Sure thing, your RAG approach sounds intriguing, especially since you're sidestepping vector databases. But doesn't the input context length cap affect it? (chatgpt plus at 32K [0] or gpt4 via open ai at 128K [1]) Seems like those cases would be pretty rare though.
[0]: https://openai.com/chatgpt/pricing#:~:text=8K-,32K,-32K
[1]: https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turb...