Understanding Embeddings, Vector Databases, and Cosine Similarity in RAG-LLM Systems
A Technical Research Paper on Semantic Search, Vector Representation, Retrieval-Augmented Generation, and the Foundations of GraphRAG
Abstract
Retrieval-Augmented Generation (RAG) has become an important architecture for building Large Language Model (LLM) applications that require access to private, specialized, current, or domain-specific information.
At the heart of many RAG systems is a deceptively simple idea:
Convert information into numerical vectors, store those vectors in a vector database, and retrieve information whose vectors are mathematically similar to the user's query.
This process involves several important technologies and mathematical concepts, including embeddings, vector spaces, high-dimensional representations, similarity metrics, cosine similarity, nearest-neighbor search, indexing, chunking, vector databases, and retrieval pipelines.
Understanding these concepts is essential for engineers developing RAG applications using platforms such as RAGFlow, Neo4j, LlamaIndex, LangChain, Haystack, OpenSearch, Elasticsearch, and dedicated vector databases.
This paper explains these concepts progressively, beginning with ordinary text and ending with a complete RAG retrieval pipeline. It also explains why cosine similarity is commonly used, how embeddings represent semantic meaning, how vector databases perform similarity search, why chunking matters, and what limitations exist when semantic similarity alone is insufficient.
The paper then establishes the foundation for understanding hybrid RAG, GraphRAG, and Agentic GraphRAG.
1. Introduction
Large Language Models such as GPT-class models can generate remarkably sophisticated responses. However, an LLM does not automatically know the contents of an organization's private documents, internal databases, service manuals, product catalogs, engineering reports, or recently created information.
A common solution is Retrieval-Augmented Generation.
A simplified RAG system works as follows:
User Question
↓
Convert question into an embedding
↓
Search a vector database
↓
Retrieve relevant document chunks
↓
Give retrieved context to the LLM
↓
Generate an answer
The critical question therefore becomes:
How does a computer determine that two pieces of text are semantically similar?
The answer begins with embeddings.
2. What Is an Embedding?
An embedding is a numerical representation of an object.
In RAG systems, the object is commonly:
- a word
- a sentence
- a paragraph
- a document chunk
- a question
- an image
- source code
- a product description
- a database record
An embedding converts that object into a vector of numbers.
For example, a simplified embedding might look like:
"engine temperature is too high" → [0.21, -0.13, 0.87, 0.42, -0.31, ...]
Real embedding vectors generally contain hundreds or thousands of dimensions.
For example:
text ↓ embedding model ↓ [0.12, -0.43, 0.77, 0.05, ...] ↑ hundreds/thousands of dimensions
The individual numbers generally do not have simple human-readable meanings such as:
dimension 1 = engine dimension 2 = temperature dimension 3 = vehicle
Instead, the vector represents patterns learned by the embedding model.
3. Why Convert Text into Numbers?
Computers can compare numerical vectors mathematically.
Suppose we have:
Query A: "engine temperature is high" Query B: "the engine is overheating"
Although the words are different, humans recognize that the two statements are semantically related.
A good embedding model attempts to place these two texts relatively close together in vector space.
By contrast:
Query C: "the company's accounting department is hiring"
would generally occupy a different region of semantic space.
Therefore:
Similar meaning ↓ Similar vector representation ↓ Small vector distance / high similarity
This is the fundamental idea behind semantic search.
4. Vector Representation
A vector is an ordered list of numerical values.
For example:
A = [2, 4, 6]
and
B = [1, 3, 5]
These are three-dimensional vectors.
In a real embedding system, the vectors may have hundreds or thousands of dimensions.
For example:
Embedding: [0.182, -0.442, 0.731, 0.119, ... 0.038]
If the embedding has 768 dimensions, then each piece of text is represented by a point in a 768-dimensional vector space.
Humans cannot visualize 768 dimensions directly, but mathematics can operate on them.
5. The Vector Space
Think of a simple two-dimensional coordinate system:
y ↑ | B | A | ------------------+----------------→ x | C | |
Each object is represented by a point.
In an embedding space, however, we might have:
768 dimensions 1024 dimensions 1536 dimensions 3072 dimensions
depending on the embedding model.
The important concept is not visualizing all dimensions.
The important concept is:
Semantically related objects tend to have useful geometric relationships in the embedding space.
6. Embedding Models
An embedding model converts input data into vectors.
The general pipeline is:
Text ↓ Tokenizer ↓ Neural Network ↓ Semantic Representation ↓ Embedding Vector
For example:
"How do I diagnose P0171?" ↓ Embedding Model ↓ [0.17, -0.42, 0.83, ...]
The embedding model is different from the LLM that eventually generates the answer.
This distinction is important.
Embedding model
Purpose:
Convert information into numerical representations suitable for retrieval or similarity comparison.
Generative LLM
Purpose:
Generate natural-language responses based on instructions and context.
A RAG system therefore commonly contains at least two AI components:
┌───────────────┐ │ Embedding │ │ Model │ └───────┬───────┘ │ ↓ Vector Search │ ↓ Retrieved Context │ ↓ ┌───────────────┐ │ Generative │ │ LLM │ └───────────────┘
7. Document Embedding
Suppose an organization has a service manual.
The original document might contain:
Chapter 4 Fuel System Diagnostics A lean condition may be caused by: - vacuum leaks - insufficient fuel pressure - injector problems - inaccurate airflow measurement
A RAG system generally does not embed the entire book as one giant vector.
Instead, it performs document chunking.
For example:
Document ↓ Chunk 1 Chunk 2 Chunk 3 Chunk 4 ... Chunk N
Each chunk is then converted into an embedding.
Chunk 1 → Vector 1 Chunk 2 → Vector 2 Chunk 3 → Vector 3 ... Chunk N → Vector N
These vectors can then be stored in a vector database.
8. Why Chunking Matters
Chunking is one of the most important design decisions in RAG.
Suppose a 500-page service manual is treated as one piece of text.
Its embedding may represent a very broad mixture of concepts.
The retrieval system may then have difficulty identifying the exact section required to answer a question.
Instead:
500-page manual ↓ meaningful chunks ↓ individual embeddings ↓ precise retrieval
Good chunking attempts to preserve enough context while keeping individual chunks focused.
Possible strategies include:
- fixed-size chunks
- paragraph-based chunks
- sentence-based chunks
- semantic chunks
- heading-aware chunks
- document-structure-aware chunks
- overlapping chunks
9. The User Query Also Becomes a Vector
Suppose the user asks:
"What can cause an engine to run too lean?"
The query is passed through the same or compatible embedding model.
User Question ↓ Embedding Model ↓ Query Vector
For example:
Q = [0.21, -0.44, 0.76, ...]
The vector database now needs to find document vectors that are most similar to Q.
This is where similarity metrics become important.
10. What Is Similarity?
Similarity is a mathematical measurement indicating how closely two vectors are related according to a particular metric.
Several approaches are commonly used:
- cosine similarity
- Euclidean distance
- Manhattan distance
- dot product / inner product
For many text-embedding applications, cosine similarity is particularly useful.
11. Cosine Similarity
Cosine similarity measures the angle between two vectors rather than primarily measuring their absolute length.
The formula is:
cosine similarity(A,B)=A⋅B∥A∥∥B∥\text{cosine similarity}(A,B) = \frac{A\cdot B} {\|A\|\|B\|}
where:
- AA = vector A
- BB = vector B
- A⋅BA\cdot B = dot product
- ∥A∥\|A\| = magnitude of vector A
- ∥B∥\|B\| = magnitude of vector B
The fundamental idea is:
If two vectors point in approximately the same direction, their cosine similarity is high.
12. Dot Product
The dot product is calculated by multiplying corresponding components and adding the results.
For:
A=[a1,a2,a3]A=[a_1,a_2,a_3]
and:
B=[b1,b2,b3]B=[b_1,b_2,b_3]
the dot product is:
A⋅B=a1b1+a2b2+a3b3A\cdot B=a_1b_1+a_2b_2+a_3b_3
For example:
A=[1,2,3]A=[1,2,3] B=[4,5,6]B=[4,5,6]
Then:
A⋅B=(1)(4)+(2)(5)+(3)(6)A\cdot B = (1)(4)+(2)(5)+(3)(6) =4+10+18=4+10+18 =32=32
The dot product is therefore 32.
13. Vector Magnitude
The magnitude, or length, of a vector is:
∥A∥=a12+a22+a32\|A\|=\sqrt{a_1^2+a_2^2+a_3^2}
For:
A=[1,2,3]A=[1,2,3]
we get:
∥A∥=12+22+32\|A\|=\sqrt{1^2+2^2+3^2} =14=\sqrt{14}
The magnitude is approximately:
3.7423.742
14. Calculating Cosine Similarity
Consider:
A=[1,2,3]A=[1,2,3]
and:
B=[4,5,6]B=[4,5,6]
We already calculated:
A⋅B=32A\cdot B=32
The magnitudes are:
∥A∥=14\|A\|=\sqrt{14}
and:
∥B∥=77\|B\|=\sqrt{77}
Therefore:
cosine similarity=321477\text{cosine similarity} = \frac{32} {\sqrt{14}\sqrt{77}}
which is approximately:
0.9750.975
A value close to 1 indicates that the vectors point in very similar directions.
15. Geometric Meaning of Cosine Similarity
Cosine similarity can be understood geometrically.
For two vectors:
B / / / θ / /________ A
The angle θ\theta between the vectors is important.
The relationship is:
cos(θ)\cos(\theta)
If:
θ=0∘\theta=0^\circ
then:
cos(0)=1\cos(0)=1
The vectors point in exactly the same direction.
If:
θ=90∘\theta=90^\circ
then:
cos(90∘)=0\cos(90^\circ)=0
The vectors are orthogonal.
If:
θ=180∘\theta=180^\circ
then:
cos(180∘)=−1\cos(180^\circ)=-1
The vectors point in opposite directions.
For modern embedding systems, the practical interpretation of scores depends on the embedding model and implementation. A score should not automatically be interpreted as a universal percentage of semantic similarity.
16. Why Cosine Similarity Is Useful for Text
Consider two vectors:
A = [10, 20, 30] B = [1, 2, 3]
They point in exactly the same direction.
Their lengths are very different, but their orientation is identical.
Cosine similarity therefore gives:
11
This can be useful because the semantic direction represented by an embedding can matter more than its absolute magnitude.
17. Normalized Embeddings
Many systems normalize vectors to unit length.
For a normalized vector:
∥A∥=1\|A\|=1
and:
∥B∥=1\|B\|=1
The cosine similarity becomes:
A⋅BA\cdot B
This means that cosine similarity and dot product can become mathematically equivalent when vectors are normalized.
This is an important implementation detail when configuring vector databases.
18. Vector Database
A vector database is a database designed to store and retrieve high-dimensional vectors efficiently.
Instead of primarily asking:
SELECT * FROM documents WHERE title LIKE '%engine%';
a vector database can perform:
Find the vectors most similar to this query vector.
A simplified database record might contain:
ID Document ID Chunk ID Text Embedding Vector Metadata
For example:
Document: ServiceManual.pdf Chunk: 428 Text: "A lean condition can result..." Vector: [0.17, -0.42, ...] Metadata: vehicle = Subaru system = Fuel chapter = Diagnostics
The vector and metadata can be stored together or in coordinated data structures.
19. Vector Search
The retrieval process looks like:
User Query ↓ Embedding Model ↓ Query Vector ↓ Vector Database ↓ Similarity Search ↓ Top-K Results
Suppose the system requests:
Top K = 5
The vector database returns the five document chunks with the highest similarity according to the selected metric.
For example:
|
Rank |
Chunk |
Similarity |
|---|---|---|
|
1 |
Fuel-system diagnosis |
0.91 |
|
2 |
Vacuum-leak diagnosis |
0.88 |
|
3 |
Fuel-pressure testing |
0.84 |
|
4 |
Injector diagnosis |
0.81 |
|
5 |
Airflow measurement |
0.78 |
These numbers are illustrative rather than universal thresholds.
20. Approximate Nearest Neighbor Search
A major challenge is scale.
Suppose a knowledge base contains:
10 million document chunks
and each embedding contains:
1,536 dimensions
Comparing the query against every vector can become computationally expensive.
Vector databases therefore use indexing techniques for efficient approximate nearest-neighbor search.
Common approaches include:
- HNSW
- IVF
- Product Quantization
- Disk-based ANN methods
- specialized vector indexes
The objective is:
Find highly relevant vectors without exhaustively comparing the query against every stored vector.
21. HNSW
HNSW stands for:
Hierarchical Navigable Small World
It creates a graph-like structure that allows efficient navigation through vector space.
Conceptually:
Layer 2 A ----------- D \ \ \ E \ / Layer 1 A -- B -- C -- D -- E -- F
Instead of examining every vector, the search navigates through the index toward promising candidates.
This is one reason vector databases can perform similarity searches over very large datasets efficiently.
22. Metadata Filtering
Vector similarity should not necessarily be the only retrieval mechanism.
Suppose a database contains manuals for:
- Toyota
- Honda
- Subaru
- Ford
- BMW
The user asks about a Subaru.
The system can combine:
Semantic similarity + Metadata filter
For example:
vehicle_make = Subaru vehicle_year = 2010 engine = 2.5L
Then similarity search operates on a more appropriate subset.
This is called filtered vector search.
23. Hybrid Search
Vector search is powerful, but exact keywords can also be important.
Consider:
"P0171"
A vector search may understand the general concept of a lean condition.
However, exact matching of the diagnostic code P0171 is also valuable.
A hybrid retrieval architecture can combine:
Keyword Search + Vector Search + Metadata Filtering + Optional Graph Search
This is often more robust than relying on only one retrieval mechanism.
24. Embeddings Do Not "Understand" Documents Like Humans
It is important not to misunderstand embeddings.
An embedding is not a compressed copy of a document in the ordinary sense.
It is a learned numerical representation.
The model has learned statistical relationships between patterns in its training data.
Consequently:
Embedding ≠ complete database of facts
This has important consequences for RAG.
A vector database should not be treated as a replacement for the original documents.
The original source should remain available for:
- evidence
- citations
- verification
- auditing
- traceability
25. The Complete RAG Pipeline
A typical RAG system can now be understood as a sequence.
Phase 1 — Ingestion
PDF DOCX HTML Web pages Database Email Manual Code
↓
Phase 2 — Parsing
Extract text and document structure.
↓
Phase 3 — Chunking
Document ↓ Chunk 1 Chunk 2 Chunk 3 ...
↓
Phase 4 — Embedding
Chunk ↓ Embedding Model ↓ Vector
↓
Phase 5 — Storage
Vector Database + Metadata + Original Document Reference
↓
Phase 6 — Query
User Question ↓ Query Embedding
↓
Phase 7 — Retrieval
Vector Similarity Search ↓ Top-K Chunks
↓
Phase 8 — Context Construction
Question + Retrieved Evidence + System Instructions
↓
Phase 9 — Generation
LLM ↓ Grounded Answer
26. Where RAGFlow Fits
RAGFlow can be viewed as a practical platform for implementing significant portions of this pipeline.
A conceptual architecture is:
┌──────────────┐ │ Documents │ └──────┬───────┘ ↓ ┌──────────────┐ │ RAGFlow │ │ Parsing │ │ Chunking │ └──────┬───────┘ ↓ ┌──────────────┐ │ Embedding │ │ Model │ └──────┬───────┘ ↓ ┌──────────────┐ │ Vector │ │ Retrieval │ └──────┬───────┘ ↓ ┌──────────────┐ │ Retrieved │ │ Context │ └──────┬───────┘ ↓ ┌──────────────┐ │ LLM │ └──────┬───────┘ ↓ Answer
For an experimental engineering environment, RAGFlow can be combined with:
- Ollama
- Hugging Face
- embedding models
- reranking models
- vector stores
- Neo4j
- Docker
- Python APIs
27. Reranking
Initial vector retrieval does not necessarily produce the final best ordering.
A RAG system can therefore use a reranker.
The process becomes:
Query ↓ Vector Search ↓ Top 20–100 candidates ↓ Reranker ↓ Top 5–10 documents ↓ LLM
The first-stage vector search is optimized for efficient retrieval.
The reranker performs a more detailed relevance assessment.
This two-stage architecture can improve retrieval quality.
28. Vector Search Has Limitations
Vector similarity is not equivalent to reasoning.
For example, two documents may be semantically similar but have opposite conclusions.
A vector database does not inherently understand:
- causality
- temporal relationships
- ownership
- dependency
- authorization
- logical consistency
- numerical correctness
- complex entity relationships
This is one reason GraphRAG becomes interesting.
29. From Vector RAG to GraphRAG
Traditional RAG:
Question ↓ Embedding ↓ Vector Search ↓ Relevant Chunks ↓ LLM
GraphRAG:
Question ↓ Entity / Concept Identification ↓ Vector Search + Graph Retrieval ↓ Relationships ↓ Relevant Evidence ↓ LLM
The graph adds an explicit representation of relationships.
For example:
[DTC] | ↓ [Subsystem] | ↓ [Component] | ↓ [Test] | ↓ [Repair Procedure]
Vector search finds relevant text.
The graph helps identify how the entities relate.
30. From GraphRAG to Agentic GraphRAG
The next architectural step is to introduce agents.
Instead of executing a fixed retrieval pipeline, an agent can decide which tools to use.
User ↓ AI Agent ├── Vector Search ├── Graph Search ├── Web Search ├── SQL ├── APIs ├── Knowledge Base └── Calculation ↓ Evidence ↓ Reasoning ↓ Verification ↓ LLM Response
This architecture is particularly useful for complex engineering and enterprise applications.
31. Example: OBD-AI
Consider an automotive diagnostic system.
The user asks:
"Why is my vehicle reporting P0171?"
The system may perform:
Question ↓ Embedding ↓ Vector Search ↓ Service Manual Chunks
The graph may contain:
P0171 ↓ Lean Condition ↓ Fuel System ├── Vacuum Leak ├── Fuel Injector ├── MAF Sensor └── Fuel Pressure
The agent could then retrieve:
Diagnostic procedure + Vehicle-specific service manual + Sensor information + Previous diagnostic results
The LLM generates the final response using the retrieved evidence.
This illustrates the progression:
Embedding → Vector Database → RAG → GraphRAG → Agentic GraphRAG
32. Important Engineering Parameters
When designing a vector RAG system, engineers should evaluate:
Embedding model
- embedding dimensions
- domain suitability
- multilingual support
- model quality
- computational requirements
Chunking
- chunk size
- overlap
- semantic boundaries
- document structure
Retrieval
- similarity metric
- top-K
- metadata filters
- hybrid search
- reranking
Vector database
- scale
- indexing
- persistence
- filtering
- distributed operation
- latency
LLM
- context window
- reasoning capability
- latency
- cost
- local versus cloud deployment
Evaluation
- retrieval precision
- retrieval recall
- answer faithfulness
- citation accuracy
- latency
- cost
33. A Practical SME Architecture
For an SME research and engineering environment, an open-source architecture could look like:
User | ↓ Application | ↓ RAG / Agent / \ / \ ↓ ↓ Vector Search Graph Search | | ↓ ↓ Vector DB Neo4j \ / \ / ↓ ↓ Context | ↓ LLM | ↓ Answer
A possible implementation stack is:
Docker + RAGFlow + Ollama + Hugging Face + Neo4j + Python + FastAPI + MCP
The exact stack should be selected according to the application's requirements rather than assuming that one product is appropriate for every use case.
34. Key Conceptual Distinction
The following distinction is fundamental:
Embedding
A numerical representation of information.
Vector
The numerical representation itself.
Vector database
A system optimized for storing and retrieving vectors.
Similarity metric
A mathematical method for comparing vectors.
Cosine similarity
A metric based on the angle between vectors.
RAG
An architecture that retrieves external information before generating an LLM response.
GraphRAG
RAG that incorporates graph-based entities and relationships into retrieval.
Agentic GraphRAG
A graph-enhanced RAG architecture in which AI agents dynamically orchestrate retrieval, tools, reasoning, and verification.
35. Conclusion
Embeddings and vector databases form one of the fundamental technical foundations of modern RAG systems.
The overall chain is:
Text
→ Embedding Model
→ Vector
→ Vector Database
→ Similarity Search
→ Relevant Context
→ LLM
→ Grounded Response
Cosine similarity provides an intuitive mathematical mechanism for comparing the orientation of embedding vectors. Vector databases extend this concept into scalable retrieval systems using specialized indexing and nearest-neighbor search techniques.
However, vector similarity alone does not solve every knowledge-retrieval problem.
As applications become more complex, additional capabilities become useful:
Vector RAG
↓
Hybrid RAG
↓
GraphRAG
↓
Agentic RAG
↓
Agentic GraphRAG
The progression reflects a broader movement from simple semantic document retrieval toward systems capable of combining semantic representations, structured relationships, tools, domain knowledge, reasoning, and autonomous retrieval strategies.
For engineering applications such as OBD-AI, enterprise knowledge systems, cybersecurity, technical documentation, and research assistants, understanding the mathematics and engineering principles behind embeddings and vector search is therefore essential before designing more advanced GraphRAG and agentic architectures.