Understanding Embeddings, Vector Databases, and Cosine Similarity in RAG-LLM Systems

A Technical Research Paper on Semantic Search, Vector Representation, Retrieval-Augmented Generation, and the Foundations of GraphRAG

Abstract

Retrieval-Augmented Generation (RAG) has become an important architecture for building Large Language Model (LLM) applications that require access to private, specialized, current, or domain-specific information.

At the heart of many RAG systems is a deceptively simple idea:

Convert information into numerical vectors, store those vectors in a vector database, and retrieve information whose vectors are mathematically similar to the user's query.

This process involves several important technologies and mathematical concepts, including embeddings, vector spaces, high-dimensional representations, similarity metrics, cosine similarity, nearest-neighbor search, indexing, chunking, vector databases, and retrieval pipelines.

Understanding these concepts is essential for engineers developing RAG applications using platforms such as RAGFlow, Neo4j, LlamaIndex, LangChain, Haystack, OpenSearch, Elasticsearch, and dedicated vector databases.

This paper explains these concepts progressively, beginning with ordinary text and ending with a complete RAG retrieval pipeline. It also explains why cosine similarity is commonly used, how embeddings represent semantic meaning, how vector databases perform similarity search, why chunking matters, and what limitations exist when semantic similarity alone is insufficient.

The paper then establishes the foundation for understanding hybrid RAG, GraphRAG, and Agentic GraphRAG.

1. Introduction

Large Language Models such as GPT-class models can generate remarkably sophisticated responses. However, an LLM does not automatically know the contents of an organization's private documents, internal databases, service manuals, product catalogs, engineering reports, or recently created information.

A common solution is Retrieval-Augmented Generation.

A simplified RAG system works as follows:

User Question

↓

Convert question into an embedding

↓

Search a vector database

↓

Retrieve relevant document chunks

↓

Give retrieved context to the LLM

↓

Generate an answer

The critical question therefore becomes:

How does a computer determine that two pieces of text are semantically similar?

The answer begins with embeddings.

2. What Is an Embedding?

An embedding is a numerical representation of an object.

In RAG systems, the object is commonly:

  • a word
  • a sentence
  • a paragraph
  • a document chunk
  • a question
  • an image
  • source code
  • a product description
  • a database record

An embedding converts that object into a vector of numbers.

For example, a simplified embedding might look like:

"engine temperature is too high" → [0.21, -0.13, 0.87, 0.42, -0.31, ...]

Real embedding vectors generally contain hundreds or thousands of dimensions.

For example:

text ↓ embedding model ↓ [0.12, -0.43, 0.77, 0.05, ...] ↑ hundreds/thousands of dimensions

The individual numbers generally do not have simple human-readable meanings such as:

dimension 1 = engine dimension 2 = temperature dimension 3 = vehicle

Instead, the vector represents patterns learned by the embedding model.

3. Why Convert Text into Numbers?

Computers can compare numerical vectors mathematically.

Suppose we have:

Query A: "engine temperature is high" Query B: "the engine is overheating"

Although the words are different, humans recognize that the two statements are semantically related.

A good embedding model attempts to place these two texts relatively close together in vector space.

By contrast:

Query C: "the company's accounting department is hiring"

would generally occupy a different region of semantic space.

Therefore:

Similar meaning ↓ Similar vector representation ↓ Small vector distance / high similarity

This is the fundamental idea behind semantic search.

4. Vector Representation

A vector is an ordered list of numerical values.

For example:

A = [2, 4, 6]

and

B = [1, 3, 5]

These are three-dimensional vectors.

In a real embedding system, the vectors may have hundreds or thousands of dimensions.

For example:

Embedding: [0.182, -0.442, 0.731, 0.119, ... 0.038]

If the embedding has 768 dimensions, then each piece of text is represented by a point in a 768-dimensional vector space.

Humans cannot visualize 768 dimensions directly, but mathematics can operate on them.

5. The Vector Space

Think of a simple two-dimensional coordinate system:

y ↑ | B | A | ------------------+----------------→ x | C | |

Each object is represented by a point.

In an embedding space, however, we might have:

768 dimensions 1024 dimensions 1536 dimensions 3072 dimensions

depending on the embedding model.

The important concept is not visualizing all dimensions.

The important concept is:

Semantically related objects tend to have useful geometric relationships in the embedding space.

6. Embedding Models

An embedding model converts input data into vectors.

The general pipeline is:

Text ↓ Tokenizer ↓ Neural Network ↓ Semantic Representation ↓ Embedding Vector

For example:

"How do I diagnose P0171?" ↓ Embedding Model ↓ [0.17, -0.42, 0.83, ...]

The embedding model is different from the LLM that eventually generates the answer.

This distinction is important.

Embedding model

Purpose:

Convert information into numerical representations suitable for retrieval or similarity comparison.

Generative LLM

Purpose:

Generate natural-language responses based on instructions and context.

A RAG system therefore commonly contains at least two AI components:

┌───────────────┐ │ Embedding │ │ Model │ └───────┬───────┘ │ ↓ Vector Search │ ↓ Retrieved Context │ ↓ ┌───────────────┐ │ Generative │ │ LLM │ └───────────────┘

7. Document Embedding

Suppose an organization has a service manual.

The original document might contain:

Chapter 4 Fuel System Diagnostics A lean condition may be caused by: - vacuum leaks - insufficient fuel pressure - injector problems - inaccurate airflow measurement

A RAG system generally does not embed the entire book as one giant vector.

Instead, it performs document chunking.

For example:

Document ↓ Chunk 1 Chunk 2 Chunk 3 Chunk 4 ... Chunk N

Each chunk is then converted into an embedding.

Chunk 1 → Vector 1 Chunk 2 → Vector 2 Chunk 3 → Vector 3 ... Chunk N → Vector N

These vectors can then be stored in a vector database.

8. Why Chunking Matters

Chunking is one of the most important design decisions in RAG.

Suppose a 500-page service manual is treated as one piece of text.

Its embedding may represent a very broad mixture of concepts.

The retrieval system may then have difficulty identifying the exact section required to answer a question.

Instead:

500-page manual ↓ meaningful chunks ↓ individual embeddings ↓ precise retrieval

Good chunking attempts to preserve enough context while keeping individual chunks focused.

Possible strategies include:

  • fixed-size chunks
  • paragraph-based chunks
  • sentence-based chunks
  • semantic chunks
  • heading-aware chunks
  • document-structure-aware chunks
  • overlapping chunks

9. The User Query Also Becomes a Vector

Suppose the user asks:

"What can cause an engine to run too lean?"

The query is passed through the same or compatible embedding model.

User Question ↓ Embedding Model ↓ Query Vector

For example:

Q = [0.21, -0.44, 0.76, ...]

The vector database now needs to find document vectors that are most similar to Q.

This is where similarity metrics become important.

10. What Is Similarity?

Similarity is a mathematical measurement indicating how closely two vectors are related according to a particular metric.

Several approaches are commonly used:

  • cosine similarity
  • Euclidean distance
  • Manhattan distance
  • dot product / inner product

For many text-embedding applications, cosine similarity is particularly useful.

11. Cosine Similarity

Cosine similarity measures the angle between two vectors rather than primarily measuring their absolute length.

The formula is:

cosine similarity(A,B)=A⋅B∥A∥∥B∥\text{cosine similarity}(A,B) = \frac{A\cdot B} {\|A\|\|B\|}

where:

  • AA = vector A
  • BB = vector B
  • A⋅BA\cdot B = dot product
  • ∥A∥\|A\| = magnitude of vector A
  • ∥B∥\|B\| = magnitude of vector B

The fundamental idea is:

If two vectors point in approximately the same direction, their cosine similarity is high.

12. Dot Product

The dot product is calculated by multiplying corresponding components and adding the results.

For:

A=[a1,a2,a3]A=[a_1,a_2,a_3]

and:

B=[b1,b2,b3]B=[b_1,b_2,b_3]

the dot product is:

A⋅B=a1b1+a2b2+a3b3A\cdot B=a_1b_1+a_2b_2+a_3b_3

For example:

A=[1,2,3]A=[1,2,3] B=[4,5,6]B=[4,5,6]

Then:

A⋅B=(1)(4)+(2)(5)+(3)(6)A\cdot B = (1)(4)+(2)(5)+(3)(6) =4+10+18=4+10+18 =32=32

The dot product is therefore 32.

13. Vector Magnitude

The magnitude, or length, of a vector is:

∥A∥=a12+a22+a32\|A\|=\sqrt{a_1^2+a_2^2+a_3^2}

For:

A=[1,2,3]A=[1,2,3]

we get:

∥A∥=12+22+32\|A\|=\sqrt{1^2+2^2+3^2} =14=\sqrt{14}

The magnitude is approximately:

3.7423.742

14. Calculating Cosine Similarity

Consider:

A=[1,2,3]A=[1,2,3]

and:

B=[4,5,6]B=[4,5,6]

We already calculated:

A⋅B=32A\cdot B=32

The magnitudes are:

∥A∥=14\|A\|=\sqrt{14}

and:

∥B∥=77\|B\|=\sqrt{77}

Therefore:

cosine similarity=321477\text{cosine similarity} = \frac{32} {\sqrt{14}\sqrt{77}}

which is approximately:

0.9750.975

A value close to 1 indicates that the vectors point in very similar directions.

15. Geometric Meaning of Cosine Similarity

Cosine similarity can be understood geometrically.

For two vectors:

B / / / θ / /________ A

The angle θ\theta between the vectors is important.

The relationship is:

cos⁡(θ)\cos(\theta)

If:

θ=0∘\theta=0^\circ

then:

cos⁡(0)=1\cos(0)=1

The vectors point in exactly the same direction.

If:

θ=90∘\theta=90^\circ

then:

cos⁡(90∘)=0\cos(90^\circ)=0

The vectors are orthogonal.

If:

θ=180∘\theta=180^\circ

then:

cos⁡(180∘)=−1\cos(180^\circ)=-1

The vectors point in opposite directions.

For modern embedding systems, the practical interpretation of scores depends on the embedding model and implementation. A score should not automatically be interpreted as a universal percentage of semantic similarity.

16. Why Cosine Similarity Is Useful for Text

Consider two vectors:

A = [10, 20, 30] B = [1, 2, 3]

They point in exactly the same direction.

Their lengths are very different, but their orientation is identical.

Cosine similarity therefore gives:

11

This can be useful because the semantic direction represented by an embedding can matter more than its absolute magnitude.

17. Normalized Embeddings

Many systems normalize vectors to unit length.

For a normalized vector:

∥A∥=1\|A\|=1

and:

∥B∥=1\|B\|=1

The cosine similarity becomes:

A⋅BA\cdot B

This means that cosine similarity and dot product can become mathematically equivalent when vectors are normalized.

This is an important implementation detail when configuring vector databases.

18. Vector Database

A vector database is a database designed to store and retrieve high-dimensional vectors efficiently.

Instead of primarily asking:

SELECT * FROM documents WHERE title LIKE '%engine%';

a vector database can perform:

Find the vectors most similar to this query vector.

A simplified database record might contain:

ID Document ID Chunk ID Text Embedding Vector Metadata

For example:

Document: ServiceManual.pdf Chunk: 428 Text: "A lean condition can result..." Vector: [0.17, -0.42, ...] Metadata: vehicle = Subaru system = Fuel chapter = Diagnostics

The vector and metadata can be stored together or in coordinated data structures.

19. Vector Search

The retrieval process looks like:

User Query ↓ Embedding Model ↓ Query Vector ↓ Vector Database ↓ Similarity Search ↓ Top-K Results

Suppose the system requests:

Top K = 5

The vector database returns the five document chunks with the highest similarity according to the selected metric.

For example:

Rank

Chunk

Similarity

1

Fuel-system diagnosis

0.91

2

Vacuum-leak diagnosis

0.88

3

Fuel-pressure testing

0.84

4

Injector diagnosis

0.81

5

Airflow measurement

0.78

These numbers are illustrative rather than universal thresholds.

20. Approximate Nearest Neighbor Search

A major challenge is scale.

Suppose a knowledge base contains:

10 million document chunks

and each embedding contains:

1,536 dimensions

Comparing the query against every vector can become computationally expensive.

Vector databases therefore use indexing techniques for efficient approximate nearest-neighbor search.

Common approaches include:

  • HNSW
  • IVF
  • Product Quantization
  • Disk-based ANN methods
  • specialized vector indexes

The objective is:

Find highly relevant vectors without exhaustively comparing the query against every stored vector.

21. HNSW

HNSW stands for:

Hierarchical Navigable Small World

It creates a graph-like structure that allows efficient navigation through vector space.

Conceptually:

Layer 2 A ----------- D \ \ \ E \ / Layer 1 A -- B -- C -- D -- E -- F

Instead of examining every vector, the search navigates through the index toward promising candidates.

This is one reason vector databases can perform similarity searches over very large datasets efficiently.

22. Metadata Filtering

Vector similarity should not necessarily be the only retrieval mechanism.

Suppose a database contains manuals for:

  • Toyota
  • Honda
  • Subaru
  • Ford
  • BMW

The user asks about a Subaru.

The system can combine:

Semantic similarity + Metadata filter

For example:

vehicle_make = Subaru vehicle_year = 2010 engine = 2.5L

Then similarity search operates on a more appropriate subset.

This is called filtered vector search.

23. Hybrid Search

Vector search is powerful, but exact keywords can also be important.

Consider:

"P0171"

A vector search may understand the general concept of a lean condition.

However, exact matching of the diagnostic code P0171 is also valuable.

A hybrid retrieval architecture can combine:

Keyword Search + Vector Search + Metadata Filtering + Optional Graph Search

This is often more robust than relying on only one retrieval mechanism.

24. Embeddings Do Not "Understand" Documents Like Humans

It is important not to misunderstand embeddings.

An embedding is not a compressed copy of a document in the ordinary sense.

It is a learned numerical representation.

The model has learned statistical relationships between patterns in its training data.

Consequently:

Embedding ≠ complete database of facts

This has important consequences for RAG.

A vector database should not be treated as a replacement for the original documents.

The original source should remain available for:

  • evidence
  • citations
  • verification
  • auditing
  • traceability

25. The Complete RAG Pipeline

A typical RAG system can now be understood as a sequence.

Phase 1 — Ingestion

PDF DOCX HTML Web pages Database Email Manual Code

↓

Phase 2 — Parsing

Extract text and document structure.

↓

Phase 3 — Chunking

Document ↓ Chunk 1 Chunk 2 Chunk 3 ...

↓

Phase 4 — Embedding

Chunk ↓ Embedding Model ↓ Vector

↓

Phase 5 — Storage

Vector Database + Metadata + Original Document Reference

↓

Phase 6 — Query

User Question ↓ Query Embedding

↓

Phase 7 — Retrieval

Vector Similarity Search ↓ Top-K Chunks

↓

Phase 8 — Context Construction

Question + Retrieved Evidence + System Instructions

↓

Phase 9 — Generation

LLM ↓ Grounded Answer

26. Where RAGFlow Fits

RAGFlow can be viewed as a practical platform for implementing significant portions of this pipeline.

A conceptual architecture is:

┌──────────────┐ │ Documents │ └──────┬───────┘ ↓ ┌──────────────┐ │ RAGFlow │ │ Parsing │ │ Chunking │ └──────┬───────┘ ↓ ┌──────────────┐ │ Embedding │ │ Model │ └──────┬───────┘ ↓ ┌──────────────┐ │ Vector │ │ Retrieval │ └──────┬───────┘ ↓ ┌──────────────┐ │ Retrieved │ │ Context │ └──────┬───────┘ ↓ ┌──────────────┐ │ LLM │ └──────┬───────┘ ↓ Answer

For an experimental engineering environment, RAGFlow can be combined with:

  • Ollama
  • Hugging Face
  • embedding models
  • reranking models
  • vector stores
  • Neo4j
  • Docker
  • Python APIs

27. Reranking

Initial vector retrieval does not necessarily produce the final best ordering.

A RAG system can therefore use a reranker.

The process becomes:

Query ↓ Vector Search ↓ Top 20–100 candidates ↓ Reranker ↓ Top 5–10 documents ↓ LLM

The first-stage vector search is optimized for efficient retrieval.

The reranker performs a more detailed relevance assessment.

This two-stage architecture can improve retrieval quality.

28. Vector Search Has Limitations

Vector similarity is not equivalent to reasoning.

For example, two documents may be semantically similar but have opposite conclusions.

A vector database does not inherently understand:

  • causality
  • temporal relationships
  • ownership
  • dependency
  • authorization
  • logical consistency
  • numerical correctness
  • complex entity relationships

This is one reason GraphRAG becomes interesting.

29. From Vector RAG to GraphRAG

Traditional RAG:

Question ↓ Embedding ↓ Vector Search ↓ Relevant Chunks ↓ LLM

GraphRAG:

Question ↓ Entity / Concept Identification ↓ Vector Search + Graph Retrieval ↓ Relationships ↓ Relevant Evidence ↓ LLM

The graph adds an explicit representation of relationships.

For example:

[DTC] | ↓ [Subsystem] | ↓ [Component] | ↓ [Test] | ↓ [Repair Procedure]

Vector search finds relevant text.

The graph helps identify how the entities relate.

30. From GraphRAG to Agentic GraphRAG

The next architectural step is to introduce agents.

Instead of executing a fixed retrieval pipeline, an agent can decide which tools to use.

User ↓ AI Agent ├── Vector Search ├── Graph Search ├── Web Search ├── SQL ├── APIs ├── Knowledge Base └── Calculation ↓ Evidence ↓ Reasoning ↓ Verification ↓ LLM Response

This architecture is particularly useful for complex engineering and enterprise applications.

31. Example: OBD-AI

Consider an automotive diagnostic system.

The user asks:

"Why is my vehicle reporting P0171?"

The system may perform:

Question ↓ Embedding ↓ Vector Search ↓ Service Manual Chunks

The graph may contain:

P0171 ↓ Lean Condition ↓ Fuel System ├── Vacuum Leak ├── Fuel Injector ├── MAF Sensor └── Fuel Pressure

The agent could then retrieve:

Diagnostic procedure + Vehicle-specific service manual + Sensor information + Previous diagnostic results

The LLM generates the final response using the retrieved evidence.

This illustrates the progression:

Embedding → Vector Database → RAG → GraphRAG → Agentic GraphRAG

32. Important Engineering Parameters

When designing a vector RAG system, engineers should evaluate:

Embedding model

  • embedding dimensions
  • domain suitability
  • multilingual support
  • model quality
  • computational requirements

Chunking

  • chunk size
  • overlap
  • semantic boundaries
  • document structure

Retrieval

  • similarity metric
  • top-K
  • metadata filters
  • hybrid search
  • reranking

Vector database

  • scale
  • indexing
  • persistence
  • filtering
  • distributed operation
  • latency

LLM

  • context window
  • reasoning capability
  • latency
  • cost
  • local versus cloud deployment

Evaluation

  • retrieval precision
  • retrieval recall
  • answer faithfulness
  • citation accuracy
  • latency
  • cost

33. A Practical SME Architecture

For an SME research and engineering environment, an open-source architecture could look like:

User | ↓ Application | ↓ RAG / Agent / \ / \ ↓ ↓ Vector Search Graph Search | | ↓ ↓ Vector DB Neo4j \ / \ / ↓ ↓ Context | ↓ LLM | ↓ Answer

A possible implementation stack is:

Docker + RAGFlow + Ollama + Hugging Face + Neo4j + Python + FastAPI + MCP

The exact stack should be selected according to the application's requirements rather than assuming that one product is appropriate for every use case.

34. Key Conceptual Distinction

The following distinction is fundamental:

Embedding

A numerical representation of information.

Vector

The numerical representation itself.

Vector database

A system optimized for storing and retrieving vectors.

Similarity metric

A mathematical method for comparing vectors.

Cosine similarity

A metric based on the angle between vectors.

RAG

An architecture that retrieves external information before generating an LLM response.

GraphRAG

RAG that incorporates graph-based entities and relationships into retrieval.

Agentic GraphRAG

A graph-enhanced RAG architecture in which AI agents dynamically orchestrate retrieval, tools, reasoning, and verification.

35. Conclusion

Embeddings and vector databases form one of the fundamental technical foundations of modern RAG systems.

The overall chain is:

Text

→ Embedding Model

→ Vector

→ Vector Database

→ Similarity Search

→ Relevant Context

→ LLM

→ Grounded Response

Cosine similarity provides an intuitive mathematical mechanism for comparing the orientation of embedding vectors. Vector databases extend this concept into scalable retrieval systems using specialized indexing and nearest-neighbor search techniques.

However, vector similarity alone does not solve every knowledge-retrieval problem.

As applications become more complex, additional capabilities become useful:

Vector RAG

↓

Hybrid RAG

↓

GraphRAG

↓

Agentic RAG

↓

Agentic GraphRAG

The progression reflects a broader movement from simple semantic document retrieval toward systems capable of combining semantic representations, structured relationships, tools, domain knowledge, reasoning, and autonomous retrieval strategies.

For engineering applications such as OBD-AI, enterprise knowledge systems, cybersecurity, technical documentation, and research assistants, understanding the mathematics and engineering principles behind embeddings and vector search is therefore essential before designing more advanced GraphRAG and agentic architectures.