Doc2LoRA provides decodable representations of scientific ideas

We construct a map of scientific papers where every point is a large language model, open to questions about the encoded ideas, including regions with no existing papers.

Chand Sahil Mansuri*, Joel Zachariah*, Sadamori Kojaku
School of Systems Science and Industrial Engineering, Binghamton University
*Equal contribution

Every point becomes a model

We use Doc-to-LoRA, a hypernetwork that reads a document and creates a LoRA adapter: a small module that changes the output of a language model as if it had read the document.

While large language models have many parameters, the embedding is compact as it contains only the information about the document. It offers idea arithmetic: any point—single document, group average, or an interpolation—can be turned into a LoRA adapter and loaded into a language model for direct question answering.

Doc2LoRA pipeline: text is encoded into an embedding, reshaped by a reversible function, and decoded to a LoRA adapter. Blended embeddings create blended adapters.
Figure 1. Doc2LoRA encodes a document to an embedding. A reversible transform g reshapes embeddings for search. Any embedding, including blends, is decoded to a LoRA adapter for the language model.

For example: we embedded two recipes, Italian cacio e pepe and Japanese kake udon. The Italian adapter produces “Pecorino Romano” as the cheese. Halfway between the two, the model describes “Kake Udon with Creamy Parmesan Sauce”—mixing noodles from one and cheese from the other.

A bowl of cacio e pepe: thin pasta in a glossy Pecorino and black pepper sauce.Cacio e pepe
A bowl of thick udon noodles coated in a creamy Parmesan sauce, topped with grated cheese.Halfway: Kake Udon with Creamy Parmesan Sauce
A bowl of kake udon: thick udon noodles in a clear dashi broth with sliced green onion.Kake udon

Decoded at the midpoint, verbatim. Prompt: “Write out this recipe: give the dish a name, list the ingredients, and describe the preparation steps.”

Recipe Name: Kake Udon with Creamy Parmesan Sauce Ingredients: - 200 g fresh udon noodles - 2 tablespoons unsalted butter - 2 tablespoons mirin - 2 tablespoons soy sauce - 1 teaspoon wasabi paste - 1/2 cup grated Parmesan cheese - 1/2 cup heavy cream - 1 spring onion, thinly sliced - Optional: a pinch of freshly ground black pepper Preparation Steps: 1. Cook the fresh udon noodles according to the package instructions. Drain and rinse with cold water. 2. In a deep skillet or pan, melt the unsalted butter over medium heat. 3. Add the mirin and cook fsr 1 minute. 4. Stir in the soy sauce and wasabi paste. 5. Gradually add the grated Parmesan cheese, stirring constantly to prevent lumps from forming. 6. Pour in the heavy cream and bring to a simmer. 7. Cook for 2 minutes, stirring occasionally, until the sauce has thickened. 8. Add the cooked udon noodles to the skillet and toss to coat them evenly with the sauce. 9. Remove from heat and top with the sliced spring onion. 10. Optional: finish with a pinch of freshly ground black pepper. 11. Serve immediately.
The two recipes and the dish decoded halfway between them. The pictures are illustrations generated with Gemini from each dish's description. Doc2LoRA produces the text, not the pictures.

Map generation results

We evaluated on American Physical Society physics papers labeled by PACS, a hierarchical topic scheme for organizing fields and subfields.

Naming groups of papers

For each topic, we averaged paper embeddings and asked the model to generate a 2-3 word label. We compared these outputs to official PACS names and five baselines (including keyword extraction and vector-to-text decoders). Doc2LoRA's labels best matched the official names by word overlap and were preferred in language model panel comparisons.

Violin plot of word overlap with the official topic names for seven methods. Doc2LoRA averages .64, ahead of ICAE .57, in-context prompting .50, KeyLLM .31, BERTopic .30, vec2text .30 and T2L .28. Heat map of head-to-head win rates between the seven methods, judged by five language models. Doc2LoRA has the highest row mean, .76.
Figure 2. Naming 28 physics topics. Left: word overlap between each method's label and the official name (1 = identical). Each dot is a topic; diamond is the mean. Right: panel win rate: how often five language models prefer row over column; 0.5 = tie. D2L: Doc2LoRA; In-ctx: language model receives papers directly; ICAE; KeyLLM; BERTopic; vec2text; T2L: Text-to-LoRA.

Shorter vectors, broader ideas

Averaging divergent papers shortens the embedding vector. Shorter vectors tend to be more generic, broader categories, and the level of generality can be controlled by shrinking the vector length (right). Shrinking a Wikipedia article's vector shifts its label from Citric acid cycle metabolism to Cellular respiration. Near zero, the model defaults to its base answer.

Averaged vectors of physics topics placed by vector length, in three rows for field, division and subdivision, with decoded labels such as Particle physics, Condensed matter physics and Quantum information. Four Wikipedia articles with their vectors shrunk step by step. Decoded labels become broader as the vector shortens, for example from Citric acid cycle metabolism to Cellular respiration.
Figure 3. Shorter vectors make labels broader. Left: Each row shows average vectors for physics topics by hierarchy level; circle size shows number of papers; orange labels are decoded, grey are official where different. Right: Four Wikipedia article vectors are progressively shrunk; labels update at each step. In the shaded region, the base model output appears.

Reading the space between two papers

We interpolated between two physics papers and, at each point, asked the model to write an abstract. Move the slider to shift the weight from paper A to paper B. Doc2LoRA's abstract follows the weight and combines terms from both papers near the switch, while a language model given both papers and the weight in its prompt blends them at every weight.

Paper A Paper B Weight on paper B: 5/12
Doc2LoRA
16 paper-A words · 13 paper-B words

This research investigates the learning of quantum gates in a unitary model of quantum computation, where a quantum system evolves under a unitary transformation and is trained using a quantum neural network. The problem involves determining whether a quantum system can learn a target unitary transformation—possibly imperfect—by iteratively adjusting a set of quantum gates, with the goal of achieving a high fidelity of the learned transformation. The study employs a statistical mechanics approach, specifically the replica method and annealing techniques, to analyze the learning process in the large-N limit, where N represents the number of qubits in the system. The key finding is that the learning process is governed by a critical temperature T, and the asymptotic error rate scales as 1/N² for smooth unitaries, with a universal scaling law that depends on the smoothness of the unitary transformation. For nonsmooth unitaries, the error rate can exhibit a more complex, potentially divergent behavior, and the learning process may fail to converge to a high-fidelity state.

In-context
29 paper-A words · 14 paper-B words

This research investigates the statistical-mechanical behavior of quantum neural networks under the influence of both stochastic training and quantum decoherence, blending the framework of learning from examples (A) with the principles of quantum error correction (B) in a 58:42 proportion. The central problem is to understand how the generalization performance of a quantum neural network—trained on random examples—degrades or improves in the presence of arbitrary decoherence affecting up to t qubits, and whether the statistical-mechanical predictions of generalization curves (e.g., inverse power laws or discontinuous transitions) remain valid in a quantum setting. Using replica theory and the annealed approximation, we analyze the average generalization error as a function of training examples and network size, while incorporating quantum error-correcting codes that protect the encoded information against decoherence. Our findings reveal that for realizable learning tasks, the high-temperature and annealed approximations still provide accurate descriptions of generalization, even when quantum noise is present, but for unrealizable rules, the system exhibits a phase transition to a spin-glass-like state with degenerate minima, analogous to classical perceptrons, yet now stabilized by quantum error correction. We propose a classification of asymptotic learning curves in quantum neural networks, showing that the…

Table 1. Abstracts decoded along the path from paper A, Statistical mechanics of learning from examples, to paper B, Good quantum error-correcting codes exist, two papers from the same physics subtopic. Each cell shows the full decoded abstract; “…” marks where in-context prompting reached the length limit. Words that appear in only one of the two source abstracts are highlighted: paper A’s in blue, paper B’s in orange. The bar above each column shows their ratio. In-context: a language model given both papers and the mixing weight in its prompt.

Adding g greatly improves search

Doc2LoRA is designed for text generation, not ranking. Adding a small reversible transform, g, trained on 42,332 citation pairs from OpenAlex, improves its search results to the level of the commonly used text embeddings. g is the only trained part and preserves reversibility, so all points can be decoded as before.

We evaluated Doc2LoRA (with and without g) against five common text encoders on 14 benchmarks covering four tasks: next-paper prediction, topic classification, collaboration prediction, and author disambiguation. Without g, Doc2LoRA ranks last. With g, it performs on par with EmbeddingGemma and GTE, ahead of Instructor and SPECTER2, and behind SBERT.

Doc2LoRA without g Doc2LoRA with g

Swipe sideways to see the whole chart.

Next-paper prediction (AUC) gain from g .80 .85 .90 .95 Economics SBERT: .937 GTE: .910 EmbeddingGemma: .906 Instructor: .877 SPECTER2: .903 Doc2LoRA without g: .813 Doc2LoRA with g: .910 +.097 Psychology SBERT: .937 GTE: .912 EmbeddingGemma: .910 Instructor: .888 SPECTER2: .908 Doc2LoRA without g: .799 Doc2LoRA with g: .923 +.124 Physics SBERT: .956 GTE: .962 EmbeddingGemma: .947 Instructor: .908 SPECTER2: .941 Doc2LoRA without g: .880 Doc2LoRA with g: .958 +.078 Topic classification (macro-F1) .20 .30 .40 .50 .60 Economics SBERT: .418 GTE: .398 EmbeddingGemma: .362 Instructor: .379 SPECTER2: .344 Doc2LoRA without g: .242 Doc2LoRA with g: .374 +.132 Psychology SBERT: .366 GTE: .350 EmbeddingGemma: .358 Instructor: .348 SPECTER2: .319 Doc2LoRA without g: .217 Doc2LoRA with g: .340 +.123 Physics SBERT: .573 GTE: .587 EmbeddingGemma: .588 Instructor: .574 SPECTER2: .598 Doc2LoRA without g: .590 Doc2LoRA with g: .590 +.000 Collaboration prediction (AUC) .60 .70 .80 Economics SBERT: .655 GTE: .627 EmbeddingGemma: .633 Instructor: .620 SPECTER2: .632 Doc2LoRA without g: .596 Doc2LoRA with g: .629 +.033 Psychology SBERT: .691 GTE: .668 EmbeddingGemma: .666 Instructor: .649 SPECTER2: .665 Doc2LoRA without g: .602 Doc2LoRA with g: .687 +.085 Physics SBERT: .832 GTE: .839 EmbeddingGemma: .828 Instructor: .854 SPECTER2: .826 Doc2LoRA without g: .792 Doc2LoRA with g: .795 +.003 Author-name disambiguation (B³ F1) .70 .80 .90 zbMATH SBERT: .944 GTE: .939 EmbeddingGemma: .934 Instructor: .936 SPECTER2: .934 Doc2LoRA without g: .933 Doc2LoRA with g: .932 −.001 QIAN SBERT: .865 GTE: .866 EmbeddingGemma: .832 Instructor: .832 SPECTER2: .832 Doc2LoRA without g: .787 Doc2LoRA with g: .850 +.063 ArnetMiner SBERT: .698 GTE: .698 EmbeddingGemma: .706 Instructor: .679 SPECTER2: .669 Doc2LoRA without g: .678 Doc2LoRA with g: .708 +.030 PubMed SBERT: .832 GTE: .792 EmbeddingGemma: .822 Instructor: .801 SPECTER2: .739 Doc2LoRA without g: .693 Doc2LoRA with g: .815 +.122 KISTI SBERT: .793 GTE: .785 EmbeddingGemma: .771 Instructor: .754 SPECTER2: .746 Doc2LoRA without g: .708 Doc2LoRA with g: .767 +.059
Figure 4. Scores on each benchmark; higher is better. Grey markers are the five text encoders, and the grey band spans the worst to the best of them. The line from the open to the filled circle is what g adds, and the number at the right is the size of that gain. Thin whiskers on the Doc2LoRA points show one bootstrap standard deviation and are mostly hidden behind the circles. Each axis is cropped to the range of the scores.
All scores
TaskField / datasetDoc2LoRAText encoders
without gwith gSBERTGTEEmbeddingGemmaInstructorSPECTER2
Next-paper predictionAUCEconomics.813±.001.910±.001.937±.001.910±.001.906±.001.877±.001.903±.001
Psychology.799±.001.923±.001.937±.001.912±.001.910±.001.888±.001.908±.001
Physics.880±.001.958±.001.956±.001.962±.001.947±.001.908±.001.941±.001
Topic classificationmacro-F1Economics.242±.015.374±.022.418±.022.398±.023.362±.021.379±.021.344±.021
Psychology.217±.010.340±.020.366±.023.350±.019.358±.020.348±.026.319±.020
Physics.590±.012.590±.009.573±.011.587±.010.588±.010.574±.010.598±.007
Collaboration predictionAUCEconomics.596±.011.629±.012.655±.011.627±.012.633±.012.620±.012.632±.011
Psychology.602±.012.687±.010.691±.011.668±.011.666±.011.649±.011.665±.011
Physics.792±.004.795±.004.832±.003.839±.003.828±.003.854±.003.826±.004
Author-name disambiguationB³ F1zbMATH.933±.001.932±.001.944±.001.939±.001.934±.002.936±.001.934±.001
QIAN.787±.005.850±.004.865±.004.866±.004.832±.004.832±.004.832±.004
ArnetMiner.678±.004.708±.005.698±.005.698±.004.706±.004.679±.004.669±.004
PubMed.693±.007.815±.006.832±.006.792±.006.822±.006.801±.006.739±.006
KISTI.708±.002.767±.002.793±.002.785±.002.771±.002.754±.002.746±.002
Mean ± standard deviation over 1,000 bootstrap resamples. Bold marks the best score in each row. For author-name disambiguation, SPECTER replaces SPECTER2.

Where this goes next

Traditional maps of science show only published work. Here, we built a new map that also covers unexplored areas—spaces between fields and combinations of ideas that haven't been tried. We can target gaps directly: What problem would a paper here address? What methods and experiments could bridge the divide?

This enables new questions: Can we trace a path across disciplines and see how concepts connect? Can we limit the map to papers up to a certain date, generate ideas for the "knowledge holes," and later see if those predictions appear in new research?

We must emphasize that our map reveals potential ideas, but verifying their validity is left for future investigation. The map does not assess how promising an idea is. That judgment is up to people. Yet, we believe that the Doc2LoRA embedding is a powerful tool for generating novel research questions and invites scientific curiosity.

Citation

@misc{mansuri2026doc2lora,
  title         = {Doc2LoRA Provides Decodable Representations of Scientific Ideas},
  author        = {Mansuri, Chand Sahil and Zachariah, Joel and Kojaku, Sadamori},
  year          = {2026},
  eprint        = {2609.38374},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2609.38374}
}