Method and device for improving RAG recall rate in voice question and answer scene

By cleaning and compressing the speech recognition results, and combining a multi-model embedding system and semantic fidelity evaluation, the optimal embedding model was selected and a hot word list was constructed. This solved the problem of insufficient recall in the speech question answering scenario and improved the recall rate and matching accuracy of professional vocabulary in speech question answering.

CN120913552APending Publication Date: 2025-11-07ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511040109.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In voice question answering scenarios, issues such as speech transcription errors, semantic ambiguity, and entity mismatches can lead to a decrease in the semantic similarity between the query vector and related documents in the knowledge base, affecting recall performance. In particular, when using professional terms, the model is prone to identifying professional words as general words or homonyms, resulting in insufficient recall accuracy.

Method used

By semantically cleaning and compressing the speech recognition results, high-quality data input is generated. Multiple candidate embedding vectors are used to generate models to calculate semantic fidelity scores, select the optimal model, and construct a hot word list by combining word frequency and semantic contribution to guide the question answering module to perform professional semantic matching.

Benefits of technology

It significantly improves the RAG recall rate and knowledge matching accuracy in voice question answering scenarios, solves the problems of entity mismatch and semantic drift caused by speech transcription errors, and enhances the model's recall ability and robustness in expressing professional vocabulary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913552A_ABST
    Figure CN120913552A_ABST
Patent Text Reader

Abstract

The invention provides an RAG recall rate improving method and device in a voice question and answer scene, and relates to the technical field of data processing.The method comprises the steps that semantic cleaning processing is conducted on an original corpus containing a voice recognition result, semantic compression is conducted on the cleaned original corpus, and the original corpus is obtained; utilizing the plurality of candidate embedded vector generation models to respectively execute vector generation operation, and outputting word vectors; for each word vector, calculating a semantic fidelity score; evaluating the plurality of semantic fidelity scores, and selecting a target embedding vector generation model with the optimal semantic fidelity score from the plurality of candidate embedding vector generation models; calculating a word frequency value and an inverse document frequency value of each word according to data input, judging whether the words are professional hot words or not, and screening out the professional hot words to construct a hot word list; and jointly inputting an embedded vector output by the target embedded vector generation model and the hot word list into a question and answer module, and outputting a target answer text. According to the invention, the RAG recall rate in the voice question and answer scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a RAG recall rate improving method and device in a voice question and answer scenario. BACKGROUND

[0002] In a voice question and answer scenario, a retrieval-augmentation-generation (RAG) technology performs vector retrieval and document matching to generate an answer by taking an automatic speech recognition result as input and combining an external knowledge base, but due to problems such as speech transcription errors, semantic ambiguity and entity mismatch, the semantic similarity between the input query vector and the real relevant document in the knowledge base decreases, thereby significantly affecting the recall rate performance; therefore, it is necessary to jointly optimize from multiple dimensions such as transcription correction, semantic representation enhancement, context preservation and professional word recognition to improve the matching degree between the query vector and the target document, thereby fundamentally enhancing the knowledge retrieval capability and recall accuracy of the RAG model in the voice question and answer.

[0003] In the prior art, due to the presence of a large amount of redundant information such as intonation words, repeated words and colloquial language in the voice input, the speech recognition model is prone to semantic drift or context misplacement in the transcription process, thereby masking or misjudging the key information, especially when it comes to professional language, the model is more likely to recognize it as a general word or a homonym, which significantly reduces the restoration degree of the professional words in the transcribed text, and ultimately causes the query vector generated based on the transcribed result to have insufficient accuracy in semantic representation, which cannot accurately match the corresponding professional hot words in the knowledge base, thereby seriously affecting the integrity and accuracy of knowledge recall in the voice question and answer.

[0004] Therefore, a method is needed to improve the RAG recall rate in the voice question and answer scenario. SUMMARY

[0005] The present application provides a RAG recall rate improving method and device in a voice question and answer scenario, which can improve the RAG recall rate in the voice question and answer scenario.

[0006] In a first aspect of the present application, a RAG recall rate improving method in a voice question and answer scenario is provided, the method comprising: performing semantic cleaning processing on original corpus containing a speech recognition result, and performing semantic compression on the cleaned original corpus to generate data input; based on the data input, performing vector generation operation by using a plurality of candidate embedding vector generation models to output word vectors; for each of the word vectors, calculating a semantic fidelity score; using a gradient ranking preference optimization algorithm to evaluate a plurality of the semantic fidelity scores, and selecting a target embedding vector generation model with the optimal semantic fidelity score from the plurality of the candidate embedding vector generation models; Calculate the term frequency value and the inverse document frequency value of each word based on the data input, determine whether the word is a professional hot word, and filter out the professional hot words to construct a hot word table; Input the embedding vector output by the target embedding vector generation model and the hot word table into a question and answer module, and output a target answer text.

[0007] On the basis of the above technical solutions, preferably, the gradient ranking preference optimization algorithm is used to evaluate a plurality of semantic fidelity scores, and a target embedding vector generation model with the optimal semantic fidelity score is selected from a plurality of candidate embedding vector generation models, specifically including: Based on the semantic fidelity score, a model evaluation sample set is constructed, the evaluation sample set is indexed by the unique identifier of the data input and the name of the candidate embedding vector generation model as the category identifier, and the average semantic fidelity score corresponding to the word vector output by each candidate embedding vector generation model for the same data input is recorded as the model performance value; The evaluation sample set is passed through a pair-wise comparison unit to sort and compare the fidelity scores of any two embedding vector generation models under the same data input, and a ranking preference label is constructed based on the score difference; According to the ranking preference label, a differentiable loss function is defined and the ranking function parameters are iteratively optimized based on the backpropagation mechanism, so that the model gradually learns the true ranking relationship represented by the fidelity score difference; After optimization and convergence, the ranking function is used to uniformly rank the overall semantic fidelity performance of all candidate embedding vector generation models under different data inputs, and the candidate embedding vector generation model with the highest ranking stability and ranked first under most data inputs is selected as the target embedding vector generation model.

[0008] On the basis of the above technical solutions, preferably, after the gradient ranking preference optimization algorithm is used to evaluate a plurality of semantic fidelity scores and select a target embedding vector generation model with the optimal semantic fidelity score from a plurality of candidate embedding vector generation models, the method further includes: Based on the data input, a training sample set containing proper noun enhancement training is constructed, and the training sample set includes a plurality of labeled proper noun phrases, each proper noun phrase is combined with a context semantic compression representation to form a training pair; inputting the training sample set into the target embedding vector generation model, constructing a bidirectional context prediction task and a word vector clustering consistency task as joint training targets, the bidirectional context prediction task being used to constrain the target embedding vector generation model to retain the complete context dependency structure of the proper noun when generating the word vector, and the word vector clustering consistency task being used to cluster the word vector positions corresponding to the same proper noun in different semantic compression representations in the embedding space; During the training process, a small batch gradient descent mechanism is used to optimize the encoding parameters, position encoding parameters and word vector generation weights in the target embedding vector generation model. After the training converges, the model effect is jointly evaluated based on the proper noun recognition accuracy and the semantic fidelity variation on the validation set, and if the preset precision improvement threshold is met, the fine-tuned target embedding vector generation model is solidified and used in the subsequent RAG recall stage.

[0009] On the basis of the above technical solutions, preferably, for each word vector, a semantic fidelity score is calculated, specifically including: Based on the original semantic compression representation in the data input corresponding to each word vector, a word-level semantic comparison unit is constructed, each comparison unit including a word vector and a corresponding semantic context fragment; A reference semantic vector is generated according to the semantic context fragment; For each word vector and the corresponding reference semantic vector, a matching calculation based on cosine similarity is performed, and the calculation result is taken as the initial semantic fidelity score of the word vector; The context information of adjacent words is introduced, a weighted sliding window is constructed, adjacent word vectors are dynamically aggregated to obtain an enhanced semantic vector, and the enhanced semantic vector is used to replace the original word vector for similarity calculation to obtain the semantic fidelity score.

[0010] On the basis of the above technical solutions, preferably, the frequency value and the inverse document frequency value of each word are calculated for the data input to determine whether the word is a professional hot word, and the professional hot word list is constructed by screening out the professional hot words, specifically including: Based on the data input, a unified segmentation operation is performed, and a segmenter consistent with the target embedding vector generation model is used to perform segmentation processing on each semantic compression representation to obtain a set of semantically consistent words; Based on the word set, the total frequency of each word appearing in all semantic compression representations is counted to obtain the frequency value, and the number of documents in which each word appears in different semantic compression representations is counted at the same time, and the corresponding inverse document frequency value is calculated in combination with the total amount of the data input; Sort all the inverse document frequency values in descending order, and construct a candidate word set for words with inverse document frequency values higher than a preset threshold value; For each target word in the candidate word set, perform semantic contribution calculation in combination with the context; If the target word meets both the inverse document frequency value being in the top preset percentage and the semantic contribution exceeding a preset contribution, mark the target word as the professional hotword; Collect all the professional hotwords and remove duplicates to construct the hotword list.

[0011] On the basis of the above technical solutions, preferably, based on the data input, a plurality of candidate embedding vector generation models are used to respectively perform vector generation operations and output word vectors, specifically including: The data input is input into a plurality of candidate embedding vector generation models as unified semantic input content, wherein each embedding vector generation model adopts different model architectures or base parameter configurations, including but not limited to a semantic embedding model based on a Transformer structure, a semantic matching model optimized by contrast learning, and a word vector encoding model enhanced for proper nouns; A word vector set corresponding to each token unit of the data input is obtained, wherein each embedding vector generation model performs semantic modeling operations based on an internal token mechanism and a context modeling mechanism after receiving the data input, constructs a corresponding context-related encoding representation after dividing the semantic compressed representation into a plurality of token units, and outputs a word vector corresponding to each token unit.

[0012] On the basis of the above technical solutions, preferably, the original corpus containing the speech recognition result is subjected to semantic cleaning processing, and the cleaned original corpus is subjected to semantic compression to generate the data input, specifically including: The original corpus is subjected to cleaning processing based on rule matching and character screening mechanisms, and stop words, mood words, non-semantic punctuation symbols and character sequences with a confidence lower than a preset threshold value in the original corpus are deleted to obtain a structured cleaned corpus; The structured cleaned corpus is input into a pre-trained language model to perform generative compression operations based on a context consistency maintaining mechanism, and output a semantic compressed representation, wherein the semantic compressed representation expresses the semantic core information retained in the original corpus in the form of a question-answer pair or a central event expression; The semantic compressed representation is taken as the data input.

[0013] The device is used for executing the RAG recall rate improving method in a voice question and answer scene, and comprises an acquisition module, a processing module, and an output module. The acquisition module is configured to perform semantic cleaning processing on original corpus containing voice recognition results, and perform semantic compression on the cleaned original corpus to generate data input. The processing module is configured to perform vector generation operation on the data input by using a plurality of candidate embedding vector generation models to output word vectors. The processing module is configured to calculate a semantic fidelity score for each word vector. The processing module is configured to evaluate a plurality of semantic fidelity scores by using a gradient ranking preference optimization algorithm, and select a target embedding vector generation model with the optimal semantic fidelity score from the plurality of candidate embedding vector generation models. The processing module is configured to calculate the term frequency value and the inverse document frequency value of each word for the data input, determine whether the word is a professional hot word, and filter out the professional hot word to construct a hot word table. The output module is configured to input the embedding vector output by the target embedding vector generation model and the hot word table into a question and answer module, and output a target answer text.

[0014] Preferably, the processing module is configured to construct a model evaluation sample set based on the semantic fidelity scores, the evaluation sample set is indexed by a unique identifier of the data input and classified by the name of the candidate embedding vector generation model, and records the average semantic fidelity score of the word vectors output by each candidate embedding vector generation model for the same data input as a model performance value. The processing module is configured to perform pair comparison on the evaluation sample set by constructing a pair comparison unit, compare the fidelity scores of any two embedding vector generation models under the same data input, and construct a ranking preference label based on the score difference. The processing module is configured to define a differentiable loss function according to the ranking preference label, and iteratively optimize the ranking function parameters based on a back propagation mechanism, so that the model gradually learns the real ranking relationship represented by the fidelity score difference. The output module is configured to perform unified ranking on the overall semantic fidelity performance of all the candidate embedding vector generation models under different data inputs by using the ranking function after optimization and convergence, and select the candidate embedding vector generation model with the highest ranking stability and ranking first under the majority of data inputs as the target embedding vector generation model.

[0015] On the basis of the above technical solutions, preferably, the processing module is configured to construct a training sample set containing proprietary name enhancement training based on the data input, the training sample set including a plurality of proprietary name phrases with annotations, and each proprietary name phrase is combined with a context semantic compression representation to form a training pair; The processing module is configured to input the training sample set into the target embedding vector generation model, and construct a bidirectional context prediction task and a word vector clustering consistency task as joint training targets, the bidirectional context prediction task is configured to constrain the target embedding vector generation model to retain the complete context dependency structure of the proprietary name when generating the word vector, and the word vector clustering consistency task is configured to cluster the word vector positions corresponding to the same proprietary name in different semantic compression representations in the embedding space; The processing module is configured to use a small batch gradient descent mechanism to optimize the encoding parameters, position encoding parameters, and word vector generation weights in the target embedding vector generation model during the training process. The processing module is configured to, after the training converges, jointly evaluate the model effect based on the proprietary name recognition accuracy and the semantic fidelity change amount on a validation set, and if a preset precision improvement threshold is met, the fine-tuned target embedding vector generation model is solidified and used in the subsequent RAG recall stage.

[0016] On the basis of the above technical solutions, preferably, the acquisition module is configured to construct a word-level semantic comparison unit based on the original semantic compression representation in the data input corresponding to each word vector, and each comparison unit includes one word vector and a corresponding semantic context fragment. The processing module is configured to generate a reference semantic vector according to the semantic context fragment. The processing module is configured to perform matching calculation based on cosine similarity on each word vector and the corresponding reference semantic vector, and the calculation result is used as the initial semantic fidelity score of the word vector. The processing module is configured to introduce the context information of adjacent words, construct a weighted sliding window, dynamically aggregate adjacent word vectors, obtain an enhanced semantic vector, and replace the original word vector with the enhanced semantic vector for similarity calculation to obtain the semantic fidelity score.

[0017] On the basis of the above technical solutions, preferably, the processing module is configured to perform a unified segmentation operation based on the data input, and a segmenter consistent with the target embedding vector generation model is used to perform segmentation processing on each semantic compression representation to obtain a semantic-consistent word set. The processing module is configured to calculate a total frequency of each word in all semantic compression representations based on the set of words, to obtain a term frequency value, and to calculate a number of documents in which each word appears in different semantic compression representations, and to calculate a corresponding inverse document frequency value in combination with a total amount of the data input; The processing module is configured to sort all the inverse document frequency values in descending order, and to construct a candidate word set of words corresponding to inverse document frequency values higher than a preset threshold value. The processing module is configured to calculate a semantic contribution of each target word in the candidate word set in combination with a context. The processing module is configured to mark the target word as the professional hot word if the target word meets both the inverse document frequency value being located in a front preset percentage and the semantic contribution exceeding a preset contribution. The processing module is configured to collect and deduplicate all the professional hot words to construct the hot word table.

[0018] On the basis of the above technical solutions, preferably, the processing module is configured to input the data input as a unified semantic input content into a plurality of candidate embedding vector generation models, wherein each embedding vector generation model adopts a different model architecture or base parameter configuration, including but not limited to a semantic embedding model based on a Transformer structure, a semantic matching model optimized by contrast learning, and a word vector encoding model enhanced for proper nouns. The acquisition module is configured to acquire a set of word vectors corresponding to each segmented unit of the data input, wherein each embedding vector generation model performs semantic modeling operation based on an internal segmentation mechanism and a context modeling mechanism after receiving the data input, constructs a context-related encoding representation after segmenting the semantic compression representation into a plurality of segmented units, and outputs a word vector corresponding to each segmented unit.

[0019] On the basis of the above technical solutions, preferably, the processing module is configured to perform cleaning processing on the original corpus based on a rule matching and character screening mechanism, delete stop words, mood words, non-semantic punctuation symbols, and character sequences with a confidence lower than a preset threshold in the original corpus, and obtain a structured cleaned corpus. The processing module is configured to input the structured cleaned corpus into a pre-trained language model, perform generative compression operation based on a context consistency maintaining mechanism, and output a semantic compression representation, wherein the semantic compression representation expresses semantic core information retained in the original corpus in the form of a question-answer pair or a central event expression. The processing module is configured to input the semantic compression representation as the data input.

[0020] In a third aspect of the present application, an electronic device is provided, comprising a processor, a memory, a user interface and a network interface, the memory being configured to store instructions, the user interface and the network interface each being configured to communicate with other devices, and the processor being configured to execute the instructions stored in the memory to cause the electronic device to perform the method according to any one of the preceding aspects.

[0021] In a fourth aspect of the present application, a computer-readable storage medium is provided, which stores instructions that, when executed, perform the method according to any one of the preceding aspects.

[0022] In summary, the one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. The present application constructs a full-process semantic enhancement mechanism oriented to the characteristics of voice input, removes redundant information in the voice recognition text from the source and compresses it into a unified semantic structure, ensures that the subsequent vector generation has a high-quality semantic basis, introduces a multi-model embedding candidate system, and selects the optimal embedding representation through a semantic fidelity measurement and a gradient sorting preference optimization algorithm, significantly improves the semantic accuracy of the query vector, and at the same time, combines statistical word frequency and semantic contribution to filter professional hot words and construct a hot word table, guides the question and answer module to strengthen the matching and generation of professional semantic areas, thereby effectively solving the entity mismatch and semantic drift problems caused by voice transcription errors, and comprehensively improving the recall ability and knowledge matching accuracy of the RAG model in the voice question and answer scene.

[0023] 2. The gradient sorting preference optimization algorithm is introduced to sort and learn the semantic fidelity scores of multiple embedding vector generation models, so as to automatically select the embedding model with stable performance and optimal fidelity in the majority of data input in the multi-model candidate structure, significantly improve the fitting degree of query semantic representation and real context, effectively avoid the semantic expression deviation caused by improper model architecture selection, and ensure the vector retrieval accuracy of RAG from the source.

[0024] 3. The specific term enhancement training is performed on the target embedding vector generation model, and the bidirectional context prediction and word vector clustering consistency mechanism are jointly introduced, so that the model can retain the complete semantic structure of the specific term and maintain cross-context consistency when generating the word vector, thereby enhancing the expression robustness and discrimination ability of the model to the domain vocabulary, and effectively improving the recall coverage and entity positioning ability of RAG when processing professional questions.

[0025] 4. Construct a word-level semantic comparison unit and perform cosine similarity calculation based on context reference vector, further introduce a weighted sliding window to generate an enhanced semantic vector, so that the semantic fidelity evaluation process fully considers the context dependence of word vectors in the semantic compression structure, improves the accuracy and resolution of semantic evaluation, provides a high reliability evaluation basis for subsequent embedding model screening, and thus improves the stability of the overall semantic modeling.

[0026] 5. Perform TF-IDF calculation on the semantic compression representation and joint screening mechanism of semantic contribution degree, accurately identify professional hot words with high information gain for knowledge matching, and construct a hot word table with semantic recognition ability, further as a guide signal input in the speech recognition and question answering stage, improve the restoration degree of professional field core words and the attention weight of the generation model, and significantly enhance the retrieval ability of RAG to professional semantic fragments.

[0027] 6. Input data into multiple embedding vector generation models with different structures respectively, and output word-level context encoding representation, construct a multi-view semantic embedding set, provide expression basis for subsequent semantic fidelity analysis and model fine-tuning, can fully capture the modeling bias and semantic granularity difference of different models on semantic compression representation, provide structural support for optimizing semantic representation quality and model selection strategy, and improve the expression consistency and representation coverage of semantic vectors. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1 is a flowchart of a RAG recall rate improvement method in a voice question and answer scenario disclosed by an embodiment of the present application; Figure 2 is a module schematic diagram of a RAG recall rate improvement device in a voice question and answer scenario disclosed by an embodiment of the present application; Figure 3 is a structural schematic diagram of an electronic device disclosed by an embodiment of the present application.

[0029] Explanation of reference signs: 201, acquisition module; 202, processing module; 203, output module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION

[0030] In order for those skilled in the art to better understand the technical solutions in the specification, the technical solutions in the specification will be described clearly and completely in the specification, with reference to the drawings in the specification. Obviously, the described embodiments are only part of the embodiments of the present application, not all embodiments.

[0031] In the description of the embodiments of the present application, the words such as "for example" or "for instance" are used to represent an example, illustration or description. Any embodiment or design scheme described as "for example" or "for instance" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words such as "for example" or "for instance" are intended to present the relevant concept in a specific manner.

[0032] In the description of the embodiments of the present application, the term "a plurality of" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are used only for description purposes and should not be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Therefore, the features defined with "first" and "second" can be explicitly or implicitly included one or more of the features. The terms "include", "contain", "have" and their variants mean "include but are not limited to", unless otherwise specifically emphasized.

[0033] In the voice question and answer scene, the retrieval enhancement generation technology (RAG) is limited by problems such as speech transcription errors, semantic drift and entity recognition errors, resulting in a decrease in semantic similarity between query vectors and knowledge base documents, which seriously affects the integrity and accuracy of knowledge recall; in the prior art, the voice input often contains redundant information such as adverbs, repeated words and colloquial language, which is easy to cause semantic misplacement of the transcription, especially the restoration ability of professional domain vocabulary is insufficient, and then causes the decrease of query semantic representation precision and the failure of professional hot word matching, therefore, it is urgent to jointly optimize from multiple aspects such as transcription correction, semantic compression, context preservation and professional hot word identification, to systematically improve the recall performance of RAG in the voice question and answer.

[0034] The embodiment discloses a RAG recall rate improvement method in a voice question and answer scene, referring to Figure 1 , comprising the following steps S110-S160: S110, performing semantic cleaning processing on original corpus containing voice recognition results, and performing semantic compression on the cleaned original corpus to generate data input.

[0035] The RAG recall rate improvement method in a voice question and answer scene disclosed by the embodiment of the present application is applied to a server. The server includes but is not limited to electronic devices such as mobile phones, tablet computers, wearable devices, PC (Personal Computer), and the like, and can also be a background server running a RAG recall rate improvement method in a voice question and answer scene. The server can be realized by an independent server or a server cluster composed of multiple servers.

[0036] In a possible implementation, the original corpus containing the speech recognition result is subjected to semantic cleaning processing, and the cleaned original corpus is subjected to semantic compression to generate data input, specifically including: performing cleaning processing on the original corpus based on rule matching and character screening mechanism, deleting stop words, mood words, non-semantic punctuation marks and character sequences with confidence lower than a preset threshold in the original corpus to obtain structured cleaning corpus; inputting the structured cleaning corpus into a pre-trained language model to perform generative compression operation based on context consistency maintenance mechanism, and outputting semantic compression representation, wherein the semantic compression representation expresses the semantic core information retained in the original corpus in the form of a question and answer pair or a central event expression; taking the semantic compression representation as data input.

[0037] Specifically, first, the initial text output by the speech recognition model is input as the original corpus, which often contains non-informational content such as mood words “hmm”, “ah”, “this” and the like in speech interaction, as well as repeated words, colloquial expressions and low-confidence character sequences due to background noise or unclear speech. Based on rule matching and character screening mechanism, a multi-dimensional cleaning rule set is constructed, including: a stop word list, a mood word list, a regular punctuation recognition rule and a confidence determination logic. After performing word segmentation processing, each word segmentation unit is traversed, if the word segmentation unit matches a stop word or a mood word, or its recognition confidence is lower than a set threshold (such as 0.85), it is removed from the corpus; and the sequence of punctuation marks containing continuous punctuation marks or not conforming to the standard of semantic expression (such as “?!、、、”) is subjected to merging or removing processing, and finally the structured cleaning corpus is output, which has stable syntactic structure and semantic coherence. For example, in the transcription result “this student, hmm, is studying computer in that very powerful, hmm, university”, “this”, “hmm” and “that very powerful” are deleted, and the core structure “student studying computer in university” is retained, constituting a structured expression.

[0038] First, a pre-trained language model with strong context modeling capability and generative semantic extraction capability, such as GPT, T5 or FLAN series model, is selected. The structured cleaning corpus is input as an input to perform encoding and semantic extraction process. The context consistency maintenance mechanism refers to the language model retaining the logical relationship, reference relationship and event main line between sentences or substructures during compression to prevent important semantic chain from breaking or information jumping. In actual operation, a compression prompt such as "compress the following sentences into a question and answer form" or "extract the event description points" is configured for each structured cleaning corpus to guide the language model to focus on the semantic backbone during compression. A semantic compression representation is output by the generative language model, that is, a summary expression of the original content, which is shorter in length but complete in semantics. For example, the semantic compression representation generated for "a student studies computer science and artificial intelligence and wins a prize in a related competition" can be "a student studies computer science and artificial intelligence and wins a prize", which retains the core event and subject relationship.

[0039] The semantic compression representation generated by the foregoing language model is used as a standardized input format for all subsequent modules, including embedding vector generation, semantic fidelity evaluation, professional hot word statistics, and vector retrieval entry of the RAG question and answer module. The semantic compression representation has less redundant information, higher semantic density and clearer entity relationship than the original corpus, which can better adapt to the tokenizer of the embedding vector generation model, improve the semantic concentration of the word vector, and also help to identify statistically significant and semantically key proper nouns and keywords in subsequent word frequency statistics, thereby providing a unified and high-fidelity semantic input basis for improving the recall rate of RAG in the voice question and answer scenario.

[0040] S120, based on the data input, performing vector generation operation by using multiple candidate embedding vector generation models respectively, and outputting word vectors.

[0041] In one possible implementation, based on the data input, performing vector generation operation by using multiple candidate embedding vector generation models respectively, and outputting word vectors, specifically including: inputting the data input as a unified semantic input content into the multiple candidate embedding vector generation models respectively, wherein each embedding vector generation model adopts different model architecture or base parameter configuration, including but not limited to a semantic embedding model based on a Transformer structure, a semantic matching model optimized by contrastive learning, and a word vector encoding model enhanced for proper nouns; obtaining a set of word vectors corresponding to each token unit of the data input, wherein each embedding vector generation model performs semantic modeling operation based on its token mechanism and context modeling mechanism after receiving the data input, constructs a corresponding context-related encoding representation after dividing the semantic compression representation into multiple token units, and outputs a word vector corresponding to each token unit.

[0042] Specifically, based on the aforementioned semantic compression representation, an input data batch is constructed, and each semantic compression representation is sent as an independent input sample to multiple candidate embedding vector generation models for processing. The embedding vector generation models include semantic representation models with different structural designs or training mechanisms, such as semantic embedding models using a multi-layer Transformer structure and trained on large-scale unsupervised corpus, which use a multi-head self-attention mechanism to capture intra-sentence and inter-sentence dependencies.

[0043] The semantic matching model trained using the contrastive learning method narrows the embedding distance between semantic similar samples by constructing positive and negative sample pairs; and a word vector encoding model is constructed for proper nouns and optimized by named entity recognition to determine the word boundary, thereby strengthening the stability and discrimination of the model when processing domain-specific entity expressions. The structural differences between each model are reflected in their word segmentation strategies, context windows, training objective functions, and parameter optimization paths, thereby forming a multi-perspective semantic representation system for subsequent semantic fidelity evaluation.

[0044] In each candidate embedding vector generation model, after receiving data input, the built-in word segmenter of the model is first called to perform word segmentation. Different models may use Byte Pair Encoding (BPE), WordPiece, or SentencePiece word segmentation mechanisms, which are responsible for dividing the semantic compression representation into finer-grained word units and determining the relative position of each word unit in the context. Then, the model constructs a context-aware word vector representation for each word unit through its context modeling mechanism, usually a multi-layer Transformer encoder. This means that when generating the current word vector, not only the information of the word itself is used, but also the representations of other words in its context are fused, thereby capturing long-distance dependencies and syntactic structures. The final output of the word vector set is a two-dimensional tensor structure, where each row represents a context semantic representation of a word unit. The tensor dimension is nxd, where n is the number of word units and d is the embedding dimension. Taking "the student won the first prize in the artificial intelligence competition" as an example of semantic compression representation input, the model can output seven groups of word vectors corresponding to "the", "student", "won", "artificial", "intelligence", "competition", and "first prize". Each group of vectors maintains a uniform embedding dimension, which is used for subsequent semantic fidelity calculation and model evaluation.

[0045] S130, for each word vector, calculate the semantic fidelity score.

[0046] In a possible implementation, for each word vector, a semantic fidelity score is calculated, specifically comprising: based on the original semantic compression representation in the data input corresponding to each word vector, a word-level semantic comparison unit is constructed, each comparison unit including a word vector and a corresponding semantic context segment; a reference semantic vector is generated according to the semantic context segment; for each word vector and the corresponding reference semantic vector, a cosine similarity-based matching calculation is performed, and the calculation result is taken as the initial semantic fidelity score of the word vector; the context information of adjacent words is introduced, a weighted sliding window is constructed, adjacent word vectors are dynamically aggregated to obtain an enhanced semantic vector, and the enhanced semantic vector is used to replace the original word vector for similarity calculation to obtain the semantic fidelity score.

[0047] Specifically, for each candidate embedding vector generated by the word vector, the source semantic compression representation corresponding to the word vector is retrieved, and the start and end positions of the word in the sentence are extracted as a semantic context segment with a fixed length of context window content as the center; each semantic context segment and the center word vector form a word-level semantic comparison unit, which is used to evaluate whether the word vector accurately expresses the semantic core in the context. For example, in the semantic compression representation "students win awards in artificial intelligence experiments", the word vector generated for the word "artificial", the context segment is composed of the two words "in" "students" and "intelligence" "experiment" on the left and right of the word, and the word vector corresponding to "artificial" is combined to form a comparison unit for subsequent similarity evaluation.

[0048] The semantic context segment in each comparison unit is input into a pre-trained language understanding model such as BERT or RoBERTa to obtain the overall semantic representation of the segment through a context encoding process; the reference semantic vector is the context embedding result obtained by the language understanding model after compressing the entire context segment on the syntactic and semantic levels, which is used as a judgment standard. To ensure that the reference semantic vector has a global perspective, the [CLS] site output vector is used as the context aggregation representation, and the maximum pooling mechanism is used to enhance its sensitivity to local semantic changes, ensuring that the semantic consistency between the center word and its context is included in the representation.

[0049] Each pair of word vectors and reference semantic vectors is input into a cosine similarity function for similarity calculation, and the formula is as follows: wherein represents the word vector, represents the reference semantic vector, and the symbol "·" represents the dot product operation, denotes L2 norm; the score is used to measure the degree of expression of the current word vector in the semantic position of the context fragment, the value range is [-1, 1], the higher the score, the more accurate the word vector can convey the semantic information required by the semantic position.

[0050] With the position of the current word vector in the semantic compression representation as the center, a sliding window of a preset length (such as ±2) is extended forward and backward, all word vectors in the window are collected and weighted summation is performed thereon; the weight distribution is performed according to the relative position index attenuation strategy, the word vector closer to the center word has a larger weight, and an enhanced semantic vector is constructed: wherein denotes the word vector of the i-th word, denotes the weight coefficient thereof, denotes the window radius, is an attenuation factor; the enhanced semantic vector more fully expresses the semantic role of the word in its semantic environment. The enhanced semantic vector

[0051] is compared with the corresponding reference semantic vector to perform again cosine similarity calculation, and the output result is used as the final semantic fidelity score: The final semantic fidelity score reflects the accuracy of the word vector in expressing the semantic position of the semantic under the condition of context enhancement, and is used for subsequent ranking optimization and screening process of the embedding vector generation model.

[0052] S140, a gradient ranking preference optimization algorithm is used to evaluate a plurality of semantic fidelity scores, and a target embedding vector generation model with the optimal semantic fidelity score is selected from a plurality of candidate embedding vector generation models.

[0053] ​In a possible implementation, the gradient ranking preference optimization algorithm is adopted to evaluate the plurality of semantic fidelity scores, and a target embedding vector generation model with the optimal semantic fidelity score is selected from the plurality of candidate embedding vector generation models, and specifically comprises: constructing a model evaluation sample set based on the semantic fidelity scores, the evaluation sample set is indexed by a unique identifier of data input and classified by a name of the candidate embedding vector generation model, and records an average semantic fidelity score corresponding to the word vector output by the candidate embedding vector generation model for the same data input as a model performance value; the evaluation sample set is passed through a pair-wise comparison unit to compare and sort the fidelity scores of any two embedding vector generation models under the same data input, and a score difference is used to construct a ranking preference label; according to the ranking preference label, a differentiable loss function is defined and the ranking function parameters are iteratively optimized based on a back propagation mechanism, so that the model gradually learns the real ranking relationship represented by the fidelity score difference; after optimization and convergence, the ranking function is used to uniformly rank the overall semantic fidelity performance of all candidate embedding vector generation models under different data inputs, and the candidate embedding vector generation model with the highest ranking stability and ranking first under most data inputs is selected as the target embedding vector generation model.

[0054] Specifically, first, for each data input sample, a set of word vectors output by all candidate embedding vector generation models under the input condition is obtained, and the semantic fidelity score of each group of word vectors is calculated based on the foregoing calculation process; the average fidelity score of all word segmentation units is calculated for each group of word vectors corresponding to the embedding vector generation model, and the average value is taken as the performance indicator of the model under the current data input. Each data input sample has a unique identifier (such as UUID) for indexing the performance of the corresponding sample under different models; the name of the candidate embedding vector generation model is used as the category identifier to distinguish the source of the fidelity score. The finally constructed model evaluation sample set is in the form of a structured table, each row containing the data input identifier, the model name and the corresponding semantic fidelity average value, which is used to support the subsequent ranking training process.

[0055] Under each data input, any two candidate embedding vector generation models are randomly or exhaustively selected to form a model pair , and the corresponding fidelity average values are compared; if , a ranking preference label is generated, indicating that model is better than model , otherwise indicates the opposite ranking relationship. The pair-wise comparison unit is in the form of a triple , where represents the fidelity score difference, which is the input feature of the ranking learning model, and is used to train the ranking function to identify the mapping relationship between the score difference and the real preference.

[0056] The RankNet loss function based on the Logistic loss is used to construct the ranking optimization objective, and the loss function is defined as follows: wherein is a Sigmoid function, represents the difference between the two model fidelity values, is a ranking preference label. By minimizing the loss function and updating the ranking function parameters based on the back propagation mechanism, the ranking model gradually learns the functional relationship between the score difference and the preference label, thereby enhancing its ability to distinguish the pros and cons of semantic fidelity. Training uses optimization algorithms such as Adam or SGD to iteratively optimize the linear weights of the ranking function in batch update mode until the validation set ranking accuracy reaches a stable or the loss function converges.

[0057] On all data input samples, the trained ranking function is used to perform scoring and ranking operations on the mean fidelity of the model output generated by each candidate embedding vector, and the average ranking and ranking fluctuation standard deviation of each model on all samples are calculated; select the embedding vector generation model that ranks first in most data input samples and has the smallest cross-sample ranking fluctuation as the target embedding vector generation model. This model has stable expression ability and global performance advantage in the dimension of semantic fidelity, and is the only output source of the final semantic representation generation module, input to the subsequent professional hot word indexing, question and answer vector retrieval and answer generation stage, to ensure that the semantic input in the RAG process has consistency, high fidelity and robustness.

[0058] In a possible implementation, after the plurality of candidate embedding vector generation models are evaluated by using the gradient ranking preference optimization algorithm, and the target embedding vector generation model with the optimal semantic fidelity score is selected from the plurality of candidate embedding vector generation models, the method further includes: constructing a training sample set containing proper noun enhancement training based on the data input, the training sample set including a plurality of labeled proper noun phrases, each proper noun phrase being combined with a context semantic compression representation to form a training pair; inputting the training sample set into the target embedding vector generation model, constructing a bidirectional context prediction task and a word vector clustering consistency task as a joint training target, the bidirectional context prediction task being used to constrain the target embedding vector generation model to retain the complete context dependency structure of the proper noun when generating the word vector, and the word vector clustering consistency task being used to cluster the word vector positions corresponding to the same proper noun in different semantic compression representations in the embedding space; during the training process, a small batch gradient descent mechanism is used to optimize the encoding parameters, the position encoding parameters, and the word vector generation weight in the target embedding vector generation model; after the training converges, the model effect is jointly evaluated based on the proper noun recognition accuracy and the semantic fidelity change on the validation set, and if the preset precision improvement threshold is met, the fine-tuned target embedding vector generation model is solidified and used in the subsequent RAG recall stage.

[0059] Specifically, first, the proper noun phrases in the data input that have completed semantic compression processing are recognized by using a named entity recognition model (such as BERT-NER or RoBERTa-NER), and the proper noun phrases include but are not limited to names of persons, places, organizations, products, and technical terms, and are combined with the context fragments thereof in the compressed corpus to form training sample pairs. Each training sample pair is formed by a semantic compression representation and a proper noun phrase contained in the representation, and the start and end boundaries of the proper noun are accurately labeled for supervised learning. For example, in the semantic compression representation “XX University develops a new artificial intelligence chip”, “XX University” and “artificial intelligence chip” are recognized as two independent proper noun phrases, and the context structure thereof is retained as the model training input by constructing the corresponding training pairs.

[0060] The bidirectional context prediction task is a mechanism for enhancing the semantic modeling capability of a language model, which aims to make the model consider the context information before and after a certain word when generating the vector of the word, thereby improving the sensitivity of the word vector to changes in context. In specific implementation, part of the context information in the training sample is masked, and the model is asked to predict the embedding representation of the missing word position. The cosine similarity or L2 distance between the true embedding and the predicted embedding is used as the loss function for optimization. The word vector clustering consistency task is used to ensure that the word vector representations of the same proper noun in different semantic compression representations remain consistent, i.e., close to each other in the embedding space. In specific implementation, the intra-class distance of the word vectors of the proper nouns with the same label is calculated using a center clustering mechanism, and the Euclidean distance or KL divergence between the word vectors of the proper nouns in different contexts is minimized, thereby achieving semantic aggregation. The joint training mechanism can simultaneously improve the context expression capability and the stability of cross-context entity recognition of the target embedding vector generation model.

[0061] The training samples are sent into the model in batches, each batch including a number of semantic compression representations and their corresponding proper noun labels. During the forward propagation process, the model generates embedding representations based on the current parameters and calculates the loss value. Subsequently, the gradient is updated for the multi-head attention parameters in the Transformer encoding layer, the position encoding parameters (used to model the relative order information of each word in the sequence), and the projection weights in the output layer for generating embedding vectors based on the gradient backpropagation algorithm. The update method is adjusted using the Adam or SGD optimizer, and the iteration is performed until the loss converges or the preset number of rounds is reached. For example, if the learning rate is 1e-4, each round of training includes 5000 batches, and each batch contains 128 training samples, the parameter convergence can usually be completed within 5 to 10 rounds of training.

[0062] In the validation set, semantic compression representations containing a large number of labeled proper nouns are selected and input into the fine-tuned target embedding vector generation model. The recognition accuracy at each proper noun position is calculated, i.e., whether the model can stably output semantically consistent word vectors in different contexts. At the same time, the difference in semantic fidelity score before and after fine-tuning is compared, and the average improvement is calculated. If the recognition accuracy exceeds the preset threshold (e.g., 90%) and the semantic fidelity improvement value exceeds the set proportion (e.g., 5%), it is considered that the fine-tuning effect meets the performance standard.

[0063] If the above evaluation results meet the preset precision improvement threshold, the fine-tuned target embedding vector generation model is solidified, i.e., the training weights are saved and the model structure is locked, which is used for vector generation operations in the subsequent RAG recall stage. In order to ensure that professional entities can be represented more stably and accurately and maintain semantic consistency when facing real voice question and answer applications, thereby improving the knowledge recall coverage and accuracy of the retrieval enhancement generation mechanism in the voice question and answer scenario.

[0064] S150: Calculate the word frequency and inverse document frequency of each word for the data input, determine whether the word is a professional hot word, and filter out professional hot words to build a hot word list.

[0065] In one possible implementation, the word frequency (Word Frequency) and inverse document frequency (InVF) values ​​of each word are calculated for the data input. The method determines whether a word is a professional hot word and selects professional hot words to construct a hot word list. Specifically, this includes: performing a unified word segmentation operation based on the data input; using a word segmenter consistent with the target embedding vector generation model to perform word segmentation processing on each semantically compressed representation to obtain a semantically consistent word set; based on the word set, calculating the total frequency of each word in all semantically compressed representations to obtain the Word Frequency value, and simultaneously calculating the number of documents in which each word appears in different semantically compressed representations, and calculating the corresponding InVF value based on the total amount of data input; sorting all InVF values ​​in descending order, and constructing a candidate word set for words with InVF values ​​higher than a preset threshold; calculating the semantic contribution of each target word in the candidate word set based on the context; if a target word simultaneously satisfies that its InVF value is in the top preset percentile and its semantic contribution exceeds a preset contribution, then the target word is marked as a professional hot word; collecting all professional hot words and removing duplicates to construct a hot word list.

[0066] Specifically, each compressed text is read from the semantic compression representation set, and word segmentation is performed using the word segmenter corresponding to the target embedding vector generation model (such as BPE, WordPiece, or SentencePiece). This ensures that the segmentation granularity, word boundaries, and sub-word structure remain completely consistent with the subsequent embedding vector generation process, avoiding representational bias introduced by segmentation differences. All processed words are stored in the order they appear in the compressed representation to form a word set, which is used for subsequent word frequency statistics and inverse document frequency calculation. For example, for the semantic compression representation "intelligent chips improve image recognition accuracy," the WordPiece word segmenter produces a word set such as "intelligent," "chip," "improve," "image," "recognition," and "accuracy."

[0067] For each word The total number of times it appears in all semantic compression representations is denoted as . Count how many different semantic compression representations the word appears in, and denote it as . The total number of data entries is denoted as The inverse document frequency (IVF) value is calculated using the following formula: The "+1" is used to avoid a denominator of zero. The word frequency value reflects the density of word usage in the data, while the inverse document frequency value reflects the scarcity of words and their information gain capability, which is subsequently used to filter high-weight words with domain characteristics.

[0068] Sort all inverse document frequency values ​​in descending order, and construct a candidate word set for words with inverse document frequency values ​​higher than a preset threshold. Specifically, this includes: Based on this, all words are sorted and the top 100% are selected. The first 10% of words are selected as the initial candidate set, and a lower threshold is set (e.g., ...). Further filtering out commonly used words with insufficient information. This screening strategy focuses the candidate set on words that appear less frequently in semantically compressed representations but have a significant effect on discriminative expression, and have the potential to become professional hot words.

[0069] The context window of each candidate word in the semantically compressed representation is extracted as an input fragment and fed into a pre-trained language model. The influence of the word on the overall semantic representation vector of the fragment is calculated. The semantic contribution is defined as the semantic shift between the context vector after the word is masked and the original context vector, and is represented by cosine distance. in Represents the original context semantic vector, This represents the semantic vector of the context after the target word is masked. The higher the contribution, the stronger the role of the word in expressing the semantic meaning of the context.

[0070] Check each target word separately Is it before sorting? If a word is selected based on its semantic contribution (within a certain percentage) and whether it exceeds a threshold (e.g., 0.2), it is identified as a professional hot word. This dual screening mechanism ensures that the selected words are both informationally scarce and semantically crucial, effectively distinguishing them from high-frequency general words and non-core vocabulary.

[0071] All qualified professional hot words are uniformly summarized and duplicate terms are removed to form a unique set. The weight information of each hot word, such as TF-IDF value or semantic contribution score, can be added as an additional field to construct a standard hot word vocabulary. This hot word vocabulary can be provided as an external prior input to the speech recognition model to adjust its language modeling priority and recognition dictionary strategy, thereby improving the recognition accuracy and semantic restoration ability of professional words in the speech input stage, and further enhancing the semantic matching ability and recall accuracy based on the retrieval enhancement generation mechanism.

[0072] S160: Input the embedding vector output by the target embedding vector generation model and the hot word vocabulary into the question answering module to output the target answer text.

[0073] First, the query embedding vector, processed by the target embedding vector generation model, is obtained. This embedding vector originates from the set of full-sequence word vectors obtained by performing vector generation operations on the semantic compression representation. A retrieval vector representing the overall query semantics is generated through average pooling or CLS site extraction. Simultaneously, a pre-constructed hot word vocabulary is loaded. Each term in the hot word vocabulary carries its corresponding word frequency value, inverse document frequency value, and normalized semantic contribution score, forming a multi-dimensional weight vector. A hot word enhancement mechanism transforms this vocabulary into an attention-guided matrix for the query vector. This involves adding semantic weights or adjusting the embedding position of word segments containing hot words in the original embedding vector. This enhances the query vector's responsiveness to the target semantic region during the vector matching stage, thereby improving its coverage breadth and matching accuracy in knowledge base retrieval.

[0074] Next, the question-answering module receives the embedded vector and the contextual attention matrix guided by hot words as joint input, and performs a document matching process based on vector retrieval. The document matching selects the document embedding vector that is closest to the query embedding vector in the knowledge base based on a similarity measurement algorithm (such as cosine similarity or inner product similarity). Each document slice constitutes a candidate knowledge fragment set. To improve recall stability, target terms from the hot word list are introduced as keyword filters during the matching process, retaining only knowledge fragments containing multiple professional hot words or semantically similar words, avoiding the introduction of irrelevant information due to transcription ambiguity or semantic dilution. This filtering mechanism is based on a hot word matching rate scoring function, calculated as follows: in, For matching rate, Indicates the first Candidate document fragments, This represents the hot word list, and the matching rate reflects the degree of coverage between document fragments and the hot word set.

[0075] After document matching and hot word-guided filtering, the selected knowledge fragments are input as contextual conditions into the generative model. This model is typically a large-scale autoregressive language model, which performs language generation based on the joint encoding of query embeddings and knowledge fragments to produce the final target answer text. To enhance the semantic correspondence between the generated content and the query, the language model maintains soft alignment with hot word terms during decoding, prioritizing the generation of answer statements containing high-weight hot words to improve the professionalism and relevance of the answer. The final output target answer text is the result of the search enhancement generation mechanism based on hot word reinforcement, possessing higher semantic relevance, coverage of professional terminology, and contextual consistency.

[0076] This embodiment also discloses a device for improving RAG recall in a voice question-and-answer scenario, referring to... Figure 2, comprising an acquisition module 201, a processing module 202, and an output module 203, the device is used to execute the RAG recall rate improvement method in any one of the above voice question and answer scenarios, wherein: The acquisition module 201 is used to perform semantic cleaning processing on the original corpus containing the voice recognition result, and perform semantic compression on the cleaned original corpus to generate data input.

[0077] The processing module 202 is used to perform vector generation operations on the data input by using a plurality of candidate embedding vector generation models respectively, and output word vectors.

[0078] The processing module 202 is used to calculate a semantic fidelity score for each word vector.

[0079] The processing module 202 is used to evaluate a plurality of semantic fidelity scores by using a gradient ranking preference optimization algorithm, and select a target embedding vector generation model with the optimal semantic fidelity score from the plurality of candidate embedding vector generation models.

[0080] The processing module 202 is used to calculate the term frequency value and the inverse document frequency value of each word for the data input, determine whether the word is a professional hot word, and filter out professional hot words to construct a hot word table.

[0081] The output module 203 is used to input the embedding vector output by the target embedding vector generation model and the hot word table into a question and answer module, and output a target answer text.

[0082] In a possible implementation, the processing module 202 is used to construct a model evaluation sample set based on the semantic fidelity scores, the evaluation sample set is indexed by a unique identifier of the data input and classified by the name of the candidate embedding vector generation model, and the average value of the semantic fidelity scores corresponding to the word vectors output by each candidate embedding vector generation model for the same data input is recorded as a model performance value.

[0083] The processing module 202 is used to pass the evaluation sample set through a pair-wise comparison unit to sort and compare the fidelity scores of any two embedding vector generation models under the same data input, and construct a ranking preference label based on the score difference.

[0084] The processing module 202 is used to define a differentiable loss function according to the ranking preference label and iteratively optimize the ranking function parameters based on a back propagation mechanism, so that the model gradually learns the real ranking relationship represented by the fidelity score difference.

[0085] The output module 203 is configured to, after optimization convergence, uniformly rank the overall semantic fidelity performances of all candidate embedding vector generation models under different data inputs using the ranking function, and select a candidate embedding vector generation model with the highest ranking stability and ranking first under most data inputs as the target embedding vector generation model.

[0086] In a possible implementation, the processing module 202 is configured to construct a training sample set containing proper noun enhancement training based on the data input, and the training sample set includes a plurality of labeled proper noun phrases, and each proper noun phrase is combined with the context semantic compression representation to form a training pair.

[0087] The processing module 202 is configured to input the training sample set into the target embedding vector generation model, and construct a bidirectional context prediction task and a word vector clustering consistency task as a joint training target, the bidirectional context prediction task is used to constrain the target embedding vector generation model to retain the complete context dependency structure of the proper noun when generating the word vector, and the word vector clustering consistency task is used to cluster the word vector positions corresponding to the same proper noun in different semantic compression representations in the embedding space.

[0088] The processing module 202 is configured to use a small batch gradient descent mechanism to optimize the encoding parameters, position encoding parameters and word vector generation weights in the target embedding vector generation model during the training process.

[0089] The processing module 202 is configured to, after training convergence, jointly evaluate the model effect based on the proper noun recognition accuracy and the semantic fidelity change amount on the validation set, and if a preset precision improvement threshold is met, the fine-tuned target embedding vector generation model is solidified and used in the subsequent RAG recall stage.

[0090] In a possible implementation, the acquisition module 201 is configured to construct a word-level semantic comparison unit based on the original semantic compression representation in the data input corresponding to each word vector, and each comparison unit includes a word vector and a corresponding semantic context segment.

[0091] The processing module 202 is configured to generate a reference semantic vector according to the semantic context segment.

[0092] The processing module 202 is configured to perform a matching calculation based on cosine similarity on each word vector and the corresponding reference semantic vector, and the calculation result is used as the initial semantic fidelity score of the word vector.

[0093] The processing module 202 is configured to introduce the context information of adjacent words, construct a weighted sliding window, dynamically aggregate adjacent word vectors to obtain an enhanced semantic vector, and replace the original word vector with the enhanced semantic vector for similarity calculation to obtain the semantic fidelity score.

[0094] In a possible implementation, the processing module 202, configured to perform the unified tokenization operation based on the data input, performs tokenization processing on each semantic compression representation by using a tokenizer consistent with the target embedding vector generation model, to obtain a set of semantically consistent words.

[0095] The processing module 202 is configured to, based on the set of words, count a total frequency of each word appearing in all semantic compression representations to obtain a term frequency value, and simultaneously count a number of documents in which each word appears in different semantic compression representations, and calculate a corresponding inverse document frequency value in combination with a total amount of the data input.

[0096] The processing module 202 is configured to sort all inverse document frequency values in descending order, and construct a candidate word set of words corresponding to inverse document frequency values higher than a preset threshold value.

[0097] The processing module 202 is configured to, for each target word in the candidate word set, perform semantic contribution degree calculation in combination with the context.

[0098] The processing module 202 is configured to, if the target word simultaneously satisfies that the inverse document frequency value is located in a front preset percentage and the semantic contribution degree exceeds a preset contribution degree, mark the target word as a professional hot word.

[0099] The processing module 202 is configured to collect and deduplicate all professional hot words to construct a hot word table.

[0100] In a possible implementation, the processing module 202 is configured to input the data input as unified semantic input content into a plurality of candidate embedding vector generation models respectively, wherein each embedding vector generation model adopts different model architectures or base parameter configurations, including but not limited to a semantic embedding model based on a Transformer structure, a semantic matching model optimized by contrastive learning, and a word vector encoding model enhanced for proper nouns.

[0101] The acquisition module 201 is configured to acquire a set of word vectors corresponding to each tokenization unit of the data input, wherein each embedding vector generation model, after receiving the data input, performs semantic modeling operation based on an internal tokenization mechanism and a context modeling mechanism, constructs a corresponding context-related encoding representation after cutting the semantic compression representation into a plurality of tokenization units, and outputs a word vector corresponding to each tokenization unit.

[0102] In a possible implementation, the processing module 202 is configured to perform cleaning processing on the original corpus based on a rule matching and character screening mechanism, delete stop words, modal words, non-semantic punctuation marks and character sequences with a confidence lower than a preset threshold in the original corpus, and obtain a structured cleaned corpus.

[0103] The processing module 202 is configured to input the structured cleaning corpus into a pre-trained language model, perform a generative compression operation based on a context consistency maintenance mechanism, and output semantic compression representation, wherein the semantic compression representation expresses semantic core information retained in the original corpus in the form of a question-answer pair or a central event expression.

[0104] The processing module 202 is configured to input the semantic compression representation as data.

[0105] It should be noted that the device provided in the above embodiment is only used as an example to divide the above functional modules to achieve its functions, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is described in the method embodiment, which will not be described here.

[0106] The embodiment also discloses an electronic device, which refers to Figure 3 The electronic device can include at least one processor 301, at least one communication bus 302, a user interface 303, a network interface 304, and at least one memory 305.

[0107] The communication bus 302 is configured to realize the connection and communication between the components.

[0108] The user interface 303 can include a display screen (Display) and a camera (Camera), and the optional user interface 303 can further include a standard wired interface and a wireless interface.

[0109] The network interface 304 can optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0110] The processor 301 can include one or more processing cores. The processor 301 connects various parts within the server through various interfaces and lines, performs various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 305, and calling data stored in the memory 305. Alternatively, the processor 301 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 301 can integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes operating systems, user interfaces, and application programs. The GPU is responsible for rendering and drawing the content to be displayed on the display screen. The modem is used to process wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 301, but can be realized by a separate chip.

[0111] The memory 305 can include a random access memory (RAM) and a read-only memory (ROM). Alternatively, the memory includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The data storage area can store data related to the above-mentioned various method embodiments, etc. The memory 305 can also be at least one storage device located away from the aforementioned processor 301. As a computer storage medium, the memory 305 can include an operating system, a network communication module, a user interface 303 module, and an application program of the RAG recall rate improvement method in a voice question and answer scenario.

[0112] In Figure 3In the electronic device shown, the user interface 303 is mainly used to provide an interface for the user to input, and obtain data input by the user. The processor 301 can be used to invoke an application program of the RAG recall rate improvement method in a voice question and answer scenario stored in the memory 305, and when executed by one or more processors 301, the electronic device performs the method of one or more of the above embodiments.

[0113] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the order of the described actions, because according to the present application, certain steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0114] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0115] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the units is only a logical function division. There can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different parts can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical or other forms.

[0116] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0117] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0118] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable memory. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory 305 and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned memory 305 includes: a U disk, a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0119] The present application also discloses a computer readable storage medium, which stores instructions. When executed by one or more processors 301, the instructions cause an electronic device to perform one or more methods as described in the above embodiments.

[0120] The above are only exemplary embodiments of the present disclosure, and cannot limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon considering the specification and practicing the disclosure. The present application is intended to cover any variations, uses or adaptive changes of the present disclosure that follow the general principles of the present disclosure and include common knowledge or conventional techniques in the art that are not described in the present disclosure. The specification and examples are only considered to be exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A method for improving RAG recall rate in a voice question and answer scenario, characterized in that, The method comprises: performing semantic cleaning processing on original corpus containing voice recognition results, and performing semantic compression on the cleaned original corpus to generate data input; based on the data input, using multiple candidate embedding vector generation models to respectively perform vector generation operations, and outputting word vectors; for each word vector, calculating a semantic fidelity score; using a gradient ranking preference optimization algorithm to evaluate multiple semantic fidelity scores, and selecting a target embedding vector generation model with the optimal semantic fidelity score from the multiple candidate embedding vector generation models; calculating the term frequency value and inverse document frequency value of each word for the data input, determining whether the word is a professional hot word, and screening out the professional hot words to construct a hot word table; inputting the embedding vector output by the target embedding vector generation model and the hot word table into a question and answer module to output a target answer text.

2. The RAG recall rate improvement method in a voice question and answer scene according to claim 1, characterized in that, The method further comprises: based on the semantic fidelity score, constructing a model evaluation sample set, the evaluation sample set being indexed by the unique identifier of the data input and classified by the name of the candidate embedding vector generation model, recording the average semantic fidelity score of the word vectors output by each candidate embedding vector generation model for the same data input as the model performance value; using a pair-wise comparison unit to sort and compare the fidelity scores of any two embedding vector generation models under the same data input, and constructing a ranking preference label based on the score difference; based on the ranking preference label, defining a differentiable loss function and iteratively optimizing the ranking function parameters based on a backpropagation mechanism, so that the model gradually learns the true ranking relationship represented by the fidelity score difference; after optimization convergence, using the ranking function to uniformly rank the overall semantic fidelity performance of all candidate embedding vector generation models under different data inputs, and selecting the candidate embedding vector generation model with the highest ranking stability and ranking first under most data inputs as the target embedding vector generation model.

3. The RAG recall rate improvement method in a voice question and answer scene according to claim 2, characterized in that, After the method of using a gradient ranking preference optimization algorithm to evaluate multiple semantic fidelity scores and selecting a target embedding vector generation model with the optimal semantic fidelity score from multiple candidate embedding vector generation models, the method further comprises: based on the data input, constructing a training sample set containing proper noun enhancement training, the training sample set including multiple labeled proper noun phrases, and each proper noun phrase being combined with context semantic compression representation to form a training pair; inputting the training sample set into the target embedding vector generation model, constructing a bidirectional context prediction task and a word vector clustering consistency task as joint training targets, the bidirectional context prediction task being used to constrain the target embedding vector generation model to retain the complete context dependency structure of the proper noun when generating the word vector, and the word vector clustering consistency task being used to cluster the word vector positions corresponding to the same proper noun in different semantic compression representations in the embedding space; adopting a small batch gradient descent mechanism to optimize the encoding parameters, position encoding parameters and word vector generation weights in the target embedding vector generation model during the training process; after the training converges, jointly evaluating the model effect based on the proper noun recognition accuracy and the semantic fidelity variation on the validation set, and if the preset accuracy improvement threshold is met, then solidifying the fine-tuned target embedding vector generation model and using it in the subsequent RAG recall stage.

4. The RAG recall rate improvement method in a voice question and answer scene according to claim 1, characterized in that, The semantic fidelity score is calculated for each word vector, specifically including: Based on the original semantic compression representation in the data input corresponding to each word vector, a word-level semantic comparison unit is constructed, each comparison unit including a word vector and a corresponding semantic context fragment; A reference semantic vector is generated according to the semantic context fragment; For each word vector and the corresponding reference semantic vector, a cosine similarity-based matching calculation is performed, and the calculation result is used as the initial semantic fidelity score of the word vector; The context information of adjacent words is introduced to construct a weighted sliding window, and adjacent word vectors are dynamically aggregated to obtain an enhanced semantic vector, and the enhanced semantic vector is used to replace the original word vector for similarity calculation to obtain the semantic fidelity score.

5. The RAG recall rate improvement method in a voice question and answer scene according to claim 1, characterized in that, The term frequency value and the inverse document frequency value of each word in the data input are calculated to determine whether the word is a professional hot word, and the professional hot words are filtered to construct a hot word list, specifically including: Based on the data input, a unified segmentation operation is performed, and a segmenter consistent with the target embedding vector generation model is used to perform segmentation processing on each semantic compression representation to obtain a set of semantically consistent words; Based on the word set, the total frequency of each word appearing in all semantic compression representations is counted to obtain the term frequency value, and the number of documents in which each word appears in different semantic compression representations is also counted, and the corresponding inverse document frequency value is calculated based on the total amount of data input; All the inverse document frequency values are sorted in descending order, and the words corresponding to the inverse document frequency values higher than the preset threshold are constructed into a candidate word set; For each target word in the candidate word set, semantic contribution is calculated in combination with the context; If the target word meets the conditions that the inverse document frequency value is in the top preset percentage and the semantic contribution exceeds the preset contribution, the target word is marked as the professional hot word; All the professional hot words are collected and de-duplicated to construct the hot word list.

6. The RAG recall rate improvement method in a voice question and answer scene according to claim 1, characterized in that, The data input is used to perform vector generation operations using multiple candidate embedding vector generation models to output word vectors, specifically including: Input the data input as unified semantic input content into a plurality of candidate embedding vector generation models respectively, wherein each embedding vector generation model adopts a different model architecture or base parameter configuration, including but not limited to a semantic embedding model based on a Transformer structure, a semantic matching model optimized by contrast learning, and a word vector encoding model enhanced for proper nouns; Obtain a word vector set corresponding to each segmented unit of the data input, wherein each embedding vector generation model performs semantic modeling operations based on an internal segmentation mechanism and a context modeling mechanism after receiving the data input, constructs a corresponding context-related encoding representation after segmenting the semantic compressed representation into a plurality of segmented units, and outputs a word vector corresponding to each segmented unit.

7. The RAG recall rate improvement method in a voice question and answer scene according to claim 1, characterized in that, The method comprises the following steps: Performing cleaning processing on the original corpus based on a rule matching and character screening mechanism, deleting stop words, mood words, non-semantic punctuation marks, and character sequences with a confidence lower than a preset threshold in the original corpus, and obtaining a structured cleaned corpus; Inputting the structured cleaned corpus into a pre-trained language model to perform generative compression operations based on a context consistency maintenance mechanism, and outputting a semantic compressed representation, wherein the semantic compressed representation expresses the semantic core information retained in the original corpus in the form of a question-answer pair or a central event expression; The semantic compressed representation is used as data input.

8. A device for improving RAG recall rate in a voice question and answer scenario, characterized in that, The device is used to perform the RAG recall rate improvement method in a voice question and answer scene as claimed in any one of claims 1-7, and the device comprises an acquisition module (201), a processing module (202), and an output module (203), wherein: The acquisition module (201) is used to perform semantic cleaning processing on the original corpus containing the voice recognition result, and perform semantic compression on the cleaned original corpus to generate data input; The processing module (202) is used to perform vector generation operations based on the data input using a plurality of candidate embedding vector generation models respectively, and output word vectors; The processing module (202) is used to calculate a semantic fidelity score for each word vector; The processing module (202) is used to evaluate a plurality of semantic fidelity scores using a gradient ranking preference optimization algorithm, and select a target embedding vector generation model with the best semantic fidelity score from the plurality of candidate embedding vector generation models; The processing module (202) is used to calculate the term frequency value and the inverse document frequency value of each word for the data input, determine whether the word is a professional hot word, and screen out the professional hot word to construct a hot word table; The output module (203) is used to input the embedding vector output by the target embedding vector generation model and the hot word table into a question and answer module, and output a target answer text.

9. An electronic device, comprising: The electronic device comprises a processor (301), a communication bus (302), a user interface (303), a network interface (304), and a memory (305) for storing instructions, the user interface (303) and the network interface (304) are both used for communicating with other devices, the communication bus (302) is used for realizing the connection communication between components in the electronic device, and the processor (301) is used for executing the instructions stored in the memory (305) to make the electronic device execute the method in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores instructions, when the instructions are executed, the method in any one of claims 1-7 is executed.

Citation Information

Cited By

  • Enterprise-level unstructured knowledge governance-oriented method and storage medium

    CN121597848A