A Chinese semantic matching enhancement method and system based on a BAAI-bge model

By combining the BAAI-bge model with multi-source heterogeneous data and auxiliary libraries, the shortcomings of Chinese semantic matching technology in processing complex sentences and polysemous words are solved, efficient semantic matching is achieved in scenarios with limited resources or high real-time requirements, and the cross-domain adaptability of the model and user experience are improved.

CN120578756BActive Publication Date: 2025-10-10YICHUANG JINGYUN DIGITAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511094190.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-10-10
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

Existing Chinese semantic matching technology has shortcomings in deeply understanding complex sentences, polysemous words and metaphors, and is easily affected by noise and ambiguity. The model relies on high-quality annotated data, consumes a lot of resources, and is difficult to deploy in scenarios with limited resources or high real-time requirements. It also has poor cross-domain adaptability, resulting in recommendation bias and a decline in user experience.

Method used

Through the BAAI-bge model, combined with multi-source heterogeneous data and auxiliary libraries, semantic enrichment and matching enhancement are carried out, and auxiliary libraries are built using professional knowledge graphs and knowledge bases to split and reorganize structured data, perform semantic enhancement and result optimization, adapt to different model states, and improve the accuracy and robustness of semantic matching.

Benefits of technology

Improve the accuracy and cross-domain adaptability of Chinese semantic matching in scenarios with limited resources or high real-time requirements, improve the results of question-answering and retrieval tasks, provide richer result information, and enhance user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578756B_ABST
    Figure CN120578756B_ABST
Patent Text Reader

Abstract

The application discloses a Chinese semantic matching enhancement method and system based on a BAAI-bge model, and relates to the technical field of semantic matching. The method comprises the following steps: requirement judgment: judging the training state of the model to obtain a judgment result; data acquisition: acquiring multi-source heterogeneous data used for model training to obtain to-be-processed data, and acquiring search information when the model is used to obtain to-be-searched data. Through the judgment of the training state of the model, the semantic richness of the search data is carried out by using an auxiliary enhancement method for the model after training, and the training data is enriched by using a semantic enhancement method for the model before training, so that the semantic matching of the subsequent model when coping with the search data is enhanced. The two methods complement each other, are suitable for different application scenarios, help to improve the accuracy, robustness and cross-field adaptability of matching, and make the system perform better when processing complex semantics and professional knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semantic matching technology, and in particular to a Chinese semantic matching enhancement method and system based on a BAAI-bge model. Background Art

[0002] Chinese semantic matching aims to determine the semantic similarity between two Chinese texts or whether they have similar meanings. It is widely used in scenarios such as question-answering systems, information retrieval, and knowledge question-answering. Chinese semantic expression is highly diverse and flexible. The same sentence can express the same meaning using different vocabulary and sentence structures. Some sentences are compact, while others are long and complex. Identifying these differences places higher demands on the model.

[0003] Patent publication number CN112966524A is a Chinese sentence semantic matching method and system based on a multi-granularity twin network. It uses Word2Vec to obtain pre-trained word vectors, and the input Chinese sentence sequence will be converted into a vector representation through the embedding layer; secondly, it enters the multi-granularity encoding layer to capture the complex semantic features of the sentence from the perspective of characters and words respectively; then, the feature vector output by the previous layer is input into the semantic interaction layer for semantic interaction; finally, the semantic interaction result is sent to the output layer to obtain the result of whether the sentence semantics are similar. The present invention proposes a new multi-granularity encoding method, which captures richer semantic information in the sentence from the perspectives of characters and words and obtains more features. The twin structure adopted by the present invention theoretically reduces the number of parameters, so that the model can achieve faster training speed.

[0004] Current semantic matching technology still has shortcomings in deeply understanding complex sentences, polysemous words and metaphors, and is easily affected by noise and ambiguity, leading to misjudgment or understanding deviation. In addition, the model is highly dependent on a large amount of high-quality annotated data, consumes a lot of resources, and is difficult to deploy in scenarios with limited resources or high real-time requirements. Moreover, the model performs poorly in cross-domain or multimodal information fusion, has insufficient generalization ability, and is prone to understanding deviation or incorrect matching, especially in applications involving professional terminology, polysemous words or non-standard expressions. These drawbacks can lead to recommendation bias, information loss, decreased user experience and other problems in actual applications, limiting the widespread implementation and development of semantic matching technology, and thus the present invention is proposed. Summary of the Invention

[0005] The purpose of the present invention is to provide a Chinese semantic matching enhancement method and system based on the BAAI-bge model to solve the problems raised in the above background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a Chinese semantic matching enhancement method based on the BAAI-bge model, comprising:

[0007] Demand determination: determine the model training state to obtain a determination result;

[0008] Data acquisition: acquire multi-source heterogeneous data used for model training to obtain to-be-processed data, and acquire retrieval information when the model is used to obtain to-be-retrieved data;

[0009] Data processing: pre-process the to-be-processed data and the to-be-retrieved data to obtain processed data and retrieval data;

[0010] It is characterized in that it comprises:

[0011] Determination result analysis: the determination result comprises a model training completion state and a model training incomplete state;

[0012] Semantic enrichment: based on the processed data, an auxiliary construction method is used to construct a semantic library for enriching data to obtain an auxiliary library, based on the auxiliary library and the model training incomplete state, a semantic enhancement method is used to enrich the semantics of the processed data to obtain enriched data, and based on the enriched data, semantic matching enhancement is performed on the model training;

[0013] Matching enhancement: the retrieval data is split to obtain a plurality of structured data, the splitting process is recorded to obtain data numbers, based on the structured data and the auxiliary library, an auxiliary enhancement method is used to enrich the semantics of the structured data and reorganize them to obtain a plurality of matching data;

[0014] Result optimization: based on the matching data, a result optimization method is used to obtain result information based on the model in the training completion state to complete semantic matching enhancement.

[0015] Further, the auxiliary construction method comprises: acquiring the working field of the model to obtain a target field, acquiring professional field knowledge graphs and knowledge information in the target field to obtain enhancement information, presetting hypothetical data, splitting the processed data and the hypothetical data to obtain first sub-data and second sub-data, acquiring field information of the first sub-data and the second sub-data to obtain first and second fields, traversing the enhancement information based on the first and second fields to obtain first and second related fields, matching the information of the first and second sub-data in the first and second related fields to obtain first and second related data, integrating the first related data and the first sub-data to obtain first enriched data, integrating the second related data and the second sub-data to obtain second enriched data, and integrating the first and second enriched data to obtain the auxiliary library.

[0016] Furthermore, the first sub-data and the second sub-data are target data, the source of the target data is recorded, keywords are extracted from the target data to obtain target keywords, homophones, synonyms and antonyms are obtained based on the target keywords to obtain added keywords, alternative words for the target keywords are obtained in real time to obtain new keywords, the added keywords and the new keywords are added to the target data accordingly to obtain derivative data, and the derivative data is stored in the auxiliary library according to the source of the target data to complete the real-time update of the auxiliary library.

[0017] Furthermore, the semantic enhancement method includes: extracting data corresponding to the processed data in the auxiliary library to obtain first rich data, the first rich data includes several data points, presetting hypothetical factors, adding hypothetical factors to the data points and obtaining hypothetical results to obtain derived rich data, expanding the field based on the field information of the data points to obtain derived fields, obtaining relevant information of the data points in the derived fields to obtain second rich data, and integrating the first rich data and the second rich data to obtain rich data.

[0018] Furthermore, the process of presetting hypothetical factors is: obtaining variable information of several data points to obtain target factors, obtaining synonyms, antonyms and synonyms of the target factors to obtain derived factors, and integrating all derived factors to obtain hypothetical factors. Hypothetical factors are multidimensional factors.

[0019] Furthermore, the auxiliary enhancement method includes: obtaining a data retrieval direction based on retrieval data, extracting variable information from a number of structured data to obtain target structured data, extracting information corresponding to the target structured data in the auxiliary library to obtain structure-enriched data, reorganizing the retrieval data based on the data number to obtain a number of reorganized data, comparing the retrieval data and the reorganized data based on the data retrieval direction to determine the rationality of the reorganized data to obtain a rationality result, and eliminating unreasonable reorganized data from the number of reorganized data based on the rationality result to obtain matching data.

[0020] Furthermore, the result optimization method includes: presetting usage requirements, the usage requirements include expansion requirements and efficiency requirements, obtaining target requirements, determining the final requirements based on the target requirements and the usage requirements, determining the data closest to the search data among several matching data to obtain the optimal data, when the final requirement is an expansion requirement, importing several matching data and search data into the model to obtain result information, and when the final requirement is an efficiency requirement, importing the optimal data and search data into the model to obtain result information.

[0021] A Chinese semantic matching enhancement system based on the BAAI-bge model uses the above-mentioned Chinese semantic matching enhancement method based on the BAAI-bge model.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] The Chinese semantic matching enhancement method and system based on the BAAI-bge model determines the training status of the model to obtain the training status result or the untrained status result. For the model when the training is completed, the auxiliary enhancement method is used to semantically enrich the retrieval data to facilitate the model's reasoning and matching. For the model when the training is not completed, the semantic enhancement method is used to enrich the training data to enhance the semantic matching of the subsequent model when dealing with the retrieval data. At the same time, after the training is completed with the training data, the above operations for the trained model can also be added to enhance the model's processing effect of Chinese semantic matching on the retrieval data. The two methods are used to target models in different states to achieve deployment in scenarios with limited resources or high real-time requirements. The two methods complement each other and are suitable for different application scenarios. They help to improve the accuracy, robustness and cross-domain adaptability of matching, so that the system performs better in processing complex semantics and professional knowledge, and improves the effects of question answering, retrieval and reasoning tasks.

[0024] At the same time, the first rich data and the second rich data are integrated to obtain an auxiliary library, that is, the two rich data sets are integrated to generate an auxiliary library, and an overview resource with domain knowledge depth and information breadth is obtained to facilitate subsequent information enrichment processing. In order to further improve the efficiency of model output, the hypothetical data in the auxiliary construction method is the preset retrieval data. By presetting, the efficiency of subsequent enrichment of the retrieval data is improved. The training data and the hypothetical data are placed in the auxiliary library at the same time, which can increase the information richness of the auxiliary library and facilitate the subsequent full enrichment of the required retrieval data. After binding the derived data with the corresponding source information of the source data, it is dynamically stored in the auxiliary library, so that the auxiliary library maintains real-time, rich and diverse semantic expressions, supports subsequent semantic matching and knowledge retrieval, and enhances the diversity and expression ability of the data through extensions such as synonyms and antonyms. After training with the enriched data, the model is not limited to the original word form when matching.

[0025] At the same time, through the setting of semantic enhancement methods, multi-level and multi-angle knowledge supplementation is reflected, the semantic expression ability of training data is enhanced, and the semantic depth and breadth are improved. The rich semantic information helps the model to better identify similar semantics and enhance the robustness of the model. The domain expansion mechanism enables the model to adapt to different or emerging fields and enhance the generalization ability. Through the setting of auxiliary enhancement methods, the structured retrieval data is reorganized after corresponding enriched semantics to obtain several reorganized data, and the rationality of the reorganized data is judged. This can increase the scope of the retrieval data imported into the model, improve the richness of the data in the imported model, thereby enhancing Chinese semantic matching, providing users with richer result information, and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1It is a schematic diagram of the overall process structure of the present invention;

[0027] Figure 2 This is a schematic structural diagram of the auxiliary construction method of the present invention;

[0028] Figure 3 This is a schematic diagram of the retrieval data splitting structure of the present invention;

[0029] Figure 4 Schematic diagram of the semantic enhancement structure of the model in different states of the present invention. DETAILED DESCRIPTION

[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0031] The BAAI-bge model is a series of general semantic vector models developed by the Beijing Academy of Artificial Intelligence (BAAI) to improve the accuracy and efficiency of text retrieval systems. Based on a cross-encoder architecture, this model can re-rank the top documents returned by the retrieval system to optimize search results. Its core features include multilingual support (Chinese and English), efficient semantic understanding, and excellent performance in processing large amounts of text data. Enhanced Chinese semantic matching refers to a series of methods and techniques that improve the model's ability to understand and judge the semantic relationships between Chinese texts in natural language processing (NLP) tasks. Its core goal is to enable the model to more accurately and robustly determine the semantic similarity or correlation between two Chinese texts.

[0032] like Figures 1-4 As shown, the present invention provides a technical solution: a Chinese semantic matching enhancement method based on the BAAI-bge model, comprising:

[0033] Demand judgment: judge the model training status and obtain the judgment result;

[0034] It should be noted that by judging the training status of the model, the status results of "training completed" or "not trained completed" are obtained. When training is completed, it mainly relies on the trained model and auxiliary enhancement methods to semantically enrich the retrieval data to facilitate model reasoning and matching; when training is not completed, it relies on semantic enhancement methods to enrich the training data to enhance the semantic matching of subsequent models when dealing with retrieval data.

[0035] Data acquisition: Acquire multi-source heterogeneous data used in model training to obtain the data to be processed, and obtain the retrieval information when the model is used to obtain the data to be retrieved;

[0036] It should be noted that obtaining the data to be processed means obtaining the multi-source heterogeneous data used for model training, that is, including structured, semi-structured or unstructured training source data (such as text, databases, knowledge graphs, sensor data, etc.); obtaining the data to be retrieved means obtaining the retrieval information in the model usage phase, that is, the target data to be retrieved. Multi-source heterogeneous data means that the data has differences in format, source, content, etc., such as social media text, industry databases, patent information, etc. This link provides a data basis for subsequent processing, enrichment and matching.

[0037] Data processing: pre-process the data to be processed and the data to be retrieved to obtain the processed data and the retrieved data;

[0038] It should be noted that the preprocessing process includes but is not limited to text cleaning and denoising, with the aim of unifying data formats, reducing noise, and providing a basis for semantic enrichment and matching.

[0039] The invention is characterized by comprising:

[0040] Judgment result analysis: Judgment results include the state where model training is completed and the state where model training is not completed;

[0041] It should be noted that the untrained state also includes the need to input new data into the model for subsequent training.

[0042] Semantic enrichment: Based on the processed data and the retrieved data, an auxiliary construction method is used to construct a semantic library for data enrichment to obtain an auxiliary library. Based on the auxiliary library and the untrained state of the model, a semantic enhancement method is used to semantically enrich the processed data to obtain enriched data. The model is trained based on the enriched data to complete semantic matching enhancement;

[0043] It should be noted that an auxiliary construction method is adopted to use domain knowledge (knowledge graphs, professional vocabulary, etc.) to build an auxiliary library for data enrichment. When the model training is not yet completed (for example, in the early stage of the model or the online adjustment stage), the processed data is enhanced by using the auxiliary library through the "semantic enhancement method". In the specific implementation process, knowledge entities and relationships are introduced into the text or data existing in the processed data to expand the semantic expression so that the data "reflects the potential semantics more richly and in a deeper level". The model is trained with enriched data to improve the semantic understanding ability and matching effect of the model, so as to improve the ability of subsequent models to cope with retrieval data. This part realizes the "knowledge supplement" and "semantic expansion" of the data before training. It makes full use of external domain knowledge to make up for the lack of knowledge when the model training is not completed, and guides the model to learn deep semantics early by enhancing the data, so as to achieve a certain enhanced matching effect before the training is completed.

[0044] Matching enhancement: Split the search data to obtain several structured data, record the splitting process to obtain data numbers, and use the auxiliary enhancement method based on the structured data to semantically enrich the structured data and reorganize it to obtain several matching data;

[0045] It should be noted that the retrieval data is split into several structured sub-data (such as entities, tags, paragraphs), and the numbers of each sub-data in the splitting process are recorded for tracking and reorganization. The structured sub-data are semantically enhanced and reorganized using the auxiliary library to form matching data. This link realizes the combination of "structuring + knowledge", integrates scattered information into knowledge-structured fragments, improves the semantic depth and breadth of matching, and records the data number, that is, the memory split number, to ensure that it can be accurately reorganized after enrichment, maintaining the integrity and traceability of the data.

[0046] Result optimization: Based on the matching data, the result optimization method is used in conjunction with the trained model to obtain the result information to complete semantic matching enhancement.

[0047] It should be noted that the number of matching data may be multiple. The matching data obtained by the set result optimization method is combined with the user's usage needs to be imported into the model to generate result information that meets the user's needs. The semantic matching is enhanced by importing the enriched data into the model. By introducing the professional knowledge base or knowledge graph into the trained model, the semantic expression of the data to be retrieved is enhanced by using rich entity and relationship information, thereby achieving deeper semantic matching. At the same time, in the process of model training, the training data is enriched with auxiliary knowledge to improve the model's ability to understand professional knowledge and implicit relationships, such as Figure 4 As shown in (a) and (b), the process of semantic matching enhancement is carried out for both trained and untrained models under different training states. These two methods complement each other and are suitable for different application scenarios. They help improve the accuracy, robustness and cross-domain adaptability of matching, making the system perform better when processing complex semantics and professional knowledge, and improving the results of question answering, retrieval and reasoning tasks.

[0048] like Figure 2As shown, the auxiliary construction method includes: obtaining the working domain of the model to obtain the target domain, obtaining the professional domain knowledge graph and knowledge information in the target domain to obtain enhanced information, splitting the processed data and the retrieved data to obtain the first sub-data and the second sub-data, obtaining the domain information of the first sub-data and the second sub-data to obtain the first domain and the second domain, traversing the enhanced information based on the first domain and the second domain to obtain the first related domain and the second related domain, matching the information of the first sub-data and the second sub-data in the first related domain and the second related domain to obtain the first related data and the second related data, integrating the first related data and the first sub-data to obtain the first enriched data, integrating the second related data and the second sub-data to obtain the second enriched data, and integrating the first enriched data and the second enriched data to obtain the auxiliary library.

[0049] It should be noted that the process of obtaining the working domain of the model is to obtain the dominant domain of the current application or training of the model. After identifying the target domain, as the basis for subsequent knowledge and data processing, in the target domain, through professional domain knowledge graphs, related databases and other resources, rich structured and unstructured knowledge information is extracted. This information constructs enhanced information and provides deep knowledge support for subsequent data processing. The processing data and the retrieval data are split to obtain the first sub-data and the second sub-data, that is, multiple data in the processing data and the retrieval data are split into single independent data, and the domain information of the first sub-data and the second sub-data are analyzed respectively to obtain the first domain and the second domain. This step is to clarify the professional domain to which each piece of data belongs, and to provide a basis for subsequent domain traversal and matching. Based on the first domain and the second domain, the relationship network of the knowledge graph is used to traverse and expand the first related domain and the second related domain, including those closely related to the original domain, Potentially related or overlapping secondary fields, this process ensures a more comprehensive knowledge base and enriches the semantics and structure of the data used for model training and for retrieval. In the first related field and the second related field, the information in the first sub-data and the second sub-data is combined to perform a matching operation. The goal is to find the first related data and the second related data, that is, the matching result that is more in line with the domain semantics of the original data. The first enriched data and the second enriched data are integrated to obtain an auxiliary library, that is, the two enriched data sets are integrated to generate an auxiliary library, and an overview resource with domain knowledge depth and information breadth is obtained to facilitate subsequent information enrichment processing. In order to further improve the efficiency of model output, the hypothetical data in the auxiliary construction method is the pre-set retrieval data. By presetting, the efficiency of subsequent enrichment of the retrieval data is improved. By placing the training data and the hypothetical data in the auxiliary library at the same time, the information richness of the auxiliary library can be increased, which is convenient for the subsequent full enrichment of the required retrieval data.

[0050] The first sub-data and the second sub-data are target data, which record the source of the target data, perform keyword extraction on the target data to obtain target keywords, obtain homophones, synonyms and antonyms based on the target keywords to obtain added keywords, obtain alternative words for the target keywords in real time to obtain newly added keywords, add the added keywords and the newly added keywords to the target data accordingly to obtain derivative data, and store the derived data in the auxiliary library in accordance with the source of the target data to complete the real-time update of the auxiliary library.

[0051] It should be noted that, during the processing, by recording the "source" for each target data, that is, the original source from which the data is extracted, the traceability of the data and the integrity of the source information are guaranteed, and the core keywords (such as entity names, professional terms, key phrases) are automatically extracted from the target data as the basis for expansion and matching. In the specific implementation process, keyword extraction algorithms (TF-IDF, TextRank, NER entity recognition, etc.) can be used to implement it, and word vectors (such as Word2Vec, GloVe), part-of-speech networks or knowledge bases (such as synonyms, professional thesaurus) are used to find words that are similar in meaning, pronunciation or semantics to the target keywords. For example, the target keyword "car" can be expanded to "car", "sedan", and "vehicle"; the process of obtaining alternative words for the target keyword in real time to obtain the newly added keywords is to retrieve the corresponding keywords on the Internet in real time. Close the replacement words to avoid lag, generate multiple versions with similar semantics but different expressions through keyword replacement, juxtaposition, expansion, etc., improve data diversity, bind the derived data with the corresponding source information of the source data, and dynamically store it in the auxiliary library to keep the auxiliary library in real-time, rich and diverse semantic expressions, support subsequent semantic matching and knowledge retrieval, and enhance data diversity and expression capabilities through synonyms, antonyms and other extensions. After training with enriched data, the model is not restricted to the original word form when matching. The derivative expansion of data reduces matching failures caused by differences in word expression, improves the fault tolerance of the model and system, updates the auxiliary library in real time, forms a closed loop, continuously enriches the knowledge graph and semantic support, and provides a more complete knowledge background. In different applications, there is no need to retrain the model, and more accurate matching can be achieved based on the dynamically expanded keywords and knowledge base.

[0052] like Figure 1 As shown, the semantic enhancement method includes: extracting data corresponding to the processed data in the auxiliary library to obtain first rich data, the first rich data includes several data points, presetting hypothetical factors, adding hypothetical factors to the data points and obtaining hypothetical results to obtain derived rich data, expanding the field based on the field information of the data points to obtain a derived field, obtaining relevant information of the data points in the derived field to obtain second rich data, and integrating the first rich data and the second rich data to obtain rich data.

[0053] It should be noted that obtaining the first enriched data means filtering out knowledge points or data segments related to the currently processed data from the auxiliary library. In the specific implementation process, keywords, entities or semantic features can be used to match knowledge items in the auxiliary library. At the same time, similarity calculation (such as vector similarity) or rule matching can also be used to filter out relevant content, lay the foundation for rich semantics, and provide knowledge support for subsequent hypothetical reasoning. On the basis of the original data points, based on preset hypothetical factors (such as "if", "maybe", "in a certain situation"), hypothetical reasoning is performed to generate multi-dimensional and multi-angle derivative enriched data, add hypothetical words ("if", "perhaps", "possible") to the data points, simulate different conditions or inference scenarios, and generate corresponding inference results or inference information based on the derived enriched data obtained by adding hypothetical factors, enrich the data content, enhance the reasoning dimension of the data, and enable the model to process the corresponding data. It is more flexible when it comes to information that is similar but in different situations. It uses the domain labels or domain features of data points and generates derivative domains through domain migration or expansion technology. In the specific implementation process, the association relationships in the knowledge graph can be used to search for other domains related to the current domain. Through this process, the coverage of knowledge and information is expanded, and the system's adaptability to unknown or cross-domain problems is enhanced. The process of obtaining relevant information in the derivative domain is consistent with the way of obtaining enhanced information. When the derivative domain is located in the target domain, the keywords of the derivative domain are used to retrieve relevant content from the auxiliary library. Through the setting of the semantic enhancement method, multi-level and multi-angle knowledge supplementation is reflected, the semantic expression ability of the training data is enhanced, and the semantic depth and breadth are improved. The rich semantic information helps the model better identify similar semantics and enhance the robustness of the model. The domain expansion mechanism enables the model to adapt to different or emerging fields and enhance generalization capabilities.

[0054] The process of presetting hypothetical factors is: obtaining variable information of several data points to obtain the target factor, obtaining synonyms, antonyms and synonyms of the target factor to obtain the derived factor, and integrating all the derived factors to obtain the hypothetical factor. The hypothetical factor is a multidimensional factor.

[0055] It should be noted that the process of extracting variable information can be obtained by training a dedicated variable extraction model, and the variable information is the changeable information in the data points that does not affect the original directional meaning of the data points, such as month, etc. The process of obtaining the derived factors of the target factor can use natural language processing (NLP) dictionaries, word libraries, word vector models (such as Word2Vec, GloVe) and other resources to search for target factors, and integrate all related derived factors into hypothetical factors. This is a set operation that includes multiple dimensions (such as time and space), forming a multi-angle, multi-variable hypothesis expression. Rich hypothetical factors provide multi-dimensional, multi-angle reasoning clues for the model, and multiple synonymous, antonymous, and synonymous word expansions ensure consistent understanding in different contexts and expressions. Multi-dimensional hypotheses can handle complex or variable scenarios, improving the flexibility of the model in matching, and rich hypothesis factors help the system consider more possibilities, making more comprehensive and reasonable judgments.

[0056] As shown in Figure 1 and Figure 3 , the auxiliary enhancement method includes: obtaining data retrieval direction based on retrieval data, extracting variable information in a plurality of structured data to obtain target structured data, extracting information corresponding to the target structured data in the auxiliary library to obtain structured rich data, reorganizing the retrieval data based on the data number to obtain a plurality of reorganized data, comparing the retrieval data and the reorganized data based on the data retrieval direction to determine the rationality of the reorganized data to obtain a rationality result, and eliminating unreasonable reorganized data in the plurality of reorganized data based on the rationality result to obtain matching data.

[0057] It should be noted that the data retrieval direction is a specific directional feature, such as forward retrieval and reverse retrieval. The specific role is to avoid the combination of reorganized data and antonyms of variables, which can cause the reorganized data to have an antonym bias with the retrieval data. By setting the auxiliary enhancement method, the retrieval data processed by the structured method is reorganized after corresponding semantic enrichment, a plurality of reorganized data are obtained, and the rationality of the reorganized data is judged. This can increase the range of importing the retrieval data into the model, improve the richness of the data in the imported model, and enhance the semantic matching of the noon, thereby providing users with more rich result information and improving the user experience.

[0058] As shown in Figure 1 , the result optimization method includes: presetting usage requirements, the usage requirements including expansion requirements and efficiency requirements, obtaining target requirements, determining final requirements based on the target requirements and usage requirements, determining the data closest to the retrieval data in the plurality of matching data to obtain optimal data, and when the final requirement is the expansion requirement, importing the plurality of matching data and the retrieval data into the model to obtain result information, and when the final requirement is the efficiency requirement, importing the optimal data and the retrieval data into the model to obtain result information.

[0059] It should be noted that the expansion demand is the result information after the user needs to expand the retrieved data, the efficiency demand is the result information after the user needs and the retrieved data is not expanded, the target demand is obtained through the user, and the final demand is determined by combining the target demand with the usage demand. By importing several matching data and retrieval data into the model, the result information after the retrieval data is expanded can be obtained. By importing the optimal data and retrieval data into the model, the data processing sample is increased and the richness of the model processing data is improved, so as to find the results that meet the user needs and realize the enhancement of Chinese semantic matching.

[0060] Although embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is limited by the accompanying embodiments and their equivalents.

Claims

1. A Chinese semantic matching enhancement method based on the BAAI-bge model, comprising: Demand judgment: judge the model training status and obtain the judgment result; Data acquisition: Acquire multi-source heterogeneous data used in model training to obtain the data to be processed, and obtain the retrieval information when the model is used to obtain the data to be retrieved; Data processing: pre-process the data to be processed and the data to be retrieved to obtain the processed data and the retrieved data; The invention is characterized by comprising: Judgment result analysis: Judgment results include the state where model training is completed and the state where model training is not completed; Semantic enrichment: Based on the processed data, an auxiliary construction method is used to build a semantic library for enriching data to obtain an auxiliary library. Based on the auxiliary library and the untrained state of the model, a semantic enhancement method is used to semantically enrich the processed data to obtain enriched data. Model training is performed based on the enriched data to complete semantic matching enhancement; Matching enhancement: Split the search data to obtain several structured data, record the splitting process to obtain data numbers, and use the auxiliary enhancement method based on the structured data to semantically enrich the structured data and reorganize it to obtain several matching data; Result optimization: Based on the matching data, the result optimization method is used in conjunction with the trained model to obtain the result information and complete the semantic matching enhancement; The auxiliary construction method includes: obtaining the working domain of the model to obtain the target domain, obtaining the professional domain knowledge graph and knowledge information in the target domain to obtain enhanced information, presetting hypothetical data, splitting the processed data and the hypothetical data to obtain the first sub-data and the second sub-data, obtaining the domain information of the first sub-data and the second sub-data to obtain the first domain and the second domain, traversing the enhanced information based on the first domain and the second domain to obtain the first related domain and the second related domain, matching the information of the first sub-data and the second sub-data in the first related domain and the second related domain to obtain the first related data and the second related data, integrating the first related data and the first sub-data to obtain the first enriched data, integrating the second related data and the second sub-data to obtain the second enriched data, and integrating the first enriched data and the second enriched data to obtain the auxiliary library.

2. The Chinese semantic matching enhancement method based on the BAAI-bge model according to claim 1 is characterized by: The first sub-data and the second sub-data are target data, which record the source of the target data, perform keyword extraction on the target data to obtain target keywords, obtain homophones, synonyms and antonyms based on the target keywords to obtain added keywords, obtain alternative words for the target keywords in real time to obtain new keywords, add the added keywords and the new keywords to the target data accordingly to obtain derivative data, and store the derived data in the auxiliary library in accordance with the source of the target data to complete the real-time update of the auxiliary library.

3. The Chinese semantic matching enhancement method based on the BAAI-bge model according to claim 1 is characterized in that: The semantic enhancement method includes: extracting data corresponding to the processed data in the auxiliary library to obtain first enriched data, the first enriched data includes several data points, presetting hypothetical factors, adding hypothetical factors to the data points and obtaining hypothetical results to obtain derived enriched data, expanding the field based on the field information of the data points to obtain derived fields, obtaining relevant information of the data points in the derived fields to obtain second enriched data, and integrating the first enriched data and the second enriched data to obtain enriched data.

4. The Chinese semantic matching enhancement method based on the BAAI-bge model according to claim 3 is characterized by: The process of presetting hypothetical factors is: obtaining variable information of several data points to obtain the target factor, obtaining synonyms, antonyms and synonyms of the target factor to obtain the derived factor, and integrating all the derived factors to obtain the hypothetical factor. The hypothetical factor is a multidimensional factor.

5. The Chinese semantic matching enhancement method based on the BAAI-bge model according to claim 1 is characterized in that: The auxiliary enhancement method includes: obtaining a data retrieval direction based on retrieval data, extracting variable information from a plurality of structured data to obtain target structured data, extracting information corresponding to the target structured data in an auxiliary library to obtain structure-enriched data, reorganizing the retrieval data based on the data number to obtain a plurality of reorganized data, comparing the retrieval data and the reorganized data based on the data retrieval direction to determine the rationality of the reorganized data to obtain a rationality result, and eliminating unreasonable reorganized data from the plurality of reorganized data based on the rationality result to obtain matching data.

6. The Chinese semantic matching enhancement method based on the BAAI-bge model according to claim 1 is characterized in that: The result optimization method includes: presetting usage requirements, where the usage requirements include expansion requirements and efficiency requirements, obtaining target requirements, determining final requirements based on the target requirements in combination with usage requirements, determining the data closest to the search data among a number of matching data to obtain optimal data, and when the final requirement is an expansion requirement, importing a number of matching data and search data into the model to obtain result information; and when the final requirement is an efficiency requirement, importing the optimal data and search data into the model to obtain result information.

7. A Chinese semantic matching enhancement system based on the BAAI-bge model, characterized by: A Chinese semantic matching enhancement method based on the BAAI-bge model as described in any one of claims 1-6 is used.

Citation Information

Patent Citations

  • Chinese sentence semantic matching method and system based on multi-granularity twin network

    CN112966524A

  • Emotion recognition method and device, equipment and storage medium

    CN115050077A

  • Urban building safety risk knowledge graph construction method

    CN119378661A