Long text matching method and system based on large model, terminal and medium

By extracting key semantic information of long text based on the big model and performing similarity calculation, the problems of low efficiency and low accuracy of long text similarity calculation in the prior art are solved, and efficient and accurate long text matching is achieved.

CN119940336AActive Publication Date: 2025-05-06ZHENJUE TECH (SHANGHAI) CO LTD

Patent Information

Application Number
CN202510415233.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process the similarity calculation of long texts, especially when the text is very long. The training data set has high requirements for device video memory, and traditional methods cannot understand semantics and have low accuracy.

Method used

Using a long text matching method based on a large model, the key semantic information of the long text is extracted through the first large model and similarity calculation is performed based on these key information, including dividing the key semantic information into keywords for matching, or extracting the key semantic information of the second long text for similarity calculation.

Benefits of technology

It solves the problem that traditional models are difficult to deal with long text, reduces the requirements for video memory, improves the efficiency and accuracy of calculations, and can effectively retain semantic information for matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940336A_ABST
    Figure CN119940336A_ABST
Patent Text Reader

Abstract

The invention provides a large model-based long text matching method and system, a terminal and a medium. The method comprises the following steps of: extracting first key semantic information of a first long text by utilizing a first large model; performing matching similarity calculation with the second long text based on the first key semantic information; wherein matching similarity calculation is carried out based on the first key semantic information and the second long text, and the method further comprises any one of the following steps: dividing the first key semantic information into corresponding keywords, and carrying out matching similarity calculation on the keywords and the second long text; extracting second key semantic information of the second long text by utilizing the first large model; and performing matching similarity calculation on the first key semantic information and the second key semantic information. According to the method, dimension reduction of the long text is assisted through the large model, the problem of text processing is firstly solved, key information in the original long text can be well reserved, and meanwhile the problem that a traditional method occupies high video memory is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text processing technology, and in particular to a long text matching method and system based on a large model, as well as a corresponding computer terminal and a computer-readable storage medium. Background Art

[0002] The similarity calculation of long texts is still a difficult problem in this field. In actual business scenarios, long texts are not necessarily processed and unambiguous pure texts (that is, data used in academia, data that can be directly trained). Actual long texts may contain a lot of noise, such as various symbols, spaces, line breaks, etc. with unknown meanings. First, the obtained data must be cleaned to obtain clean pure text information, and then the similarity of the long texts can be calculated. At present, the methods for calculating the similarity of long texts can be roughly divided into two categories. One is the deep learning method. At present, the better deep learning models for calculating the similarity of long texts are Sentence-BERT and Longformer models. However, for such deep learning models, it is difficult to process the similarity between long texts of thousands or tens of thousands of words, because when the text is very long, the training data set has very high requirements for the device video memory. Another method is the SimHash algorithm. The main work of the SimHash algorithm is to reduce the dimension of the text and generate a SimHash value. The SimHash values ​​of different texts are compared by the Hamming distance to determine the similarity of two texts. The characteristic is that when the text content is long, the accuracy of SimHash is higher, but the accuracy of SimHash for short text content is often not guaranteed. The calculation using this method is more complicated and cannot understand the semantics.

[0003] After searching, it was found that the Chinese invention patent application "A text matching method for long texts" with publication number CN117828028A adopts a dual-tower Longformer model to build text information for matching text pairs. As mentioned above, the Longformer model used in this method has high requirements for video memory and is difficult to process texts with tens of thousands of characters. Summary of the invention

[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a long text matching method, system, terminal and medium based on a large model.

[0005] According to one aspect of the present invention, a long text matching method based on a large model is provided, comprising: Extracting first key semantic information of the first long text using the first large model; Based on the first key semantic information, performing matching similarity calculation with the second long text; The step of performing matching similarity calculation with the second long text based on the first key semantic information further includes any one of the following: - dividing the first key semantic information into corresponding keywords, and performing matching similarity calculation between the keywords and the second long text; - Using the first large model to extract second key semantic information of the second long text; and performing matching similarity calculation on the first key semantic information and the second key semantic information.

[0006] According to a second aspect of the present invention, a long text matching system based on a large model is provided, comprising: A first key information extraction module, which uses a first large model to extract first key semantic information of a first long text; A similarity matching module, which performs matching similarity calculation with the second long text based on the first key semantic information; Wherein, the similarity matching module further includes any one of the following: - a direct matching unit, which is used to divide the first key semantic information into corresponding keywords and perform matching similarity calculation between the keywords and the second long text; - A second key information extraction unit and a similarity calculation unit; wherein: The second key information extraction unit extracts second key semantic information of the second long text using the first large model; The similarity calculation unit is used to perform matching similarity calculation on the first key semantic information and the second key semantic information.

[0007] According to a third aspect of the present invention, another long text matching method based on a large model is provided, comprising: Use the second largest model to perform dimensionality reduction on several known long text pairs with annotated matching degrees to construct a training data set; Using the training data set to train a deep learning model; The new long text pair is used as input to the trained deep learning model and the matching score is output.

[0008] According to a fourth aspect of the present invention, another large model-based long text matching system is provided, comprising: A training sample construction module, which uses the second largest model to perform dimensionality reduction on several known long text pairs with annotated matching degrees to construct a training data set; A model training module, which uses the training data set to train a deep learning model; The matching module is used to take new long text pairs as input to the trained deep learning model and output a matching score.

[0009] According to a fifth aspect of the present invention, there is provided a computer terminal comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, can be used to execute the method described above in the present invention, or to run the system described above in the present invention.

[0010] According to a sixth aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can be used to execute the method described above in the present invention, or to run the system described above in the present invention.

[0011] Due to the adoption of the above technical solution, the present invention has at least one of the following beneficial effects compared with the prior art: The present invention uses a large model to generate keywords, solving the problem that traditional models are difficult to perform secondary generation of long texts.

[0012] The present invention matches the keywords of the first long text with the second long text, thereby solving the problem that the traditional model can only perform classification matching according to preset topics. These keywords are closer to the original text information than the preset topics, thus avoiding the problem of low matching degree in the future.

[0013] The present invention performs matching similarity calculation on the first key semantic information and the second key semantic information, thereby solving the problem of text matching with different text contents but the same semantics.

[0014] The present invention adopts a technology based on a large model to process long texts, which solves the problem that general long text matching technology requires high device video memory, and ensures accuracy while retaining semantic information matching.

[0015] The present invention uses a large model to help reduce the dimension of long texts. It first solves the problem of text processing and can better retain the key information in the original long text. At the same time, it solves the problem of high video memory occupancy of traditional methods and ensures accuracy while retaining semantic information matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 The present invention is a flowchart of a first large model-based long text matching method in one embodiment of the present invention.

[0017] Figure 2This is a flowchart of the first matching similarity calculation in a preferred embodiment of the present invention.

[0018] Figure 3 This is a flow chart of the second matching similarity calculation in a preferred embodiment of the present invention.

[0019] Figure 4 The figure is a schematic diagram of the components of the first large model-based long text matching system in one embodiment of the present invention.

[0020] Figure 5 The flowchart is a second large model-based long text matching method in one embodiment of the present invention.

[0021] Figure 6 The figure is a schematic diagram of component modules of a second large model-based long text matching system in one embodiment of the present invention.

[0022] Figure 7 This is a workflow diagram of a second large model-based long text matching system in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following is a detailed description of the embodiments of the present invention: This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and a specific operation process are given. It should be pointed out that for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.

[0024] The existing long text matching technology has high requirements for video memory, is difficult to process texts with tens of thousands of characters, is relatively complex to calculate and cannot understand semantics. In view of the above problems, an embodiment of the present invention provides a long text matching method based on a large model. The method uses a large model to help reduce the dimension of long texts. First, the problem of text processing is solved, and the key information in the original long text can be better preserved. At the same time, the problem of high video memory usage of traditional methods is solved. On the basis of retaining semantic information matching, the accuracy rate is guaranteed.

[0025] Specifically, Figure 1 As shown, the large model-based long text matching method provided in this embodiment may include: S1, extracting the first key semantic information of the first long text using the first large model; S2, based on the first key semantic information, performing matching similarity calculation with the second long text; The method of performing matching similarity calculation with the second long text based on the first key semantic information may further include any one or more of the following: S21, dividing the first key semantic information into corresponding keywords, and performing matching similarity calculation between the keywords and the second long text; S22, using the first large model to extract second key semantic information of the second long text; and performing matching similarity calculation on the first key semantic information and the second key semantic information.

[0026] In some preferred embodiments, in the above S11, the first large model can use the following pre-trained large models: any non-encoder-only model. Further, any large model with inductive generation capabilities can be used as the first large model. The large model can be trained by a self-constructed data set, or other trained open source large models can be used.

[0027] In some preferred embodiments, Figure 2 As shown, the above S21, dividing the first key semantic information into corresponding keywords, and performing matching similarity calculation between the keywords and the second long text, may further include: S211, dividing the key semantic information generated by the first large model into corresponding keywords, further preferably: By setting the corresponding prompt words, or even using Chain of Thought (CoT) technology to improve the reasoning ability of large language models, the required keyword information can be generated. In addition, according to grammatical rules, regular expressions can be used for secondary splitting to obtain the desired keywords.

[0028] S222, comparing the generated keywords with the second long text, the number of matched keywords is the similarity between the two texts, further preferably, including any one of the following methods: - Identify and count the frequency of occurrence of the generated keywords (i.e. the number of keywords) in the second long text, assign corresponding weight scores according to the importance of the keywords, and give a weighted score to the number of keywords, which is the similarity between the two texts; - You can pre-build a word list containing keyword synonyms, antonyms or specific abbreviations. When these words appear in the second long text, it is considered a successful match. The number of matched words is the similarity between the two texts.

[0029] In the above S21, the more keywords matched, the higher the direct match between the two texts. The matching similarity calculation method provided in this step is applicable to long texts where the two texts are very long and contain a lot of professional information. By extracting these key semantic information for comparison, the relevance between the two long texts can be effectively determined. For example, if both texts are academic papers, the method provided in this step is more effective.

[0030] In some preferred embodiments, Figure 3 As shown, the above S22, using the first large model to extract the second key semantic information of the second long text; matching the first key semantic information and the second key semantic information to calculate the similarity, can further include: S221, extracting key semantic information from the second long text through the first large model to obtain second key semantic information; S222, performing similarity calculation on the first key semantic information and the second key semantic information, including: Pearson correlation coefficient, Euclidean distance, Cosine similarity and Manhattan distance calculation methods.

[0031] In the above S22, the two key semantic information are similarly calculated to determine whether the two long texts are related. The matching similarity calculation method provided in this step is more suitable for common texts with less professional content.

[0032] Based on the same inventive concept, an embodiment of the present invention provides a long text matching system based on a large model.

[0033] Specifically, Figure 4 As shown, the large model-based long text matching system provided in this embodiment may include: A first key information extraction module, which uses a first large model to extract first key semantic information of a first long text; A similarity matching module, which performs matching similarity calculation with the second long text based on the first key semantic information; The similarity matching module may further include any one or more of the following: - A direct matching unit, which is used to divide the first key semantic information into corresponding keywords and perform matching similarity calculation between the keywords and the second long text; - A second key information extraction unit and a similarity calculation unit; wherein: A second key information extraction unit, extracting second key semantic information of the second long text using the first large model; The similarity calculation unit is used to perform matching similarity calculation on the first key semantic information and the second key semantic information.

[0034] The working contents of the functional modules constituting the large model-based long text matching system provided in the above embodiment of the present invention are further described in detail below.

[0035] In this system, the long text obtained is summarized by a large model to obtain key semantic information. At present, most mainstream large models have good ability to extract key semantic information from long texts. Of course, for specific business scenarios, the model with corresponding fine-tuning will have better results. The key semantic information obtained by summarization generally includes the key semantics of the long text. Next, the first key information extraction module and any functional unit of the similarity matching module can be used to implement matching similarity calculation with another long text.

[0036] Adopt direct matching unit; first, the key semantic information generated by the first large model needs to be divided into corresponding keywords through the first key information extraction module, and then the generated keywords are compared with another long text. The number of matched keywords is the similarity between the two texts. The more matched keywords, the higher the direct match between the two texts. This method is suitable for two very long texts containing a lot of professional information. Extracting these key information for comparison can effectively determine the relevance between the two texts. For example, when both texts are academic papers, this method works well.

[0037] The second key information extraction unit and the similarity calculation unit are used; first, the first key semantic information of a long text that has passed through the first large model is extracted through the first key information extraction module, and the second key semantic information of another long text that has passed through the first large model is extracted through the second key information extraction unit; then the two key semantic information are similarly calculated through the similarity calculation unit, and the similarity calculation includes Pearson correlation coefficient, Euclidean distance, Cosine similarity, Manhattan distance and other methods, which are relatively common methods in machine learning. This method is suitable for judging whether two texts are related. Because these methods require that the text cannot be too long, the information extracted by the large model can be understood as the titles of the two texts. This method is more suitable for ordinary texts with less professionalism.

[0038] An embodiment of the present invention also provides another long text matching method based on a large model.

[0039] Specifically, Figure 5 As shown, the large model-based long text matching method provided in this embodiment may include: M1, using the second largest model to reduce the dimensionality of several known long text pairs with annotated matching degrees to construct a training data set; M2, train a deep learning model using the training data set; M3 takes the new long text pair as the input of the trained deep learning model and outputs the matching score.

[0040] In some preferred embodiments, the above-mentioned M1, the second largest model can adopt the following pre-trained model: any non-encoder-only model.

[0041] In some preferred implementations, the above M1 uses the second largest model to perform dimensionality reduction processing on a number of known long text pairs with annotated matching degrees to construct a training data set, and may further include: M11, using the second largest model, extracts key information of several known long text pairs with annotated matching degrees, reduces the dimension of the known long text, and constructs new text; M12, combines new text with the annotated matches to form a training dataset.

[0042] Based on the same inventive concept, an embodiment of the present invention also provides another long text matching system based on a large model.

[0043] Specifically, Figure 6 As shown, the large model-based long text matching system provided in this embodiment may include: A training sample construction module, which uses the second largest model to perform dimensionality reduction on several known long text pairs with annotated matching degrees to construct a training data set; A model training module, which trains a deep learning model using a training data set; The matching module is used to take new long text pairs as input to the trained deep learning model and output a matching score.

[0044] In some preferred implementations, the above training sample construction module may further include: Long text dimension reduction unit, which uses the second largest model to extract key information of several known long text pairs with marked matching degrees, and reduces the dimension of the known long text to form a new text; A dataset construction unit is used to combine new text with annotated matches to form a training dataset.

[0045] The working contents of the functional modules constituting the large model-based long text matching system provided in the above embodiment of the present invention are further described in detail below.

[0046] In this system, a training data set is constructed to train the deep learning model to obtain the matching scores of unknown long text pairs, thereby judging the similarity between long texts and achieving long text matching.

[0047] like Figure 7As shown, further, several known long text pairs are obtained, and the known long text pairs have been annotated with the matching degree of the text pairs. According to the video memory of the existing equipment, a pre-trained large model is used to reduce the data dimension of the known long text pairs. After experimental testing, a text pair of about 5,000 words occupies about 800m of video memory when trained using the Longformer model. According to the existing equipment, the length of the text data is reduced to the range available to the equipment for deep learning model training. Finally, the deep learning model trained can be used to judge the similarity between texts. Here, the large model mainly plays a role in removing the original text noise, retaining useful information, reducing the text length, and reducing the demand for equipment. The deep learning model can use a Transformer model, or it can be a SentenceBert model and a Longformer model based on the Transformer architecture.

[0048] An embodiment of the present invention further provides a computer terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the processor can be used to execute any method of the above-mentioned embodiments of the present invention, or to run any system of the above-mentioned embodiments of the present invention.

[0049] Optionally, the memory is used to store programs; the memory may include volatile memory (English: volatile memory), such as random access memory (English: random-access memory, abbreviated: RAM), such as static random access memory (English: static random-access memory, abbreviated: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviated: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory is used to store computer programs (such as applications, functional modules, etc. that implement the above method), computer instructions, etc., and the above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.

[0050] The processor is used to execute the computer program stored in the memory to implement the various steps of the method or various modules of the system involved in the above embodiments. For details, please refer to the relevant descriptions in the above method and system embodiments.

[0051] The processor and the memory may be independent structures or integrated structures. When the processor and the memory are independent structures, the memory and the processor may be coupled and connected via a bus.

[0052] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it can be used to execute any method of the above embodiments of the present invention, or to run any system of the above embodiments of the present invention.

[0053] Among them, computer-readable media include computer storage media and communication media, wherein the communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general or special-purpose computer. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in a user device. Of course, the processor and the storage medium can also be present in a communication device as discrete components.

[0054] The large-model-based long text matching method, system, terminal and medium provided in the above-mentioned embodiments of the present invention adopt a technology based on a large model to process long texts, which solves the problem that general long text matching technology requires high device video memory, and ensures accuracy while retaining semantic information matching; by using a large model to help reduce the dimension of long texts, the problem of text processing is first solved, and the key information in the original long text can be better retained. At the same time, it solves the problem of high video memory occupation of traditional methods, and ensures accuracy while retaining semantic information matching.

[0055] All matters not covered in the above embodiments of the present invention are well known in the art.

[0056] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A long text matching method based on a large model, characterized in that: include: Extracting first key semantic information of the first long text using the first large model; Based on the first key semantic information, performing matching similarity calculation with the second long text; The step of performing matching similarity calculation with the second long text based on the first key semantic information further includes any one of the following: - dividing the first key semantic information into corresponding keywords, and performing matching similarity calculation between the keywords and the second long text; - Using the first large model to extract second key semantic information of the second long text; and performing matching similarity calculation on the first key semantic information and the second key semantic information.

2. The long text matching method based on a large model according to claim 1 is characterized in that: The dividing the first key semantic information into corresponding keywords includes: By using the set prompt words and / or thought chain technology, the key semantic information generated by the first large model is divided into corresponding first keywords; based on the grammatical rules, the first keywords are split twice using regular expressions to generate the final desired keywords; The matching similarity calculation between the keyword and the second long text includes any one of the following: - Compare the generated keywords with the second long text and count the number of matched keywords; assign weight scores according to the importance of the corresponding keywords, and calculate the weighted score of the number of keywords, which is the similarity between the two texts; - Using the generated keywords, construct a vocabulary containing keywords, keyword synonyms, keyword antonyms and / or specific abbreviations related to the keywords; comparing the vocabulary with the second long text, and the number of matched words is the similarity between the two texts.

3. The long text matching method based on a large model according to claim 1 is characterized in that: extracting second key semantic information of the second long text by using the first large model; The matching similarity calculation of the first key semantic information and the second key semantic information includes: Extracting second key semantic information from the second long text through the first large model; The first key semantic information and the second key semantic information are subjected to similarity calculation, including: Pearson correlation coefficient, Euclidean distance, Cosine similarity and Manhattan distance calculation methods.

4. A long text matching system based on a large model, characterized in that: include: A first key information extraction module, which uses a first large model to extract first key semantic information of a first long text; A similarity matching module, which performs matching similarity calculation with the second long text based on the first key semantic information; Wherein, the similarity matching module further includes any one of the following: - a direct matching unit, which is used to divide the first key semantic information into corresponding keywords and perform matching similarity calculation between the keywords and the second long text; - A second key information extraction unit and a similarity calculation unit; wherein: The second key information extraction unit extracts second key semantic information of the second long text using the first large model; The similarity calculation unit is used to perform matching similarity calculation on the first key semantic information and the second key semantic information.

5. A long text matching method based on a large model, characterized in that: include: Use the second largest model to perform dimensionality reduction on several known long text pairs with annotated matching degrees to construct a training data set; Using the training data set to train a deep learning model; The new long text pair is used as input to the trained deep learning model and the matching score is output.

6. The large model-based long text matching method according to claim 5, characterized in that: The second largest model is used to perform dimensionality reduction processing on a number of known long text pairs with annotated matching degrees to construct a training data set, including: Using the second largest model, extract key information of several known long text pairs with marked matching degrees, reduce the dimension of the known long texts, and construct new texts; The new text is combined with the marked matching degree to form a training data set.

7. A long text matching system based on a large model, characterized in that: include: A training sample construction module, which uses the second largest model to perform dimensionality reduction on several known long text pairs with annotated matching degrees to construct a training data set; A model training module, which uses the training data set to train a deep learning model; The matching module is used to take new long text pairs as input to the trained deep learning model and output a matching score.

8. The large model-based long text matching system according to claim 7, characterized in that: The training sample construction module includes: A long text dimension reduction unit, which uses the second largest model to extract key information of several known long text pairs with marked matching degrees, and reduces the dimension of the known long text to form a new text; A data set construction unit is used to combine the new text with the marked matching degree to form a training data set.

9. A computer terminal comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it can be used to execute the method described in any one of claims 1-3, 5-6, or run the system described in any one of claims 4, 7-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it can be used to execute the method described in any one of claims 1 to 3, 5 to 6, or to run the system described in any one of claims 4, 7 to 8.

Citation Information

Patent Citations

  • Project duplicate checking method, device and equipment and storage medium

    CN110377886A

  • Long text semantic similarity calculation method based on comparative learning

    CN114707516A

  • Feature category determination method and device, electronic equipment and storage medium

    CN114722162A

  • Long text abstract generation method, system and equipment and storage medium

    CN116050397A

  • Long text similarity calculation method and device based on semantic progressive fusion

    CN117113094A

Cited By

  • Strong association control method for long text streaming conversion

    CN120524919A

  • A strong correlation control method for long text streaming conversion

    CN120524919B