Text matching method and device based on probability distribution, equipment and storage medium
By obtaining semantic feature distribution from a professional knowledge base, calculating the probability distribution of user input text, and using similarity distance to determine text matching results, the problem of low text matching accuracy under noise influence in existing technologies is solved, and stable semantic matching in high-noise environments is achieved.
Patent Information
- Application Number
- CN202511101607.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-08-07
AI Technical Summary
When user input is noisy, existing text matching methods have low fault tolerance and cannot effectively improve matching accuracy.
By obtaining semantic feature distribution from a professional knowledge base, calculating the probability distribution of user input text, and using similarity distance calculation to determine text matching results, a probability distribution model is used to capture the core intent of user input, thus solving the problem of traditional methods being sensitive to noise.
It significantly improves the fault tolerance and matching accuracy of text matching, and can stably retrieve semantically relevant professional knowledge in high-noise environments.
Smart Images

Figure CN121029964A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of natural language processing, and particularly relates to a text matching method and device based on probability distribution, equipment and a storage medium. BACKGROUND
[0002] With the application of large language models based on RAG (Retrieval-augmented Generation) methods, in order to improve the generation effect of the model in a specific field or professional task, the most relevant content will be retrieved from the professional knowledge base after the user input text, and submitted to the large language model for processing together with the user input text.
[0003] However, when there are errors (such as spelling errors, key term errors, and grammatical confusion) or noise (such as irrelevant information and ambiguous expressions) in the user input, the matching accuracy of such methods will often decrease significantly.
[0004] Therefore, how to improve the fault tolerance of text matching when there is noise in the user input information is a problem that needs to be solved at present. SUMMARY
[0005] The main purpose of the present application is to provide a text matching method and device based on probability distribution, equipment and a storage medium, which aims to solve the technical problem of low fault tolerance of text matching when there is noise in the user input information.
[0006] To achieve the above purpose, the present application provides a text matching method based on probability distribution, which comprises: Obtaining the semantic feature distribution of each professional knowledge text from the professional knowledge base to obtain a knowledge probability distribution set; Obtaining the user input text and calculating the corresponding user text probability distribution; Calculating the similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set; Determining the text matching result according to the minimum value of the similarity distance.
[0007] In an embodiment, the step of obtaining the semantic feature distribution of each professional knowledge text from the professional knowledge base to obtain a knowledge probability distribution set comprises: Obtaining each professional knowledge text from the professional knowledge base and calculating the word frequency or word vector of the professional knowledge text; Calculating the knowledge text probability distribution according to the word frequency or word vector to obtain the knowledge probability distribution set.
[0008] In an embodiment, the step of calculating the knowledge text probability distribution according to the word frequency or the word vector to obtain the knowledge probability distribution set comprises: If the knowledge text probability distribution is calculated according to the word vector, each knowledge text in the professional knowledge base is numbered in order to obtain a knowledge number set, and the knowledge number set comprises a mapping relationship between each knowledge text and each knowledge number; A target knowledge text is selected, and the target knowledge text is subjected to vectorization processing to obtain a word vector; Based on a probability distribution formula, the probability distribution of each dimension of the word vector is calculated to obtain the probability distribution of the target knowledge text; A mapping relationship between each knowledge number and the probability distribution of the target knowledge text is established to obtain the knowledge probability distribution set.
[0009] In an embodiment, the step of selecting the target knowledge text comprises: If it is the first selection, the knowledge text corresponding to the starting number in the knowledge number set is selected as the target knowledge text; If it is not the first selection, the current number is sequentially increased to obtain a target number, and the knowledge text corresponding to the target number is selected as the target knowledge text.
[0010] In an embodiment, the step of establishing a mapping relationship between each knowledge number and the probability distribution of the target knowledge text to obtain the knowledge probability distribution set comprises: The knowledge number set is traversed, and if there is an unprocessed knowledge text, the unprocessed knowledge text is defined as the target knowledge text; The step of vectorizing the target knowledge text to obtain a word vector is returned.
[0011] In an embodiment, the step of calculating the knowledge text probability distribution according to the word frequency or the word vector to obtain the knowledge probability distribution set comprises: If the knowledge text probability distribution is calculated according to the word frequency, each professional knowledge text is obtained from the professional knowledge base, and the professional vocabulary of each professional knowledge text is extracted to obtain the word frequency of the professional vocabulary; Based on the word frequency of the professional vocabulary, the professional vocabulary weight is calculated by introducing the inverse document frequency to obtain a weighted calculation result; The weighted calculation result is subjected to normalization processing to obtain the knowledge probability distribution set.
[0012] In an embodiment, the step of calculating the similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set comprises: The similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set is calculated by a Bhattacharyya distance formula. The Bayes distance formula is expressed as:
[0013] wherein, denotes the probability distribution of the i-th knowledge text, denotes the probability distribution of the user input text, denotes the similarity distance between the probability distribution of the i-th knowledge text and the probability distribution of the user input text.
[0014] In addition, to achieve the above object, the present application further proposes a text matching device based on probability distribution, which comprises: A first probability calculation module is configured to obtain semantic feature distribution of each professional knowledge text from a professional knowledge base, and obtain a knowledge probability distribution set. A second probability calculation module is configured to obtain a user input text and calculate a corresponding user text probability distribution. A similarity distance calculation module is configured to calculate the similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set. A text matching determination module is configured to determine a text matching result according to the minimum value of the similarity distance.
[0015] In addition, to achieve the above object, the present application further proposes a text matching device based on probability distribution, which comprises a memory, a processor and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the text matching method based on probability distribution as described above.
[0016] In addition, to achieve the above object, the present application further proposes a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the text matching method based on probability distribution as described above.
[0017] The one or more technical solutions proposed by the present application have at least the following technical effects: The semantic feature distribution of each professional knowledge text is obtained from the professional knowledge base to obtain a knowledge probability distribution set, which solves the problem that the semantic uncertainty cannot be captured due to the use of fixed word vectors. The internal features of the text are abstracted by the probability distribution to realize the standardized representation of the knowledge base.
[0018] The user input text is obtained, and the corresponding user text probability distribution is calculated. The same means for calculating the user text probability distribution of the user input text is adopted, which solves the limitation of the loss of word vectors when the user input is incorrect, and captures the core intent of the query through the distribution probability model instead of the surface words, which significantly improves the fault tolerance of text matching.
[0019] By calculating the similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set, the problem of sensitivity to small changes in vector similarity calculation is solved, and the overall semantic similarity is measured through distribution distance instead of local matching, so that the comparison process can tolerate distribution deviation.
[0020] Finally, the minimum value of the similarity distance is used to determine the text matching result, which solves the problem of low matching accuracy in traditional technology, realizes the effect of stable retrieval of semantically related professional knowledge under high noise, and significantly improves the fault tolerance and matching accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0021] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application.
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0023] Figure 1 The flowchart of the first embodiment of the text matching method based on probability distribution of the present application is shown. Figure 2 The sub-flowchart of calculating the user text probability distribution in the embodiment of the present application is shown. Figure 3 The flowchart of the second embodiment of the text matching method based on probability distribution of the present application is shown. Figure 4 The flowchart of obtaining the knowledge probability distribution set based on the BERT model in the embodiment of the present application is shown. Figure 5 The sub-flowchart of step S40 in the embodiment of the present application is shown. Figure 6 The module structure diagram of the text matching device based on probability distribution in the embodiment of the present application is shown. Figure 7 The device structure diagram of the hardware running environment involved in the text matching method based on probability distribution in the embodiment of the present application is shown.
[0024] The objectives, functional characteristics and advantages of the present application will be further illustrated in conjunction with the embodiments, with reference to the accompanying drawings. DETAILED DESCRIPTION
[0025] It should be understood that the specific embodiments described herein are merely intended to explain the technical solutions of the present application, and are not intended to limit the present application.
[0026] In order to better understand the technical solutions of the present application, the specific embodiments will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0027] In the application of large language models based on RAG, in order to improve the generation effect of the model in a specific field or professional task, a key strategy is to retrieve the most relevant content fragments (knowledge entries) from a professional knowledge base (such as documents, knowledge graphs, etc.) when processing user input text. Then, the retrieved professional content is spliced (or combined in other ways) with the original user input text and submitted to the large language model for processing, thereby guiding and constraining the generation process with professional knowledge.
[0028] The core of the current mainstream retrieval method is to calculate the similarity (such as cosine similarity) between the user query vector and all knowledge entry vectors, and select one or more entries with the highest similarity as the result. However, this method has a significant limitation: its matching accuracy is highly dependent on the accuracy of the user input text. When there are errors (such as spelling errors, key term errors, and syntax confusion) or noise (such as irrelevant information, ambiguous expressions) in the user input, the generated query vector often deviates from its true semantic intent. Since the vector similarity calculation is extremely sensitive to such semantic shifts, the "most relevant" professional knowledge entries retrieved may be significantly different from the content the user actually needs, or even completely irrelevant. This phenomenon is particularly common in large-scale application scenarios, as the quality of user input varies greatly and it is difficult to ensure accuracy.
[0029] Therefore, how to improve the fault tolerance of text matching when there is noise in the user input information is a problem that needs to be solved at present.
[0030] Based on this, the embodiments of the present application provide a text matching method based on probability distribution, which can be applied to a terminal and can be executed by hardware or software in the terminal. The terminal includes but is not limited to a mobile phone or a tablet computer with a touch-sensitive surface (such as a touch screen display and / or a touchpad) and other portable communication devices. It should also be understood that in some embodiments, the terminal can not be a portable communication device, but a desktop computer with a touch-sensitive surface (such as a touch screen display and / or a touchpad).
[0031] Reference Figure 1 ,Figure 1 Flowchart of the first embodiment of the text matching method based on probability distribution of the present application.
[0032] In this embodiment, the text matching method based on probability distribution includes steps S10-S40: Step S10, obtain the semantic feature distribution of each professional knowledge text from the professional knowledge base to obtain a knowledge probability distribution set.
[0033] It should be noted that the professional knowledge text can be understood as the specific text content stored in the professional knowledge base, such as professional documents or materials, covering structured knowledge in a specific field. The semantic feature distribution can be understood as representing the text as a probability distribution through a probability model (such as embedding space mapping) to capture its inherent semantic uncertainty. The knowledge probability distribution set can be a set composed of semantic feature distributions extracted from all professional knowledge texts in the professional knowledge base, used to represent the abstract features of knowledge.
[0034] Step S20, obtain the user input text and calculate the corresponding user text probability distribution.
[0035] It should be noted that the user input text is the input query text provided by the user, which may contain errors or noise. The user text probability distribution refers to the semantic feature distribution calculated by applying the same probability model to the user input text to represent the semantic intent of the query. For example, regardless of the method used, the probability distribution of the professional knowledge text and the user input text can be calculated. The user input text is processed in the same way as the professional knowledge to obtain its probability distribution , where is consistent with the vocabulary set of professional knowledge, and the subflowchart is referred to Figure 2 , including: (1) Obtain user input: after the user completes the input of all text information, the system obtains the text content input by the user in real time, and prepares to submit it to the large language model for processing. After completion, go to step (2).
[0036] (2) Calculate word vector U: use the same BERT model as professional knowledge processing to calculate the word vector U of the user input text. After completion, go to step (3).
[0037] (3) Calculate probability distribution : based on the word vector U of the user input, calculate the probability distribution of each dimension uj through the probability distribution formula, and finally obtain the probability distribution of the user input.
[0038] Step S30, calculate the similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set.
[0039] It should be noted that the distance metric for quantifying the difference between two probability distributions is used to evaluate semantic similarity. The similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set can be calculated by similarity distance measurement methods, such as Hellinger distance or Wasserstein distance.
[0040] Step S40, determining the text matching result according to the minimum value of the similarity distance.
[0041] It should be noted that the best matching professional knowledge text determined by the minimum value of the similarity distance is used as the final output.
[0042] In this embodiment, the technical means of extracting semantic feature distribution from professional knowledge base is adopted, and the professional knowledge is represented as a knowledge probability distribution set, which solves the problem of not being able to capture semantic uncertainty due to the use of fixed word vectors. By abstracting the internal features of the text through probability distribution, the standardized representation of the knowledge base is realized. The same means of calculating the user text probability distribution of the user input text is adopted, which solves the limitation of the word vector being easily lost when the user input is incorrect. By capturing the core intent of the query rather than the surface words through the distribution probability model, the fault tolerance of the text matching is significantly improved. The technical means of calculating the similarity distance is adopted, which solves the problem of sensitivity to small changes in vector similarity calculation. By measuring the overall semantic similarity through distribution distance rather than local matching, the comparison process can tolerate distribution deviation. Finally, the means of determining the matching result according to the minimum similarity distance solves the problem of low matching accuracy of traditional technology, realizes the effect of stable retrieval of semantically related professional knowledge under high noise, and significantly improves the fault tolerance and matching accuracy.
[0043] Reference Figure 3 , Figure 3 The flowchart of the second embodiment of the text matching method based on probability distribution of the present application is shown in the figure. Based on the first embodiment shown in the above Figure 1 The second embodiment of the text matching method based on probability distribution of the present application is proposed.
[0044] In the second embodiment, the step S10 comprises: Step S101, obtaining each professional knowledge text from the professional knowledge base, and calculating the word frequency or word vector of the professional knowledge text.
[0045] It should be noted that when dealing with texts with a relatively significant vocabulary distribution, the word frequency of the specialized knowledge text can be calculated. When it is necessary to capture the deep semantic features of the text, the word vectors of the specialized knowledge text can be calculated. Word frequency can be understood as quantifying the text content by statistically analyzing the frequency of word occurrences in the text. Words can be mapped to dense vectors using a pre-trained model to capture semantic information. The pre-trained model can be a Bidirectional Encoder Representations from Transformers (BERT) model.
[0046] Step S102: Calculate the probability distribution of knowledge text based on word frequency or word vector to obtain a set of knowledge probability distributions.
[0047] It should be noted that the probability distribution of knowledge text can be a semantic feature distribution generated by using a probability calculation model based on word frequency or word vectors, which represents the potential semantic structure of knowledge text.
[0048] For example, when calculating the probability distribution based on the BERT model, each piece of professional knowledge text is transformed into a k-dimensional word vector. ,in( = [ , ,…, ,…, ]), ( k =768). Next, the word vectors are normalized into a probability distribution. Its calculation formula is expressed as formula (1).
[0049] (1) When calculating the probability distribution based on word frequency, a set of specialized vocabulary within the field is compiled based on all standard documents. },in The set is a series of ordinal numbers for specialized terms, with a size of [size missing]. k Then, the word frequency in each professional knowledge text was statistically analyzed. Technical terms Frequency of occurrence in the text. If a word does not appear, its frequency is set to 0. Probability distribution based on word frequency. The calculation formula is expressed as formula (2).
[0050] (2) In this embodiment, the technical means of calculating the word frequency or word vector of professional knowledge text is adopted to solve the problem that the traditional technology needs to rely on accurate input to generate effective vectors. The word frequency or word vector itself can tolerate local lexical variation, providing a robust foundation for subsequent probability distribution calculation. The technical means of generating knowledge text probability distribution is adopted to solve the limitation that the traditional vector similarity is sensitive to input noise. Finally, a knowledge probability distribution set is formed, which realizes the conversion of the professional knowledge base into a robust distribution representation, provides bottom support for fault-tolerant matching, and significantly improves the noise resistance of subsequent semantic similarity comparison.
[0051] In one embodiment, step S102 includes: if the knowledge text probability distribution is calculated according to the word vector, numbering each knowledge text in the professional knowledge base in order to obtain a knowledge number set, the knowledge number set includes the mapping relationship between each knowledge text and each knowledge number; selecting a target knowledge text and performing vectorization processing on the target knowledge text to obtain a word vector; based on the probability distribution formula, the probability distribution of each dimension of the word vector is calculated to obtain the probability distribution of the target knowledge text; and the mapping relationship between each knowledge number and the probability distribution of the target knowledge text is established to obtain a knowledge probability distribution set.
[0052] It should be noted that the knowledge number set can be understood as an index that assigns a unique number to all texts in the professional knowledge base in order. The mapping relationship between the text and the number is established. The target knowledge text can be understood as the professional knowledge text currently selected for probability distribution calculation, which can be dynamically specified by the number. Vectorization processing can be understood as a process of converting the target knowledge text into a dense vector sequence through a pre-trained model (such as BERT) to capture semantic features.
[0053] Specifically, the step of selecting the target knowledge text includes: if it is the first selection, selecting the knowledge text corresponding to the starting number in the knowledge number set as the target knowledge text; if it is not the first selection, the current number is sequentially incremented as the target number, and the knowledge text corresponding to the target number is selected as the target knowledge text.
[0054] It should be noted that the mechanism of controlling the traversal process by the number sequence - the first processing from the starting number, and the sequential traversal is automatically implemented by incrementing the number for non-first selection. Specifically, the step of establishing the mapping relationship between each knowledge number and the probability distribution of the target knowledge text to obtain the knowledge probability distribution set includes: traversing the knowledge number set, if there is an unprocessed knowledge text, defining the unprocessed knowledge text as the target knowledge text; returning to the step of vectorizing the target knowledge text to obtain the word vector.
[0055] It should be noted that each text is processed in the order of the number to ensure that the knowledge base is fully covered and there is no omission.
[0056] An example of using the BERT model to convert each knowledge point into a word vector and obtain the specific process of the knowledge probability distribution set is shown in Figure 4 . Specifically, it can include the following steps: (1) Organize professional knowledge: Collect and organize professional knowledge in a certain field through artificial means, ensuring that the collected knowledge covers the field comprehensively. After completion, proceed to step (2).
[0057] (2) Number professional knowledge: According to the order from 1, uniformly number all professional knowledge texts to ensure that each knowledge point has a unique number, and finally form a numbered set of professional knowledge. After completion, proceed to step (3).
[0058] (3) Select knowledge points: If it is the first selection, select the knowledge point numbered 1; if not, select the knowledge point corresponding to the next number of the current knowledge point. After completion, proceed to step (4).
[0059] (4) Calculate word vector : Use the BERT model to convert the current knowledge point into a word vector . After completion, proceed to step (5).
[0060] (5) Calculate probability distribution : Based on the probability distribution formula, calculate the probability distribution of each dimension of the word vector . After completing the calculation of all dimensions, the probability distribution of the knowledge point is obtained . After completion, proceed to step (6).
[0061] (6) Whether all knowledge is traversed: Check whether all knowledge points have been processed and the probability distribution is calculated. If there are still unprocessed knowledge points, return to step (3); otherwise, proceed to step (7).
[0062] (7) Form a probability distribution set: Combine the probability distributions of all knowledge points into a set for subsequent use.
[0063] In this embodiment, the technical means of sequentially numbered mapping knowledge text (generating knowledge number set) is adopted, which solves the problem of memory overflow caused by loading the whole library at one time in traditional methods, and realizes the scalable processing of large-scale knowledge base through step-by-step indexing. Through dynamic selection of target knowledge text and incremental numbering, the limitation of low efficiency of traditional batch calculation is solved. Based on the vectorization and probability distribution calculation of the target text, the problem that the traditional word vector direct matching is sensitive to local noise is solved. Through the probability distribution calculation of each dimension of the word vector based on the probability distribution formula, the probability distribution of the target knowledge text is obtained, so that the representation has statistical fault tolerance. Finally, by traversing to establish the mapping relationship between the number and the distribution, compared with the traditional static vector library, the dynamic construction and structured storage of the probability distribution of the knowledge base are realized, which provides a foundation support for the subsequent high robustness similarity comparison, and significantly improves the stability and matching noise resistance of the system when processing massive knowledge.
[0064] In one embodiment, step S102 further comprises: if the knowledge text probability distribution is calculated according to the word frequency, obtaining each professional knowledge text from the professional knowledge base, extracting the professional vocabulary of each professional knowledge text to obtain the word frequency of the professional vocabulary; introducing the inverse document frequency to calculate the weight of the professional vocabulary based on the word frequency of the professional vocabulary, to obtain the weighted calculation result; and normalizing the weighted calculation result to obtain the knowledge probability distribution set.
[0065] Exemplarily, the domain-specific vocabulary can be screened from the professional knowledge text through a domain dictionary or a term recognition model. The word frequency of the professional vocabulary is the frequency of the professional vocabulary appearing in the current knowledge text. The inverse document frequency is an index quantifying the scarcity of the vocabulary. The multiplication of the term frequency (TF) and the inverse document frequency (IDF) obtains the term frequency-inverse document frequency (TF-IDF) to amplify the distinguishing weight of the domain key term. The weighted calculation result is normalized, such as scaled to the interval [0, 1] (such as divided by the vector length), to obtain the knowledge probability distribution set.
[0066] Exemplarily, the calculation formula of TF-IDF is represented as formula (3).
[0067] (3) In formula (3), represents the term frequency-inverse document frequency of the professional vocabulary represents the word frequency of the professional vocabulary represents the inverse document frequency of the professional vocabulary and The calculation manners of the two are respectively represented as formula (4) and formula (5).
[0068] (4) (5) In formula (4) and formula (5), the word frequency represents the professional vocabulary appearing in the current knowledge text, is the total number of documents, is the number of documents containing the word . Through normalization processing, the probability distribution is finally obtained, and the calculation formula is represented as formula (6).
[0069] (6)
[0070] In the embodiment, the technical means of extracting professional vocabulary and counting word frequency is adopted, the problem of semantic ambiguity caused by the inclusion of general words in traditional word frequency statistics is solved, and the core feature expression is strengthened by focusing on the field key terms; further, the inverse document frequency weighting calculation is introduced, the problem of over-sensitivity of pure word frequency to common terms is solved, and the field discrimination degree of distribution representation is improved by suppressing high-frequency general words and highlighting low-frequency professional words. Finally, the probability distribution set is generated through normalization, the text semantics is converted into standardized probability distribution, the subsequent similarity distance calculation is immune to the interference of text length difference, and the robustness and fault tolerance of semantic matching in professional scenarios are significantly improved.
[0071] In one embodiment, step S30 comprises: calculating the similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set by the Bhattacharyya distance formula; wherein the Bhattacharyya distance formula is represented as formula (7).
[0072] (7)
[0073] In formula (7), represents the probability distribution of the i-th knowledge text, represents the probability distribution of the user input text, represents the similarity distance between the probability distribution of the i-th knowledge text and the probability distribution of the user input text.
[0074] In this embodiment, the Bhattacharyya distance between the user input text and all professional knowledge texts is calculated, and the professional knowledge text with the smallest distance is found from the calculation result. The professional knowledge corresponding to this text is closest in semantics to the user input text, and therefore can be considered as the most matched professional knowledge with the user input. The matched professional knowledge text and the user input text are submitted to the large language model together, so that the model generates a more accurate answer or classification result. This is suitable for text classification, intelligent question answering, information retrieval and other tasks, and can significantly improve the robustness and accuracy of matching.
[0075] In an embodiment, the sub-process of step S40 is as shown in Figure 5 . Specifically, it can include the following steps: (1) Initialize MinD and BestIndex: define a variable MinD for storing the smallest similarity distance and a variable BestIndex for storing the sequence number of the corresponding professional knowledge. The initial values are set as MinD = 1000 and BestIndex = -1. After completion, go to step (2).
[0076] (2) Select professional knowledge: select a probability distribution of a professional knowledge from the set of professional knowledge probability distributions for processing. If it is the first selection, select the first record in the set; otherwise, select the next probability distribution of the current record. After completion, go to step (3).
[0077] (3) Get probability distribution Pi(wj): extract the probability distribution of the current professional knowledge and its corresponding sequence number . After completion, go to step (4).
[0078] (4) Calculate Bhattacharyya distance : use the Bhattacharyya distance formula to calculate the similarity distance Di between the user input probability distribution Q(uj) and the current professional knowledge probability distribution Pi(wj). After completion, go to step (5).
[0079] (5) <MinD: compare and the current minimum distance MinD. If <MinD, it indicates that the current distance is smaller, go to step (6); otherwise, go directly to step (7).
[0080] (6) Smaller MinD and BestIndex: update MinD to the current smaller distance , and update BestIndex to the sequence number of the current professional knowledge. After completion, go to step (7).
[0081] (7) Whether to traverse completion: check whether all the probability distributions of the professional knowledge have been calculated the similarity distance. If yes, go to step (8); otherwise, return to step (2) to continue processing the untraversed professional knowledge.
[0082] (8) Determine the best match: according to the professional knowledge corresponding to the BestIndex, determine the best professional knowledge matched by the user input, and end the process.
[0083] In this embodiment, by calculating the similarity of the user input and the probability distribution of the professional knowledge, the professional knowledge most matched with the user input is selected in combination with the Bhattacharyya distance. In combination with the RAG technology, the user input and the matched professional knowledge are submitted to the large language model together, the model reasons the user input and generates the most appropriate professional knowledge answer, which significantly improves the robustness and accuracy of the matching.
[0084] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the text matching method based on probability distribution of the present application. More forms of simple transformation based on this technical concept are within the protection scope of the present application.
[0085] The present application also provides a text matching device based on probability distribution, please refer to Figure 6 , the text matching device based on probability distribution comprises: The first probability calculation module 10 is used for obtaining the semantic feature distribution of each professional knowledge text from the professional knowledge base, and obtaining a knowledge probability distribution set; The second probability calculation module 20 is used for obtaining the user input text and calculating the corresponding user text probability distribution; The similarity distance calculation module 30 is used for calculating the similarity distance of the user text probability distribution and each distribution in the knowledge probability distribution set; The text matching determination module 40 is used for determining the text matching result according to the minimum value of the similarity distance. The text matching device based on probability distribution provided by the present application adopts the text matching method based on probability distribution in the above embodiment, and can solve the technical problem of low fault tolerance of text matching in the case that the user input information has noise. Compared with the prior art, the text matching device based on probability distribution provided by the present application has the same beneficial effects as the text matching method based on probability distribution provided by the above embodiment, and other technical features in the text matching device based on probability distribution are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0086] The application provides a probability distribution based text matching device, the probability distribution based text matching device comprises: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the probability distribution based text matching method in the above embodiment one.
[0087] Reference will now be made to the following description Figure 7 , which shows a structural schematic diagram of a probability distribution based text matching device suitable for being used to implement the embodiments of the application. The probability distribution based text matching device in the embodiments of the application can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 7 The probability distribution based text matching device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the application.
[0088] As shown in Figure 7 , the probability distribution based text matching device can include a processing apparatus 1001 (for example, a central processor, a graphic processor, and the like), which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or loaded from a storage apparatus 1003 to a random access memory (RAM) 1004. In the RAM 1004, various programs and data required for operation of the probability distribution based text matching device are also stored. The processing apparatus 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: input apparatuses 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, and the like; output apparatuses 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; the storage apparatus 1003 including, for example, a magnetic tape, a hard disk, and the like; and a communication apparatus 1009. The communication apparatus 1009 can allow the probability distribution based text matching device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7The probability distribution based text matching apparatus with various systems is shown, but it should be understood that not all of the shown systems are required to be implemented or present. More or fewer systems can alternatively be implemented or present.
[0089] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program codes for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through a communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0090] The probability distribution based text matching apparatus provided by the present application adopts the probability distribution based text matching method in the above-mentioned embodiments, and can solve the technical problem of low fault tolerance of text matching in the case of noise in user input information. Compared with the prior art, the probability distribution based text matching apparatus provided by the present application has the same beneficial effects as the probability distribution based text matching method provided by the above-mentioned embodiments, and other technical features in the probability distribution based text matching apparatus are the same as the features disclosed in the above-mentioned embodiments, which will not be described here.
[0091] It should be understood that parts of the present application can be realized in hardware, software, firmware, or a combination thereof. In the description of the above-mentioned embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0092] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0093] The present application provides a computer readable storage medium having stored thereon computer readable program instructions (i.e. computer program) for executing the probability distribution based text matching method in the above-mentioned embodiments.
[0094] The computer readable storage medium provided in the present application may, for example, be a U disk, but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, system, or device, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection with one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present embodiment, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer readable storage medium can be transmitted in any suitable medium, including but not limited to electrical wires, optical cables, RF (Radio Frequency), and the like, or any suitable combination of the above.
[0095] The modules described in the embodiments of the present application can be implemented in software or hardware. In some cases, the name of the module does not limit the unit itself.
[0096] The readable storage medium provided in the present application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer programs) for executing the above-mentioned text matching method based on probability distribution, and can solve the technical problem of low fault tolerance of text matching in the case of noise in user input information. Compared with the prior art, the computer readable storage medium provided in the present application has the same beneficial effects as the text matching method based on probability distribution provided in the above-mentioned embodiments, and will not be described here.
[0097] The above only describes some embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structural transformation made by using the contents of the present application specification and drawings, or direct / indirect application in other related technical fields is included in the patent protection scope of the present application.
Claims
1. A text matching method based on probability distribution, characterized in that, The method includes: Semantic feature distributions of various professional knowledge texts are obtained from the professional knowledge base to obtain a set of knowledge probability distributions; Obtain the user input text and calculate the corresponding user text probability distribution; Calculate the similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set; The text matching result is determined based on the minimum similarity distance.
2. The method as described in claim 1, characterized in that, The step of obtaining the semantic feature distribution of each professional knowledge text from the professional knowledge base to obtain the knowledge probability distribution set includes: Obtain various professional knowledge texts from the professional knowledge base, and calculate the word frequency or word vector of the professional knowledge texts; The probability distribution of knowledge text is calculated based on word frequency or word vectors to obtain a set of knowledge probability distributions.
3. The method as described in claim 2, characterized in that, The step of calculating the probability distribution of knowledge text based on word frequency or word vectors to obtain a set of knowledge probability distributions includes: If the probability distribution of knowledge texts is calculated based on word vectors, and each knowledge text in the professional knowledge base is numbered sequentially, a set of knowledge numbers is obtained. The set of knowledge numbers includes the mapping relationship between each knowledge text and each knowledge number. Select the target knowledge text and vectorize it to obtain word vectors; Based on the probability distribution formula, the probability distribution of the target knowledge text is calculated for each dimension of the word vector. Establish a mapping relationship between each knowledge number and the probability distribution of the target knowledge text to obtain a set of knowledge probability distributions.
4. The method as described in claim 3, characterized in that, The step of selecting the target knowledge text includes: If this is the first time selecting, the knowledge text corresponding to the starting number in the knowledge number set will be selected as the target knowledge text. If this is not the first selection, the current number will be incremented sequentially to become the target number, and the knowledge text corresponding to the target number will be selected as the target knowledge text.
5. The method as described in claim 3, characterized in that, The step of establishing the mapping relationship between each knowledge number and the probability distribution of the target knowledge text to obtain the knowledge probability distribution set includes: Traverse the knowledge ID set; if there is unprocessed knowledge text, define the unprocessed knowledge text as the target knowledge text. Return to the step of vectorizing the target knowledge text to obtain word vectors.
6. The method as described in claim 1, characterized in that, The step of calculating the probability distribution of knowledge text based on word frequency or word vectors to obtain a set of knowledge probability distributions includes: If the probability distribution of knowledge texts is calculated based on word frequency, and each knowledge text is obtained from the knowledge base, the professional vocabulary of each knowledge text is extracted to obtain the word frequency of the professional vocabulary. Based on the word frequency of the aforementioned professional terms, inverse document frequency is introduced to calculate the weight of the professional terms, and a weighted calculation result is obtained; The weighted calculation results are normalized to obtain a set of knowledge probability distributions.
7. The method as described in claim 1, characterized in that, The step of calculating the similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set includes: The similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set is calculated using the Bach distance formula. The Bach distance formula is expressed as follows: in, Indicates the first The probability distribution of a knowledge text. This represents the probability distribution of the user's input text. Indicates the first The similarity distance between the probability distribution of a knowledge text and the probability distribution of the user input text.
8. A text matching device based on probability distribution, characterized in that, The device includes: The first probability calculation module is used to obtain the semantic feature distribution of each professional knowledge text from the professional knowledge base and obtain a set of knowledge probability distributions. The second probability calculation module is used to obtain user input text and calculate the corresponding user text probability distribution. The similarity distance calculation module is used to calculate the similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set; The text matching determination module is used to determine the text matching result based on the minimum similarity distance.
9. A text matching device based on probability distribution, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the probability distribution-based text matching method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the text matching method based on probability distribution as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Text error correction method and device, equipment and storage medium
CN119443087A
Multi-modal knowledge question and answer retrieval method and system for specific professional field
CN119621921A
Knowledge question-answering method and system based on large model and storage medium
CN120162401A
Local knowledge base RAG method and device based on Bayesian reasoning
CN120218247A
Expert knowledge recommendation method and apparatus, computer device, and storage medium
WO2020119063A1
Cited By
Team optimization matching method and system based on bidirectional matching mechanism
CN121920798A