Personnel information consistency calculation method based on mixed text similarity

Through the mixed text similarity calculation method, combined with editing distance and synonym word forest and sentence vector model, the accuracy and efficiency of the data consistency calculation of multi-source heterogeneous personnel information data is solved, and more comprehensive search results are achieved.

CN120579529APending Publication Date: 2025-09-02AEROSPACE SCI & IND INTELLIGENT OPERATION RES & INFORMATION SECURITY RES INST (WUHAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411956531.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-29
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

When calculating the consistency of multi-source heterogeneous personnel information data, the data is redundant and ambiguity, and a single similarity method is difficult to ensure accuracy and performance, resulting in incomplete search results.

Method used

A mixed text similarity calculation method is adopted, combined with editing distance, synonym forest and sentence vector model, the similarity calculation method is selected based on the field item attributes, and the calculation amount is reduced through similarity transferability, and the manual checksum knowledge base expansion is supported.

Benefits of technology

It improves the accuracy and efficiency of data consistency calculations, ensures the comprehensiveness and accuracy of search results, and reduces the consumption of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579529A_ABST
    Figure CN120579529A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information similarity judgment, and particularly relates to a personnel information consistency calculation method based on mixed text similarity, which is used for calculating text occurrence frequency and searching a matching process by calculating similarity among texts and judging whether the texts are consistent or not, so that the statistical process and the matching process are more comprehensive, and the matching efficiency is improved. And omission of some texts with identical semantics due to individual character inconsistency is avoided. And based on the consistency calculation result, the retrieval word is expanded in the retrieval process, so that the retrieval result is more comprehensive. According to the method, the transitivity of the similarity is utilized, when the similarity of a text a and a text c is known and the text a and the text b are completely consistent, the similarity of the text b and the text c does not need to be calculated, the similarity of the text a and the text c is directly obtained, and the calculation amount is reduced. According to the method, when the similarity is calculated based on the character strings, text pairs needing to be calculated are deleted and selected according to the length difference of the character strings and the proportion of the same characters, and the calculation amount is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information similarity judgment, and in particular relates to a method for calculating personnel information consistency based on mixed text similarity. Background Art

[0002] Extracting the identity information of a single individual from multi-source, heterogeneous personnel information data can be challenging. Due to the varying quality of different data sources, inconsistent information representation languages, and varying data formats, even after undergoing common data cleaning and conversion rules, the data stored in the database still suffers from redundancy and ambiguity, making it difficult for users to obtain a more accurate description of the individual from this complex set of identity data. In this system, it is necessary to calculate the most reliable identity information for an individual. Data consistency is a key indicator of data credibility. The more frequently data values ​​from different sources appear, the more accurately they describe the individual's identity. When calculating data frequency, determining whether two data points are consistent is a key challenge. Furthermore, when users search for personnel information, inconsistent search terms and synonyms can lead to incomplete search results. Therefore, calculating consistency across information significantly impacts the accuracy of the system's related functions.

[0003] Currently, the main methods for calculating the similarity between texts include string-based methods, knowledge ontology or classification systems, and statistics and calculations based on large-scale corpora. String-based methods are simple and convenient, and can quickly and intuitively reflect the differences between texts. However, they only calculate similarity based on the differences in the text itself and cannot reflect the degree of semantic similarity. Corpus-based methods can objectively reflect the morphology, syntax, and semantics of texts and discover effective connections between implicit characters, but they are highly dependent on the corpus and suffer from data sparsity and noise interference. Knowledge base-based methods are based on human world knowledge, and the calculation results reflect the similarity information contained in the knowledge ontology. However, they are easily influenced by human subjective knowledge, and the meaning of words may change over time and in different environments.

[0004] When calculating information consistency, the issue of consistency and inconsistency must be addressed. Similarity is an approximate value between 0 and 1, so two texts with high similarity do not necessarily meet consistency. For personal identity information, the correlation between similarity and consistency varies depending on the dimension. Using the aforementioned single similarity method for consistency determination is difficult to guarantee accuracy and performance. Summary of the Invention

[0005] (1) Technical issues to be resolved

[0006] The technical problem to be solved by the present invention is: how to provide a method for calculating the consistency of personnel information based on mixed text similarity.

[0007] (2) Technical solution

[0008] To solve the above technical problems, the present invention provides a method for calculating the consistency of personnel information based on mixed text similarity, the method comprising the following steps:

[0009] Step 1: Aggregate and integrate various data from multiple data sources based on related fields, and pre-process the data values ​​to obtain all personal identity information;

[0010] Step 2: Obtain all field values ​​for the field items to be calculated, construct field value text pairs, and query the pre-maintained field value similarity to see whether the data value pair or the data value pair that is consistent with it has been calculated. If so, directly obtain the similarity value. If the similarity between the two data values ​​has not been calculated, proceed to the next step of calculation.

[0011] Step 3: Determine whether the field item attribute to be calculated is a field item with no multi-word synonyms in the field value. If so, proceed to step 4 to calculate the similarity between the field values; otherwise, proceed to step 5 to calculate the similarity;

[0012] The case where the field item does not have multiple synonyms in the field value includes: containing standard dictionary values ​​and having no semantic meaning;

[0013] Step 4: Use edit distance to calculate similarity and calculate the similarity of field values ​​without multiple synonyms;

[0014] Step 5: Calculate the similarity between the two texts using the semantic similarity method based on the synonym dictionary and sentence vector model;

[0015] Step 6: Calculate the similarity between the field values ​​in all field items using the method in step 5. Select an appropriate threshold based on different calculation methods. If the value is greater than the threshold, it is considered consistent; if the value is less than the threshold, it is considered inconsistent.

[0016] Step 7: For the similarity results, a user adjustment interface is provided. If the user identifies a text pair as consistent, the similarity is set to 1;

[0017] Step 8: For similarity values ​​greater than 0.5, the text pairs and their similarity values ​​are memorized to reduce the amount of subsequent calculations. At the same time, the text pairs with the highest similarity values ​​greater than the threshold are maintained in the synonym dictionary, gradually supplementing and expanding industry knowledge and improving the applicability of the dictionary to professional fields.

[0018] Step 9: Calculate the frequency of occurrence of each field value in the personnel identity field; count the number of consistent texts in each field and calculate the ratio to the total number of field values, so that the system can measure the credibility of the field values;

[0019] Step 10: During the information retrieval process, the search terms are expanded based on the consistency calculation results, and the search results are displayed in order according to the similarity of the search terms. After the search terms are entered, the search terms are searched with relevant and consistent texts. If satisfactory data still cannot be matched in the database, the search is carried out in order according to the similarity until a suitable result is retrieved or there are no similar words.

[0020] In step 4, the edit distance is a method for measuring the literal similarity between two character strings, which is defined as the minimum number of operations required to convert one character string into another. The smaller the edit distance, the more similar the two character strings are. The operations include: deletion, insertion, and replacement.

[0021] The calculation process of step 4 is as follows:

[0022] For two strings S and T, with lengths n and m respectively, their edit distance is as follows, where lev S,T (i, j) represents the edit distance between the first i and first j characters of strings S and T:

[0023]

[0024] The similarity between texts is inversely proportional to the edit distance. The formula for calculating the similarity Sim(S,T) is as follows:

[0025]

[0026] To judge the consistency of text, a high degree of similarity is required. To save computing resources, a quick pruning judgment is performed on the distance calculation. When the length difference of the two strings is large or there are few identical characters, it can be judged in advance that the similarity between the two strings will not be too high, and the probability of the text being consistent is very small. The judgment calculation formula is as follows:

[0027]

[0028] Where SSL(S,T) represents the number of identical characters in two strings, θ1 represents the length difference threshold, and θ2 represents the identical character threshold. If the length difference is higher than θ1 or the identical character ratio is lower than θ2, the similarity is not calculated.

[0029] A similarity of 1 indicates that the two strings are completely identical. The closer the similarity is to 1, the more identical the strings are.

[0030] Wherein, in step 5, the calculation process is:

[0031] First, the text is segmented and stop words are removed. The following two methods are selected according to the size of the text vocabulary:

[0032] (1) After word segmentation, the field value of some field items consists of only one or two words. For such short texts, the synonym knowledge base in the Synonymous Cilin is used to solve the morphological problem of Chinese to a certain extent. The similarity calculation formula based on the Synonymous Cilin is as follows:

[0033]

[0034] s w and t w Represented as a word segmentation vocabulary of strings S and T, where LCP(s w ,t w ) is expressed as s w and t w The longest common path, where Depth(LCP(s w ,t w )) represents the depth distance between the two words’ nearest common parent nodes, Path(s w ,t w ) represents the shortest path between two words, α is the depth adjustment parameter, and β is the path adjustment parameter;

[0035] Convert the similarity between two words into the similarity between two short texts. First, calculate the similarity of a single word. P and Q represent the number of word segments in the strings S and T respectively. The similarity of the pth word in the string S is as follows:

[0036]

[0037] Calculate the weighted sum of the similarities of each word to obtain the similarity of the final text. p and δ q are the similarity weights of two text words respectively. The weight is higher for core words;

[0038]

[0039] (2) For some complex texts such as long texts with varying lengths, similarity is measured by generating sentence vectors and then calculating the distance between the vectors;

[0040] The vector representation of text maps the text into a low-dimensional vector space, so that semantically similar text is closer in the vector space. A pre-trained sentence vector generation model is used to obtain the vector representation of text sentences. Pre-trained models usually have good generalization capabilities and can effectively represent sentences as vectors with semantic information. The similarity is measured using cosine similarity, which is determined by calculating the cosine value of the angle between two vectors. -1 indicates complete dissimilarity, 0 indicates no similarity, and 1 indicates complete similarity.

[0041] Assume that two texts S and T generate two d-dimensional vectors, represented as [S1, S2, ..., S d ] and [T1,T2,...,T d ], the calculation formula is as follows:

[0042]

[0043] (3) Beneficial effects

[0044] Compared with the prior art, the key technical innovations of the present invention are:

[0045] (1) By calculating the similarity between texts, we can determine whether the texts are consistent. This is used to count the frequency of text occurrences and the search and matching process, making the statistical and matching processes more comprehensive and avoiding missing texts that are semantically identical due to inconsistencies in individual characters. Based on the consistency calculation results, the search terms are expanded during the search process, making the search results more comprehensive.

[0046] (2) By utilizing the transitivity of similarity, when the similarity between text a and text c is known, and text a is completely consistent with text b, there is no need to calculate the similarity between text b and text c. The similarity between text a and text c can be directly taken to reduce the amount of calculation.

[0047] (3) When calculating similarity, select a similarity calculation method based on the attributes of the field item and the characteristics of the field value. When the field value is a standard dictionary item or a meaningless field, select string-based similarity calculation; for fields containing multiple synonyms, when the vocabulary of the text pair is small and the difference is not large, select Cilin-based similarity calculation, and calculate the similarity between the texts based on the similarity between the vocabulary; for longer texts or when the vocabulary difference is large, generate sentence vectors based on the pre-trained model and calculate the similarity.

[0048] (4) When calculating similarity based on character strings, the text pairs that need to be calculated are selected based on the difference in string length and the proportion of identical characters to reduce the amount of calculation.

[0049] (5) When calculating similarity based on Cilin, the similarity between texts is the weighted sum of the similarities between words. Considering the importance of words, core words have a higher weight.

[0050] (6) The consistency judgment results support manual verification, and can be maintained, the knowledge base expanded, and fed back to the consistency calculation system, gradually improving the calculation efficiency and accuracy of the system's consistency judgment.

[0051] Compared with the prior art, the advantages of the present invention are:

[0052] (1) By calculating the similarity between texts, we can determine whether the texts are consistent. This is used to count the frequency of text occurrences and search and match processes, making the statistical and matching processes more comprehensive and avoiding missing texts that are semantically identical due to inconsistencies in individual characters.

[0053] (2) The present invention selects different similarity calculation methods according to the attributes of the field items and the characteristics of the field values, which greatly improves the calculation efficiency while ensuring the accuracy of the semantic similarity calculation.

[0054] (3) Using similarity transferability can help reduce the amount of calculation.

[0055] (4) When calculating based on string similarity, characters with large differences are removed by judging the difference in string length and the proportion of identical characters in advance, thereby reducing the amount of calculation and improving efficiency.

[0056] (5) When similarity calculation is performed based on Cilin, the importance of words is considered when converting the similarity between words into the similarity between texts, and core words are given higher weights, thereby improving the accuracy of similarity calculation.

[0057] (6) Based on the similarity calculation results and manual verification results, maintain and expand the knowledge base, continuously improve the applicability of the knowledge base to the industry, and make the similarity calculation results more in line with reality. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is the overall flow chart of the present invention.

[0059] Figure 2 This is a flowchart of the similarity calculation method. The left side is based on string similarity calculation, and the right side is based on word forest and sentence vector similarity calculation.

[0060] Figure 3 Flowchart of applying similarity results to information retrieval. DETAILED DESCRIPTION

[0061] In order to make the purpose, content, and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below with reference to the accompanying drawings and examples.

[0062] The present invention proposes a hybrid similarity calculation method for judging the consistency of personnel information. This method uses different similarity calculation methods and different thresholds for consistency judgment according to the characteristics of the dimensions for different dimensions of personnel identity information. In addition, the similarity can be transferred. If the field value a is consistent with the field value b, then the similarity between a and c is equal to the similarity between b and c. This helps to reduce the amount of calculation and improve the efficiency of system operation. The judgment of consistency can be used to measure whether two data values ​​are the same to determine whether the data values ​​are credible. For two completely consistent data, the knowledge base is supplemented and maintained, and the knowledge base with industry characteristics is gradually enriched to improve the calculation accuracy. The judgment of consistency can also be used for information retrieval, synonymous replacement of search terms, and more comprehensive retrieval.

[0063] After data from different sources has been cleaned and normalized, some fields with standard dictionary values, such as education and gender, have uniform data values. For non-semantic fields, such as names and nicknames, the similarity differences between fields with standard dictionary values ​​or non-semantic values ​​in the system are not multi-word synonyms. Using knowledge base or corpus-based methods, the similarity differences are not obvious and the calculation is more complex. Therefore, judging consistency based on the literal differences between strings is more efficient and accurate. For other fields, such as home address, work experience, and rewards and punishments, the data expressions from different sources may be inconsistent. The literal differences alone cannot reflect the consistency of two records. Therefore, it is necessary to judge consistency based on knowledge and the implicit relationships between the text.

[0064] To solve the above technical problems, the present invention provides a method for calculating the consistency of personnel information based on mixed text similarity, the method comprising the following steps:

[0065] Step 1: Aggregate and integrate various data from multiple data sources based on related fields, and pre-process the data values ​​to obtain all personal identity information;

[0066] Step 2: Obtain all field values ​​for the field items to be calculated, construct field value text pairs, and query the pre-maintained field value similarity to see whether the data value pair or the data value pair that is consistent with it has been calculated. If so, directly obtain the similarity value. If the similarity between the two data values ​​has not been calculated, proceed to the next step of calculation.

[0067] Step 3: Determine whether the field item attribute to be calculated is a field item with no multi-word synonyms in the field value. If so, proceed to step 4 to calculate the similarity between the field values; otherwise, proceed to step 5 to calculate the similarity;

[0068] The case where the field item does not have multiple synonyms in the field value includes: containing standard dictionary values ​​and having no semantic meaning;

[0069] Step 4: Use edit distance to calculate similarity and calculate the similarity of field values ​​without multiple synonyms;

[0070] Step 5: Calculate the similarity between the two texts using the semantic similarity method based on the synonym dictionary and sentence vector model;

[0071] Step 6: Calculate the similarity between the field values ​​in all field items using the method in step 5. Select an appropriate threshold based on different calculation methods. If the value is greater than the threshold, it is considered consistent; if the value is less than the threshold, it is considered inconsistent.

[0072] Step 7: For the similarity results, a user adjustment interface is provided. If the user identifies a text pair as consistent, the similarity is set to 1;

[0073] Step 8: For similarity values ​​greater than 0.5, the text pairs and their similarity values ​​are memorized to reduce the amount of subsequent calculations. At the same time, the text pairs with the highest similarity values ​​greater than the threshold are maintained in the synonym dictionary, gradually supplementing and expanding industry knowledge and improving the applicability of the dictionary to professional fields.

[0074] Step 9: Calculate the frequency of occurrence of each field value in the personnel identity field; count the number of consistent texts in each field and calculate the ratio to the total number of field values, so that the system can measure the credibility of the field values;

[0075] Step 10: During the information retrieval process, the search terms are expanded based on the consistency calculation results, and the search results are displayed in order according to the similarity of the search terms. After the search terms are entered, the search terms are searched with relevant and consistent texts. If satisfactory data still cannot be matched in the database, the search is carried out in order according to the similarity until a suitable result is retrieved or there are no similar words.

[0076] In step 4, the edit distance is a method for measuring the literal similarity between two character strings, which is defined as the minimum number of operations required to convert one character string into another. The smaller the edit distance, the more similar the two character strings are. The operations include: deletion, insertion, and replacement.

[0077] The calculation process of step 4 is as follows:

[0078] For two strings S and T, with lengths n and m respectively, their edit distance is as follows, where lev S,T (i, j) represents the edit distance between the first i and first j characters of strings S and T:

[0079]

[0080] The similarity between texts is inversely proportional to the edit distance. The formula for calculating the similarity Sim(S,T) is as follows:

[0081]

[0082] To judge the consistency of text, a high degree of similarity is required. To save computing resources, a quick pruning judgment is performed on the distance calculation. When the length difference of the two strings is large or there are few identical characters, it can be judged in advance that the similarity between the two strings will not be too high, and the probability of the text being consistent is very small. The judgment calculation formula is as follows:

[0083]

[0084] Where SSL(S,T) represents the number of identical characters in two strings, θ1 represents the length difference threshold, and θ2 represents the identical character threshold. If the length difference is higher than θ1 or the identical character ratio is lower than θ2, the similarity is not calculated.

[0085] A similarity of 1 indicates that the two strings are completely identical. The closer the similarity is to 1, the more identical the strings are.

[0086] Wherein, in step 5, the calculation process is:

[0087] First, the text is segmented and stop words are removed. The following two methods are selected according to the size of the text vocabulary:

[0088] (1) After word segmentation, the field value of some field items consists of only one or two words. For such short texts, the synonym knowledge base in the Synonymous Cilin is used to solve the morphological problem of Chinese to a certain extent. The similarity calculation formula based on the Synonymous Cilin is as follows:

[0089]

[0090] s w and t w Represented as a word segmentation vocabulary of strings S and T, where LCP(s w ,t w ) is expressed as s w and t w The longest common path, where Depth(LCP(s w ,t w )) represents the depth distance between the two words’ nearest common parent nodes, Path(s w ,t w ) represents the shortest path between two words, α is the depth adjustment parameter, and β is the path adjustment parameter;

[0091] Convert the similarity between two words into the similarity between two short texts. First, calculate the similarity of a single word. P and Q represent the number of word segments in the strings S and T respectively. The similarity of the pth word in the string S is as follows:

[0092]

[0093] Calculate the weighted sum of the similarities of each word to obtain the similarity of the final text. p and δ q are the similarity weights of two text words respectively. The weight is higher for core words;

[0094]

[0095] (2) For some complex texts such as long texts with varying lengths, similarity is measured by generating sentence vectors and then calculating the distance between the vectors;

[0096] The vector representation of text maps the text into a low-dimensional vector space, so that semantically similar text is closer in the vector space. A pre-trained sentence vector generation model is used to obtain the vector representation of text sentences. Pre-trained models usually have good generalization capabilities and can effectively represent sentences as vectors with semantic information. The similarity is measured using cosine similarity, which is determined by calculating the cosine value of the angle between two vectors. -1 indicates complete dissimilarity, 0 indicates no similarity, and 1 indicates complete similarity.

[0097] Assume that two texts S and T generate two d-dimensional vectors, represented as [S1, S2, ..., S d ] and [T1,T2,...,T d ], the calculation formula is as follows:

[0098]

[0099] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for calculating consistency of personnel information based on mixed text similarity, characterized in that: The method comprises the following steps: Step 1: Aggregate and integrate various data from multiple data sources based on related fields, and pre-process the data values ​​to obtain all personal identity information; Step 2: Obtain all field values ​​for the field items to be calculated, construct field value text pairs, and query the pre-maintained field value similarity to see whether the data value pair or the data value pair that is consistent with it has been calculated. If so, directly obtain the similarity value. If the similarity between the two data values ​​has not been calculated, proceed to the next step of calculation. Step 3: Determine whether the field item attribute to be calculated is a field item with no multi-word synonyms in the field value. If so, proceed to step 4 to calculate the similarity between the field values; otherwise, proceed to step 5 to calculate the similarity; The case where the field item does not have multiple synonyms in the field value includes: containing standard dictionary values ​​and having no semantic meaning; Step 4: Use edit distance to calculate similarity and calculate the similarity of field values ​​without multiple synonyms; Step 5: Calculate the similarity between the two texts using the semantic similarity method based on the synonym dictionary and sentence vector model; Step 6: Calculate the similarity between the field values ​​in all field items using the method in step 5. Select an appropriate threshold based on different calculation methods. If the value is greater than the threshold, it is considered consistent; if the value is less than the threshold, it is considered inconsistent. Step 7: For the similarity results, a user adjustment interface is provided. If the user identifies a text pair as consistent, the similarity is set to 1; Step 8: For similarity values ​​greater than 0.5, the text pairs and their similarity values ​​are memorized to reduce the amount of subsequent calculations. At the same time, the text pairs with the highest similarity values ​​greater than the threshold are maintained in the synonym dictionary, gradually supplementing and expanding industry knowledge and improving the applicability of the dictionary to professional fields. Step 9: Calculate the frequency of occurrence of each field value in the personnel identity field; count the number of consistent texts in each field and calculate the ratio to the total number of field values, so that the system can measure the credibility of the field values; Step 10: During the information retrieval process, the search terms are expanded based on the consistency calculation results, and the search results are displayed in order according to the similarity of the search terms. After the search terms are entered, the search terms are searched with relevant and consistent texts. If satisfactory data still cannot be matched in the database, the search is carried out in order according to the similarity until a suitable result is retrieved or there are no similar words.

2. The method for calculating consistency of personnel information based on mixed text similarity according to claim 1, characterized in that: In step 4, the edit distance is a method for measuring the literal similarity between two character strings, and is defined as the minimum number of operations required to convert one character string into another. The smaller the edit distance, the more similar the two character strings are. The operations include: deletion, insertion, and replacement.

3. The method for calculating consistency of personnel information based on mixed text similarity according to claim 2, characterized in that: The calculation process of step 4 is as follows: For two strings S and T, with lengths n and m respectively, their edit distance is as follows, where lev S,T (i, j) represents the edit distance between the first i and first j characters of strings S and T: The similarity between texts is inversely proportional to the edit distance. The formula for calculating the similarity Sim(S,T) is as follows: To judge the consistency of text, a high degree of similarity is required. To save computing resources, a quick pruning judgment is performed on the distance calculation. When the length difference of the two strings is large or there are few identical characters, it can be judged in advance that the similarity between the two strings will not be too high, and the probability of the text being consistent is very small. The judgment calculation formula is as follows: Where SSL(S,T) represents the number of identical characters in two strings, θ1 represents the length difference threshold, and θ2 represents the identical character threshold. If the length difference is higher than θ1 or the identical character ratio is lower than θ2, the similarity is not calculated. A similarity of 1 indicates that the two strings are completely identical. The closer the similarity is to 1, the more identical the strings are.

4. The method for calculating consistency of personnel information based on mixed text similarity according to claim 3, characterized in that: In step 5, the calculation process is: First, the text is segmented and stop words are removed. The following two methods are selected according to the size of the text vocabulary: (1) After word segmentation, the field value of some field items consists of only one or two words. For such short texts, the synonym knowledge base in the Synonymous Cilin is used to solve the morphological problem of Chinese to a certain extent. The similarity calculation formula based on the Synonymous Cilin is as follows: s w and t w Represented as a word segmentation vocabulary of strings S and T, where LCP(s w ,t w ) is expressed as s w and t w The longest common path, where Depth(LCP(s w ,t w )) represents the depth distance between the two words’ nearest common parent nodes, Path(s w ,t w ) represents the shortest path between two words, α is the depth adjustment parameter, and β is the path adjustment parameter; Convert the similarity between two words into the similarity between two short texts. First, calculate the similarity of a single word. P and Q represent the number of word segments in the strings S and T respectively. The similarity of the pth word in the string S is as follows: Calculate the weighted sum of the similarities of each word to obtain the similarity of the final text. p and δ q are the similarity weights of two text words respectively. The weight is higher for core words; (2) For some complex texts such as long texts with varying lengths, similarity is measured by generating sentence vectors and then calculating the distance between the vectors; The vector representation of text maps the text into a low-dimensional vector space, so that semantically similar text is closer in the vector space. A pre-trained sentence vector generation model is used to obtain the vector representation of text sentences. Pre-trained models usually have good generalization capabilities and can effectively represent sentences as vectors with semantic information. The similarity is measured using cosine similarity, which is determined by calculating the cosine value of the angle between two vectors. -1 indicates complete dissimilarity, 0 indicates no similarity, and 1 indicates complete similarity. Assume that two texts S and T generate two d-dimensional vectors, represented as [S1, S2, ..., S d ] and [T1,T2,...,T d ], the calculation formula is as follows:

5. The method for calculating consistency of personnel information based on mixed text similarity according to claim 4, characterized in that: In step 5, the field values ​​of some field items are only one or two words after word segmentation, including occupation and place of origin.

6. The method for calculating consistency of personnel information based on mixed text similarity according to claim 3, characterized in that: This method calculates similarity between texts to determine whether they are consistent. This is used to count text occurrences and perform search and matching, making the counting and matching processes more comprehensive and avoiding missing texts that are semantically identical despite individual character inconsistencies. Based on the consistency calculation results, search terms are expanded during the search process, resulting in more comprehensive search results.

7. The method for calculating consistency of personnel information based on mixed text similarity according to claim 3, characterized in that: The method utilizes the transitivity of similarity. When the similarity between text a and text c is known and text a is completely consistent with text b, there is no need to calculate the similarity between text b and text c. The similarity between text a and text c can be directly taken to reduce the amount of calculation.

8. The method for calculating consistency of personnel information based on mixed text similarity according to claim 3, characterized in that: When calculating similarity, the method selects a similarity calculation method based on the attributes of the field item and the characteristics of the field value; When the field value is a standard dictionary item or a meaningless field, select string-based similarity calculation. For fields containing multiple synonyms and when the text pairs have a small vocabulary and small differences, select Cilin-based similarity calculation to calculate the similarity between the texts based on the similarity between the words. For longer texts or when the vocabulary varies greatly, generate sentence vectors based on the pre-trained model and calculate similarity.

9. The method for calculating consistency of personnel information based on mixed text similarity according to claim 3, characterized in that: The method uses Cilin to perform similarity calculations. When converting inter-lexical similarity into inter-text similarity, the method considers the importance of the lexical elements and gives higher weight to core lexical elements, thereby improving the accuracy of similarity calculations.

10. The method for calculating consistency of personnel information based on mixed text similarity according to claim 3, characterized in that: The method maintains and expands the knowledge base based on similarity calculation results and manual verification results, continuously improves the applicability of the knowledge base to the industry, and makes the similarity calculation results more in line with reality.