Method, device and equipment for determining personal information by same person
By combining the similarity judgment methods of structured, semi-structured and unstructured texts, the problem of inefficient judgment of candidates in corporate recruitment is solved, efficient and accurate fan judgment is achieved, and the efficiency of talent pool management is improved.
Patent Information
- Application Number
- CN202510280048.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, corporate recruitment specialists are inefficient and prone to errors when judging whether candidates already exist in the internal talent pool, especially because the vector characterization method is poor due to the lengthy content of resume text and structured features.
Using a combination of structured, semi-structured and unstructured text, structured text is extracted through large language models, semi-structured text is generated in blocks, and target similarity is generated using text similarity model and state-aware hybrid expert model to comprehensively judge the similarity of personal information.
It realizes efficient and accurate judgment of whether personal information is duplicated, reduces the possibility of misjudgment, improves the efficiency of talent pool management and saves manual comparison time.
Smart Images

Figure CN120257965A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more specifically, to a method, device, and equipment for determining whether personal information belongs to the same person. Background Art
[0002] In modern recruitment processes, enterprise recruitment specialists need to obtain a large number of candidate resumes on various third-party recruitment websites, add the candidates who meet the requirements to the enterprise's internal talent pool, so that the enterprise's internal talent pool accumulates a large number of candidate resumes; when the enterprise recruitment specialist uses the third-party website to select candidates again, it is impossible to quickly determine whether the candidate already exists in the internal talent pool, resulting in the recruitment specialist spending a lot of time and effort manually comparing the resumes obtained from the third party with the resumes recorded in the internal talent pool to check for duplicates. This not only has low efficiency but also is prone to comparison errors.
[0003] In the prior art, a vector representation method is used to judge the similarity of resumes. However, the text content of resumes is usually quite long and often exceeds the maximum text length that the encoder of the vector representation method can handle. Moreover, the resume text has specific structural features, making the above vector representation method less effective in judging whether resumes are duplicates. The above resumes are a form of personal information.
[0004] In summary, how to efficiently and accurately judge whether personal information is duplicated is a problem that needs to be solved currently. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method, device, and equipment for determining whether personal information belongs to the same person, which can efficiently and accurately judge whether personal information is duplicated.
[0006] First aspect, an embodiment of the present invention provides a method for determining personal information of the same person. The method includes: obtaining first personal information and second personal information; respectively determining at least one structured text, at least one semi-structured text, and unstructured text corresponding to each of them according to the first personal information and the second personal information, where the structured text is structured information extracted from the first personal information and the second personal information, the semi-structured text is chunked information extracted from the first personal information and the second personal information, and the unstructured text is an unstructured representation of the first personal information and the second personal information; determining at least one dense similarity and at least one sparse similarity according to the at least one structured text corresponding to the first personal information and the second personal information respectively, determining at least one statistical feature similarity according to the at least one semi-structured text corresponding to the first personal information and the second personal information respectively, and determining a global similarity according to the unstructured text corresponding to the first personal information and the second personal information respectively; inputting the at least one dense similarity, the at least one sparse similarity, the at least one statistical feature similarity, and the global similarity into a pre-trained state-aware mixture of experts model to generate a target similarity, where the target similarity is used to represent the similarity degree between the first personal information and the second personal information.
[0007] Optionally, the method further includes:
[0008] In response to the target similarity being greater than or equal to a set threshold, determining that the first personal information and the second personal information are personal information of the same person.
[0009] Optionally, the step of respectively determining at least one structured text corresponding to each of them according to the first personal information and the second personal information specifically includes:
[0010] Inputting the first personal information and the second personal information into a large language model respectively, and automatically extracting at least one structured text corresponding to each of them through the field extractor of the large language model.
[0011] Optionally, the step of respectively determining at least one semi-structured text corresponding to each of them according to the first personal information and the second personal information specifically includes:
[0012] Chunking the first personal information and the second personal information respectively according to set rules to generate at least one semi-structured text corresponding to each of them.
[0013] Optionally, determining at least one dense similarity and at least one sparse similarity according to the at least one structured text respectively corresponding to the first personal information and the second personal information specifically includes:
[0014] Input the at least one structured text respectively corresponding to the first personal information and the second personal information into a pre-set text similarity model, and output at least one dense similarity and at least one sparse similarity.
[0015] Optionally, the text similarity model is an open-source general semantic vector model fine-tuned by a contrastive learning loss function.
[0016] Optionally, determining at least one statistical feature similarity according to the at least one semi-structured text respectively corresponding to the first personal information and the second personal information specifically includes:
[0017] Input the at least one semi-structured text respectively corresponding to the first personal information and the second personal information into a multi-level statistical feature extractor to generate at least one set of statistical sub-feature similarities, where the statistical sub-feature similarities include Jaccard similarity, locality-sensitive hashing similarity, and Levenshtein distance similarity;
[0018] Input the at least one set of statistical sub-feature similarities into a multi-layer perceptron linear layer, and output at least one statistical feature similarity.
[0019] Optionally, determining a global similarity according to the unstructured texts respectively corresponding to the first personal information and the second personal information specifically includes:
[0020] Perform a vector inner product on the unstructured text corresponding to the first personal information and the unstructured text corresponding to the second personal information to generate a global similarity.
[0021] Optionally, the process of training the text similarity model specifically includes:
[0022] Obtain historical data, where the historical data includes positive sample data and negative sample data, the positive sample data is similar personal information of the same person, and the negative sample is different similar personal information;
[0023] Construct a data set according to the historical data;
[0024] Train the open-source general semantic vector model through the data set and the contrastive learning loss function to generate the text similarity model.
[0025] Optionally, the constructing the data set according to the historical data specifically includes:
[0026] Perform random noise addition processing on the historical data to construct the data set.
[0027] Optionally, the process of training the state-aware mixture-of-experts model specifically includes:
[0028] Determine the weight coefficients of each expert model in the state-aware mixture-of-experts model according to the gating network, and then determine the state-aware mixture-of-experts model.
[0029] In a second aspect, an embodiment of the present invention provides a personal information same-person determination device, the device includes: an acquisition unit, configured to acquire first personal information and second personal information; a first determination unit, configured to respectively determine at least one structured text, at least one semi-structured text, and unstructured text corresponding to each of them according to the first personal information and the second personal information, where the structured text is structured information extracted from the first personal information and the second personal information, the semi-structured text is chunk information extracted from the first personal information and the second personal information, and the unstructured text is an unstructured representation of the first personal information and the second personal information; a second determination unit, configured to determine at least one dense similarity and at least one sparse similarity according to the at least one structured text respectively corresponding to the first personal information and the second personal information, determine at least one statistical feature similarity according to the at least one semi-structured text respectively corresponding to the first personal information and the second personal information, and determine a global similarity according to the unstructured text respectively corresponding to the first personal information and the second personal information; a generation unit, configured to input the at least one dense similarity, the at least one sparse similarity, the at least one statistical feature similarity, and a global similarity into a pre-trained state-aware mixture-of-experts model to generate a target similarity, where the target similarity is used to represent the similarity degree between the first personal information and the second personal information.
[0030] Optionally, the device further includes:
[0031] A determination unit, in response to the target similarity being greater than or equal to a set threshold, is configured to determine that the first personal information and the second personal information are personal information of the same person.
[0032] Optionally, the first determination unit is specifically configured to:
[0033] Input the first personal information and the second personal information into a large language model respectively, and automatically extract at least one structured text corresponding to each of them through the field extractor of the large language model.
[0034] Optionally, the first determination unit is specifically configured to:
[0035] Chunk the first personal information and the second personal information respectively according to a set rule to generate at least one semi-structured text corresponding to each of them.
[0036] Optionally, the second determination unit is specifically configured to:
[0037] Input the at least one structured text corresponding to the first personal information and the second personal information respectively into a pre-set text similarity model, and output at least one dense similarity and at least one sparse similarity.
[0038] Optionally, the text similarity model is an open-source general semantic vector model fine-tuned by a contrastive learning loss function.
[0039] Optionally, the second determination unit is specifically configured to:
[0040] Input the at least one semi-structured text corresponding to the first personal information and the second personal information respectively into a multi-level statistical feature extractor to generate at least one set of statistical sub-feature similarities, where the statistical sub-feature similarities include Jaccard similarity, locality-sensitive hashing similarity, and Levenshtein distance similarity;
[0041] Input the at least one set of statistical sub-feature similarities into a linear layer of a multi-layer perceptron to output at least one statistical feature similarity.
[0042] Optionally, the second determination unit is specifically configured to:
[0043] Perform a vector inner product on the unstructured text corresponding to the first personal information and the unstructured text corresponding to the second personal information to generate a global similarity.
[0044] Optionally, in the process of training the text similarity model, the obtaining unit is further configured to:
[0045] Obtain historical data, where the historical data includes positive sample data and negative sample data, the positive sample data is similar personal information of the same person, and the negative sample is different similar personal information;
[0046] The device further includes a construction unit for constructing a data set according to the historical data;
[0047] The generating unit is further configured to train an open-source general semantic vector model through the data set and a contrastive learning loss function to generate the text similarity model.
[0048] Optionally, the construction unit is specifically configured to:
[0049] Perform random noise addition processing on the historical data to construct the data set.
[0050] Optionally, during the process of training the state-aware mixture-of-experts model, the apparatus further includes:
[0051] A model determination unit, configured to determine the weight coefficients of each expert model in the state-aware mixture-of-experts model according to a gating network, and further determine the state-aware mixture-of-experts model.
[0052] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method described in any one of the first aspect or any possible implementation of the first aspect.
[0053] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when executed by a processor, implement the method described in any one of the first aspect or any possible implementation of the first aspect.
[0054] In the embodiments of the present invention, by obtaining the first personal information and the second personal information; respectively determining at least one structured text, at least one semi-structured text, and unstructured text corresponding to each of them according to the first personal information and the second personal information, where the structured text is structured information extracted from the first personal information and the second personal information, the semi-structured text is chunk information extracted from the first personal information and the second personal information, and the unstructured text is an unstructured representation of the first personal information and the second personal information; determining at least one dense similarity and at least one sparse similarity according to the at least one structured text corresponding to the first personal information and the second personal information respectively, and determining at least one statistical feature similarity according to the at least one semi-structured text corresponding to the first personal information and the second personal information respectively, and determining a global similarity according to the unstructured text corresponding to the first personal information and the second personal information respectively; inputting the at least one dense similarity, the at least one sparse similarity, the at least one statistical feature similarity, and the one global similarity into a pre-trained state-aware mixture-of-experts model to generate a target similarity, where the target similarity is used to represent the similarity degree between the first personal information and the second personal information. Through the above method, it is possible to efficiently and accurately determine whether the first personal information and the second personal information are personal information of the same person. Description of the Drawings
[0055] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0056] Figure 1 is a flowchart of a method for determining the same person for personal information in an embodiment of the present invention;
[0057] Figure 2 is a flowchart of another method for determining the same person for personal information in an embodiment of the present invention;
[0058] Figure 3 is a flowchart of a method for training a text similarity model in an embodiment of the present invention;
[0059] Figure 4 is a schematic diagram of a device for determining the same person for personal information in an embodiment of the present invention;
[0060] Figure 5 is a schematic diagram of an electronic device in an embodiment of the present invention. Detailed Embodiments
[0061] The following describes the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. In order to avoid obscuring the essence of the present application, well-known methods, processes, flows, components, and circuits are not described in detail.
[0062] In addition, those of ordinary skill in the art should understand that the accompanying drawings provided here are for illustrative purposes only, and the drawings are not necessarily drawn to scale.
[0063] Unless the context clearly requires otherwise, words such as "including" and "comprising" in the entire application document should be interpreted as having an inclusive meaning rather than an exclusive or exhaustive meaning; that is, it is the meaning of "including but not limited to".
[0064] In the description of the present application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0065] In the prior art, when an enterprise recruitment specialist uses a third-party website to select candidates, the third-party website usually only displays partial personal information of the candidates for privacy protection reasons, resulting in incomplete personal information. Even some important fields in the personal information may be blurred. Moreover, a large amount of personal information has accumulated in the internal talent pool, and the enterprise recruitment specialist cannot quickly determine whether a candidate already exists in the internal talent pool, resulting in the recruitment specialist spending a lot of time and energy manually comparing whether the resume obtained from the third party is duplicate with the resume recorded in the internal talent pool. Also, it will lead to the problem of over-expansion of the internal talent pool. The increase in redundant data not only wastes storage resources but may also affect the management efficiency of the talent pool. Among them, the resume can be a form of manifestation of personal information, and personal information can also include other forms of manifestation.
[0066] Currently, a vector representation method is used to judge the similarity of resumes. For example, an open-source general semantic vector model is adopted. The open-source general semantic vector model can specifically be the General Embedding (BGE) model of the Beijing Academy of Artificial Intelligence. However, since the text content of resumes is usually quite long and often exceeds the maximum text length that the BGE encoder can handle, and also, there are deficiencies in the importance of structured fields, scarcity of samples, imbalance between global and local judgments, and uncertainties brought about by missing fields when using the vector representation method. The above reasons make the above vector representation method have poor effect when judging whether resumes are duplicate, which may lead to misjudgment or even incorrect identification of the same person. Therefore, how to efficiently and accurately judge whether personal information is duplicate is a problem that needs to be solved currently.
[0067] In an embodiment of the present invention, to solve the above problems, a method for determining the same person of personal information is proposed, specifically as Figure 1 shown. The method includes:
[0068] Step S101, obtain a first personal information and a second personal information.
[0069] Specifically, the first personal information and the second personal information may include multiple contents such as basic information, educational background, work experience, hobbies, etc. Among them, the basic information includes name, age, and ethnicity; the educational background includes the time of studying in university and the name of the school, etc.; the work experience includes the working time, the name of the company where one has worked, etc., which is specifically determined according to the actual situation, and only exemplary descriptions are given here.
[0070] Step S102, respectively determine at least one structured text, at least one semi-structured text, and unstructured text corresponding to each of them according to the first personal information and the second personal information.
[0071] Specifically, the structured text is the structured information extracted from the first person information and the second person information, the semi-structured text is the chunked information extracted from the first person information and the second person information, and the unstructured text is the unstructured representation of the first person information and the second person information.
[0072] In a possible implementation manner, the step of respectively determining at least one corresponding structured text according to the first person information and the second person information specifically includes: inputting the first person information and the second person information into a large language model respectively, and automatically extracting at least one corresponding structured text for each of them through the Fields Extractor of the large language model (LLM), where different structured texts belong to different structural fields; wherein, the large language model can also be referred to as a large model, an Artificial Intelligence (AI) model or a large language model, and the large language model is a deep learning model based on a transformer architecture, capable of processing and generating natural language text, usually trained on a large amount of text data, having the ability to understand and generate language, and widely applied to dialogue systems, text generation, text extraction and other natural language processing tasks.
[0073] For example, taking the first person information as an example, assume that the first person information includes multiple items such as basic information, educational background, work experience, hobbies, etc. Through the large language model, automatic extraction is performed on the first person information, and multiple important fields such as basic information, educational background, work experience, hobbies, etc. are extracted, and each important field is a structured text; through the above method, the first person information is accurately parsed to ensure that the information of the important fields in the first person information is fully utilized.
[0074] In a possible implementation manner, the step of respectively determining at least one corresponding semi-structured text according to the first person information and the second person information specifically includes: chunking the first person information and the second person information respectively according to set rules to generate at least one corresponding semi-structured text for each of them; wherein, the above set rules can be taking each paragraph as a chunk, chunking by specific punctuation marks, taking a set number of sentences as a chunk, or taking a set text length as a chunk, which is specifically determined according to the actual situation.
[0075] For example, taking the first person information as an example, assume that every 10 sentences in the first person information are taken as a chunk, and each chunk is a semi-structured text.
[0076] In a possible implementation, the unstructured text is an unstructured representation of the first personal information and the second personal information. For example, the first personal information and the second personal information are represented by vectors, that is, the above unstructured representation.
[0077] Step S103: Determine at least one dense similarity and at least one sparse similarity according to the at least one structured text respectively corresponding to the first personal information and the second personal information, determine at least one statistical feature similarity according to the at least one semi-structured text respectively corresponding to the first personal information and the second personal information, and determine a global similarity according to the unstructured text respectively corresponding to the first personal information and the second personal information.
[0078] In a possible implementation, the determining at least one dense similarity and at least one sparse similarity according to the at least one structured text respectively corresponding to the first personal information and the second personal information specifically includes: inputting the at least one structured text respectively corresponding to the first personal information and the second personal information into a pre-set text similarity model, and outputting at least one dense similarity and at least one sparse similarity; wherein, the above text similarity model is an open-source general semantic vector model fine-tuned by a contrastive learning loss function, and the open-source general semantic vector model can be the BGE model, the BGE-M3 model, or other general models.
[0079] Illustrating with an example, assuming that three structured texts are extracted from the first personal information and three structured texts are also extracted from the second personal information, then each structured text extracted from the first personal information and the structured text in the same structural domain in the second personal information are input into the fine-tuned BGE-M3 model, and one dense similarity and one sparse similarity are output, so that the three pairs of structured texts correspond to three dense similarities and three sparse similarities.
[0080] In a possible implementation, determining at least one statistical feature similarity according to the at least one semi-structured text respectively corresponding to the first person information and the second person information specifically includes: inputting the at least one semi-structured text respectively corresponding to the first person information and the second person information into a multi-level statistical feature extractor, generating at least one set of statistical sub-feature similarities, where the statistical sub-feature similarities include Jaccord Similarity, SimHash Similarity, and Levenshtein Similarity; inputting the at least one set of statistical sub-feature similarities into a multi-layer perceptron linear layer, and outputting at least one statistical feature similarity.
[0081] For example, assume that four semi-structured texts are determined from the first person information, and four structured texts are also determined from the second person information. Then, input each semi-structured text determined from the first person information and the corresponding semi-structured text of the second person information into the multi-level statistical feature extractor, outputting a set of Jaccord Similarity, SimHash Similarity, and Levenshtein Similarity. Inputting the above set of Jaccord Similarity, SimHash Similarity, and Levenshtein Similarity into the multi-layer perceptron linear layer outputs one statistical feature similarity. Then, four pairs of semi-structured texts generate four statistical feature similarities.
[0082] In a possible implementation, determining a global similarity according to the unstructured texts respectively corresponding to the first person information and the second person information specifically includes: performing a vector inner product on the unstructured text corresponding to the first person information and the unstructured text corresponding to the second person information, generating a global similarity.
[0083] Step S104: Input the at least one dense similarity, the at least one sparse similarity, the at least one statistical feature similarity, and the one global similarity into a pre-trained state-aware mixture-of-experts model, generating a target similarity, where the target similarity is used to represent the similarity degree between the first person information and the second person information.
[0084] Specifically, during the process of training the State-aware Mixture of Experts (MOE) model, the missing fields in the first person information and the second person information are used to guide the determination of the weight coefficients of each expert model in the State-aware Mixture of Experts model according to the gating network. After determining the weight coefficients of each expert model, the weight coefficients are multiplied by the corresponding expert models, and the products are added together to form the State-aware Mixture of Experts model.
[0085] In the embodiments of the present invention, different levels of similarity such as global similarity and local similarity are respectively determined by the above method. Among them, the local similarity includes dense similarity, sparse similarity, and statistical feature similarity, realizing the balanced determination of global information and local information, providing a more hierarchical same-person determination result, and reducing the possibility of misjudgment. And, according to the missing situation of the fields in the personal information, the gating network is guided to set different weight values for different expert models, making up for the uncertainty brought by the missing fields, and being able to make a reasonable judgment even in the case of incomplete information, improving the robustness and accuracy of the State-aware Mixture of Experts model.
[0086] In a possible implementation manner, after the step S104, there are also other steps, specifically as Figure 2 shown, including:
[0087] Step S105: In response to the target similarity being greater than or equal to the set threshold, determine that the first person information and the second person information are personal information of the same person.
[0088] For example, assume that the set threshold is 0.7. In response to the target similarity being 0.8, since the target similarity 0.8 is greater than the set threshold, it is determined that the first person information and the second person information are personal information of the same person. If the target similarity is 0.6, and the target similarity 0.6 is less than the set threshold, it is determined that the first person information and the second person information are not personal information of the same person. The set threshold can also be selected as other thresholds, which are specifically determined according to the actual situation.
[0089] In a possible implementation manner, the process of training the text similarity model is specifically as Figure 3 shown, including the following steps:
[0090] Step S301: Obtain historical data.
[0091] Specifically, the historical data includes positive sample data and negative sample data. The positive sample data is similar personal information of the same person, and the negative sample is different similar personal information. Among them, the negative sample is a sample with similar information such as name, education information, and work experience, but in fact, it is not the same person.
[0092] For example, the positive sample data comes from the same-person data precipitated in business history and reflects a relatively real same-person matching scenario.
[0093] Step S302: Construct a data set according to the historical data.
[0094] Specifically, perform random noise addition processing on the historical data to construct the data set.
[0095] In a possible implementation manner, since personal information on third-party websites may be missing or hidden, it is necessary to perform random noise addition processing on the positive samples. For example, randomly modify information such as name, educational background, and work experience in the personal information of the positive samples to enhance the diversity and generalization ability of the data.
[0096] Step S303: Train an open-source general semantic vector model through the data set and a contrastive learning loss function to generate the text similarity model.
[0097] Specifically, the contrastive loss function includes a dense contrastive learning loss function and a sparse contrastive learning loss function. Train the open-source general semantic vector model based on the dense contrastive learning loss function and the sparse contrastive learning loss function to generate the text similarity model. Therefore, the text similarity model can output two types of similarities: dense similarity and sparse similarity.
[0098] Through the above embodiments, through the above-mentioned random noise addition processing method, element-level data augmentation can be achieved, the diversity of sample data can be enriched, the robustness of the text similarity model can be enhanced, the problem of sample scarcity can be effectively addressed, the generalization ability of the text similarity model can be improved, and its judgment accuracy in complex scenarios can be enhanced.
[0099] In the embodiments of the present invention, an apparatus for determining the same person of personal information is provided, such as Figure 4As shown in the figure, it specifically includes: an acquisition unit 401, a first determination unit 402, a second determination unit 403, and a generation unit 404; among them, the acquisition unit 401 is used to acquire the first personal information and the second personal information; the first determination unit 402 is used to respectively determine at least one structured text, at least one semi-structured text, and unstructured text corresponding to each of them according to the first personal information and the second personal information, where the structured text is the structured information extracted from the first personal information and the second personal information, the semi-structured text is the chunked information extracted from the first personal information and the second personal information, and the unstructured text is the unstructured representation of the first personal information and the second personal information; the second determination unit 403 is used to determine at least one dense similarity and at least one sparse similarity according to the at least one structured text corresponding to the first personal information and the second personal information respectively, determine at least one statistical feature similarity according to the at least one semi-structured text corresponding to the first personal information and the second personal information respectively, and determine a global similarity according to the unstructured text corresponding to the first personal information and the second personal information respectively; the generation unit 404 is used to input the at least one dense similarity, the at least one sparse similarity, the at least one statistical feature similarity, and a global similarity into a pre-trained state-aware mixture-of-experts model to generate a target similarity, where the target similarity is used to represent the similarity degree between the first personal information and the second personal information.
[0100] Furthermore, the device further includes:
[0101] A determination unit, in response to the target similarity being greater than or equal to a set threshold, is used to determine that the first personal information and the second personal information are personal information of the same person.
[0102] Furthermore, the first determination unit is specifically used for:
[0103] Input the first personal information and the second personal information into a large language model respectively, and automatically extract at least one structured text corresponding to each of them through the field extractor of the large language model.
[0104] Furthermore, the first determination unit is specifically used for:
[0105] Chunk the first personal information and the second personal information respectively according to set rules to generate at least one semi-structured text corresponding to each of them.
[0106] Furthermore, the second determination unit is specifically used for:
[0107] Input the at least one structured text corresponding to the first personal information and the second personal information into a pre-set text similarity model, and output at least one dense similarity and at least one sparse similarity.
[0108] Further, the text similarity model is an open-source general semantic vector model fine-tuned by a contrastive learning loss function.
[0109] Further, the second determination unit is specifically configured to:
[0110] Input the at least one semi-structured text corresponding to the first personal information and the second personal information into a multi-level statistical feature extractor to generate at least one set of statistical sub-feature similarities, where the statistical sub-feature similarities include Jaccard similarity, locality-sensitive hashing similarity, and Levenshtein distance similarity;
[0111] Input the at least one set of statistical sub-feature similarities into a multi-layer perceptron linear layer to output at least one statistical feature similarity.
[0112] Further, the second determination unit is specifically configured to:
[0113] Perform a vector inner product on the unstructured text corresponding to the first personal information and the unstructured text corresponding to the second personal information to generate a global similarity.
[0114] Further, in the process of training the text similarity model, the obtaining unit is further configured to:
[0115] Obtain historical data, where the historical data includes positive sample data and negative sample data, the positive sample data is similar personal information of the same person, and the negative sample is different similar personal information;
[0116] The device further includes a construction unit configured to construct a data set according to the historical data;
[0117] The generating unit is further configured to train an open-source general semantic vector model through the data set and a contrastive learning loss function to generate the text similarity model.
[0118] Further, the construction unit is specifically configured to:
[0119] Perform random noise addition processing on the historical data to construct the data set.
[0120] Further, during the process of training the state-aware mixture-of-experts model, the device further includes:
[0121] A model determination unit is configured to determine the weight coefficients of the respective expert models in the state-aware mixture-of-experts model according to a gating network, and further determine the state-aware mixture-of-experts model.
[0122] Figure 5 is a schematic structural diagram of the electronic device in an embodiment of the present invention. As Figure 5 shown, it includes a general computer hardware structure, which at least includes a processor 501 and a memory 502. The processor 501 and the memory 502 are connected through a bus 503. The memory 502 is adapted to store instructions or programs executable by the processor 501. The processor 501 may be an independent microprocessor or a set of one or more microprocessors. Thus, by executing the instructions stored in the memory 502, the processor 501 executes the method flow of the embodiment of the present invention as described above to implement data processing and control of other devices. The bus 503 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to a display controller 504, a display device, and an input / output (I / O) device 505. The input / output (I / O) device 505 may be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices well known in the art. Typically, the input / output device 505 is connected to the system through an input / output (I / O) controller 506.
[0123] Among them, the instructions stored in the memory 502 are executed by at least one processor 501 to implement: obtaining first personal information and second personal information; respectively determining at least one structured text, at least one semi-structured text, and unstructured text corresponding to each of them according to the first personal information and the second personal information, where the structured text is structured information extracted from the first personal information and the second personal information, the semi-structured text is chunk information extracted from the first personal information and the second personal information, and the unstructured text is an unstructured representation of the first personal information and the second personal information; determining at least one dense similarity and at least one sparse similarity according to the at least one structured text corresponding to the first personal information and the second personal information respectively, determining at least one statistical feature similarity according to the at least one semi-structured text corresponding to the first personal information and the second personal information respectively, and determining a global similarity according to the unstructured text corresponding to the first personal information and the second personal information respectively; inputting the at least one dense similarity, the at least one sparse similarity, the at least one statistical feature similarity, and the one global similarity into a pre-trained state-aware mixture-of-experts model to generate a target similarity, where the target similarity is used to represent the similarity degree between the first personal information and the second personal information.
[0124] Specifically, the electronic device includes one or more processors 501 and a memory 502. Figure 5 Taking one processor 501 as an example. The processor 501 and the memory 502 can be connected through a bus or other means. Figure 5 Taking the connection through a bus as an example. The memory 502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The processor 501 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in the memory 502, that is, implementing the above method for determining personal information and person determination.
[0125] The memory 502 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store an option list, etc. In addition, the memory 502 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 502 optionally includes a memory remotely set relative to the processor 501, and these remote memories can be connected to an external device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0126] One or more modules are stored in the memory 502 and, when executed by one or more processors 501, implement the method for determining personal information and person determination in any of the above method embodiments.
[0127] As those skilled in the art will realize, various aspects of the embodiments of the present invention can be implemented as a system, a method, or a computer program product. Therefore, various aspects of the embodiments of the present invention can take the following forms: a completely hardware implementation, a completely software implementation (including firmware, resident software, microcode, etc.), or an implementation that combines software aspects and hardware aspects, which is generally referred to as "circuit", "module", or "system" in this article. In addition, various aspects of the embodiments of the present invention can take the following form: a computer program product implemented in one or more computer-readable media, the computer-readable media having computer-readable program code implemented thereon.
[0128] Any combination of one or more computer-readable media may be utilized. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of embodiments of the present invention, the computer-readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0129] The computer-readable signal medium may include a propagated digital signal having computer-readable program code embodied therein, either as in a baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including but not limited to: electromagnetic, optical, or any suitable combination thereof. The computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0130] Any suitable medium may be used to transmit the program code embodied on the computer-readable medium, including but not limited to wireless, wireline, fiber optic cable, RF, etc., or any suitable combination of the foregoing.
[0131] The computer program code for performing operations in connection with aspects of the embodiments of the present invention may be written in any combination of one or more programming languages, including: object-oriented programming languages such as Java, Smalltalk, C++; and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet service provider).
[0132] The flowchart illustrations and / or block diagrams of the method, apparatus (system) and computer program product according to the embodiments of the present invention described above depict various aspects of the embodiments of the present invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and the combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions (executed via the processor of the computer or other programmable data processing device) create a means for implementing the functions / actions specified in the flowchart and / or block diagram block or blocks.
[0133] These computer program instructions can also be stored in a computer-readable medium that can direct a computer, other programmable data processing device, or other apparatus to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in the flowchart and / or block diagram block or blocks.
[0134] The computer program instructions can also be loaded onto a computer, other programmable data processing device, or other apparatus to cause a series of operational steps to be performed on the computer, other programmable device, or other apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in the flowchart and / or block diagram block or blocks.
[0135] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse. When a user refuses to process personal information other than the necessary information required for basic functions, it will not affect the user's use of basic functions.
Claims
1. A method for determining personal information and its counterparts, characterized in that, The method includes: Obtain the first person information and the second person information; Respectively determine at least one structured text, at least one semi-structured text, and unstructured text corresponding to each of them according to the first person information and the second person information, where the structured text is the structured information extracted from the first person information and the second person information, the semi-structured text is the chunked information extracted from the first person information and the second person information, and the unstructured text is the unstructured representation of the first person information and the second person information; Determine at least one dense similarity and at least one sparse similarity respectively according to the at least one structured text corresponding to the first person information and the second person information, determine at least one statistical feature similarity according to the at least one semi-structured text corresponding to the first person information and the second person information respectively, and determine a global similarity according to the unstructured text corresponding to the first person information and the second person information respectively; Input the at least one dense similarity, the at least one sparse similarity, the at least one statistical feature similarity, and the one global similarity into a pre-trained state-aware mixture-of-experts model to generate a target similarity, where the target similarity is used to represent the similarity degree between the first person information and the second person information.
2. The method according to claim 1, wherein The method further includes: In response to the target similarity being greater than or equal to a set threshold, determine that the first person information and the second person information are personal information of the same person.
3. The method according to claim 1, wherein The step of respectively determining at least one structured text corresponding to each of them according to the first person information and the second person information specifically includes: Input the first person information and the second person information into a large language model respectively, and automatically extract at least one structured text corresponding to each of them through the field extractor of the large language model.
4. The method according to claim 1, wherein The step of respectively determining at least one semi-structured text corresponding to each of them according to the first person information and the second person information specifically includes: Chunk the first person information and the second person information respectively according to set rules to generate at least one semi-structured text corresponding to each of them.
5. The method according to claim 1, characterized in that, The step of respectively determining at least one dense similarity and at least one sparse similarity according to the at least one structured text corresponding to the first person information and the second person information specifically includes: Input the at least one structured text corresponding to the first person information and the second person information into a pre-set text similarity model, and output at least one dense similarity and at least one sparse similarity.
6. The method according to claim 5, wherein The text similarity model is an open-source general semantic vector model fine-tuned by a contrastive learning loss function.
7. The method according to claim 1, characterized in that, Determine at least one statistical feature similarity according to the at least one semi-structured text corresponding to the first person information and the second person information respectively, specifically including: Input the at least one semi-structured text corresponding to the first personal information and the second personal information into a multi-level statistical feature extractor to generate at least one set of statistical sub-feature similarities, where the statistical sub-feature similarities include Jaccard similarity, locality-sensitive hashing similarity, and Levenshtein distance similarity; Input the at least one set of statistical sub-feature similarities into a multi-layer perceptron linear layer to output at least one statistical feature similarity.
8. The method according to claim 1, wherein Determine a global similarity according to the unstructured texts corresponding to the first personal information and the second personal information respectively, specifically including: Perform a vector inner product on the unstructured text corresponding to the first personal information and the unstructured text corresponding to the second personal information to generate a global similarity.
9. The method according to claim 6, wherein The process of training the text similarity model specifically includes: Obtain historical data, where the historical data includes positive sample data and negative sample data, the positive sample data is similar personal information of the same person, and the negative sample is different similar personal information; Construct a data set according to the historical data; Train an open-source general semantic vector model through the data set and a contrastive learning loss function to generate the text similarity model.
10. The method according to claim 9, characterized in that, The constructing the data set according to the historical data specifically includes: Perform random noise addition processing on the historical data to construct the data set.
11. The method according to claim 1, characterized in that, The process of training the state-aware mixture-of-experts model specifically includes: Determine the weight coefficients of each expert model in the state-aware mixture-of-experts model according to a gating network, and further determine the state-aware mixture-of-experts model.
12. An apparatus for determining personal information doppelgangers, characterized in that, The device includes: An acquisition unit, configured to acquire first personal information and second personal information; A first determination unit, configured to respectively determine at least one structured text, at least one semi-structured text, and unstructured text corresponding to each of them according to the first personal information and the second personal information, where the structured text is structured information extracted from the first personal information and the second personal information, the semi-structured text is chunk information extracted from the first personal information and the second personal information, and the unstructured text is an unstructured representation of the first personal information and the second personal information; A second determination unit, configured to determine at least one dense similarity and at least one sparse similarity according to the at least one structured text corresponding to the first personal information and the second personal information respectively, determine at least one statistical feature similarity according to the at least one semi-structured text corresponding to the first personal information and the second personal information respectively, and determine a global similarity according to the unstructured texts corresponding to the first personal information and the second personal information respectively; A generation unit, configured to input the at least one dense similarity, the at least one sparse similarity, the at least one statistical feature similarity, and a global similarity into a pre-trained state-aware mixture-of-experts model to generate a target similarity, where the target similarity is used to represent the similarity degree between the first personal information and the second personal information.
13. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1-11 is implemented.