Cross-language named entity recognition method based on big and small model collaborative reasoning

By adopting the collaborative reasoning method of size and model in cross-language named entity recognition technology, combining the inference results of small models and large models, the problem of poor performance of the existing technology in complex contexts is solved, and higher recognition accuracy and interpretability are achieved.

CN120218069APending Publication Date: 2025-06-27NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510282078.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing cross-language named entity recognition technology performs poorly in complex contexts, and is limited by the lack of multilingual understanding capabilities, resulting in misjudgment in the model when identifying named entities.

Method used

The method based on size and model collaborative reasoning is adopted, and the pre-trained model is fine-tuned, combined with the inference results of small models and large models, and the multi-step inference and reconfirmation mechanism of large models is used to integrate the naming entity recognition results of both.

Benefits of technology

Improves the accuracy of cross-language named entity recognition, especially in complex contexts, reduces error rates, optimizes user experience, and provides better interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218069A_ABST
    Figure CN120218069A_ABST
Patent Text Reader

Abstract

The invention provides a cross-language named entity recognition method based on big and small model collaborative reasoning. The method comprises the following steps: step 1, preparing a training set and a test set; 2, using the training set to finely adjust the pre-training model to obtain a cross-language named entity recognition model; step 3, using a cross-language named entity recognition model to perform reasoning on the data in the test set to obtain a small model side reasoning result; 4, reasoning the data in the test set by using the large model to obtain a large model side reasoning result; and 5, fusing the small model side reasoning result and the large model side reasoning result to obtain a final recognition result of the cross-language named entity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cross - language named entity recognition method, in particular to a cross - language named entity recognition method based on collaborative inference of large and small models. Background Art

[0002] The information provided in this part is only background information related to the present disclosure, and it is not necessarily prior art.

[0003] Information Extraction (IE) is one of the branches in the field of Natural Language Processing (NLP), aiming to enable machines to extract structured text information from unstructured text. For example, it extracts many information such as the event subject, entity type, event time, etc. from news text. As an important task in natural language processing, it plays an important role in achieving precise understanding of text, summarizing text semantics, analyzing entity relationships, etc., helping people quickly and accurately extract effective information from massive data.

[0004] Named Entity Recognition (NER) is one of the basic tasks in the field of information extraction, mainly responsible for identifying named entities in text and classifying them into different predefined entity types. This task is an important part of many downstream tasks and is applied in many industrial products. For example, in commercial Web search engines such as Microsoft Bing, NER is an important module for Query understanding and question answering. For voice assistants such as Siri and Cortana, NER is a key module for achieving Spoken Language Understanding (SLU).

[0005] In the prior art, Cross - Lingual Transfer (CLT) aims to transfer high - resource languages with rich entity labels to low - resource languages. The goal of the cross - language NER task is to use the labeled data of the source language to achieve the NER task on low - resource languages.

[0006] Currently, the mainstream cross - language named entity recognition technology is mainly achieved by fine - tuning the multi - language pre - trained model of the encoder structure. The multi - language pre - trained model is an artificial intelligence model trained on a large - scale multi - language text for specific tasks. For example, the mBERT model based on the Transformer structure has certain multi - language understanding ability. Through fine - tuning with specific task data, it can adapt to different downstream tasks. Regarding the cross - language feature, existing work trains the model by designing different supervision signals, enabling the model to have the ability to perform NER tasks in the target language. Existing work is mainly divided into three categories, namely feature - based methods, translation - based methods, and self - training methods.

[0007] The working principle of the cross - language NER model is to judge the probability that each word in the text belongs to different categories of named entity labels, and select the one with the highest probability as the prediction result. Existing methods all achieve this task based on fine - tuning the multi - language pre - trained model.

[0008] However, although these solutions can all improve the effect of the cross - language NER task to a certain extent, in some relatively complex contexts, their performance is still not satisfactory. The multi - language understanding ability of the language model itself is an important factor restricting the performance of these solutions. Limited by the model size and architecture, although the multi - language pre - trained model with the traditional encoder architecture has certain multi - language understanding ability, in some relatively complex contexts, its performance is still poor. For example, for the text "Ein Spieler stand im Mittelpunkt des Interesses, der vor der Saison von Wicker zum Zweitligisten TV Gelnhausen gewechselt Ralph" . Among them, Wicker is the name of a team, and the named entity type is organization (ORG). However, due to the insufficient multi - language understanding ability of the above - mentioned model, it usually misjudges it as a person's name (PER). The weak multi - language understanding ability causes these models to perform poorly on some samples that require context understanding.

[0009] It should be noted that the information disclosed in the above background technology section is only used to strengthen the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0010] Object of the Invention: The technical problem to be solved by the present invention is to provide a cross - language named entity recognition method based on collaborative reasoning of large and small models in view of the deficiencies of the prior art.

[0011] To solve the above technical problems, the present invention discloses a cross - language named entity recognition method based on collaborative inference of large and small models, including the following steps:

[0012] Step 1, prepare a training set and a test set;

[0013] Step 2, use the training set to fine - tune the pre - trained model to obtain a cross - language named entity recognition model;

[0014] Step 3, use the cross - language named entity recognition model to infer the data in the test set to obtain the inference result on the small - model side;

[0015] Step 4, use the large model to infer the data in the test set to obtain the inference result on the large - model side;

[0016] Step 5, fuse the inference results on the small - model side and the large - model side to obtain the final recognition result of cross - language named entities.

[0017] Further, the training set in Step 1 is a set of training set data Train Set, including: source - language labeled data and target - language unlabeled data; among them, the source - language labeled data is the source - language text that has been manually annotated, and each source - language text is annotated with entities of specific types, and the target - language unlabeled data is the target - language text that has not been annotated;

[0018] The test set is a set of test set data Test Set, and the test set data Test Set is the target - language labeled data, which is the target - language text that has been manually annotated, and each target - language text is annotated with entities of specific types.

[0019] Further, obtaining the cross - language named entity recognition model in Step 2 includes:

[0020] Step 2 - 1, use the source - language labeled data in the training set to train the pre - trained model;

[0021] Step 2 - 2, use the pre - trained model to infer the target - language unlabeled data in the training set to obtain pseudo - labels on the target - language unlabeled data;

[0022] Step 2 - 3, add the target - language unlabeled data with pseudo - labels to the training process, and train the pre - trained model by gradient descent to obtain a cross - language named entity recognition model.

[0023] Further, obtaining the inference result on the small - model side in Step 3 includes:

[0024] Step 3 - 1, input each target - language text in the test set into the cross - language named entity recognition model;

[0025] Step 3-2: The cross-lingual named entity recognition model outputs the named entities corresponding to different words in the target language text as the recognition inference result.

[0026] Step 3-3: Repeat Step 3-1 to Step 3-2, and summarize the recognition inference results of each target language text in the test set to obtain the inference result on the small model side.

[0027] Furthermore, when inferring the data in the test set in Step 4, a multi-step inference method is adopted.

[0028] Furthermore, the multi-step inference method described in Step 4 includes:

[0029] Step 4-1: Provide the target language text in the test set to the large model and specify the named entity type TYPE.

[0030] Step 4-2: The large model performs inference to obtain the overall paraphrase Expl sent and the word-level paraphrase Expl word of the input target language text, specifically including:

[0031] By inputting prompting words to the large model, the large model gives the explanation of the target language text and the explanation of each word in the target language text, and obtains the overall paraphrase Expl sent and the word-level paraphrase Expl word .

[0032] Step 4-3: Through the inference of the large model, obtain the potential entities Entity p in the target language text, specifically including:

[0033] For the specified named entity type TYPE, prompts are respectively constructed, and the large model infers to obtain the entities of the corresponding type included in the target language text of the corresponding target language in the test set, that is, the potential entities Entity p ;

[0034] Step 4-4: Through the inference of the large model, obtain the potential entity paraphrase Expl p corresponding to the potential entity Entity phra in the target language text;

[0035] Step 4-5: Through the inference of the large model, comprehensively consider the overall paraphrase Expl sent of the target language text and the entity paraphrase Expl phrase to confirm the named entity type EntityType p of the entity.

[0036] Further, the named entity type EntityType of the confirmed entity described in steps 4-5 p , includes:

[0037] Design the discrimination of the named entity type TYPE corresponding to the potential entity Entity p as a multiple-choice question. In the prompt words of the large model, provide the target language text, the corresponding paraphrase Expl of the potential entity phras and the potential entity Entity p . Each option of the multiple-choice question corresponds to 1 named entity type TYPE. Use the large model to select the correct named entity type to obtain the final named entity type EntityType corresponding to the potential entity Entity p . p .

[0038] Further, for the fusion of the inference results on the small model side and the inference results on the large model side described in step 5, a reconfirmation mechanism is adopted.

[0039] Further, the reconfirmation mechanism described in step 5 includes:

[0040] Step 5-1: Compare the inference results on the small model side and the inference results on the large model side, and find the conflicting parts, specifically including:

[0041] Let the set of inference results on the large model side be P LLM , and the set of inference results on the small model side be P SLM . Then the conflicting part is the difference between the above two sets P LLM -P SLM and P SLM -P LLM ; the non-conflicting part is the common part P LLM ∩P SLM ;

[0042] Step 5-2: For the conflicting parts, make a judgment based on the large model, specifically including:

[0043] For each piece of test set data Test Set containing conflicting results, construct the prompt words of the large model. The prompt words include the target language text and the conflicting parts, and use the large model to judge and output the correct result P judge ;

[0044] Step 5-3: Combine the non-conflicting part P LLM ∩P SLM and the correct result P judge obtained in step 5-2, that is, take the union of the above two to obtain the named entity recognition result P Union, as the final recognition result of cross - language named entities.

[0045] Furthermore, the pre - trained model described in step 2 is a BERT model.

[0046] Beneficial effects:

[0047] 1. By designing explicit multi - step reasoning of the large model, the present invention makes full use of the strong multi - language understanding ability of the large model, so that the large model is adapted to the named entity recognition task.

[0048] 2. By designing a re - confirmation mechanism to process the reasoning results of the large model and the small model, and using the reasoning ability of the large model to eliminate the contradictory entities in different results, the named entity recognition results of the large model and the small model are fused.

[0049] 3. The present invention only needs to fine - tune the small model, and the fine - tuning cost is relatively low compared with that of the large model, and the effect is good.

[0050] 4. The present invention can better make up for the deficiency of the model language understanding ability in traditional methods, thereby reducing the error rate of the cross - language named entity recognition system and optimizing the user experience.

[0051] 5. The present invention has better interpretability, can provide corresponding explanations and judgment bases for the text and the entities in the text, and helps to improve the user's understanding of the recognized entities. Brief description of the drawings

[0052] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above - mentioned and / or other advantages of the present invention will become clearer.

[0053] Figure 1 is a schematic diagram of the collaborative reasoning algorithm process of the large and small models in the present invention.

[0054] Figure 2 is a schematic diagram of the multi - step reasoning process of the large model in the present invention.

[0055] Figure 3 is a schematic diagram of the re - confirmation mechanism process in the present invention. Specific embodiments

[0056] The overall idea of the present invention is as follows: A cross-lingual named entity recognition method based on collaborative reasoning of a large model and a small model is proposed. When using the large model to achieve named entity recognition, by designing explicit multi-step reasoning, the large model is adapted to this task, improving the extraction effect of named entity recognition of the large model. When merging the results, by designing a reconfirmation mechanism, the contradictory entities existing in different results are resolved, so as to fuse the reasoning results of the large model and the small model, alleviate the deficiency of the multi-language understanding ability of the small model, and improve the recognition rate of named entities in complex scenarios.

[0057] The specific technical solution of the present invention is as follows:

[0058] The overall process of the cross-lingual named entity recognition method based on collaborative reasoning of a large model and a small model is as Figure 1 shown:

[0059] Step 101, prepare named entity recognition labels, training sets, and test sets; among them, the training set is a set of training set data Train Set, including: source language labeled data and target language unlabeled data; the test set is a set of test set data Test Set, including multiple pieces of target language text manually annotated, and each piece of data is annotated with specific types of entities, which is used to evaluate the effect of cross-lingual named entity recognition. When fine-tuning a cross-lingual named entity recognition model based on a pre-trained model (Bidirectional Encoder Representation from Transformers, BERT), specific types of named entities need to be specified. According to the required named entity type TYPE, the labeled data on the source language with corresponding annotations and the text of the target language unlabeled data are provided as training data for the cross-lingual named entity recognition model to learn.

[0060] Step 102, use the source language labeled data and target language unlabeled data in the training set to fine-tune based on the pre-trained model to obtain a cross-lingual named entity recognition model (such as the TCS cross-lingual named entity recognition model). For the specific fine-tuning process, refer to the TCS model. Use the source language labeled data in the training set to train the pre-trained model, and use this model to reason about the target language unlabeled data to obtain pseudo-labels on the target language unlabeled data. Then, add the target language unlabeled data with pseudo-labels to the training process and train the model by means of gradient descent to obtain a cross-lingual named entity recognition model.

[0061] Step 103: Use the cross-lingual named entity recognition model obtained through fine-tuning to perform inference on the test set Test Set. Each target language text in the test set is input into the cross-lingual named entity recognition model, and the model outputs the named entity recognition inference results corresponding to different words in the target language text. The results of each text in the test set are aggregated to obtain the named entity recognition inference results of the small model side for the test set.

[0062] Step 104: Use a large model (such as ChatGPT, LLaMA) to perform multi-step inference on each test set text in the test set Test Set to obtain the inference results of the large model side for the target language text in the test set. The large model has strong multi-language understanding ability, and this method makes full use of its understanding ability and designs a multi-step inference scheme. The specific multi-step inference process is as Figure 2 shown:

[0063] Step 401: Input the test set text in the corresponding target language and the named entity type TYPE. Inference based on the large model requires providing the large model with the test set text in the target language and specifying the named entity type for further inference in subsequent steps.

[0064] Step 402: Through large model inference, obtain the overall paraphrase Expl sent and word-level paraphrase Expl word . Different from other natural language understanding tasks, named entity recognition requires further analysis of the fine-grained word-level information of the text. Therefore, in addition to the overall interpretation of the text, it is also necessary to construct an interpretation of different words in the text. By inputting prompts to require the large model to give an interpretation of the target language test set text and each word in the text, the overall paraphrase and word-level paraphrase of the text can be obtained.

[0065] Step 403: Through large model inference, obtain the potential entities Entity p in the text. Specifically, for the specified named entity type, prompts need to be constructed separately to ask the large model about the entities of the corresponding type contained in the target language text in the target language test set, so as to obtain the preliminary results of different types of entities in the target language text. However, since this result only undergoes single-step inference, there is usually a lot of noise. Therefore, it is necessary to further judge the type of the preliminary results of different types of entities obtained to reduce the noise. These preliminary results, that is, the entities whose final types have not been determined, are called potential entities. Through this step, the entity situations of different types in the text can be initially obtained, but the true categories of these potential entities Entity p need to be further judged.

[0066] Step 404: Through large model reasoning, obtain the potential entity Entity found in the target language test set text in the previous step p The corresponding entity interpretation Expl in this text phrase . Specifically, construct a corresponding prompt, and provide the target language test set text, the corresponding overall interpretation, and the potential entity to the large model in the prompt, and require the large model to output the interpretation of the potential entity.

[0067] Step 405: Through large model reasoning, synthesize the overall interpretation Expl of the text sent and the entity interpretation Expl phras to confirm the named entity type EntityType of the entity p . Specifically, here the discrimination of the named entity type corresponding to the potential entity is designed in the form of a multiple-choice question. In the prompt, provide the target language text, the interpretation of the corresponding potential entity, and the potential entity. Each option of the multiple-choice question corresponds to a specific named entity type, and require the model to select the correct named entity type. The large model inputs the above prompt content and outputs the option of the correct named entity type, so as to obtain the final named entity type EntityType corresponding to the potential entity p .

[0068] Step 105: Through the reconfirmation mechanism, fuse the reasoning results of the large model and the small model. The specific process of the reconfirmation mechanism is as Figure 3 shown:

[0069] Step 501: Input the reasoning results of the large model and the small model for subsequent steps.

[0070] Step 502: Compare the reasoning results of the large model and the small model, and find the parts where there are conflicts in the reasoning results of the large model and the small model. The conflicting parts are the parts where the predictions of the two are inconsistent. Let the set of the reasoning results of the large model be P LLM , and the set of the reasoning results of the small model be P SLM , and the conflicting part is the difference between the two sets P LLM -P SLM and P SLM -P LLM . The non-conflicting part is the common part P of the prediction results of the two LLM ∩P SLM .

[0071] Step 503: For the conflicting results, make an inference based on the large model to judge which result is more correct. Specifically, for each test set data TestSet containing conflicting results, construct a prompt for the large model to make an inference. The prompt contains the target language test set text and the conflicting part P LLM -PSLM and P SLM -P LLM , require the large language model to judge and output which inference result is more reasonable, so as to obtain the correct result P judged by the model judge .

[0072] Step 504, merge the results of the common part and the conflicting part, that is, the common result P of each target language text in the test set LLM ∩P SLM and the correct result P judged in the previous step judge take the union to obtain the final merged named entity recognition result P Union .

[0073] The high-performance large language model mentioned in the present invention refers to a large language model with a parameter quantity greater than or equal to 70 billion, while the small model mentioned refers to a large language model with a parameter quantity less than 10 billion.

[0074] Example:

[0075] Step 101, input named entity recognition tags, such as PER (person name), LOC (address name), ORG (organization name), named entity marking data in English (such as for the English text EU rejects German call to boycott British lamb., where the named entities include EU, marked as ORG), and unmarked named entity recognition data in German (such as the German text Den Grünen müssen derweil die Ohren geklungen haben.). Taking the TCS model as an example, it is necessary to specify entities of the PER, LOC, and ORG types, and provide corresponding English marking data and German unmarked data of these three types of named entity types for model learning.

[0076] Step 102, use the source language marking data and target language unmarked data in the training set, fine-tune based on the pre-trained model BERT, and train using the training method in the TCS model to obtain a named entity recognition model in German.

[0077] Step 103, use the trained TCS German named entity recognition model to perform inference on the German test set text (such as Wicker siegte gegen den TVG mit 11:8(7:5)), and obtain the named entity recognition inference result of the TCS model for this text. The inference result of the model is: Wicker[PER], TVG[ORG].[[]]

[0078] Step 104, use the ChatGPT model to perform multi-step reasoning on the text in the target language test set, and obtain the reasoning results of the large model for the target language text. The specific multi-step reasoning process is as Figure 2 shown:

[0079] Step 401, input the text of the target language test set and the named entity recognition tags. For the reasoning of the large model, the text of the target language (such as Wicker siegte gegen den TVG mit 11:8(7:5)) and specific named entity types (such as PER, LOC, ORG) need to be provided to the model for further inference by the large model in the subsequent steps.

[0080] Step 402, through the reasoning of the large model, obtain the overall interpretation and word-level interpretation of the text in the target language test set. For the target language text Wicker siegte gegen den TVG mit 11:8(7:5), the overall interpretation obtained through the large model is:

[0081] This is likely referring to a sports match where Wicker(probably a team or a club) defeated TVG(another team or club). The final score was 11:8, and the halftime score(or score at a specific point in the match) was 7:5.

[0082] The word-level interpretation needs to explain each word in the text separately. For example, its word-level interpretation of TVG is:

[0083] another team or club that participated in the sports match and was defeated by Wicker.

[0084] Step 403, through the reasoning of the large model, obtain the potential entities in the text. For example, for the type PER, ask the model the following prompt:

[0085] Based on the explanation for each word in the sentence and the sentence meaning above, then give each person's name (without any adjective or title) mentioned in the sentence. Give the name of person in list form like [NAME1, NAME2] using a single row in german. If there's no name of person, output 'no person's name in the sentence'.

[0086] Through the above prompts, the potential entity of the PER type obtained by the model is Wicker.

[0087] Step 404, through large model reasoning, obtain the specific meaning of Wicker in the text, and its prompt is:

[0088] Please explain the meaning of the entity in the sentence based on the explanation of sentence meaning.

[0089] Sentence: Wicker siegte gegen den TVG mit 11:8(7:5)

[0090] Explanation: This is likely referring to a sports match where Wicker (probably a team or a club) defeated TVG (another team or club). The final score was 11:8, and the halftime score (or score at a specific point in the match) was 7:5.

[0091] Based on the explanation for the sentence, the entity 'Wicker' in this sentence refers to

[0092] Through the above prompts, the model outputs the explanation of 'Wicker' in the text:

[0093] a team or a club that participated in the sports match and won against TVG.

[0094] Step 405, through large model reasoning, comprehensively combine the overall interpretation and entity interpretation of the text to confirm the specific type of the entity. For 'Wicker', the input prompts are as follows:

[0095] The following is a multiple-choice question. Please choose the most reasonable one from the following options.

[0096] Question: In the Sentence: Wicker siegte gegen den TVG mit 11:8(7:5)

[0097] the Proper Noun 'Wicker' may means a team or a club that participated in the sports match and won against TVG. in this sentence. Which type entity is it most likely to be in this sentence?

[0098] A. Organization

[0099] B. Location

[0100] C. Person

[0101] Answer:

[0102] The large model outputs option A according to the prompts.

[0103] Step 105, through the reconfirmation mechanism, integrate the reasoning results of the large model and the small model. The specific process of the reconfirmation mechanism is as followsFigure 3 As shown below:

[0104] Step 501: Input the inference results of the large model (Wicker[ORG], TVG[ORG]) and the inference results of the small model (Wicker[PER], TVG[ORG]) for subsequent steps.

[0105] Step 502: Compare the inference results of the large model and the small model, and find the parts where the inference results of the large model and the small model conflict, which are Wicker[ORG] and Wicker[PER], and the non-conflicting part is TVG[ORG].

[0106] Step 503: For the conflicting inference results, namely Wicker[ORG] and Wicker[PER], make an inference based on the large model to determine which result is more correct. For the above conflict, the prompt input to the large model is as follows:

[0107] Instruction: You are a linguist proficient in named entity recognition. Two individuals have disagreed on the entity annotations for the same sentence. Please determine which annotation is more reasonable, focusing only on persons, locations, and organizations.

[0108] For the sentence: SENT

[0109] Person A thinks 'Wicker' is the name of an organization in this sentence.

[0110] Person B thinks 'Wicker' is the name of a person in this sentence.

[0111] Their opinions are contradictory. Which label do you think is right?

[0112] Finally, the output of the model is:

[0113] Based on the context of the sentence, it is more reasonable to consider 'Wicker' as the name of an organization rather than a person. This is because the verb'siegte' (won) suggests that 'Wicker' is the subject of the sentence. Since TVG (likely short for TV Gelnhausen) is a sports club, it is reasonable to assume that 'Wicker' is also a team or organization, as sports teams or clubs typically compete against each other.

[0114] Thus, it is determined that Wicker[ORG] is correct.

[0115] In step 504, the results of the common part and the conflicting part are merged, that is, the common result TVG[ORG] in each sentence is combined with the correct result Wicker[ORG] determined in the previous step to obtain the final merged named entity recognition result, that is, Wicker is predicted as an entity of the ORG type, and TVG is also predicted as an entity of the ORG type.

[0116] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the content of the invention of a cross-language named entity recognition method based on collaborative inference of large and small models and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0117] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a computer program, that is, a software product. This computer program software product can be stored in a storage medium, including several instructions to enable a device including a data processing unit (which can be a personal computer, server, single-chip microcomputer, MCU or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0118] The present invention provides an idea and method for a cross-language named entity recognition method based on collaborative inference of large and small models. There are many specific methods and ways to implement this technical solution. The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.

Claims

1. A cross-language named entity recognition method based on large and small model collaborative reasoning, characterized in that: The following steps are involved: Step 1: Prepare training set and test set; Step 2: Use the training set to fine-tune the pre-trained model to obtain a cross-language named entity recognition model; Step 3: Use the cross-language named entity recognition model to infer the data in the test set and obtain the inference results on the small model side; Step 4: Use the large model to infer the data in the test set and obtain the inference results on the large model side; Step 5: Fuse the inference results of the small model side and the large model side to obtain the final recognition result of the cross-language named entity.

2. According to the cross-language named entity recognition method based on large and small model collaborative reasoning according to claim 1, it is characterized in that: The training set described in step 1 is a set of training set data Train Set, including: source language labeled data and target language unlabeled data; wherein, the source language labeled data is manually annotated source language text, each source language text is annotated with a specific type of entity, and the target language unlabeled data is an unlabeled target language text; The test set is a collection of test set data Test Set, and the test set data Test Set is target language labeled data, which is manually annotated target language text, and each target language text is annotated with a specific type of entity.

3. According to the cross-language named entity recognition method based on large and small model collaborative reasoning according to claim 2, it is characterized in that: The cross-language named entity recognition model obtained in step 2 includes: Step 2-1, using the source language labeled data in the training set to train the pre-trained model; Step 2-2, use the pre-trained model to infer the target language unlabeled data in the training set to obtain pseudo labels on the target language unlabeled data; In step 2-3, unlabeled data of the target language containing pseudo labels are added to the training process, and the pre-trained model is trained by gradient descent to obtain a cross-language named entity recognition model.

4. According to the cross-language named entity recognition method based on large and small model collaborative reasoning according to claim 3, it is characterized in that: The step 3 of obtaining the inference result on the small model side includes: Step 3-1, input each target language text in the test set into the cross-language named entity recognition model; Step 3-2, the cross-language named entity recognition model outputs the named entities corresponding to different words in the target language text as the recognition reasoning result; Step 3-3, repeat steps 3-1 to 3-2, and summarize the recognition and reasoning results of each target language text in the test set to obtain the reasoning results on the small model side.

5. According to the cross-language named entity recognition method based on large and small model collaborative reasoning according to claim 4, it is characterized in that: The data in the test set described in step 4 is inferred using a multi-step reasoning method.

6. The cross-language named entity recognition method based on large and small model collaborative reasoning according to claim 5, characterized in that: The multi-step reasoning method described in step 4 includes: Step 4-1, provide the target language text in the test set to the large model and specify the named entity type TYPF; Step 4-2: The large model performs reasoning to obtain the overall interpretation of the input target language text Expl sent and word level interpretation Expl word , specifically including: By inputting prompt words into the big model, the big model gives the interpretation of the target language text and the interpretation of each word in the target language text, and obtains the overall interpretation of the target language text. sent and word level interpretation Expl word ; Step 4-3: Obtain the potential entities in the target language text through large model reasoning p , specifically including: For the specified named entity type TYPE, a prompt is constructed respectively, and the large model infers the corresponding type of entity contained in the target language text in the test set of the target language, that is, the potential entity Entity p ; Step 4-4: Get the potential entity through large model reasoning p The corresponding potential entity interpretation in the target language text Expl phras ; Step 4-5: Through large model reasoning, the overall interpretation of the target language text is synthesized. sent and entity interpretation Expl phrase , confirm the named entity type EntityType of the entity p .

7. The cross-language named entity recognition method based on large and small model collaborative reasoning according to claim 6, characterized in that: The named entity type EntityType of the confirmed entity described in step 4-5 p ,include: The potential entity Entity p The corresponding named entity type TYPE is designed as a multiple-choice question. In the prompt words of the large model, the target language text and the corresponding potential entity interpretation Expl phrase and potential entity p Each option in the multiple-choice question corresponds to a named entity type TYPE. The large model is used to select the correct named entity type and obtain the potential entity Entity. p The corresponding final named entity type EntityType p .

8. The cross-language named entity recognition method based on large and small model collaborative reasoning according to claim 7, characterized in that: The fusion of the small model side reasoning results and the large model side reasoning results described in step 5 adopts a reconfirmation mechanism.

9. The cross-language named entity recognition method based on large and small model collaborative reasoning according to claim 8, characterized in that: The reconfirmation mechanism described in step 5 includes: Step 5-1: Compare the inference results of the small model side with the inference results of the large model side and find the conflicting parts, including: Suppose the set of inference results on the large model side is P LLM , the set of inference results on the small model side is P SLM , then the conflicting part is the difference P between the above two sets LLM -P SLM and P SLM -P LLM ; The non-conflicting part is the common part P of the above two sets LLM ∩P SLM ; Step 5-2: For the conflicting parts, make judgments based on the big model, including: For each test set data containing conflicting results, a large model prompt word is constructed. The prompt word includes the target language text and the conflicting part. The large model is used to judge and output the correct result P judge ; Step 5-3, merge the non-conflicting parts P LLM ∩P SLM and the correct result P obtained in step 5-2 judge , that is, take the union of the above two and obtain the named entity recognition result P Union , as the final recognition result of cross-language named entities.

10. The cross-language named entity recognition method based on large and small model collaborative reasoning according to claim 1, characterized in that: The pre-trained model described in step 2 is the BERT model.