Entity processing method and device based on large language model, equipment and storage medium

By analyzing the cluster effect of the large language model and building correlation arrays and filtering areas, the accuracy problem of the large language model when recognizing entities in the vocabulary is solved, and the accuracy and universality of named entity recognition are improved.

CN120579019APending Publication Date: 2025-09-02BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510654261.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

When identifying entities in the vocabulary, the accuracy of large language models may be reduced compared to special discriminant probability models. The existing methods rely on complex instant designs, limiting compatibility and versatility.

Method used

By analyzing the main error types of named entity recognition in large language models, using cluster effects to design filters, construct correlation arrays and filter areas, blocking the wrong entities in type recognition, and improving the accuracy of entity recognition.

Benefits of technology

It effectively reduces the errors in entity type recognition of large language models, improves the accuracy and compatibility of named entity recognition, and enhances the adaptability of the model and the universality of rapid engineering methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579019A_ABST
    Figure CN120579019A_ABST
Patent Text Reader

Abstract

The invention provides an entity processing method and device based on a large language model, equipment and a storage medium, and relates to the technical fields of large language models, natural language processing, named entity recognition and the like. The method comprises the following steps: identifying an initial entity, an identification type to which the initial entity belongs and correlation between the initial entity and a target type from a corpus through a large language model; based on the marked real type, extracting an error entity with a type identification error from the initial entity; constructing a correlation array according to the identification type to which the error entity belongs and the correlation between the error entity and the target type, and determining a filtering area from the correlation array; the filtering area is used for shielding entities with wrong type identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to large language models, natural language processing, and other technical fields, and specifically to a method, apparatus, device, and storage medium for entity processing based on a large language model. Background Art

[0002] The Generative Large Language Model (GLLM) has a large number of parameters and training data. By utilizing more parameters and larger datasets for training, its performance and sample efficiency on various downstream tasks are effectively improved.

[0003] Large language models have the ability to recognize out-of-vocabulary (OOV) entities. However, despite their broad knowledge coverage, large language models may have a certain degree of accuracy decline when recognizing in-vocabulary entities compared to specially designed discriminative probability models. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, device, and storage medium for entity processing based on a large language model.

[0005] According to one aspect of the present disclosure, a method for entity processing based on a large language model is provided, comprising:

[0006] Identifying an initial entity, the identification type to which the initial entity belongs, and the relevance between the initial entity and the target type from a corpus using a large language model;

[0007] Extracting erroneous entities with type recognition errors from the initial entities based on the marked true types;

[0008] A correlation array is constructed according to the identification type to which the erroneous entity belongs and the correlation between the erroneous entity and the target type, and a filtering area is determined from the correlation array; the filtering area is used to shield entities with type identification errors.

[0009] According to another aspect of the present disclosure, a method for entity processing based on a large language model is provided, comprising:

[0010] Identifying a target entity, an identification type to which the target entity belongs, and a correlation between the target entity and the target type from a target text using a large language model;

[0011] Whether the target entity belongs to the target type is determined according to the identification type to which the target entity belongs, the correlation between the target entity and the target type, and a predetermined filtering area.

[0012] According to one aspect of the present disclosure, there is provided an entity processing apparatus based on a large language model, comprising:

[0013] An initial entity module, configured to identify an initial entity, the identification type to which the initial entity belongs, and the relevance between the initial entity and the target type from a corpus using a large language model;

[0014] An error entity module, configured to extract error entities with type recognition errors from the initial entities based on the annotated true types;

[0015] The filtering determination module is used to construct a correlation array based on the identification type to which the erroneous entity belongs and the correlation between the erroneous entity and the target type, and to determine a filtering area from the correlation array; the filtering area is used to shield entities with type identification errors.

[0016] According to another aspect of the present disclosure, there is provided an entity processing apparatus based on a large language model, comprising:

[0017] a target entity module, configured to identify a target entity, an identification type to which the target entity belongs, and a correlation between the target entity and the target type from a target text using a large language model;

[0018] The target type module is used to determine whether the target entity belongs to the target type according to the identification type to which the target entity belongs, the correlation between the target entity and the target type, and a predetermined filtering area.

[0019] According to another aspect of the present disclosure, an electronic device is provided, the electronic device comprising:

[0020] at least one processor; and

[0021] a memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by any embodiment of the present disclosure.

[0023] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method provided by any embodiment of the present disclosure.

[0024] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements the method provided according to any embodiment of the present disclosure.

[0025] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of an entity processing method based on a large language model provided according to an embodiment of the present disclosure;

[0027] Figure 2a is a flowchart of another entity processing method based on a large language model provided according to an embodiment of the present disclosure;

[0028] Figure 2b is a schematic diagram of a public correlation array provided according to an embodiment of the present disclosure;

[0029] Figure 2c is a schematic diagram of a public filtering area provided according to an embodiment of the present disclosure;

[0030] Figure 3a is a flowchart of another entity processing method based on a large language model provided according to an embodiment of the present disclosure;

[0031] Figure 3b is a schematic diagram of a single correlation array provided according to an embodiment of the present disclosure;

[0032] Figure 3c is a schematic diagram of a single filtering area provided according to an embodiment of the present disclosure;

[0033] Figure 4 is a flowchart of another entity processing method based on a large language model provided according to an embodiment of the present disclosure;

[0034] Figure 5 is a structural diagram of an entity processing device based on a large language model provided according to an embodiment of the present disclosure;

[0035] Figure 6 is a structural diagram of another entity processing device based on a large language model provided according to an embodiment of the present disclosure;

[0036] Figure 7 It is a block diagram of an electronic device used to implement the entity processing method based on a large language model according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0037] Named entity recognition (NER) is a core natural language processing task in the field of information extraction, in which specific words or phrases explicitly specified by humans, known as named entities, are associated with a semantic context. For example, in the sentence "Paris, a city famous for its art," "Paris" is identified as an entity of type "place." Accurate NER requires precise definition of entity boundaries and accurate identification of entity types, thereby transforming unstructured data into structured information to simplify and optimize information retrieval and analysis processes. NER plays a vital role in web search, relationship extraction, and dialogue agent systems.

[0038] Large language models represent a significant breakthrough in natural language processing. Trained using massive amounts of text data and advanced neural network architectures (such as transformers), they are capable of capturing the complex structure and deep semantics of language. These models not only possess powerful text generation capabilities, capable of generating coherent and natural sentences based on context, but also demonstrate a degree of versatility and adaptability, enabling their application to a wide range of natural language processing tasks, such as text classification, sentiment analysis, and question-answering systems. With the continuous advancement of technology, large language models are finding increasingly broad application in areas such as intelligent customer service, content creation, and information retrieval, becoming a vital force driving the advancement of artificial intelligence.

[0039] Current large language models have the ability to recognize out-of-vocabulary entities. However, despite their broad knowledge base, their accuracy may be lower when recognizing in-vocabulary entities compared to specialized discriminative probability models. This difference may be attributed to the inherent generative mechanisms and universal design features of large language models. To address the challenges in named entity recognition tasks, researchers have used massive training data to enhance the adaptability of large language models. The effectiveness of this approach mainly stems from the fine-tuning of model parameters. To reduce the dependence on large amounts of training resources, researchers have cleverly designed a hint mechanism to achieve an accurate named entity recognition result. However, these methods often rely on complex on-the-fly designs, which may limit the compatibility and universality between different fast engineering methods. Therefore, how to improve the accuracy of named entity recognition of large language models is an important issue in the industry.

[0040] Figure 1 This is a flowchart of an entity processing method based on a large language model according to an embodiment of the present disclosure. The method is applicable to the case where named entity recognition is performed using a large language model and entity type recognition errors are corrected. The method can be executed by an entity processing device based on a large language model, which can be implemented in software and / or hardware and can be integrated into an electronic device. Figure 1 As shown, the entity processing method based on the large language model of this embodiment may include:

[0041] S101, identifying an initial entity, an identification type to which the initial entity belongs, and a correlation between the initial entity and a target type from a corpus using a large language model;

[0042] S102, extracting erroneous entities with type recognition errors from the initial entities based on the marked true types;

[0043] S103, constructing a correlation array according to the identification type to which the erroneous entity belongs and the correlation between the erroneous entity and the target type, and determining a filtering area from the correlation array; the filtering area is used to shield entities with type identification errors.

[0044] To improve the accuracy of entity recognition using large language models, the applicant first conducted an in-depth study of the main types of errors in named entity recognition using large language models. To this end, the applicant analyzed the named entity recognition results of various large language models with a parameter size of approximately 8 billion (8B) or more on the named entity recognition dataset released by the Conference on Computational Natural Language Learning (CoNLL) in 2003. The results primarily involved four types of errors: type recognition errors, missing entities, right boundary errors (for example, an entity has three characters, but the character on the right boundary is not recognized), and left boundary errors. Type recognition errors were the most common type of error, accounting for over 70% of the total errors.

[0045] Moreover, the applicant has conducted an in-depth study on type recognition errors and found that type recognition errors have a clustering effect. Therefore, based on the clustering effect, a filter is determined to shield entities with type recognition errors. The applicant uses a variety of large language models to perform entity recognition and correlation self-assessment, and obtains the initial entity, the recognition type to which the initial entity belongs, and the correlation between the initial entity and the target type. Exemplarily, for each large language model, the target type entity is identified from the corpus by the large language model, and the correlation between the identified entity and the target type is self-assessed to obtain the correlation between the entity and the target type; the large language model is used to identify entities of other types (or non-target types) other than the target type from the corpus, and the correlation with the target type is evaluated; the target type can be a pre-specified user type, organization type, etc. Through the above research, it is found that the correlation distribution of type recognition errors of each large language model shows a clustering effect, with a tendency to cluster. The correlation between the entity and the target type shows a higher density in a certain area, while it shows a sparser distribution in other areas.

[0046] The applicants also divided the initial entities into correct and incorrect entities based on the annotated true type. If the true type of the initial entity is the target type, the initial entity is a correct entity with a correct type identification; otherwise, the initial entity is an incorrect entity with an incorrect type identification. Based on the correlation, a correct correlation array was constructed for the correct entities, and an incorrect correlation array was constructed for the incorrect entities. The authors found that the distribution of correct and incorrect correlations showed a consistent directional trend, with the center point of the correlation distribution of correct and incorrect entities having the same directionality as the source data.

[0047] Based on the clustering effect characteristics of type recognition errors, this application designs a filtering mechanism to distinguish correct entity recognition from type error recognition. This mechanism combines the correlation between the entity and the target type to determine the filter, which is used to filter out entities with type recognition errors, thereby improving the accuracy of entity recognition of large language models.

[0048] In the process of determining the filter of the target type for any large language model, the initial entity, the recognition type to which the initial entity belongs, and the correlation between the initial entity and the target type can be identified from the corpus through the large language model. The recognition type can be the target type or other types. In order to facilitate subsequent processing, the correlation value can be a natural number of 1-10.

[0049] If the true type of the initial entity is the target type, then the initial entity is a correct entity; otherwise, the initial entity is an incorrect entity. Combined with the recognition type to which the incorrect entity belongs and the correlation between the incorrect entity and the target type, a correlation array (i.e., incorrect correlation distribution) is constructed. Based on the clustering effect, the area where the incorrect entities are clustered in the correlation array is used as a filtering area. If the large language model recognizes any entity of the target type, and the correlation between the entity and the target type falls within the filtering area, then the entity is a type recognition error and does not belong to the target type.

[0050] The technical solution provided by the embodiments of the present disclosure is based on the research finding that entity type recognition errors of large language models present a clustering effect. The large language model is used to perform entity recognition on the corpus to obtain the initial entity, the recognition type to which the initial entity belongs, and the correlation between the initial entity and the target type. Based on the true type, the erroneous entity with type recognition error is selected from the initial entity; the correlation array is constructed by combining the recognition type and the correlation to which the erroneous entity belongs, and the area where the erroneous entities are clustered in the correlation array is used as a filtering area, which can reduce the entity type recognition errors of the large language model and thus improve the accuracy of entity recognition.

[0051] Figure 2a Flowchart of another entity processing method based on a large language model provided according to an embodiment of the present disclosure. Figure 2aBased on the above embodiment, the entity processing method based on the large language model of this embodiment may include:

[0052] S201, identifying an initial entity, an identification type to which the initial entity belongs, and a correlation between the initial entity and a target type from a corpus using a large language model;

[0053] S202, extracting erroneous entities with type recognition errors from the initial entities based on the marked true types;

[0054] S203, determining a target error entity according to the identification type to which the error entity belongs; the target error entity is a common error entity or a single error entity, wherein the common error entity belongs to both the target type and other types except the target type; the single error entity belongs only to the target type and does not belong to other types;

[0055] S204: Construct a correlation array according to the correlation between the target error entity and the target type, and determine a filtering area from the correlation array.

[0056] In an embodiment of the present disclosure, an initial entity of a target type is identified from a corpus through a large language model, and the correlation between the initial entity and the target type is self-evaluated to obtain a first correlation between the initial entity and the target type; an initial entity of other types is identified from a corpus through a large language model, and a second correlation between the initial entity of other types and the target type is evaluated.

[0057] If the identification type of any initial entity includes the target type, and the actual type of the initial entity is another type, then the initial entity is an error entity. The error entities are then divided based on their identification type to obtain target error entities. Where the identification type is the target type or another type, the target error entity is either a common error entity or a single error entity. For example, if the identification type of any error entity is both the target type and another type, then the error entity is a common error entity; if the identification type of any error entity is only the target type and not another type, then the error entity is a single error entity. A correlation array is constructed based on the correlation between the target error entity and the target type, and a filtering area is determined from the correlation array. By further dividing the error entities into common error entities and single error entities based on their identification type, and constructing a correlation array for the common error entities based on the correlation between the common error entities and the target type, and determining a filtering area for the common error entities from the correlation array, and constructing a correlation array for the single error entities based on the correlation between the single error entities and the target type, and determining a filtering area for the single error entities from the correlation array, filtering areas are determined separately for common error entities and single error entities, further improving the accuracy of the filtering areas.

[0058] In an optional embodiment, constructing a correlation array based on the correlation between the target error entity and the target type, and determining a filtering area from the correlation array, includes: when the target error entity is a common error entity, constructing a common correlation array based on the first correlation and the second correlation of the common error entity; wherein the first correlation is the correlation with the target type when the identification type is the target type; the second correlation is the correlation with the target type when the identification type is other types; and determining a common filtering area from the common correlation array.

[0059] For common error entities, a first correlation is obtained by self-evaluating the correlation between the common error entity and the target type when the common error entity is identified as belonging to the target type by the large language model; and a second correlation is obtained by self-evaluating the correlation between the common error entity and the target type when the common error entity is identified as belonging to other types by the large language model; a common correlation array can be constructed with the first correlation and the second correlation as row indexes or column indexes respectively. For ease of description, the following explanation will take the first correlation as the column index and the second correlation as the row index as an example; based on the clustering effect, the area where common error entities are clustered in the common correlation array is used as a common filtering area. By constructing a common correlation array based on the first correlation and the second correlation of the common error entity, and using the area where common error entities are clustered as the common filtering area, the common error entities are shielded.

[0060] In an optional implementation, the element in the i-th column and the j-th row in the common correlation array is used to represent the number of common error entities with a first correlation of i and a second correlation of j, where i and j are both positive integers.

[0061] refer to Figure 2b , using the first correlation as the column index and the second correlation as the row index, a two-dimensional public correlation array is constructed for the public error entity. The values ​​of the first correlation and the second correlation are both natural numbers from 1 to 10. The first correlation and the second correlation of each public error entity are counted to obtain the number of public error entities with first correlation i and second correlation j. This number is used as the element a in the i-th column and j-th row of the public correlation array. ji For example, 18 The value is 44, indicating that the number of common error entities having a first correlation of 8 and a second correlation of 1 is 44.

[0062] In an optional embodiment, determining the common filtering area from the common correlation array includes: using the element values ​​in the common correlation array as weights, and weightedly averaging the column position coordinates and the row position coordinates to obtain a common center point; and determining the common filtering area from the common correlation array based on the common center point.

[0063] For example, the following formula is used to weight the column position coordinates and row position coordinates, respectively, and to obtain the average column position coordinates and the average row position coordinates:

[0064]

[0065] Among them, a jiis the value of the element in the i-th column and j-th row, I is the average column position coordinate, and J is the average row position coordinate. The column position coordinate of the common center point is I, and the row position coordinate is J.

[0066] By using the element values ​​in the public correlation array as weights, the column position coordinates and row position coordinates are weighted and averaged to obtain the average column position coordinates and average row position coordinates of the public center point, which can accurately obtain the public center point; combined with the public correlation array, the area where public error entities gather is determined as the public filtering area, which further improves the accuracy of the public filtering area.

[0067] In an optional embodiment, a common filtering area is determined from the common correlation array based on the common center point, including: determining a common abnormal area between the origin of the common correlation array and the common center point; determining a common average distance between the elements in the common correlation array and the common center point, and determining a common correction distance based on a preset common correction coefficient and the common average distance; drawing a common correction area with the common center point as the center and the common correction distance as the radius; determining a common overlapping area between the common abnormal area and the common correction area, and eliminating the common overlapping area from the common abnormal area to obtain a common filtering area.

[0068] In the public correlation array, the origin is set to the point in the lower left corner, that is, the point with the smallest row and column coordinates. Figure 2c , the origin is the point where both the row and column coordinates are 1. Exemplarily, a rectangular common anomaly region 21 is drawn with the origin and common center of the common correlation array as the lower left vertex and the upper right vertex, respectively; the Euclidean distance between each element in the common correlation array and the common center point is calculated, and the average of each Euclidean distance is obtained to obtain the common average distance, and the common correction coefficient α is multiplied by the common average distance to obtain the common correction distance; a common correction region 22 is drawn with the common center point as the center and the common correction distance as the radius, and the common overlapping region 23 between the common anomaly region 21 and the common correction region 22 is determined, and the common overlapping region 23 is removed from the common anomaly region 21 to obtain the common filtering region. The common correction coefficient α is a hyperparameter with a value range of (0,1).

[0069] By combining the origin and common center of the common relevance array to draw a rectangular common anomaly region, the common anomaly region can clearly define the range where type errors are likely to occur, thereby improving the targeted filtering. Common overlapping regions may contain some common error entities with ambiguous types or unclear type identification. The large language model may introduce noise during the recognition of these entities, resulting in deviations in the large language model's judgment of entity types. By eliminating common overlapping regions, we can focus more on common error entities with clear characteristics, thereby improving the recognition accuracy of named entity types.

[0070] The technical solution provided by the embodiment of the present disclosure further subdivides the error entity into a common error entity and a single error entity according to the identification type to which the error entity belongs, and constructs a correlation array for the correlation between the two and the target type respectively, and then determines the filtering area for each of them from the array, thereby realizing the differentiated determination of the filtering area for different error types and effectively improving the accuracy of the filtering area; moreover, the element value in the common correlation array is used as the weight, and the column position coordinates and the row position coordinates are weighted and averaged, so as to accurately locate the common center point, and then the common error entity clustering area is determined from the common correlation array in combination with the common center point as the common process area. The filtering area further strengthens the precise coverage of the public filtering area on the erroneous entities and improves its accuracy. Furthermore, by combining the origin and the public center point of the public correlation array to draw a rectangular public abnormal area, the boundary range where type recognition is prone to errors can be clearly delineated, which significantly enhances the targeted filtering. In addition, since the public overlapping area often contains public erroneous entities with ambiguous types and unclear identification, the large language model is prone to introduce noise when processing these entities, which interferes with the accurate judgment of the entity type. After eliminating the public overlapping area, the model can focus on the public erroneous entities with clear features, thereby improving the recognition accuracy of the named entity type.

[0071] Figure 3a This is a flowchart of another entity processing method based on a large language model provided according to an embodiment of the present disclosure. Figure 3a Based on the above embodiment, the filter determination for a single erroneous entity is further limited. The entity processing method based on a large language model in this embodiment may include:

[0072] S301, identifying an initial entity, an identification type to which the initial entity belongs, and a correlation between the initial entity and a target type from a corpus using a large language model;

[0073] S302, extracting erroneous entities with type recognition errors from the initial entities based on the marked true types;

[0074] S303, determining a target error entity according to the identification type to which the error entity belongs; the target error entity is a common error entity or a single error entity, wherein the common error entity belongs to both the target type and other types except the target type; the single error entity belongs only to the target type and does not belong to other types;

[0075] S304: When the target error entity is a single error entity, construct a single correlation array based on a first correlation of the single error entity; wherein the first correlation is a correlation between the target type and the target type when the identification type is the target type;

[0076] S305: Determine a single filtering area from the single correlation array.

[0077] For a single error entity, a first correlation is obtained by self-evaluating the correlation between the single error entity and the target type when the large language model identifies the single error entity as belonging to the target type. A single correlation array can be constructed using the first correlation as a row index or a column index. For ease of description, the following explanation uses the first correlation as a column index. Based on the clustering effect, the area where the single error entity is clustered in the single correlation array is used as a single filtering area. By constructing a single correlation array based on the first correlation of the single error entity and using the area where the single error entity is clustered as a single filtering area, the single error entity is shielded.

[0078] In an optional implementation manner, the elements in the kth column of the single correlation array are used to represent the number of single error entities with a first correlation of k; wherein k is a positive integer.

[0079] refer to Figure 3b , using the first correlation as the column index, construct a one-dimensional single correlation array for the single error entity. The value of the first correlation is a natural number from 1 to 10. Count the first correlation of each single error entity and get the number of single error entities with the first correlation being k. This number is used as the element a in the kth column of the single correlation array. k For example, the value of a8 is 57, which means that there are 57 single error entities with a first correlation of 8.

[0080] In an optional embodiment, determining a single filtering area from the single correlation array includes: using the element values ​​in the single correlation array as weights, averaging the column position coordinates to obtain a single center point; and determining a single filtering area from the single correlation array based on the single center point.

[0081] For example, the following formula is used to weight the column position coordinates by taking the element values ​​in a single correlation array as weights and performing weighted averaging to obtain the average column position coordinates:

[0082]

[0083] Among them, a k is the element value of the kth column, K is the average column position coordinate. The column position coordinate of the single center point is K.

[0084] By using the element values ​​in a single correlation array as weights, the column position coordinates are weighted and averaged to obtain the average column position coordinates of a single center point, which can accurately obtain a single center point; combining the single center point to determine the area where a single error entity is concentrated from a single correlation array as a single filtering area, further improving the accuracy of the single filtering area.

[0085] In an optional embodiment, determining a single filtering area from the single correlation array based on the single center point includes: determining a single abnormal area between the origin of the single correlation array and the single center point; determining a single average distance between the elements in the single correlation array and the single center point, and determining a single correction distance based on a preset single correction coefficient and the single average distance; drawing a single correction area with the single center point as the center and the single correction distance as the radius; determining a single overlapping area between the single abnormal area and the single correction area, and eliminating the single overlapping area from the single abnormal area to obtain the single filtering area.

[0086] In a single correlation array, the origin is set to the lower left corner, which is the point with the smallest column position coordinate. Figure 3c , the origin is the point with column position coordinate 1. Exemplarily, a single abnormal region 31 is drawn with the origin and the single center point of the single correlation array as vertices; the absolute distance between each element in the single correlation array and the single center point is calculated, and the average of the absolute distances is obtained to obtain a single average distance, and the single average distance is multiplied by the preset single correction coefficient β to obtain a single corrected distance; a single corrected region 32 is drawn with the single center point as the center and the single corrected distance as the radius, a single overlapping region 33 between the single abnormal region 31 and the single corrected region 32 is determined, and the single overlapping region 33 is removed from the single abnormal region 31 to obtain a single filtered region. The single correction coefficient β is a hyperparameter with a value range of (0,1).

[0087] By combining the origin and center of a single correlation array to draw a single anomaly region, the single anomaly region can clearly define the range where type errors are prone to error, thereby improving the targeted filtering. Single overlapping regions may contain single erroneous entities with ambiguous or unclear types. The large language model may introduce noise during the recognition of these entities, causing the large language model to deviate from the entity type. By eliminating single overlapping regions, we can focus more on single erroneous entities with clear characteristics, thereby improving the recognition accuracy of named entity types.

[0088] The technical solution provided by the embodiment of the present disclosure, for a single erroneous entity, by taking the element value in a single relevance array as the weight, weightedly averaging the column position coordinates, thereby accurately locating a single center point, and then determining the single erroneous entity aggregation area from the single relevance array in combination with the single center point as a single filtering area, further enhancing the precise coverage of the erroneous entity by the single filtering area and improving its accuracy; further, by drawing a single abnormal area in combination with the origin of the single relevance array and the single center point, the boundary range where type recognition is prone to errors can be clearly delineated, significantly enhancing the targeted filtering; in addition, since a single overlapping area often contains a single erroneous entity with ambiguous type and unclear identification, a large language model is prone to introduce noise when processing these entities, interfering with the accurate judgment of the entity type. After eliminating the single overlapping area, the model can focus on the single erroneous entity with clear features, thereby improving the recognition accuracy of the named entity type.

[0089] Figure 4 This is a flowchart of an entity processing method based on a large language model according to an embodiment of the present disclosure. The method is applicable to the case where named entity recognition is performed using a large language model and entity type recognition errors are corrected. The method can be executed by an entity processing device based on a large language model, which can be implemented in software and / or hardware and can be integrated into an electronic device. Figure 4 As shown, the entity processing method based on the large language model of this embodiment may include:

[0090] S401, identifying a target entity, an identification type to which the target entity belongs, and a correlation between the target entity and the target type from a target text using a large language model;

[0091] S402 : Determine whether the target entity belongs to the target type according to the identification type to which the target entity belongs, the correlation between the target entity and the target type, and a predetermined filtering area.

[0092] In the disclosed embodiment, a target entity is identified from a target text using a large language model, and the target entity's recognition type and the correlation between the target entity and the target type are obtained. If the recognition type is the target type, a first correlation between the target entity and the target type is obtained; if the recognition type is a type other than the target type, a second correlation between the target entity and the target type is obtained. The target text is the text for which named entity recognition is to be performed.

[0093] Based on the target entity's recognition type and the correlation between the target entity and the target type, the algorithm determines whether the corresponding correlation falls within a predetermined filtering region. If so, the target entity does not belong to the target type. By filtering out entities that do not belong to the target type based on the correlation between the target entity and the target type based on the predetermined filtering region, the accuracy of entity recognition using large language models is improved.

[0094] The technical solution provided by the embodiments of the present disclosure uses a large language model to perform named entity recognition on a target text to obtain a target entity, the identification type to which the target entity belongs, and the correlation between the target entity and the target type. Based on a predetermined filtering region, the technical solution determines whether the correlation corresponding to the target entity falls within the predetermined filtering region. Once it is determined that the correlation of the target entity falls within the filtering region, it is determined that the target entity does not belong to the target type, and entities that do not belong to the target type are filtered out, effectively improving the accuracy of the large language model in entity recognition tasks.

[0095] In an optional embodiment, the determining whether the target entity belongs to the target type based on the identification type to which the target entity belongs, the correlation between the target entity and the target type, and a predetermined filtering area includes: determining the entity to be identified based on the identification type to which the target entity belongs; the entity to be identified is a public entity or a single entity, and the public entity belongs to both the target type and other types; the single entity belongs only to the target type and does not belong to the other types; determining whether the target entity belongs to the target type based on the correlation between the entity to be identified and the target type, and a predetermined filtering area.

[0096] A target entity of a target type is identified from a target text through a large language model, and the correlation between the target entity and the target type is self-evaluated to obtain a first correlation between the target entity and the target type; other types of target entities are identified from the target text through a large language model, and a second correlation between the other types of target entities and the target type is evaluated.

[0097] If the identification type of any target entity is both the target type and other types, then the target entity is a public entity, and a determination is made as to whether the public entity belongs to a predetermined public filtering area. Once it is determined that the relevance of the public entity belongs to the public filtering area, it is determined that the public entity does not belong to the target type, and entities that do not belong to the target type are filtered out. If the identification type of any target entity is only the target type and not other types, then the target entity is a single entity, and a determination is made as to whether the single entity belongs to a predetermined single filtering area. Once it is determined that the relevance of the single entity belongs to the single filtering area, it is determined that the single entity does not belong to the target type, and entities that do not belong to the target type are filtered out.

[0098] By further dividing the target entities into public entities and single entities according to the identification type to which they belong, and judging whether the public entities belong to the target type based on a predetermined public filtering area, and judging whether the single entities belong to the target type based on a predetermined single filtering area, the accuracy of entity recognition for the target type is further improved.

[0099] In an optional embodiment, determining whether the target entity belongs to the target type based on the correlation between the entity to be identified and the target type, and a predetermined filtering area, includes: when the entity to be identified is a public entity, determining whether the entity to be identified belongs to the target type based on the first correlation and the second correlation of the entity to be identified, and a predetermined public filtering area; when the entity to be identified is a single entity, determining whether the entity to be identified belongs to the target type based on the first correlation of the entity to be identified, and a predetermined single filtering area.

[0100] For public entities that belong to both the target type and other types, whether the public entity belongs to a predetermined public filtering area is determined based on the first relevance and second relevance of the public entity; for single entities that only belong to the target type and not other types, whether the single entity belongs to a predetermined single filtering area is determined based on the first relevance of the single entity.

[0101] Figure 5 This is a schematic diagram of the structure of an entity processing device based on a large language model according to an embodiment of the present disclosure. The device is suitable for performing named entity recognition through a large language model and correcting entity type recognition errors. The device can be implemented in software and / or hardware and can be integrated into electronic devices. Figure 5 As shown, the entity processing device 500 based on the large language model of this embodiment may include:

[0102] An initial entity module 510 is configured to identify an initial entity, the identification type to which the initial entity belongs, and the relevance between the initial entity and the target type from a corpus using a large language model;

[0103] An error entity module 520 is configured to extract error entities with type recognition errors from the initial entities based on the marked true types;

[0104] The filtering determination module 530 is used to construct a correlation array based on the identification type to which the erroneous entity belongs and the correlation between the erroneous entity and the target type, and to determine a filtering area from the correlation array; the filtering area is used to shield entities with type identification errors.

[0105] In an optional implementation, the filtering determination module 530 includes:

[0106] A target error entity submodule is configured to determine a target error entity according to the identification type to which the error entity belongs; the target error entity is a common error entity or a single error entity, wherein the common error entity belongs to both the target type and other types except the target type; and the single error entity belongs only to the target type and does not belong to other types;

[0107] The filtering area submodule is configured to construct a correlation array according to the correlation between the target error entity and the target type, and determine a filtering area from the correlation array.

[0108] In an optional embodiment, the filtering area submodule includes:

[0109] a public correlation array unit, configured to construct a public correlation array based on a first correlation and a second correlation of the public error entity when the target error entity is a public error entity; wherein the first correlation is the correlation with the target type when the identification type is the target type; and the second correlation is the correlation with the target type when the identification type is another type;

[0110] The common filtering unit is configured to determine a common filtering area from the common relevance array.

[0111] In an optional implementation, the element in the i-th column and the j-th row in the common correlation array is used to represent the number of common error entities with a first correlation of i and a second correlation of j, where i and j are both positive integers.

[0112] In an optional embodiment, the common filtering unit includes:

[0113] A public center subunit is used to take the element values ​​in the public correlation array as weights and perform weighted averaging on the column position coordinates and the row position coordinates to obtain a public center point;

[0114] The public filtering subunit is configured to determine a public filtering area from the public correlation array according to the public center point.

[0115] In an optional implementation manner, the public filtering subunit is specifically configured to:

[0116] determining a common anomaly region between an origin of the common correlation array and the common center point;

[0117] Determining a common average distance between elements in the common correlation array and a common center point, and determining a common correction distance based on a preset common correction coefficient and the common average distance;

[0118] Draw a common correction area with the common center point as the center and the common correction distance as the radius;

[0119] A common overlapping area between the common abnormal area and the common correction area is determined, and the common overlapping area is removed from the common abnormal area to obtain a common filtering area.

[0120] In an optional embodiment, the filtering area submodule includes:

[0121] a single array unit, configured to construct a single correlation array based on a first correlation of the single error entity when the target error entity is a single error entity; wherein the first correlation is a correlation between the identified type and the target type when the identified type is the target type;

[0122] A single filtering unit is used to determine a single filtering area from the single correlation array.

[0123] In an optional implementation manner, the elements in the kth column of the single correlation array are used to represent the number of single error entities with a first correlation of k; wherein k is a positive integer.

[0124] In an optional embodiment, the single filtration unit comprises:

[0125] A single center subunit, configured to average the column position coordinates using the element values ​​in the single correlation array as weights to obtain a single center point;

[0126] A single filtering subunit is configured to determine a single filtering area from the single correlation array according to the single center point.

[0127] In an optional embodiment, the single filtering subunit is specifically configured to:

[0128] determining a single anomalous region between an origin and a single center point of the single correlation array;

[0129] Determining a single average distance between elements in the single correlation array and a single center point, and determining a single corrected distance based on a preset single correction coefficient and the single average distance;

[0130] Draw a single correction area with the single center point as the center and the single correction distance as the radius;

[0131] A single overlapping area between the single abnormal area and the single corrected area is determined, and the single overlapping area is removed from the single abnormal area to obtain the single filtered area.

[0132] a target entity module, configured to identify a target entity, an identification type to which the target entity belongs, and a correlation between the target entity and the target type from a target text using a large language model;

[0133] The target type module is used to determine whether the target entity belongs to the target type according to the identification type to which the target entity belongs, the correlation between the target entity and the target type, and a predetermined filtering area.

[0134] The technical solution provided by the embodiments of the present disclosure creatively discovers that entity type recognition errors have a clustering effect in the large language model entity recognition scenario. Based on this, a correlation array is constructed and the named entity recognition results of the large language model are filtered according to the filter of the error clustering in the correlation array, thereby improving the accuracy of named entity recognition and reducing entity type recognition errors.

[0135] Figure 6 This is a schematic diagram of the structure of another entity processing device based on a large language model according to an embodiment of the present disclosure. The device is suitable for performing named entity recognition through a large language model and correcting entity type recognition errors. The device can be implemented in software and / or hardware and can be integrated into electronic devices. Figure 6 As shown, the entity processing device 600 based on the large language model of this embodiment may include:

[0136] A target entity module 610 is configured to identify a target entity, a recognition type to which the target entity belongs, and a correlation between the target entity and the target type from a target text using a large language model;

[0137] The target type module 620 is configured to determine whether the target entity belongs to the target type according to the identification type to which the target entity belongs, the correlation between the target entity and the target type, and a predetermined filtering area.

[0138] In an optional embodiment, the target type module 620 includes:

[0139] a to-be-identified submodule, configured to determine an entity to be identified based on the identification type to which the target entity belongs; the entity to be identified is a public entity or a single entity, the public entity belonging to both the target type and other types; the single entity belonging only to the target type and not to other types;

[0140] The target type submodule is configured to determine whether the target entity belongs to the target type according to the correlation between the entity to be identified and the target type and a predetermined filtering area.

[0141] In an optional implementation, the target type submodule includes:

[0142] a public type unit, configured to determine, if the entity to be identified is a public entity, whether the entity to be identified belongs to the target type based on the first relevance and the second relevance of the entity to be identified and a predetermined public filtering area;

[0143] The single type unit is configured to determine whether the entity to be identified belongs to the target type according to the first correlation of the entity to be identified and a predetermined single filtering area when the entity to be identified is a single entity.

[0144] The technical solution provided by the embodiments of the present disclosure creatively discovers that entity type recognition errors have a clustering effect in the large language model entity recognition scenario. Based on this, a correlation array is constructed and the named entity recognition results of the large language model are filtered according to the filter of the error clustering in the correlation array, thereby improving the accuracy of named entity recognition and reducing entity type recognition errors.

[0145] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0146] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0147] Figure 7 It is a block diagram of an electronic device used to implement the entity processing method based on a large language model according to an embodiment of the present disclosure. Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0148] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the electronic device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0149] Multiple components in the electronic device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0150] The computing unit 701 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a material delivery unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the entity processing method based on the large language model. For example, in some embodiments, the entity processing method based on the large language model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the entity processing method based on the large language model described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the entity processing method based on the large language model in any other appropriate manner (for example, by means of firmware).

[0151] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0152] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the electronic device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0153] Multiple components in the electronic device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0154] The computing unit 701 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a material delivery unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the entity processing method based on the large language model. For example, in some embodiments, the entity processing method based on the large language model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the entity processing method based on the large language model described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the entity processing method based on the large language model in any other appropriate manner (for example, by means of firmware).

[0155] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0156] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0157] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0158] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, audio input, or tactile input).

[0159] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web player through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0160] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0161] Artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, audio recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0162] Cloud computing refers to a technology system that provides network access to elastically scalable shared pools of physical or virtual resources. These resources can include servers, operating systems, networks, software, applications, and storage devices, and can be deployed and managed on-demand in a self-service manner. Cloud computing technology provides efficient and powerful data processing capabilities for the application of technologies such as artificial intelligence and blockchain, as well as for model training.

[0163] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0164] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for entity processing based on a large language model, comprising: Identifying an initial entity, the identification type to which the initial entity belongs, and the relevance between the initial entity and the target type from a corpus using a large language model; Extracting erroneous entities with type recognition errors from the initial entities based on the marked true types; A correlation array is constructed according to the identification type to which the erroneous entity belongs and the correlation between the erroneous entity and the target type, and a filtering area is determined from the correlation array; the filtering area is used to shield entities with type identification errors.

2. The method according to claim 1, wherein The step of constructing a correlation array according to the identification type to which the error entity belongs and the correlation between the error entity and the target type, and determining a filtering area from the correlation array includes: Determine a target error entity according to the identification type to which the error entity belongs; the target error entity is a common error entity or a single error entity, the common error entity belongs to both the target type and other types except the target type; the single error entity belongs only to the target type and does not belong to other types; A correlation array is constructed according to the correlation between the target error entity and the target type, and a filtering area is determined from the correlation array.

3. The method according to claim 2, wherein: The step of constructing a correlation array according to the correlation between the target error entity and the target type, and determining a filtering area from the correlation array, comprises: When the target error entity is a common error entity, a common correlation array is constructed according to the first correlation and the second correlation of the common error entity; wherein the first correlation is the correlation with the target type when the identification type is the target type; and the second correlation is the correlation with the target type when the identification type is another type. A common filtering region is determined from the common relevance array.

4. The method according to claim 3, wherein: The element in the i-th column and the j-th row in the common correlation array is used to represent the number of common error entities with a first correlation of i and a second correlation of j, where i and j are both positive integers.

5. The method according to claim 3 or 4, wherein: The determining of a common filtering area from the common relevance array comprises: Taking the element values ​​in the common correlation array as weights, weightedly averaging the column position coordinates and the row position coordinates to obtain a common center point; A common filtering area is determined from the common correlation array according to the common center point.

6. The method according to claim 5, wherein: Determining a common filtering area from the common correlation array according to the common center point includes: determining a common anomaly region between an origin of the common correlation array and the common center point; Determining a common average distance between elements in the common correlation array and a common center point, and determining a common correction distance based on a preset common correction coefficient and the common average distance; Draw a common correction area with the common center point as the center and the common correction distance as the radius; A common overlapping area between the common abnormal area and the common correction area is determined, and the common overlapping area is removed from the common abnormal area to obtain a common filtering area.

7. The method according to claim 2, wherein: The step of constructing a correlation array according to the correlation between the target error entity and the target type, and determining a filtering area from the correlation array, comprises: When the target error entity is a single error entity, constructing a single correlation array according to the first correlation of the single error entity; wherein the first correlation is the correlation between the target type and the target type when the identification type is the target type; A single filter region is determined from the single relevance array.

8. The method according to claim 7, wherein: The elements in the k-th column of the single correlation array are used to represent the number of single error entities with a first correlation of k; wherein k is a positive integer.

9. The method according to claim 7 or 8, wherein The determining of a single filtering area from the single correlation array comprises: Taking the element values ​​in the single correlation array as weights, averaging the column position coordinates to obtain a single center point; A single filtering region is determined from the single correlation array based on the single center point.

10. The method according to claim 9, wherein: Determining a single filtering area from the single correlation array based on the single center point includes: determining a single anomalous region between an origin and a single center point of the single correlation array; Determining a single average distance between elements in the single correlation array and a single center point, and determining a single corrected distance based on a preset single correction coefficient and the single average distance; Draw a single correction area with the single center point as the center and the single correction distance as the radius; A single overlapping area between the single abnormal area and the single corrected area is determined, and the single overlapping area is removed from the single abnormal area to obtain the single filtered area.

11. A method for entity processing based on a large language model, comprising: Identifying a target entity, an identification type to which the target entity belongs, and a correlation between the target entity and the target type from a target text using a large language model; Whether the target entity belongs to the target type is determined according to the identification type to which the target entity belongs, the correlation between the target entity and the target type, and a predetermined filtering area.

12. The method according to claim 11, wherein The determining whether the target entity belongs to the target type according to the identification type to which the target entity belongs, the correlation between the target entity and the target type, and a predetermined filtering area includes: Determine an entity to be identified according to the identification type to which the target entity belongs; the entity to be identified is a public entity or a single entity, the public entity belongs to both the target type and other types; the single entity belongs only to the target type and does not belong to other types; Whether the target entity belongs to the target type is determined according to the correlation between the entity to be identified and the target type, and a predetermined filtering area.

13. The method according to claim 12, wherein: The determining whether the target entity belongs to the target type according to the correlation between the entity to be identified and the target type and a predetermined filtering area includes: In a case where the entity to be identified is a public entity, determining whether the entity to be identified belongs to the target type according to the first relevance and the second relevance of the entity to be identified and a predetermined public filtering area; In the case that the entity to be identified is a single entity, whether the entity to be identified belongs to the target type is determined according to the first correlation of the entity to be identified and a predetermined single filtering area.

14. An entity processing device based on a large language model, comprising: An initial entity module, configured to identify an initial entity, the identification type to which the initial entity belongs, and the relevance between the initial entity and the target type from a corpus using a large language model; An error entity module, configured to extract error entities with type recognition errors from the initial entities based on the annotated true types; The filtering determination module is used to construct a correlation array based on the identification type to which the erroneous entity belongs and the correlation between the erroneous entity and the target type, and to determine a filtering area from the correlation array; the filtering area is used to shield entities with type identification errors.

15. An entity processing device based on a large language model, comprising: a target entity module, configured to identify a target entity, an identification type to which the target entity belongs, and a correlation between the target entity and the target type from a target text using a large language model; The target type module is used to determine whether the target entity belongs to the target type according to the identification type to which the target entity belongs, the correlation between the target entity and the target type, and a predetermined filtering area.

16. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause a computer to execute the method according to any one of claims 1-13.

18. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 13.