Named entity recognition method, device and equipment based on loop iteration

By preprocessing text data in geological domain and cyclic iteration of multi-level recognition models, the shortcomings of named entity recognition methods in fine-grainedness are solved, and a higher precision geographic entity recognition is achieved.

CN120257986APending Publication Date: 2025-07-04CHINA UNIV OF GEOSCIENCES (WUHAN) +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510169612.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing naming entity recognition methods are difficult to meet the needs of fine-grained classification in the geological field, resulting in insufficient recognition accuracy.

Method used

The named entity recognition method based on loop iteration is adopted to pre-process the data to be processed, including text cleaning, truncation and deletion, part-of-speech restoration and removal of stop words, and the first entity label is determined using the first-level recognition model, and the entity label is refined step by step through the multi-level secondary recognition model to finally determine the named entity and the target entity label.

Benefits of technology

The accuracy and recall of geographic entity recognition were improved, and the F-score value of the model increased by 13.47%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257986A_ABST
    Figure CN120257986A_ABST
Patent Text Reader

Abstract

The invention provides a named entity recognition method, device and equipment based on loop iteration, and relates to the technical field of computers.The named entity recognition method comprises the steps that to-be-processed data is preprocessed, and a first processed text is obtained; the preprocessing comprises text cleaning, truncation and deletion, part-of-speech reduction, unified coding, data resampling and stop word removal; inputting the first processing text into a primary recognition model, and determining a first entity label to which a to-be-recognized entity belongs; performing secondary identification on the to-be-identified entity based on the first entity tag and a secondary identification model, and determining a named entity and a target entity tag corresponding to the to-be-identified entity; the secondary recognition model comprises at least two stages of recognition models, and the granularity of the at least two stages of recognition models is smaller than that of the primary recognition model. According to the method, entity labels of multiple levels are obtained step by step by utilizing a loop iteration method, and the accuracy of geographic entity recognition is improved from a multi-dimensional angle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method, apparatus, and device for named entity recognition based on cyclic iteration. Background Art

[0002] The main methods of named entity recognition are divided into three categories: rule- or dictionary-based methods, statistical-based methods, and machine learning-based methods. Machine learning-based methods have dominated the task of named entity recognition in recent years. The current main technologies in named entity recognition include: marking sentences in a geological corpus as instances of entity existence and relationships, and proposing a method based on a conditional random field model to extract geological entity relationships; constructing a knowledge discovery model based on literature including geological entity recognition and entity relationship recognition; extracting geological domain terms from text, comparing the extraction accuracies with different values of word lengths, and obtaining the word length with the best accuracy; designing a geological entity recognition model based on a deep belief network (DBN) on the basis of the annotation specification and corpus of geological entity information; segmenting geological texts based on an ontology model in the geological domain, and using a Bi-LSTM-CRF neural network model to carry out named entity recognition tasks in the geological domain; training a geological domain word vector model using a large amount of unlabeled data, and adding an Attention mechanism to the Bi-LSTM-CRF model to learn the global information of the text, so as to force the marking of multiple entities with the same label in a document as the same entity type; improving the word vector model by fusing the character features obtained by using CNN (Convolutional Neural Networks) and the dynamic features obtained by ELMO (Embeddings from Language Models) for the word vectors, and enhancing the performance of the LSTM entity recognition model. With the continuous improvement of the informatization level of geological data and the development of computer technology, software and hardware, the text information mined from the relatively coarse-grained named entity recognition tasks carried out in the geological field is relatively rough, which is of limited help to subsequent work such as knowledge graph construction and knowledge reasoning, and the recognized fine-grainedness can no longer meet the current classification requirements. Summary of the Invention

[0003] The main purpose of the present invention is to provide a method, apparatus, and device for named entity recognition based on cyclic iteration to solve the problem that the current fine-grainedness of named entity recognition does not meet the current classification requirements.

[0004] The technical solution of the embodiment of the present application is implemented as follows: The first aspect of the embodiment of the present application provides a method for named entity recognition based on cyclic iteration, including: Preprocess the data to be processed to obtain the first processed text; the preprocessing includes text cleaning, truncation and deletion, part-of-speech restoration, unified encoding, data resampling, and stop word removal; Input the first processed text into a primary recognition model to determine the first entity label to which the entity to be recognized belongs; Based on the first entity label and a secondary recognition model, perform secondary recognition on the entity to be recognized to determine the named entity and target entity label corresponding to the entity to be recognized; the secondary recognition model includes at least two levels of recognition models, and the granularity of the at least two levels of recognition models is smaller than that of the primary recognition model.

[0005] Optionally, the preprocessing the data to be processed to obtain the first processed text includes: Perform text cleaning on the data to be processed to filter out irrelevant information; Select the maximum sliding window to unify the length of the data to be processed, and truncate or delete the data to be processed whose length exceeds the maximum sliding window; Obtain the part of speech of English words in the data to be processed through a predefined function, and then restore the English words to their original forms according to the part of speech; Unify the character encoding of the data to be processed; Oversample the minority classes and undersample the majority classes in the data to be processed.

[0006] Optionally, the inputting the first processed text into a primary recognition model to determine the first entity label to which the entity to be recognized belongs includes: Input the first processed text into a primary recognition model to determine the material composition, spatio-temporal description, and physical and chemical conditions corresponding to the entity to be recognized.

[0007] Optionally, the secondary recognition model includes a secondary model, and the performing secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized includes: Perform secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the second entity category corresponding to the entity to be recognized; the second entity category includes minerals, rocks, chemical elements, time, location, geological structures, geological units, fluids, and geological anomalies.

[0008] Optionally, the secondary recognition model includes a tertiary model, and the performing secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized includes: Performing secondary identification on the entity to be identified based on the second entity category and the three-level recognition model, and determining the named entity and the target entity label corresponding to the entity to be identified.

[0009] Optionally, the performing secondary identification on the entity to be identified based on the second entity category and the three-level recognition model, and determining the named entity and the target entity label corresponding to the entity to be identified includes: Performing secondary identification on the entity to be identified based on the second entity category and the three-level recognition model, and determining the mineral type, element type, and time type corresponding to the entity to be identified; the mineral type includes metallic minerals and non-metallic minerals, the element type includes isotope elements and non-isotope elements, and the time type includes geological era and isotope age; Determining the named entity and the target entity label corresponding to the entity to be identified based on the mineral type, element type, and time type.

[0010] Optionally, it further includes: training the original recognition model using historical data to obtain a target recognition model; the target recognition model includes a primary recognition model and a secondary recognition model, and each level of recognition model includes multiple types of sub-models.

[0011] The second aspect of the embodiments of the present application provides a named entity recognition device based on cyclic iteration, including: a preprocessing module, a first determination module, and a second determination module, where The preprocessing module is configured to preprocess the data to be processed to obtain a first processed text; the preprocessing includes text cleaning, truncation and deletion, part-of-speech restoration, unified encoding, data resampling, and stop word removal; The first determination module is configured to input the first processed text into the primary recognition model to determine the first entity label to which the entity to be identified belongs; The second determination module is configured to perform secondary identification on the entity to be identified based on the first entity label and the secondary recognition model, and determine the named entity and the target entity label corresponding to the entity to be identified; the secondary recognition model includes at least two levels of recognition models, and the granularity of the at least two levels of recognition models is smaller than that of the primary recognition model.

[0012] The third aspect of the embodiments of the present application provides an electronic device, including a processor and a memory; the memory stores a computer program, where the computer program, when executed by the processor, implements the named entity recognition method based on cyclic iteration described in the first aspect.

[0013] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0014] Compared with the prior art, the beneficial effects brought by the technical solution provided by this application are as follows: The present invention provides a named entity recognition method, device and equipment based on cyclic iteration. By preprocessing the data to be processed, a first processed text is obtained. After text cleaning, truncation and deletion, part-of-speech reduction and stop word removal of the first processed text, the first processed text is input into a primary recognition model to determine the first entity label to which the entity to be recognized belongs. Furthermore, based on the first entity label and a secondary recognition model, secondary recognition is performed on the entity to be recognized to determine the named entity and the target entity label corresponding to the entity to be recognized. By using the cyclic iteration method to gradually obtain entity labels at multiple levels, the accuracy of geographical entity recognition is improved from multiple dimensions. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic flowchart of a named entity recognition method based on cyclic iteration provided by an embodiment of this application; Figure 2 It is a schematic diagram of the data processing process provided by an embodiment of this application; Figure 3 It is a flowchart of named entity recognition with different granularities provided by an embodiment of the present invention; Figure 4 It is a schematic structural diagram of a named entity recognition device based on cyclic iteration provided by an embodiment of this application; Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Hereinafter, embodiments of this application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of this application. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of this application. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of this application.

[0017] The terms used herein are merely for describing specific embodiments and are not intended to limit this application. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components. All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0018] Some block diagrams and / or flowcharts are shown in the accompanying drawings. It should be understood that some of the blocks or combinations thereof in the block diagrams and / or flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when executed by the processor, these instructions can create a device for implementing the functions / operations illustrated in these block diagrams and / or flowcharts.

[0019] In some embodiments, refer to Figure 1 , Figure 1 is a schematic flowchart of the named entity recognition method based on cyclic iteration provided by the embodiments of the present application; the named entity recognition method based on cyclic iteration provided by the embodiments of the present application includes: S110, preprocess the data to be processed to obtain a first processed text; the preprocessing includes text cleaning, truncating and deleting, part-of-speech restoration, unified encoding, data resampling, and removing stop words.

[0020] In some embodiments, preprocessing the data to be processed to obtain a first processed text includes: Perform text cleaning on the data to be processed to filter out irrelevant information; Select the maximum sliding window to unify the length of the data to be processed, and truncate or delete the data to be processed whose length exceeds the maximum sliding window; Obtain the part-of-speech of English words in the data to be processed through a predefined function, and then restore the English words to their original forms according to the part-of-speech; Unify the character encoding of the data to be processed; Oversample the minority classes in the data to be processed and undersample the majority classes.

[0021] In one example, refer to Figure 2 , Figure 2Schematic diagram of the data processing process provided by the embodiments of the present application; here, the reference data in the USGS global porphyry copper deposit database was sorted out, with a total of 419 references, and the publication years of the papers in the references were counted. Based on the publication time of the papers, 87 representative journal papers were selected from the 419 references as the original text data of the training dataset according to expert knowledge. Through text cleaning, the reference, acknowledgement, and corresponding chart parts were deleted, and the original text data was constructed based on this. Finally, the text content of the selected literature was obtained, with 492,952 bytes of text data. The Word2vec function module was called to construct a word vector model for the text data of porphyry copper deposits. The system clustering method was used to perform entity word clustering analysis on the correlation of geological entities in the writing structure. A total of 64 various geological entity words were selected to form a dataset, and the dataset was systematically clustered. Then, expert knowledge was used to adjust the results of the clustering analysis to complete the construction of the label system. On the basis of the original text (unstructured data), data annotation work was carried out, and the Brat tool was used for entity annotation to make the final annotation result into structured text data. On this basis, data preprocessing work was carried out, and the cleaned data was indexed based on the label system to complete the construction of the training dataset. First, the word form reduction work was carried out, that is, the part-of-speech of English words in the text statement was obtained through a predefined function, and then the words were restored to their original forms according to the part-of-speech. Common words or terms that have no decisive effect on the context semantics in the text and have no discrimination were removed. The sentence length was corrected, and the sentence length distribution in the training data was counted. The sentence lengths in the training data were mainly distributed in interval. When performing subsequent neural network training, the maximum sliding window was selected as 40, and sentences in the training database with a length exceeding 40 words were truncated or deleted through manual intervention. Data resampling was carried out, and resampling work was carried out on the labels with a small amount of data, that is, sentences with a small number of entities were repeatedly sampled to ensure that the number of each entity in the training data was relatively balanced. After data preprocessing work such as word form reduction and stop word removal on the text, 388,310 bytes of text data were finally obtained, and the construction of the training dataset was completed. The above is the data processing process for obtaining model training. Here, the corresponding means can be used to preprocess the data to be processed to obtain the first processed text.

[0022] S120, input the first processed text into a first-level recognition model to determine the first entity label to which the entity to be recognized belongs.

[0023] It should be noted that the recognition model includes multiple levels of models. After the first processed text passes through the first-level recognition model, its approximate label range can be determined, and then through other levels of models, the label range is gradually narrowed until the target entity label is determined.

[0024] In some embodiments, in S120, inputting the first processed text into a primary recognition model to determine the first entity label to which the entity to be recognized belongs, includes: Inputting the first processed text into a primary recognition model to determine the material composition, spatio-temporal description, and physical and chemical conditions corresponding to the entity to be recognized.

[0025] S130, performing secondary recognition on the entity to be recognized based on the first entity label and a secondary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized; the secondary recognition model includes at least two levels of recognition models, and the granularity of the at least two levels of recognition models is smaller than that of the primary recognition model.

[0026] In some embodiments, the secondary recognition model includes a secondary model. In S130, performing secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized, includes: Performing secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the second entity category corresponding to the entity to be recognized; the second entity category includes minerals, rocks, chemical elements, time, location, geological structures, geological units, fluids, and geological anomalies.

[0027] In some embodiments, the secondary recognition model includes a tertiary model. In S130, performing secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized, includes: Performing secondary recognition on the entity to be recognized based on the second entity category and the tertiary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized.

[0028] In some embodiments, in S130, performing secondary recognition on the entity to be recognized based on the second entity category and the tertiary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized, includes: Performing secondary recognition on the entity to be recognized based on the second entity category and the tertiary recognition model to determine the mineral type, element type, and time type corresponding to the entity to be recognized; the mineral type includes metallic minerals and non-metallic minerals, the element type includes isotope elements and non-isotope elements, and the time type includes geological epochs and isotope ages; Determining the named entity and target entity label corresponding to the entity to be recognized based on the mineral type, element type, and time type.

[0029] Exemplarily, please refer to Figure 3 , Figure 3It is a flowchart of named entity recognition with different granularities provided by an embodiment of the present invention. When performing entity recognition, first, it is necessary to perform text preprocessing work such as part-of-speech reduction and stop word removal on the data to be processed to construct the input text. Immediately afterwards, the preprocessed text is input into the model to carry out the first-level named entity recognition task to obtain the first-level tags of the named entities. The first-level tags are used as the judgment basis to carry out the second-level named entity recognition model task to obtain the second-level tags of the named entities. The second-level tags are used as the judgment basis to carry out the third-level named entity recognition task to obtain the third-level tags of the named entities. This cycle is iterated until all granularity entity tags of the named entities are obtained. Finally, the named entities and entity tags are output.

[0030] In some embodiments, it further includes: training the original recognition model with historical data to obtain a target recognition model; the target recognition model includes a first-level recognition model and a secondary recognition model, and each level of recognition model includes multiple types of sub-models.

[0031] In an example, Table 1 is a label system table, which represents the label levels and label categories of entity recognition. The constructed label system is divided into three levels. The first-level labels are used to identify three entity categories of geological entities in terms of material composition, spatio-temporal description, and physical and chemical conditions. After completing the entity recognition of the first level, the entity recognition of the second level will be carried out based on the recognition results. The second-level labels divide the labels in the first level in more detail. The material composition is refined into three entity categories of minerals, rocks, and chemical elements. The spatio-temporal description is divided into four entity categories of time, location, geological structure, and geological unit. The physical and chemical conditions are refined into two entity categories of fluids and geological anomalies. Identifying the third-level labels needs to be based on the recognition results of the second-level labels. The third-level labels further divide the three labels of minerals, chemical elements, and time in more detail to meet the knowledge discovery requirements in different scenarios.

[0032] Table 1

[0033] In another example, Table 2 is a brief introduction table of the training data for each granularity recognition model, showing the number of sentences and the number of entity labels in each model's training dataset. In this paper, two models are constructed in total. One is a named entity recognition model based on iterative loops, which consists of seven named entity recognition models with different granularities: the first-level label recognition model model1.1, used to identify three categories of entities: material composition, spatio-temporal description, and physical and chemical conditions; the second-level label recognition models model2.1, model2.2, and model2.3. Model2.1 is used to further identify three types of entities: rock, mineral, and chemical element in the material composition category. Model2.2 is used to further identify four types of entities: time, location, geological structure, and geological unit in the spatio-temporal description category. Model2.3 is used to further identify two types of entities: geological anomaly and fluid in the physical and chemical condition category; the third-level label recognition models model3.1, model3.2, and model3.3. Model3.1 is used to further identify two types of entities: metallic minerals and non-metallic minerals in the mineral category. Model3.2 is used to further identify two types of entities: isotope elements and other elements in the chemical element category. Model3.3 is used to further identify two types of entities: geological era and isotope age in the time category. The other model is the model model1.0 constructed by traditional methods, which identifies a total of 12 entity categories including rock, non-metallic mineral, metallic mineral, isotope element, other element, geological era, isotope age, location, geological structure, geological unit, geological anomaly, and fluid at one time, and is used for comparison with the named entity recognition model based on iterative loops.

[0034] Table 2

[0035] In this embodiment, the method of iterative loops is used to obtain entity labels at three levels step by step. The final output result of the model is the entity word and the entity label, realizing the recognition task of 12 different granularity entity labels. The accuracy rate and recall rate of the model's entity recognition are significantly higher than those of the traditional named entity recognition model. The F-score value of the model (i.e., the balanced F-score) is 84.48%, which is increased by 13.47% compared with the traditional named entity recognition model.

[0036] In the embodiments of the present invention, the data to be processed is preprocessed to obtain a first processed text. After text cleaning, truncation and deletion, part-of-speech restoration, and stop word removal of the first processed text, the first processed text is input into a primary recognition model to determine the first entity label to which the entity to be recognized belongs. Furthermore, based on the first entity label and a secondary recognition model, secondary recognition is performed on the entity to be recognized to determine the named entity and the target entity label corresponding to the entity to be recognized. By using the cyclic iteration method to gradually obtain entity labels at multiple levels, the accuracy of geographical entity recognition is improved from multiple dimensions.

[0037] In some embodiments, please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a named entity recognition device based on cyclic iteration provided by the embodiments of the present application; the embodiments of the present application provide a named entity recognition device 400 based on cyclic iteration, including: a preprocessing module 410, a first determination module 420, and a second determination module 430. Among them, The preprocessing module 410 is configured to preprocess the data to be processed to obtain a first processed text; the preprocessing includes text cleaning, truncation and deletion, part-of-speech restoration, unified encoding, data resampling, and stop word removal; The first determination module 420 is configured to input the first processed text into a primary recognition model to determine the first entity label to which the entity to be recognized belongs; The second determination module 430 is configured to perform secondary recognition on the entity to be recognized based on the first entity label and a secondary recognition model to determine the named entity and the target entity label corresponding to the entity to be recognized; the secondary recognition model includes at least two levels of recognition models, and the granularity of the at least two levels of recognition models is smaller than that of the primary recognition model.

[0038] In some embodiments, the preprocessing module 410 is specifically configured to: Perform text cleaning on the data to be processed to filter out irrelevant information; Select the maximum sliding window to unify the length of the data to be processed, and truncate or delete the data to be processed whose length exceeds the maximum sliding window; Obtain the part of speech of English words in the data to be processed through a predefined function, and then restore the English words to their original forms according to the part of speech; Unify the character encoding of the data to be processed; Oversample the minority classes in the data to be processed and undersample the majority classes.

[0039] In some embodiments, the first determination module 420 is specifically configured to: Input the first processed text into a primary recognition model to determine the material composition, spatio-temporal description, and physical and chemical conditions corresponding to the entity to be recognized.

[0040] In some embodiments, the secondary recognition model includes a secondary-level model, and the second determination module 430 is specifically configured to: Perform secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the second entity category corresponding to the entity to be recognized; the second entity category includes minerals, rocks, chemical elements, time, location, geological structures, geological units, fluids, and geological anomalies.

[0041] In some embodiments, the secondary recognition model includes a tertiary-level model, and the second determination module 430 is specifically configured to: Perform secondary recognition on the entity to be recognized based on the second entity category and the tertiary recognition model to determine the named entity and the target entity label corresponding to the entity to be recognized.

[0042] In some embodiments, the second determination module 430 is specifically configured to: Perform secondary recognition on the entity to be recognized based on the second entity category and the tertiary recognition model to determine the mineral type, element type, and time type corresponding to the entity to be recognized; the mineral type includes metallic minerals and non-metallic minerals, the element type includes isotope elements and non-isotope elements, and the time type includes geological eras and isotope ages; Determine the named entity and the target entity label corresponding to the entity to be recognized based on the mineral type, element type, and time type.

[0043] In some embodiments, it further includes a training module; the training module is specifically configured to: train the original recognition model using historical data to obtain a target recognition model; the target recognition model includes a primary recognition model and a secondary recognition model, and each level of recognition model includes multiple types of sub-models.

[0044] The named entity recognition device based on cyclic iteration provided by the embodiments of the present application can implement each process in the corresponding embodiments of the above-mentioned named entity recognition method based on cyclic iteration. To avoid repetition, it will not be elaborated here.

[0045] It should be noted that the named entity recognition device based on cyclic iteration provided by the embodiments of the present application and the named entity recognition method based on cyclic iteration provided by the embodiments of the present application are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the foregoing named entity recognition method based on cyclic iteration, and the repeated parts will not be elaborated.

[0046] In some embodiments, please refer to Figure 5 , Figure 5A schematic structural diagram of an electronic device provided by an embodiment of the present application. An electronic device 500 provided by an embodiment of the present application includes a processor 510 and a memory 520; the memory 520 stores a computer program, wherein the computer program, when executed by the processor, implements the above-mentioned named entity recognition method based on loop iteration.

[0047] Specifically, the processor 510 may include, for example, a general microprocessor, an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and so on. The processor 510 may also include on-board memory for caching purposes. The processor 510 may be a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present application.

[0048] The memory 520 may be, for example, any medium capable of containing, storing, transmitting, propagating, or transporting instructions. For example, the memory 520 may include, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium. Specific examples of the memory 520 include: a magnetic storage device, such as a magnetic tape or a hard disk drive (HDD); an optical storage device, such as a compact disc (CD-ROM); it may also be, such as a random access memory (RAM) or a flash memory; and / or a wired / wireless communication link.

[0049] The present application also provides a computer-readable medium, on which a computer program is stored, and when the program is executed by the processor, it implements the above-mentioned named entity recognition method based on loop iteration. The computer-readable medium may be included in the device / device / system described in the above embodiment; or it may exist separately and not be assembled into the device / device / system. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present application is implemented.

[0050] According to an embodiment of the present application, a computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, optical fiber cable, radio frequency signal, etc., or any suitable combination of the above.

[0051] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present application. In particular, without departing from the spirit and teachings of the present application, the features recited in the various embodiments and / or claims of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application. Therefore, the scope of the present application should not be limited to the above embodiments, but should be determined not only by the appended claims but also by the equivalents of the appended claims.

Claims

1. A named entity recognition method based on cyclic iteration, characterized in that It includes: Preprocess the data to be processed to obtain the first processed text; The preprocessing includes text cleaning, truncation and deletion, part-of-speech restoration, unified encoding, data resampling, and stop word removal; Input the first processed text into a primary recognition model to determine the first entity label to which the entity to be recognized belongs; Based on the first entity label and a secondary recognition model, perform secondary recognition on the entity to be recognized to determine the named entity and target entity label corresponding to the entity to be recognized; the secondary recognition model includes at least two levels of recognition models, and the granularity of the at least two levels of recognition models is smaller than that of the primary recognition model.

2. The named entity recognition method based on cyclic iteration according to claim 1, wherein The preprocessing of the data to be processed to obtain the first processed text includes: Perform text cleaning on the data to be processed to filter out irrelevant information; Select the maximum sliding window to unify the length of the data to be processed, and truncate or delete the data to be processed whose length exceeds the maximum sliding window; Obtain the part of speech of English words in the data to be processed through a predefined function, and then restore the English words to their original forms according to the part of speech; Unify the character encoding of the data to be processed; Oversample the minority classes and undersample the majority classes in the data to be processed.

3. The named entity recognition method based on cyclic iteration according to claim 1, wherein The inputting the first processed text into a primary recognition model to determine the first entity label to which the entity to be recognized belongs includes: Input the first processed text into a primary recognition model to determine the material composition, spatio-temporal description, and physical and chemical conditions corresponding to the entity to be recognized.

4. The named entity recognition method based on cyclic iteration according to claim 1, wherein The secondary recognition model includes a secondary model. The performing secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized includes: Perform secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the second entity category corresponding to the entity to be recognized; the second entity category includes minerals, rocks, chemical elements, time, location, geological structures, geological units, fluids, and geological anomalies.

5. The named entity recognition method based on cyclic iteration according to claim 4, wherein The secondary recognition model includes a tertiary model. The performing secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized includes: Perform secondary recognition on the entity to be recognized based on the second entity category and the tertiary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized.

6. The named entity recognition method based on cyclic iteration according to claim 5, wherein The performing secondary recognition on the entity to be recognized based on the second entity category and the tertiary recognition model to determine the named entity and target entity label corresponding to the entity to be recognized includes: Perform secondary recognition on the entity to be recognized based on the second entity category and the tertiary recognition model to determine the mineral type, element type, and time type corresponding to the entity to be recognized; the mineral type includes metallic minerals and non-metallic minerals, the element type includes isotope elements and non-isotope elements, and the time type includes geological eras and isotope ages; Based on the mineral type, element type, and time type, determine the named entity and target entity label corresponding to the entity to be recognized.

7. The named entity recognition method based on cyclic iteration according to claim 1, wherein It also includes: Train the original recognition model using historical data to obtain a target recognition model; the target recognition model includes a primary recognition model and a secondary recognition model, and each level of recognition model includes multiple types of sub-models.

8. An apparatus for named entity recognition based on cyclic iteration, characterized in that Including: A preprocessing module, a first determination module, and a second determination module, where The preprocessing module is configured to preprocess the data to be processed to obtain a first processed text; the preprocessing includes text cleaning, truncation and deletion, part-of-speech restoration, unified encoding, data resampling, and stop word removal; The first determination module is configured to input the first processed text into the primary recognition model to determine the first entity label to which the entity to be recognized belongs; The second determination module is configured to perform secondary recognition on the entity to be recognized based on the first entity label and the secondary recognition model to determine the named entity and the target entity label corresponding to the entity to be recognized; the secondary recognition model includes at least two levels of recognition models, and the granularity of the at least two levels of recognition models is smaller than that of the primary recognition model.

9. An electronic device, comprising a processor and a memory; the memory stores a computer program, wherein, When the computer program is executed by the processor, it implements the iterative cycle-based named entity recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.