Checking method and device, medium and electronic equipment

By building training corpus and matching corpus, a named entity recognition model and a semantic similarity calculation model are established, and automatic identification and proofreading of standard specification names and numbers in design files is realized, which solves the problem of time-consuming manual proofreading and improves the efficiency of design work.

CN120031027APending Publication Date: 2025-05-23CHINA NAT PETROLEUM CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311557823.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In engineering design work, proofreading the standard specification names and numbers in design documents consumes a lot of manual labor, reducing the efficiency of design work.

Method used

By building training corpus and matching corpus, a named entity recognition model and a semantic similarity calculation model are established to automatically identify and proofread the names and numbers of standard specifications in design files.

Benefits of technology

It improves the efficiency of design work, reduces the time and cost of manual proofreading, and ensures the correctness and consistency of standard specifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031027A_ABST
    Figure CN120031027A_ABST
Patent Text Reader

Abstract

The invention discloses a proofreading method and device, a medium and electronic device.The method comprises the steps that a training corpus, a matching corpus and a semantic similarity calculation model are constructed, and the training corpus is composed of files including standard specification names and corresponding standard specification numbers; the matching corpus is composed of standard specification names and corresponding standard specification numbers; establishing a named entity recognition model; a named entity recognition model is adopted to train the training corpus, and a trained named entity recognition model is obtained; using the training named entity recognition model to recognize the obtained to-be-corrected text to obtain a target recognition result; and based on the matching corpus, proofreading the target recognition result by adopting a semantic similarity calculation model to obtain a target proofreading result. According to the method, the standard specification name and the standard specification number in the design file can be identified, and the efficiency of design work is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In engineering design work, when making design documents, it is usually necessary to proofread the design documents, such as whether the name of the standard specification is written correctly, whether the standard specification meets the timeliness of novelty search, and whether the referenced standard specification number is correct.

[0003] At present, proofreading work often consumes a lot of man-hours, which reduces the efficiency of design work and is not conducive to the rapid completion of design work. Summary of the invention

[0004] The embodiments of the present application provide a proofreading method, device, medium and electronic device, which can identify the standard specification name and standard specification number in the design document and improve the efficiency of the design work.

[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0006] According to a first aspect of an embodiment of the present application, a proofreading method is provided, comprising:

[0007] Constructing a training corpus, a matching corpus and a semantic similarity calculation model, wherein the training corpus is composed of files including standard specification names and corresponding standard specification numbers, and the matching corpus is composed of standard specification names and corresponding standard specification numbers;

[0008] Build a named entity recognition model;

[0009] Using a named entity recognition model to train the training corpus to obtain a training named entity recognition model;

[0010] The trained named entity recognition model is used to perform the acquired text to be proofread to obtain the target recognition result;

[0011] Based on the matching corpus, the semantic similarity calculation model is used to proofread the target recognition result to obtain the target proofreading result.

[0012] In some embodiments of the present application, based on the above scheme, the named entity recognition model is a BERT-BiLSTM-CRF model, and the named entity recognition model is used to train the training corpus to obtain a training named entity recognition model, including:

[0013] The BERT model is used to train the training corpus to obtain a training set consisting of word vectors based on context information for each word;

[0014] The BiLSTM model is used to train the training set to obtain a label score of the entity type corresponding to each word;

[0015] The CRF model is used to perform label constraints on the label scores of the entity types corresponding to each word to obtain the entity annotation results.

[0016] In some embodiments of the present application, based on the above scheme, the BERT model is used to train the training corpus to obtain a word vector for each word based on context information to form a training set, including:

[0017] The words in the training corpus are input into the BERT model, and the BERT model represents the words with feature vectors to generate word vectors based on context information.

[0018] In some embodiments of the present application, based on the above scheme, the BiLSTM model is used to train the training set to obtain the label score of the entity category corresponding to each word, including:

[0019] The word vector of each word based on the context information is transformed into a matrix dimension, and the label score of the entity category corresponding to each word is calculated.

[0020] In some embodiments of the present application, based on the above scheme, the label score of each word corresponding to the entity category of the CRF model is used to perform label constraints to obtain entity annotation results, including:

[0021] Get label constraint mechanism;

[0022] The CRF model removes incorrect logical labels through a label constraint mechanism and obtains entity labeling results based on the highest label score of the entity type corresponding to each word.

[0023] In some embodiments of the present application, based on the above solution, the step of constructing a semantic similarity calculation model includes:

[0024] The semantic similarity calculation model is constructed using the FAISS algorithm and dictionary tree.

[0025] In some embodiments of the present application, based on the above scheme, the construction of the training corpus and the matching corpus includes:

[0026] Collecting design texts, performing data extraction and data cleaning on the design texts, and obtaining a training corpus;

[0027] Crawl data from industry websites to obtain a matching corpus.

[0028] According to a second aspect of an embodiment of the present application, a proofreading device is provided, including:

[0029] Construction unit, constructing training corpus, matching corpus and semantic similarity calculation model;

[0030] Establish a unit and build a named entity recognition model;

[0031] A first obtaining unit is used to train the training corpus using a named entity recognition model to obtain a training named entity recognition model;

[0032] The second obtaining unit uses the trained named entity recognition model to perform the obtained text to be proofread to obtain the target recognition result;

[0033] The third obtaining unit proofreads the target recognition result based on the matching corpus and adopts a semantic similarity calculation model to obtain a target proofreading result.

[0034] In some embodiments of the present application, based on the aforementioned solution, the named entity recognition model is a BERT-BiLSTM-CRF model, and the first obtaining unit is configured as follows:

[0035] The fourth unit uses the BERT model to train the training corpus to obtain a training set consisting of word vectors based on context information for each word;

[0036] A fifth obtaining unit uses a BiLSTM model to train the training set to obtain a label score of an entity type corresponding to each word;

[0037] The sixth obtaining unit uses the CRF model to perform label constraints on the label scores of the entity types corresponding to each word to obtain the entity labeling results.

[0038] In some embodiments of the present application, based on the above solution, the fourth obtaining unit is configured as follows:

[0039] The generating unit inputs the words of the training corpus into the BERT model, and the BERT model represents the words with feature vectors to generate word vectors based on context information.

[0040] In some embodiments of the present application, based on the above solution, the fifth obtaining unit is configured as follows:

[0041] In the seventh unit, the word vector of the context information is transformed into a matrix dimension, and the label score of the entity type corresponding to each word is calculated.

[0042] In some embodiments of the present application, based on the above solution, the sixth obtaining unit is configured as follows:

[0043] Acquisition unit, acquisition label constraint mechanism;

[0044] In the eighth obtaining unit, the CRF model removes incorrect logical labels through the label constraint mechanism, and obtains the entity labeling result according to the highest label score of the entity type corresponding to each word.

[0045] In some embodiments of the present application, based on the aforementioned scheme, the construction unit is configured as follows:

[0046] The semantic similarity calculation model is constructed using the FAISS algorithm and dictionary tree.

[0047] In some embodiments of the present application, based on the aforementioned scheme, the construction unit is configured as follows:

[0048] A ninth obtaining unit collects design texts, extracts and cleans the design texts, and obtains a training corpus;

[0049] The tenth unit crawls data from industry websites to obtain a matching corpus.

[0050] According to a third aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. The computer program includes executable instructions. When the executable instructions are executed by a processor, the method described in any embodiment of the first aspect above is implemented.

[0051] According to a fourth aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; and a memory for storing executable instructions of the processors, wherein when the executable instructions are executed by the one or more processors, the one or more processors implement the method described in any embodiment of the first aspect above.

[0052] The beneficial effects of this application are as follows:

[0053] The training corpus consisting of standard specification names and standard specification numbers is trained by a named entity recognition model to obtain a training named entity recognition model. Based on the matching corpus, the training named entity recognition model can identify and proofread the standard specification names and standard specification numbers in the design documents, effectively improving work efficiency.

[0054] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0056] Figure 1 A flowchart of a proofreading method in an embodiment of the present application is shown;

[0057] Figure 2 A schematic diagram of a named entity recognition model in an embodiment of the present application is shown;

[0058] Figure 3 A block diagram of a proofreading device in an embodiment of the present application is shown;

[0059] Figure 4 is a schematic diagram of a computer-readable storage medium according to an embodiment of the present invention;

[0060] Figure 5 FIG. 4 is a schematic diagram of a system structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0062] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0063] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0064] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0065] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the feature. In the description of this application, unless otherwise specified, "plurality" means two or more.

[0066] Figure 1 A flowchart of a proofreading method in an embodiment of the present application is shown. Figure 2 A schematic diagram of a named entity recognition model in an embodiment of the present application is shown, see Figure 1 and Figure 2 , the present application provides a proofreading method, which includes at least S1 to S4, which are described in detail as follows:

[0067] In step S1, a training corpus, a matching corpus and a semantic similarity calculation model are constructed, wherein the training corpus is composed of files including standard specification names and corresponding standard specification numbers, and the matching corpus is composed of standard specification names and corresponding standard specification numbers.

[0068] Specifically, the file including the standard specification name and the corresponding standard specification number can be a project description or a specification book such as a product specification book or a technical specification book. Taking the oil and gas field as an example, the file is a project description and specification book in the oil and gas field. The semantic similarity calculation model is used to compare and match the corpus and target recognition results.

[0069] Optionally, the constructing a training corpus includes:

[0070] Construct a data dictionary consisting of standard specification names;

[0071] Using a data dictionary to label the standard canonical names of the files to improve identification;

[0072] Get the commonalities and characteristics of standard specification numbers;

[0073] The commonalities and characteristics of the standard specification number are retrieved through regular expressions to obtain the standard specification number.

[0074] In step S2, a named entity recognition model is established.

[0075] Specifically, the named entity recognition model can be one of an LSTM+CRF model, a Dilated-CNN model, and a BERT-BiLSTM-CRF model.

[0076] In step S3, the training corpus is trained using a named entity recognition model to obtain a training named entity recognition model.

[0077] In step S4, the trained named entity recognition model is used to perform recognition on the acquired text to be proofread to obtain a target recognition result.

[0078] In step S5, based on the matching corpus, a semantic similarity calculation model is used to proofread the target recognition result to obtain a target proofreading result.

[0079] Specifically, the text to be proofread may be a design document that needs to be proofread, and the target proofreading result may be whether the standard specification name and the corresponding standard specification number are correct, wrong, or invalid.

[0080] Optionally, the named entity recognition model is a BERT-BiLSTM-CRF model, and the named entity recognition model is used to train the training corpus to obtain a training named entity recognition model, including:

[0081] The BERT model is used to train the training corpus to obtain a training set consisting of word vectors based on context information for each word. The test set can also be obtained in the same way.

[0082] The BiLSTM model is used to train the training set to obtain a label score of the entity type corresponding to each word;

[0083] The CRF model is used to perform label constraints on the label scores of the entity types corresponding to each word to obtain the entity annotation results.

[0084] Optionally, the BERT model is used to train the training corpus to obtain a word vector for each word based on context information to form a training set, including:

[0085] The words in the training corpus are input into the BERT model, and the BERT model represents the words with feature vectors to generate word vectors based on context information.

[0086] Specifically, after the BERT model represents the word with a feature vector, it annotates the standard specification name word by word to mark the standard specification name and the corresponding standard specification number in the training corpus.

[0087] Optionally, the BiLSTM model is used to train the training set to obtain a label score of an entity type corresponding to each word, including:

[0088] The word vector of each word based on the context information is transformed into a matrix dimension, and the label score of the entity category corresponding to each word is calculated.

[0089] Optionally, the labeling constraint is performed by using the label score of the entity type corresponding to each word in the CRF model to obtain the entity labeling result, including:

[0090] Get label constraint mechanism;

[0091] The CRF model removes incorrect logical labels through a label constraint mechanism and obtains entity labeling results based on the highest label score of the entity type corresponding to each word.

[0092] Optionally, the method further comprises:

[0093] Calculate the result of the loss function.

[0094] Specifically, through the addition operation after dimensionality transformation, the result of calculating the loss function, namely the CRF loss result, is obtained.

[0095] Optionally, the constructing of a semantic similarity calculation model includes:

[0096] The semantic similarity calculation model is constructed by using the FAISS algorithm and the dictionary tree. Specifically, when the target recognition result is proofread by the semantic similarity calculation model, the FAISS algorithm is used to perform similarity retrieval on the standard specification name, and the dictionary tree is used to perform a deep traversal of the standard specification number. When FAISS is used for vector similarity search, the specification with the highest score is found by comparing the similarity, which is the final standard specification name to be found. This can be done by using the inverted index and product quantization method. The Hybrid Representation can be introduced into the dictionary to encode the word semantics of each category label in the hierarchy so that different categories are semantically distinguishable to enhance the representation information. The dictionary tree can be used to perform a deep traversal of the standard specification number, thereby comparing the standard specification number in the target recognition result with the standard specification number in the matching library, completing the recognition work, and correcting the wrong or expired standard specification name and the corresponding standard specification number.

[0097] For example, if the target recognition result is A, after searching and comparing the matching specification library, the target search result A can be found in the matching specification library, and the result will be displayed in green. If A is not found after comparison, it means that there is a specification writing error in the target recognition result or this specification has been invalidated. It will be displayed in red and the correct standard specification name will be recommended for the designer's reference.

[0098] Optionally, the constructing of the training corpus and the matching corpus includes:

[0099] Collecting design texts, performing data extraction and data cleaning on the design texts, and obtaining a training corpus;

[0100] Crawl data from industry websites to obtain a matching corpus.

[0101] Specifically, taking the oil and gas field as an example, the industry websites can be Gongbiao.com and Guanzhihui.com.

[0102] Figure 3 A block diagram of a proofreading device in an embodiment of the present application is shown, see Figure 2 According to a second aspect of an embodiment of the present application, a proofreading device 100 is provided, comprising:

[0103] A construction unit 101 constructs a training corpus, a matching corpus and a semantic similarity calculation model;

[0104] Establishing unit 102, establishing a named entity recognition model;

[0105] A first obtaining unit 103 uses a named entity recognition model to train the training corpus to obtain a training named entity recognition model;

[0106] The second obtaining unit 104 uses the trained named entity recognition model to perform the obtained text to be proofread to obtain a target recognition result;

[0107] The third obtaining unit 105 , based on the matching corpus, uses a semantic similarity calculation model to proofread the target recognition result to obtain a target proofreading result.

[0108] Optionally, the named entity recognition model is a BERT-BiLSTM-CRF model, and the first obtaining unit is configured as follows:

[0109] The fourth unit uses the BERT model to train the training corpus to obtain a training set consisting of word vectors based on context information for each word;

[0110] A fifth obtaining unit uses a BiLSTM model to train the training set to obtain a label score of an entity type corresponding to each word;

[0111] The sixth obtaining unit uses the CRF model to perform label constraints on the label scores of the entity types corresponding to each word to obtain the entity labeling results.

[0112] Optionally, the fourth obtaining unit is configured as:

[0113] The generating unit inputs the words of the training corpus into the BERT model, and the BERT model represents the words with feature vectors to generate word vectors based on context information.

[0114] Optionally, the fifth obtaining unit is configured as

[0115] In the seventh unit, the word vector of the context information is transformed into a matrix dimension, and the label score of the entity type corresponding to each word is calculated.

[0116] Optionally, the sixth obtaining unit is configured as

[0117] Acquisition unit, acquisition label constraint mechanism;

[0118] In the eighth obtaining unit, the CRF model removes incorrect logical labels through the label constraint mechanism, and obtains the entity labeling result according to the highest label score of the entity type corresponding to each word.

[0119] Optionally, the construction unit is configured as:

[0120] The semantic similarity calculation model is constructed using the FAISS algorithm and dictionary tree.

[0121] Optionally, the construction unit is configured as:

[0122] A ninth obtaining unit collects design texts, extracts and cleans the design texts, and obtains a training corpus;

[0123] The tenth unit crawls data from industry websites to obtain a matching corpus.

[0124] To summarize, the training corpus consisting of standard specification names and standard specification numbers is trained by the named entity recognition model to obtain a training named entity recognition model. Based on the matching corpus, the training named entity recognition model can identify and proofread the standard specification names and standard specification numbers in the design documents, effectively improving work efficiency.

[0125] Based on the same inventive concept, as a third aspect, the present application also provides a computer-readable storage medium on which a program product capable of implementing the above-mentioned proofreading method of the present specification is stored. In some possible implementations, various aspects of the present application can also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary implementations of the present application described in the above-mentioned "Exemplary Method" section of the present specification.

[0126] refer to Figure 4As shown, a program product 200 for implementing the above method according to an embodiment of the present application is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.

[0127] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0128] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0129] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.

[0130] Program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0131] As another aspect, the present application also provides an electronic device capable of implementing the above method.

[0132] Those skilled in the art will appreciate that various aspects of the present application may be implemented as a system, method or program product. Therefore, various aspects of the present application may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to as "circuit", "module" or "system" herein.

[0133] Refer to the following Figure 5 The electronic device 300 according to this embodiment of the present application is described. Figure 5 The electronic device 300 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0134] like Figure 5 As shown, the electronic device 300 is in the form of a general computing device. The components of the electronic device 300 may include but are not limited to: at least one processing unit 310, at least one storage unit 320, and a bus 330 connecting different system components (including the storage unit 320 and the processing unit 310).

[0135] The storage unit stores program codes, which can be executed by the processing unit 310, so that the processing unit 310 executes the steps described in the above “Example Method” section of this specification according to various exemplary implementations of the present application.

[0136] The storage unit 320 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 321 and / or a cache memory unit 322 , and may further include a read-only memory unit (ROM) 323 .

[0137] The storage unit 320 may also include a program / utility 324 having a set (at least one) of program modules 325, such program modules 325 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0138] Bus 330 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0139] The electronic device 300 may also communicate with one or more external devices 400 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 300, and / or communicate with any device that enables the electronic device 300 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 350. Furthermore, the electronic device 300 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 360. As shown, the network adapter 360 communicates with other modules of the electronic device 300 via a bus 330. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0140] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium. Other examples and implementations are within the scope and spirit of the present application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hard wiring, or a combination of any of these. In addition, each functional unit may be integrated into a processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit.

[0141] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0142] The units described as separate components may or may not be physically separated, and the components of the control device may or may not be physical units, that is, they may be located in one place or distributed in multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0143] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly or all or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc., which can store program code.

[0144] The above description is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of the claims of the present application.

Claims

1. A proofreading method, It is characterized in that include: Constructing a training corpus, a matching corpus and a semantic similarity calculation model, wherein the training corpus is composed of files including standard specification names and corresponding standard specification numbers, and the matching corpus is composed of standard specification names and corresponding standard specification numbers; Build a named entity recognition model; Using a named entity recognition model to train the training corpus to obtain a training named entity recognition model; The trained named entity recognition model is used to perform the acquired text to be proofread to obtain the target recognition result; Based on the matching corpus, the semantic similarity calculation model is used to proofread the target recognition result to obtain the target proofreading result.

2. A proofreading method according to claim 1, It is characterized in that The named entity recognition model is a BERT-BiLSTM-CRF model. The named entity recognition model is used to train the training corpus to obtain a training named entity recognition model, including: The BERT model is used to train the training corpus to obtain a training set consisting of word vectors based on context information for each word; The BiLSTM model is used to train the training set to obtain a label score of the entity type corresponding to each word; The CRF model is used to perform label constraints on the label scores of the entity types corresponding to each word to obtain the entity annotation results.

3. A proofreading method according to claim 2, It is characterized in that The BERT model is used to train the training corpus to obtain a word vector for each word based on context information to form a training set, including: The words in the training corpus are input into the BERT model, and the BERT model represents the words with feature vectors to generate word vectors based on context information.

4. A proofreading method according to claim 2, It is characterized in that The BiLSTM model is used to train the training set to obtain the label score of the entity type corresponding to each word, including: The word vector of each word based on the context information is transformed into a matrix dimension, and the label score of the entity category corresponding to each word is calculated.

5. A proofreading method according to claim 2, It is characterized in that The CRF model is used to perform label constraints on the label scores of the entity types corresponding to each word to obtain entity annotation results, including: Get label constraint mechanism; The CRF model removes incorrect logical labels through a label constraint mechanism and obtains entity labeling results based on the highest label score of the entity type corresponding to each word.

6. A proofreading method according to claim 1, It is characterized in that The constructing of the semantic similarity calculation model includes: The semantic similarity calculation model is constructed using the FAISS algorithm and dictionary tree.

7. A proofreading method according to claim 1, It is characterized in that The construction of the training corpus and the matching corpus includes: Collecting design texts, performing data extraction and data cleaning on the design texts, and obtaining a training corpus; Crawl data from industry websites to obtain a matching corpus.

8. A proofreading device, It is characterized in that The pretending includes: Construction unit, constructing training corpus, matching corpus and semantic similarity calculation model; Establish a unit and build a named entity recognition model; A first obtaining unit is used to train the training corpus using a named entity recognition model to obtain a training named entity recognition model; The second obtaining unit uses the trained named entity recognition model to perform the obtained text to be proofread to obtain the target recognition result; The third obtaining unit proofreads the target recognition result based on the matching corpus and adopts a semantic similarity calculation model to obtain a target proofreading result.

9. A computer-readable storage medium having a computer program stored thereon, It is characterized in that The computer program includes executable instructions, and when the executable instructions are executed by a processor, the method described in any one of claims 1 to 7 is implemented.

10. An electronic device, It is characterized in that include: one or more processors; A memory for storing executable instructions of the processor, wherein when the executable instructions are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.