Corpus processing method, device, storage medium and electronic device
By determining the confidence of the corpus unit and extracting feature information from the neural network, combining address hierarchy information and word segmentation tags, the problem of poor corpus correction in the prior art is solved, and the accuracy and efficiency of corpus correction are improved.
Patent Information
- Application Number
- CN202111055774.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-09
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-09-09
AI Technical Summary
The existing corpus correction methods fail to effectively utilize the reliability and standardized information of the corpus itself, resulting in limited correction effects, especially when detailed address correction is not good.
By determining the confidence of the corpus unit, extracting feature information in combination with the neural network, and using address hierarchy information and word segmentation labels, the corpus feature information is corrected, including glyph features, position features and context semantic information.
It significantly improves the accuracy of corpus correction, especially in the correction of address information, and improves the accuracy and efficiency of correction results.
Smart Images

Figure CN114281930B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence, and in particular to a corpus processing method, apparatus, storage medium, and electronic device. Background Art
[0002] Errors may occur during the acquisition and dissemination of corpora, affecting the accuracy of the final corpus. Taking the acquisition process as an example, corpus acquisition can be based on text recognition, also known as optical character recognition (OCR). OCR uses optical and computer technology to extract text from printed or written materials and convert it into a format that is both computer-readable and human-understandable. While this can be used to quickly acquire corpora, the accuracy of OCR-acquired corpora is limited. To improve the accuracy of the final corpus, corpus correction is necessary. Summary of the Invention
[0003] In order to correct the corpus and improve the accuracy of the corpus, the embodiments of the present application provide a corpus processing method, device, storage medium and electronic device.
[0004] In one aspect, an embodiment of the present application provides a method for processing corpus, the method comprising:
[0005] Determining the confidence of each corpus unit in the target corpus, where the confidence of the corpus unit represents the degree of reliability of the corpus unit in correctly expressing an associated corpus unit, where the associated corpus unit is the original corpus unit corresponding to the corpus unit in the original corpus corresponding to the target corpus;
[0006] Based on the confidence of each corpus unit, feature extraction is performed on each corpus unit to obtain feature information of each corpus unit;
[0007] Obtaining corpus feature information corresponding to the target corpus based on the feature information of each corpus unit;
[0008] Performing corpus correction processing on the corpus feature information to obtain a corrected corpus corresponding to the target corpus.
[0009] On the other hand, an embodiment of the present application provides a corpus processing device, the device comprising:
[0010] A confidence determination module is used to determine the confidence of each corpus unit in the target corpus, wherein the confidence of the corpus unit represents the reliability of the corpus unit in correctly expressing the associated corpus unit, and the associated corpus unit is the original corpus unit corresponding to the corpus unit in the original corpus corresponding to the target corpus;
[0011] a unit feature determination module, configured to extract features of each corpus unit based on the confidence level of each corpus unit to obtain feature information of each corpus unit;
[0012] A corpus feature acquisition module, configured to obtain corpus feature information corresponding to the target corpus based on the feature information of each corpus unit;
[0013] The correction processing module is used to perform corpus correction processing on the corpus feature information to obtain a corrected corpus corresponding to the target corpus.
[0014] On the other hand, an embodiment of the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the above-mentioned corpus processing method.
[0015] On the other hand, an embodiment of the present application provides an electronic device, characterized in that it includes at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the at least one processor implements the above-mentioned corpus processing method by executing the instructions stored in the memory.
[0016] The embodiments of the present application provide a corpus processing method, apparatus, storage medium and equipment. In the embodiments of the present application, the corpus feature information not only includes feature information related to the confidence of the corpus unit, but also includes segmentation label information obtained according to the segmentation of the standard information, thereby significantly enhancing the characterization capability of the corpus feature information for the corpus unit itself and the context of the corpus unit, so that the corrected corpus has better accuracy. Taking the target corpus to represent address information as an example, the corpus feature information obtained in the embodiments of the present application can characterize the address information in the target corpus to a large extent, thereby obtaining a better correction result. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 This is a flow chart of a corpus processing method provided in an embodiment of the present application;
[0019] Figure 2This is a schematic diagram of a process for determining feature information of a corpus unit provided in an embodiment of the present application;
[0020] Figure 3 This is a schematic diagram of a method for acquiring characteristic information provided by an embodiment of the present application;
[0021] Figure 4 This is a schematic diagram of the corpus feature information flow provided by the embodiment of the present application;
[0022] Figure 5 This is a schematic diagram of word segmentation information tags provided in an embodiment of the present application;
[0023] Figure 6 1 is a flowchart of a neural network training method provided in an embodiment of the present application;
[0024] Figure 7 This is a block diagram of a corpus processing device provided in an embodiment of the present application;
[0025] Figure 8 This is a schematic diagram of the hardware structure of a device provided in an embodiment of the present application for implementing the method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of them. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the embodiments of the present application.
[0027] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] In order to make the purpose, technical solutions and advantages disclosed in the embodiments of the present application more clearly understood, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present application and are not intended to limit the embodiments of the present application.
[0029] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "multiple" means two or more. In order to facilitate understanding of the above-mentioned technical solutions and the technical effects produced by the embodiments of this application, the embodiments of this application first explain the relevant professional terms:
[0030] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0031] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0032] Bidirectional Encoder Representation from Transformers (BERT): A model used for pre-training language representation. It trains a general "language understanding" model based on text corpus. The BERT model can assist in performing natural language processing (NLP) tasks.
[0033] Long Short-Term Memory (LSTM) artificial neural network: a recurrent neural network suitable for capturing sequence before and after position information and predicting sequences.
[0034] Sequence-to-Sequence (Seq2Seq) model: This is a sequence-to-sequence conversion model framework used in scenarios such as machine translation and automated answering. It leverages the global information of a longer sequence and integrates the sequence context to infer a corresponding representation of the sequence.
[0035] Contextual Language Model (N-gram): N-gram is a common statistical language model in natural language processing. Its basic concept is to perform a sliding window operation on the text content, each byte in size N, to form a sequence of byte segments of length N. Each byte segment is called a gram. The frequency of all gram occurrences in a given sentence is counted. The probability of each gram occurring in the given sentence is then determined by comparing the frequency of each gram in the entire corpus. N-grams are particularly effective in determining sentence rationality, comparing sentence similarity, and performing word segmentation.
[0036] In the related art, there are relatively limited methods for correcting corpus, which can be roughly divided into three methods: rule-based correction, context model-based correction, and sequence transformation model-based correction.
[0037] Taking rule-based address correction as an example, in the address field, the province, city, and district have a certain hierarchical relationship. This relationship can be used to correct the corpus. For example, in the case of Haizhong City, Hainan Province, it is easy to get the corpus correction result as Haikou City, Hainan Province. The rule-based method uses this hierarchical inclusion relationship to correct the address. This method can effectively correct the first three levels of addresses. However, due to the large changes in detailed addresses (such as xx Street), it is difficult for the rule-based method to effectively correct detailed addresses.
[0038] Taking address correction based on a contextual model as an example, the N-gram-based approach first segmentes the input, then performs error detection to identify potentially incorrect characters. These characters are then replaced using a perplexity set, resulting in N candidate sentences. Finally, the language model is used to score the candidate sentences, and the sentence with the highest score is selected as the corrected result. The N-gram-based approach uses a hierarchical approach to the correction process, making it a relatively cumbersome process. Furthermore, the establishment of a perplexity set significantly impacts the correction results. If the correct answer does not appear in the perplexity set, the correct answer cannot be obtained.
[0039] Taking address correction based on the sequence transformation model as an example, address correction is directly modeled as a translation process. The input of this method is the corpus to be corrected, and the output is the corrected corpus. This method is simple, but because the model has a large solution space, a large amount of training data is required during model training to achieve better correction results, and obtaining training data is often very difficult.
[0040] From the above, it can be seen that the correction methods based on context models or sequence conversion models tend to correct general corpus, and do not consider the normative information that the corpus itself needs to follow, and do not utilize the relevant information on the reliability of the corpus to be corrected, thereby affecting the correction effect to a certain extent. The above-mentioned rule-based correction only utilizes some semantic rules without considering other information, and the correction effect and applicable scenarios are very limited. In view of this, the embodiment of the present application provides a corpus processing method that can not only take into account the advantages of general corpus correction, but also make full use of the information on the reliability of the corpus itself, and even the corpus-related normative information, thereby significantly improving the correction effect and improving the accuracy of the corrected corpus.
[0041] The method provided in the embodiments of the present application may relate to the field of cloud technology, for example, to the field of big data. The method provided in the embodiments of the present application can train a neural network based on big data, so that the trained neural network has the ability to perform corpus correction. Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time frame. It is a massive, high-growth, and diversified information asset that requires a new processing model to have stronger decision-making power, insight discovery, and process optimization capabilities. With the advent of the cloud era, big data has also attracted more and more attention. Big data requires special technologies to effectively process large amounts of data within a tolerable time. Technologies suitable for big data include large-scale parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.
[0042] The method provided in the embodiment of the present application may also involve a blockchain, that is, the method provided in the embodiment of the present application may be implemented based on a blockchain, or the data involved in the method provided in the embodiment of the present application may be stored based on a blockchain, or the execution subject of the method provided in the embodiment of the present application may be located in a blockchain. Blockchain is a new application model of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, and encryption algorithm. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. The blockchain may include a blockchain underlying platform, a platform product service layer, and an application service layer.
[0043] The underlying blockchain platform can include processing modules such as user management, basic services, and smart contracts. The user management module is responsible for managing the identity information of all blockchain participants, including maintaining public and private key generation (account management), key management, and maintaining the correspondence between users' real identities and blockchain addresses (authority management). It also oversees and audits transactions involving certain real identities, providing risk control rule configuration (risk control auditing), with authorization. The basic service module is deployed on all blockchain nodes to verify the validity of service requests and, after reaching consensus on valid requests, records them in storage. For a new service request, the basic service first performs interface adaptation, parsing, and authentication processing (interface adaptation). It then encrypts the service information using a consensus algorithm (consensus management). After encryption, the encrypted information is transmitted completely and consistently to the shared ledger (network communication) and recorded and stored. The smart contract module is responsible for contract registration, issuance, contract triggering, and contract execution. Developers can define contract logic in a programming language and publish it to the blockchain (contract registration). Based on the contract terms, key calls or other events trigger execution, completing the contract logic. The module also provides contract upgrade and cancellation capabilities.
[0044] The platform's product service layer provides the basic capabilities and implementation framework for typical applications. Developers can build on these basic capabilities, overlay business features, and complete the blockchain implementation of business logic. The application service layer provides application services based on blockchain solutions for business participants to use.
[0045] The embodiments of the present application can be applied to a data processing device, which can be a terminal device. The terminal device can be, for example, a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The data processing device can also be a server, which can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. Of course, the data processing device can be a terminal device and a server, that is, the two can be executed in conjunction with each other, and the terminal device and the server can be directly or indirectly connected via wired or wireless communication, and this application does not limit this.
[0046] The following describes a corpus processing method according to an embodiment of the present application. Figure 1 A flow chart of a corpus processing method provided by an embodiment of the present application is shown. The embodiment of the present application provides the method operation steps as described above in the embodiment or flow chart, but more or fewer operation steps may be included based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the order of many steps and does not represent the only execution order. When the actual system or server product is executed, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment) according to the method shown in the embodiment or the accompanying drawings. The above method may include:
[0047] S101. Determine the confidence of each corpus unit in the target corpus. The confidence of the corpus unit represents the reliability of the corpus unit in correctly expressing the associated corpus unit. The associated corpus unit is the original corpus unit corresponding to the corpus unit in the original corpus corresponding to the target corpus.
[0048] In the embodiments of the present application, the target corpus is the corpus to be corrected. The embodiments of the present application do not limit the target corpus. It can be a corpus obtained through OCR recognition or a corpus transmitted from other channels. The embodiments of the present application also do not limit the original corpus. The original corpus can be considered as the source of the target corpus. Taking OCR recognition as an example, if the target corpus is obtained by OCR recognition of the text on the road sign, then the text on the road sign is the original corpus. Taking the example of obtaining the target corpus from a certain data source, the data in the data source that corresponds to the target corpus is the original corpus.
[0049] In the embodiments of this application, a corpus unit is the smallest unit processed by a corpus model. For a corpus in Chinese, the corpus unit can be a character, and for a corpus in English, the corpus unit can be a word. Exemplarily, if the target corpus is "Building D7, Chuangxin Industrial Park, High-tech Zone, Hefei City, Anhui Province", then this target corpus includes 16 corpus units, namely '安', '徽', '省', '台', '肥', '市', '高', '新', '区', '创', '薪', '产', '业', '园', 'D', '7', '栋'.
[0050] The above target corpus can be sourced from OCR recognition. With the rapid development of technologies such as computer technology, electronic technology, and artificial intelligence, the OCR field has witnessed significant growth and has been widely applied in people's daily lives. In the information society era, a large amount of bill, form, and certificate data is generated every day. To digitize this data, OCR technology is required for extraction and entry. In these recognition scenarios, such as waybill recognition, bill recognition, and license recognition, the address field is an important recognition content, and thus a relatively high accuracy requirement is imposed. Therefore, the content of the address field can be used as the target corpus in step S101. However, in actual scenarios, due to reasons such as noise interference, for example, uneven illumination and blurred images to be recognized, the OCR recognition result may be inaccurate, which may lead to incorrect corpus units in the target corpus. Taking the original corpus as "Building D7, Innovation Industrial Park, High-tech Zone, Hefei City, Anhui Province" and the target corpus as "Building D7, Chuangxin Industrial Park, High-tech Zone, Taifei City, Anhui Province" as an example, obviously, the corpus units "台" and "薪" in the target corpus are incorrect corpus units. The purpose of correcting the target corpus is to automatically correct the incorrect corpus units into correct corpus units.
[0051] If the above target corpus is sourced from OCR recognition, then the OCR recognition result can include the confidence level of each corpus unit. Therefore, the confidence level of each corpus unit in step S101 can be directly obtained based on the OCR recognition result. Of course, if there is a situation where the confidence level of an individual corpus unit does not exist, the confidence level of this individual corpus unit can also be set to a preset value, which can be set according to the actual situation and is not limited in the embodiments of this application. In the embodiments of this application, the confidence level of a corpus unit represents the reliability of this corpus unit in correctly expressing the associated corpus unit, and the associated corpus unit is the original corpus unit corresponding to this corpus unit in the original corpus corresponding to the above target corpus. Taking the above example, the associated corpus unit of the corpus unit "台" is "合", and it is obvious that the corpus unit "台" recognized by OCR is difficult to correctly express the associated corpus unit "合". Therefore, the confidence level given by OCR to this corpus unit is likely to be small.
[0052] In other implementation scenarios, the corpus unit may not come from OCR recognition. In this case, confidence levels can be assigned to each corpus unit based on experience. This embodiment of the present application does not elaborate on this. It is only necessary to ensure that the confidence levels are consistent with the reliability of the corpus unit in correctly expressing its associated corpus units.
[0053] S102. Based on the confidence level of each corpus unit, perform feature extraction on each corpus unit to obtain feature information of each corpus unit.
[0054] Compared with related technologies, the embodiments of the present application can take the confidence of the corpus unit itself into consideration as part of the feature information of the corpus unit. The feature information of the corpus unit can also include the glyph feature information and contextual semantic information of the corpus unit. By enriching the feature information of the corpus unit, the accuracy of the correction can be improved.
[0055] In one embodiment, Figure 2 As shown, based on the confidence of each corpus unit, feature extraction is performed on each corpus unit to obtain feature information of each corpus unit, including:
[0056] S1021. Perform feature extraction based on the corpus unit on the target corpus to obtain a first embedded feature of each corpus unit, where the first embedded feature represents the glyph feature information of each corpus unit.
[0057] The first embedded feature represents a kind of information of the corpus unit itself. This information can be considered as isolated information and does not take into account the context of the corpus unit. In this application, the information that can be extracted only from the corpus unit itself is collectively referred to as the first embedded information, which may include glyph feature information and, in other embodiments, may also include semantic feature information.
[0058] S1022. Perform position-based feature extraction on the target corpus to obtain a second embedded feature of each corpus unit, where the second embedded feature represents position feature information of each corpus unit in the target corpus.
[0059] The second embedded feature represents the position information of the corpus unit in the target corpus. This position information carries the contextual semantic information of the corpus unit, which is mainly related to the position of the corpus unit in the target corpus.
[0060] In the embodiment of the present application, steps S1021-S1022 can be implemented based on a neural network with feature extraction capabilities. For example, the structure of the neural network can be formed with reference to BERT, RoBERTa (A Robustly Optimized BERTPretraining Approach, a rod-optimized bidirectional Transformer encoder), and ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately, an encoder that efficiently learns to accurately classify token replacements). The embodiment of the present application does not limit the specific structure of the neural network.
[0061] Taking BERT as an example, the core structure of the model is the Transformer model. Transformer is a new architecture proposed in May 2018 that can replace traditional recurrent neural networks and convolutional neural networks for machine learning. The structure of Transformer is divided into a left-side encoder and a right-side decoder. It not only adds multi-head attention (Multi-Head Attention), but also adds self-attention and fusion normalization (Add&Norm). Transformer learns different features from different dimensions and adds position information through positional encoding (Positional Encoding). Transformer can extract both the first embedded features and the second embedded features mentioned above. In order to achieve a higher running speed, a fully connected layer can be used as the above decoder.
[0062] Of course, in other embodiments, steps S1021-S1022 can also be performed based on other neural networks such as LSTM networks and graph neural networks. This application will not go into details about this.
[0063] S1023. For each corpus unit, obtain feature information of the corpus unit according to the confidence of the corpus unit, the first embedding feature of the corpus unit, and the second embedding feature of the corpus unit.
[0064] In a specific embodiment, obtaining the feature information of the corpus unit according to the confidence of the corpus unit, the first embedding feature of the corpus unit, and the second embedding feature of the corpus unit specifically includes:
[0065] S10231. When the confidence of the above-mentioned corpus unit is less than the confidence threshold, obtain a first target value, and obtain feature information of the above-mentioned corpus unit based on the above-mentioned first target value, the first embedded feature of the above-mentioned corpus unit, and the second embedded feature of the above-mentioned corpus unit. The above-mentioned first target value represents the degree of attention paid to the above-mentioned corpus unit during the corpus correction process when the confidence of the above-mentioned corpus unit is less than the above-mentioned confidence threshold.
[0066] S10232. When the confidence of the above-mentioned corpus unit is greater than or equal to the above-mentioned confidence threshold, a second target value is obtained, and characteristic information of the above-mentioned corpus unit is obtained according to the above-mentioned second target value, the first embedded feature of the above-mentioned corpus unit and the second embedded feature of the above-mentioned corpus unit. The above-mentioned second target value represents the degree of attention paid to the above-mentioned corpus unit during the corpus correction process when the confidence of the corpus unit is greater than or equal to the confidence threshold.
[0067] In the embodiment of the present application, i can be used to represent the serial number of a certain corpus unit, TypeEmbedding(i) represents the confidence-related feature information of the i-th corpus unit, PositionEmbedding(i) represents the second embedding feature of the i-th corpus unit, and WordEmbedding(i) can represent the first embedding feature of the i-th corpus unit. Then the feature information Embedding(i) of the i-th corpus unit can be expressed as
[0068] WordEmbedding(i)+PositionEmbedding(i)+TypeEmbedding(i)
[0069] In order to improve the overall corpus processing speed and reduce the amount of calculation, the first target value or the second target value can be determined based on the confidence level, and the second target value is different from the first target value. That is to say, in the above formula, TypeEmbedding(i) takes the first target value or the second target value. The first target value represents the degree of attention paid to the corpus unit in the process of corpus correction when the confidence level is less than the confidence threshold. The second target value represents the degree of attention paid to the corpus unit in the process of corpus correction when the confidence level is greater than or equal to the confidence threshold. The second target value can be greater than the first target value, so that the feature information of the corpus unit with high confidence level is strengthened during the subsequent corpus correction, and the feature information of the corpus unit with low confidence level is relatively weakened, thereby playing a role in feature enhancement.
[0070] The embodiment of the present application does not limit the above confidence threshold, and can be set according to actual conditions. Taking the confidence threshold as 0.95, the first target value as 0, and the second target value as 1 as an example, the formula can be obtained: If the confidence level (i) represents the confidence level in the OCR recognition result, it can generally be considered that the corpus units with a confidence level greater than 0.95 are correctly recognized corpus units. In subsequent corrections, the corpus units with a confidence level lower than 0.95 can be focused on.
[0071] Taking the target corpus "Building D7, Chuangxin Industrial Park, High-tech Zone, Taifei City, Anhui Province" as an example, the schematic diagram of the method for obtaining feature information of the above corpus unit is as follows: Figure 3 As shown, Figure 3 For each corpus unit, the first embedding feature, the second embedding feature and the corresponding TypeEmbedding (confidence-related information) are determined respectively. By fusing these three features, the feature information corresponding to each corpus unit can be obtained. The specific fusion method is not limited in the embodiment of the present application, such as accumulation, weighted accumulation, convolution, etc., which will not be elaborated in this application.
[0072] The embodiment of the present application uses the above-mentioned feature enhancement method to allow more attention to be paid to correct corpus units during the correction process, and to focus on correcting incorrect corpus units, thereby improving the accuracy of correction.
[0073] S103. Obtain the corpus feature information corresponding to the target corpus based on the feature information of each corpus unit.
[0074] In one embodiment, the feature information of each corpus unit may be concatenated to obtain the corpus feature information corresponding to the target corpus.
[0075] In another embodiment, Figure 4 As shown, the corpus feature information corresponding to the target corpus is obtained based on the feature information of each corpus unit, including:
[0076] S1031. Obtain the normative information corresponding to the target corpus, where the normative information represents the semantic normative information that the target corpus should follow.
[0077] For example, if the target corpus is address information, the standard information can be address attribution information. In a specific embodiment, an address hierarchy information table can be created. This address hierarchy information table is the standard information described above. For example, in China, the geological hierarchy information table can include four levels of address information: province information, city information within the province, district information within the city, and detailed address information within the district.
[0078] S1032. Perform word segmentation processing on the target corpus according to the standard information to obtain word segmentation label information.
[0079] By performing word segmentation, the representation granularity of the target corpus can be increased from the corpus unit level to the word level, enabling more attention to be paid to the relationships between different words during the correction process. For example, the six characters "Guangdong Province, Shenzhen City" will become two words, "Guangdong Province" and "Shenzhen City" after word segmentation. It can be seen that the context relationship between words is simpler than that between corpus units, which can effectively reduce the solution space during the correction process of corpus feature information and improve the correction speed.
[0080] Continuing with the above example, words that match the above address attribution information can be identified in the above target corpus. The above words include at least two corpus units. Determine the position of the first corpus unit in the above word in the above target corpus, and determine the word segmentation label information corresponding to the above word based on the above position.
[0081] Still taking the target corpus "Chuangxin Industrial Park, Building D7, High-tech Zone, Taifei City, Anhui Province" as an example, "Anhui Province" obviously exists in the above address hierarchy information table, so "Anhui Province" can be determined as a "word", while "Taifei City" obviously does not match the city information belonging to Anhui Province in the above address hierarchy information table, so "Taifei City" cannot be determined as a "word".至此,后续部分不必再进行识别,则可以认为目标语料“安徽省台肥市高新区创薪产业园D7栋”中只存在一个“词”,即“安徽省”。“安徽省”中的第一语料单元即为“安”,“安”的位置为1,则对应的分词标签信息可以被确定为“1”。
[0082] Please refer to Figure 5 which shows a schematic diagram of word segmentation information tags. Among them, [CLS] represents the classification marker, and [SEP] represents the short sentence marker, both of which are tag information required during the corpus processing. Before splicing the above word segmentation label information, Figure 5 the classification marker, short sentence marker, and each corpus unit in it occupy positions 0 - 19. Therefore, after the word segmentation label information is set, the position information of the position corresponding to the subsequent short sentence marker is 20.
[0083] S1033. Splice the feature information of each above corpus unit and the above word segmentation label information to obtain the above corpus feature information.
[0084] By splicing the feature information of the corpus unit and the word segmentation label information, corpus feature information can be obtained. This corpus feature information not only includes rich information extracted from the corpus unit perspective but also fully contains the context information of the corpus. This context information is represented by the word segmentation label information in the corpus feature information, further enriching the expression ability of the corpus feature information, thus improving the correction effect and effectively reducing the solution space of the correction process and improving the correction speed.
[0085] S104. Perform corpus correction processing on the corpus feature information to obtain a corrected corpus corresponding to the target corpus.
[0086] Compared with the related art, the corpus feature information in the embodiment of the present application not only includes feature information related to the confidence of the corpus unit, but also includes the word segmentation label information obtained according to the word segmentation of the standard information, thereby significantly enhancing the corpus feature information's ability to represent the corpus unit itself and the corpus unit's context. Taking the target corpus to represent address information as an example, the corpus feature information obtained in the embodiment of the present application can represent the address information in the target corpus to a large extent, thereby obtaining better correction results.
[0087] In the embodiment of the present application, the above method can be performed based on a neural network, and the above neural network includes a feature information extraction network and a correction network, such as Figure 6 As shown, the training method of the above neural network includes:
[0088] S201. Obtain sample corpus and original corpus corresponding to the sample corpus.
[0089] The relationship between the sample corpus in the embodiment of the present application and the original corpus corresponding to the sample corpus can refer to the relationship between the target corpus and the original corpus in step S101, which will not be described in detail here.
[0090] In some embodiments, considering that the reason for errors in the target corpus is often due to interference from similar characters, in order to enhance the neural network's ability to correct similar characters, in the embodiments of the present application, sample corpus can be directly obtained, or some deformations can be performed based on the above sample corpus, so that the neural network can learn more information about similar characters and enhance its ability to correct similar characters.
[0091] Specifically, a replacement operation can be performed on a target sample corpus unit to enrich the sample corpus. The target sample corpus unit is any sample corpus unit in the sample corpus. The replacement operation can be: according to a preset probability, the target sample corpus unit in the sample corpus is replaced with another corpus unit. The other corpus unit is one of the following: a random corpus unit, the target sample corpus unit itself, or a corpus unit similar to the target sample corpus unit.
[0092] Specifically, the probabilities that the aforementioned other corpus units are random corpus units, the aforementioned target sample corpus units themselves, and corpus units similar to the aforementioned target sample corpus units may also be different. The aforementioned preset probabilities, as well as the probabilities that the aforementioned other corpus units are random corpus units, the aforementioned target sample corpus units themselves, and corpus units similar to the aforementioned target sample corpus units, can be set according to specific circumstances and are not limited in the embodiments of the present application.
[0093] For example, taking the preset probability as 0.8, Word i Represents the target sample corpus unit, p(i) represents the first random number of the randomly obtained target sample corpus unit, and the first random number is between [0,1]. According to the formula The target sample corpus unit has a probability of 0.8 remaining unchanged and a probability of 0.2 being replaced by another corpus unit f(Word i ).
[0094] Furthermore, According to this formula, q(i) can be understood as the second random number of the randomly obtained target sample corpus unit, and the second random number is between [0,1]. Indicates the similar corpus unit of the target sample corpus unit, word i_random Represents a random corpus unit, Word i Represents the target sample corpus unit itself. According to the above formula, in Word i Need to be replaced with f(Word i ), there is a 0.5 probability that it will be replaced by a similar corpus unit, a 0.3 probability that it will be replaced by a random corpus unit, and a 0.2 probability that it will be replaced by itself.
[0095] Of course, the above operations can be performed on one or more target sample corpus units in the sample corpus to further enrich the sample corpus and ultimately improve the neural network's ability to correct similar characters.
[0096] S202. Determine the confidence level of each sample corpus unit in the sample corpus.
[0097] The inventive concept of the method for determining confidence is the same as that described above and will not be repeated here. The difference is that, for the target sample corpus unit described above, if the target sample corpus unit is replaced by the other corpus unit described above, the confidence of the target sample corpus unit is determined to be the first target value; if the target sample corpus unit is not replaced by the other corpus unit described above, the confidence of the target sample corpus unit is determined to be the second target value.
[0098] S203. Input the confidence of each of the above sample corpus units into the above feature information extraction network, so that the above feature information extraction network determines the feature information of each of the above sample corpus units based on the confidence of each of the above sample corpus units; and obtains the sample corpus feature information corresponding to the above sample corpus based on the feature information of each of the above sample corpus units.
[0099] S204. Input the above-mentioned sample corpus feature information into the above-mentioned correction network to perform corpus correction processing to obtain sample corrected corpus.
[0100] S205. Adjust the parameters of the feature information extraction network and the parameters of the correction network according to the differences between the original corpus and the sample correction corpus.
[0101] The execution method of steps S202-S204 in the embodiment of the present application can refer to the above-mentioned S101-S104 and will not be repeated here. The embodiment of the present application does not limit the structure of the neural network. Specifically, the network structure of the feature information extraction network and the correction network can refer to the network structure mentioned above, and other network structures and their variations in the prior art can also be used. This application does not limit this.
[0102] In the embodiments of this application, by enriching the sample corpus, the neural network's ability to correct similar characters is significantly improved, while also reducing the difficulty of obtaining sample corpus. The neural network training process not only focuses on the confidence information of the sample corpus units, but also further focuses on the contextual semantic information, thereby significantly improving the correction ability and ensuring that the corrected corpus is highly consistent with the original corpus.
[0103] The present application also discloses a corpus processing device, such as Figure 7 As shown, the above device includes:
[0104] The confidence determination module 10 is used to determine the confidence of each corpus unit in the target corpus. The confidence of the above corpus unit represents the reliability of the above corpus unit in correctly expressing the associated corpus unit. The above associated corpus unit is the original corpus unit corresponding to the above corpus unit in the original corpus corresponding to the above target corpus.
[0105] The unit feature determination module 20 is configured to extract features from each of the corpus units based on the confidence level of each of the corpus units to obtain feature information of each of the corpus units.
[0106] The corpus feature acquisition module 30 is configured to obtain corpus feature information corresponding to the target corpus based on the feature information of each corpus unit.
[0107] The correction processing module 40 is used to perform corpus correction processing on the corpus feature information to obtain a corrected corpus corresponding to the target corpus.
[0108] In one embodiment, the unit feature determination module 20 is configured to perform the following operations:
[0109] Performing feature extraction based on corpus units on the target corpus to obtain a first embedded feature of each corpus unit, wherein the first embedded feature represents glyph feature information of each corpus unit;
[0110] Performing position-based feature extraction on the target corpus to obtain a second embedded feature of each corpus unit, wherein the second embedded feature represents position feature information of each corpus unit in the target corpus;
[0111] For each corpus unit, feature information of the corpus unit is obtained according to the confidence of the corpus unit, the first embedding feature of the corpus unit, and the second embedding feature of the corpus unit.
[0112] In one embodiment, the unit feature determination module 20 is configured to perform the following operations:
[0113] When the confidence of the corpus unit is less than a confidence threshold, a first target value is obtained, and feature information of the corpus unit is obtained based on the first target value, the first embedded feature of the corpus unit, and the second embedded feature of the corpus unit, wherein the first target value represents the degree of attention paid to the corpus unit during the corpus correction process when the confidence of the corpus unit is less than the confidence threshold;
[0114] When the confidence of the above-mentioned corpus unit is greater than or equal to the above-mentioned confidence threshold, a second target value is obtained. According to the above-mentioned second target value, the first embedded feature of the above-mentioned corpus unit and the second embedded feature of the above-mentioned corpus unit, the feature information of the above-mentioned corpus unit is obtained. The above-mentioned second target value represents the degree of attention paid to the above-mentioned corpus unit during the corpus correction process when the confidence of the corpus unit is greater than or equal to the confidence threshold.
[0115] In one embodiment, the corpus feature acquisition module 30 includes:
[0116] A standard information acquisition unit is used to acquire standard information corresponding to the target corpus, wherein the standard information represents semantic standard information that the target corpus should follow;
[0117] A word segmentation unit is used to perform word segmentation processing on the target corpus according to the above-mentioned standard information to obtain word segmentation label information;
[0118] The corpus feature information acquisition unit is used to concatenate the feature information of each corpus unit and the word segmentation label information to obtain the corpus feature information.
[0119] In one embodiment, the target corpus represents address information, the standard information is address attribution information, and the word segmentation unit is used to perform the following operations:
[0120] Identifying a word in the target corpus that matches the address attribution information, wherein the word includes at least two corpus units;
[0121] The position of the first corpus unit in the word is determined in the target corpus, and the word segmentation tag information corresponding to the word is determined according to the position.
[0122] In one embodiment, a neural network training module is further included. The neural network includes a feature information extraction network and a correction network. The neural network training module is configured to perform the following operations:
[0123] Obtain sample corpus and the original corpus corresponding to the sample corpus;
[0124] Determine the confidence level of each sample corpus unit in the sample corpus;
[0125] Inputting the confidence of each sample corpus unit into the feature information extraction network, so that the feature information extraction network determines the feature information of each sample corpus unit according to the confidence of each sample corpus unit; and obtaining the sample corpus feature information corresponding to the sample corpus according to the feature information of each sample corpus unit;
[0126] Inputting the sample corpus feature information into the correction network to perform corpus correction processing to obtain a sample corrected corpus;
[0127] According to the difference between the original corpus and the sample revised corpus, the parameters of the feature information extraction network and the parameters of the revised network are adjusted.
[0128] In one embodiment, the neural network training module is configured to perform the following operations:
[0129] According to a preset probability, the target sample corpus unit in the sample corpus is replaced with another corpus unit, where the other corpus unit is one of the following: a random corpus unit, the target sample corpus unit itself, or a corpus unit similar in form to the target sample corpus unit;
[0130] The above-mentioned determination of the confidence level of each sample corpus unit in the above-mentioned sample corpus includes:
[0131] If the target sample corpus unit is replaced by the other corpus unit, the confidence of the target sample corpus unit is determined to be a first target value; if the target sample corpus unit is not replaced by the other corpus unit, the confidence of the target sample corpus unit is determined to be a second target value;
[0132] The target sample corpus unit is any sample corpus unit in the sample corpus.
[0133] Specifically, the embodiment of the present application discloses a corpus processing device and the corresponding method embodiment described above, both of which are based on the same inventive concept. For details, please refer to the method embodiment and will not be repeated here.
[0134] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned corpus processing method.
[0135] The present application also provides a computer-readable storage medium that can store a plurality of instructions. The instructions can be loaded by a processor and executed by the aforementioned corpus processing method of the present application.
[0136] Furthermore, Figure 8 A schematic diagram of the hardware structure of a device for implementing the method provided in the embodiment of the present application is shown. The above-mentioned device may participate in constituting or include the apparatus or system provided in the embodiment of the present application. Figure 8 As shown, the device 10 may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 8 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 8 More or fewer components than shown, or with Figure 8 Different configurations shown.
[0137] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the device 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0138] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the above-mentioned method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned corpus processing method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the device 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0139] Transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of device 10. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module configured to communicate with the Internet wirelessly.
[0140] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of device 10 (or mobile device).
[0141] It should be noted that the above-mentioned order of the embodiments of the present application is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above-mentioned embodiments of the present application are described in terms of specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0142] Each embodiment of the present application is described in a progressive manner. Similar portions between the embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device and server embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0143] Those skilled in the art will understand that all or part of the steps for implementing the above embodiments may be accomplished by hardware, or may be accomplished by instructing the relevant hardware through a program, and the above program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0144] The above is only a preferred embodiment of the embodiment of the present application and is not intended to limit the embodiment of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiment of the present application should be included in the scope of protection of the embodiment of the present application.
Claims
1. A corpus processing method, characterized in that: The method comprises: Determining the confidence of each corpus unit in the target corpus, where the confidence of the corpus unit represents the degree of reliability of the corpus unit in correctly expressing an associated corpus unit, where the associated corpus unit is the original corpus unit corresponding to the corpus unit in the original corpus corresponding to the target corpus; For each corpus unit, if the confidence of the corpus unit is less than a confidence threshold, a first target value is obtained, and the first target value, the first embedding feature of the corpus unit, and the second embedding feature of the corpus unit are fused to obtain feature information of the corpus unit; if the confidence of the corpus unit is greater than or equal to the confidence threshold, a second target value is obtained, and the second target value, the first embedding feature of the corpus unit, and the second embedding feature of the corpus unit are fused to obtain feature information of the corpus unit; Obtaining corpus feature information corresponding to the target corpus based on the feature information of each corpus unit; Performing corpus correction processing on the corpus feature information to obtain a corrected corpus corresponding to the target corpus; Among them, the first embedded feature and the second embedded feature respectively represent the glyph feature information and position feature information of the corpus unit, and the first target value and the second target value respectively represent the degree of attention paid to the corpus unit during the corpus correction process when the confidence is less than the confidence threshold and greater than or equal to the confidence threshold.
2. The method according to claim 1, characterized in that The method further comprises: Performing feature extraction based on corpus units on the target corpus to obtain a first embedding feature of each corpus unit; Performing position-based feature extraction on the target corpus to obtain a second embedding feature of each corpus unit.
3. The method according to claim 1 or 2, characterized in that The obtaining of the corpus feature information corresponding to the target corpus according to the feature information of each corpus unit includes: Acquiring standard information corresponding to the target corpus, wherein the standard information represents semantic standard information that the target corpus should follow; Performing word segmentation processing on the target corpus according to the specification information to obtain word segmentation label information; The feature information of each corpus unit and the word segmentation label information are concatenated to obtain the corpus feature information.
4. The method according to claim 3, characterized in that The target corpus represents address information, the standard information is address attribution information, and the word segmentation processing of the target corpus according to the standard information to obtain word segmentation label information includes: Identifying a word in the target corpus that matches the address attribution information, wherein the word includes at least two corpus units; The position of the first corpus unit in the word is determined in the target corpus, and the word segmentation tag information corresponding to the word is determined according to the position.
5. The method according to claim 1, wherein The method is implemented based on a neural network, which includes a feature information extraction network and a correction network. The training method of the neural network is as follows: Obtaining a sample corpus and an original corpus corresponding to the sample corpus; Determining the confidence level of each sample corpus unit in the sample corpus; Inputting the confidence of each sample corpus unit into the feature information extraction network, so that the feature information extraction network determines the feature information of each sample corpus unit according to the confidence of each sample corpus unit; and obtaining sample corpus feature information corresponding to the sample corpus according to the feature information of each sample corpus unit; Inputting the sample corpus feature information into the correction network to perform corpus correction processing to obtain a sample corrected corpus; According to the difference between the original corpus and the sample revised corpus, the parameters of the feature information extraction network and the parameters of the revised network are adjusted.
6. The method according to claim 5, characterized in that The obtaining of sample corpus includes: According to a preset probability, the target sample corpus unit in the sample corpus is replaced with another corpus unit, where the other corpus unit is one of the following: a random corpus unit, the target sample corpus unit itself, or a corpus unit similar in form to the target sample corpus unit; Determining the confidence of each sample corpus unit in the sample corpus includes: If the target sample corpus unit is replaced by the other corpus unit, the confidence of the target sample corpus unit is determined to be a first target value; if the target sample corpus unit is not replaced by the other corpus unit, the confidence of the target sample corpus unit is determined to be a second target value; The target sample corpus unit is any sample corpus unit in the sample corpus.
7. A corpus processing device, characterized in that: The device comprises: A confidence determination module is used to determine the confidence of each corpus unit in the target corpus, wherein the confidence of the corpus unit represents the reliability of the corpus unit in correctly expressing the associated corpus unit, and the associated corpus unit is the original corpus unit corresponding to the corpus unit in the original corpus corresponding to the target corpus; a unit feature determination module configured to, for each corpus unit, obtain a first target value when the confidence of the corpus unit is less than a confidence threshold, fuse the first target value, the first embedding feature of the corpus unit, and the second embedding feature of the corpus unit to obtain feature information of the corpus unit; and obtain a second target value when the confidence of the corpus unit is greater than or equal to the confidence threshold, fuse the second target value, the first embedding feature of the corpus unit, and the second embedding feature of the corpus unit to obtain feature information of the corpus unit; A corpus feature acquisition module, configured to obtain corpus feature information corresponding to the target corpus based on the feature information of each corpus unit; A correction processing module, configured to perform corpus correction processing on the corpus feature information to obtain a corrected corpus corresponding to the target corpus; Among them, the first embedded feature and the second embedded feature respectively represent the glyph feature information and position feature information of the corpus unit, and the first target value and the second target value respectively represent the degree of attention paid to the corpus unit during the corpus correction process when the confidence is less than the confidence threshold and greater than or equal to the confidence threshold.
8. The device according to claim 7, characterized in that The unit feature determination module is further used to: Performing feature extraction based on corpus units on the target corpus to obtain a first embedding feature of each corpus unit; Performing position-based feature extraction on the target corpus to obtain a second embedding feature of each corpus unit.
9. The device according to claim 7 or 8, characterized in that The corpus feature acquisition module is used to: Acquiring standard information corresponding to the target corpus, wherein the standard information represents semantic standard information that the target corpus should follow; Performing word segmentation processing on the target corpus according to the specification information to obtain word segmentation label information; The feature information of each corpus unit and the word segmentation label information are concatenated to obtain the corpus feature information.
10. The device according to claim 9, characterized in that The target corpus represents address information, the standard information is address attribution information, and the corpus feature acquisition module is further used to: Identifying a word in the target corpus that matches the address attribution information, wherein the word includes at least two corpus units; The position of the first corpus unit in the word is determined in the target corpus, and the word segmentation tag information corresponding to the word is determined according to the position.
11. The device according to claim 7, characterized in that It also includes a neural network training module, the neural network includes a feature information extraction network and a correction network, and the neural network training module is used to: Obtaining a sample corpus and an original corpus corresponding to the sample corpus; Determining the confidence level of each sample corpus unit in the sample corpus; Inputting the confidence of each sample corpus unit into the feature information extraction network, so that the feature information extraction network determines the feature information of each sample corpus unit according to the confidence of each sample corpus unit; and obtaining sample corpus feature information corresponding to the sample corpus according to the feature information of each sample corpus unit; Inputting the sample corpus feature information into the correction network to perform corpus correction processing to obtain a sample corrected corpus; According to the difference between the original corpus and the sample revised corpus, the parameters of the feature information extraction network and the parameters of the revised network are adjusted.
12. The device according to claim 11, characterized in that The neural network training module is further used to: According to a preset probability, the target sample corpus unit in the sample corpus is replaced with another corpus unit, where the other corpus unit is one of the following: a random corpus unit, the target sample corpus unit itself, or a corpus unit similar in form to the target sample corpus unit; Determining the confidence of each sample corpus unit in the sample corpus includes: If the target sample corpus unit is replaced by the other corpus unit, the confidence of the target sample corpus unit is determined to be a first target value; if the target sample corpus unit is not replaced by the other corpus unit, the confidence of the target sample corpus unit is determined to be a second target value; The target sample corpus unit is any sample corpus unit in the sample corpus.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement a corpus processing method according to any one of claims 1 to 6.
14. An electronic device, characterized in that: The invention comprises at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements a corpus processing method according to any one of claims 1 to 6 by executing the instructions stored in the memory.
15. A computer program product, characterized in that The computer program product includes computer instructions, and a processor of a computer device executes the computer instructions, so that the computer device executes a corpus processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text information processing method, model training method and related devices
CN110750959A