Text processing method and device, storage medium and electronic device

By using the probability values and feature vectors of multiple encoding layers in the deep learning model, and updating the model with real tags, the problems of high training complexity and high computational cost in the prior art are solved, and more efficient text processing is achieved.

CN114330239BActive Publication Date: 2025-07-08BEIJING OPPO TELECOMM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111661119.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-07-08
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

In existing deep learning scenarios, active learning query strategy functions often only consider one indicator of uncertainty, diversity or representation, resulting in complex training process and large computational volume.

Method used

By obtaining the feature vectors that reference unlabeled text, the probability values and feature vectors of multiple encoding layers are used to determine the target unlabeled text, and the reference text processing model is updated with the real label until the preset conditions are met, reducing the number and calculation amount of human labels.

Benefits of technology

It reduces the complexity and computational amount of deep learning model training, improves the efficiency and accuracy of model training, makes full use of the information of the encoding layer, and reduces the number of unlabeled text in the target.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114330239B_ABST
    Figure CN114330239B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of natural language processing, and particularly relates to a text processing method, an apparatus, a computer-readable storage medium, and an electronic device. The method includes: obtaining a reference unlabeled text, and inputting the reference unlabeled text into a pre-trained reference text processing model to obtain feature vectors of each reference unlabeled text; obtaining probability values output by at least one encoding layer, and determining a plurality of target unlabeled texts in the reference unlabeled text according to the probability values and the feature vectors; determining true labels corresponding to the target unlabeled texts, and updating the reference text processing model by using the target unlabeled texts and the true labels until the reference text processing model meets a preset condition; and using the reference text processing model that meets the preset condition to process a text to be processed to obtain a processing result. The technical solution of the embodiments of the present disclosure reduces the computational amount during text processing and reduces the complexity of the training process of the deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the development of information technology, Internet data and resources exhibit a massive characteristic. In order to effectively manage and utilize this distributed massive information, content-based information retrieval and data mining have gradually become areas of great concern. Among them, text processing technology is an important foundation for information retrieval and text mining.

[0003] Current text processing methods are mainly implemented based on deep learning. However, the active learning query strategy functions in deep learning scenarios often only consider one of the indicators of uncertainty, diversity, and representativeness in the design, and the training process of the deep learning model is complex and computationally intensive.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a text processing method, a text processing device, a computer-readable medium, and an electronic device, so as to at least to a certain extent reduce the computational amount during text processing and reduce the complexity of the training process of the deep learning model.

[0006] According to a first aspect of the present disclosure, there is provided a text processing method, including: obtaining reference unlabeled texts, and inputting the reference unlabeled texts into a pre-trained reference text processing model to obtain feature vectors of the reference unlabeled texts, wherein the reference text processing model includes a plurality of encoding layers; obtaining probability values output by at least one of the encoding layers, and determining a plurality of target unlabeled texts in the reference unlabeled texts according to the probability values and the feature vectors; determining true labels corresponding to the target unlabeled texts, and updating the reference text processing model by using the target unlabeled texts and the true labels until the reference text processing model meets a preset condition; and processing a text to be processed by using the reference text processing model that meets the preset condition to obtain a processing result.

[0007] According to a second aspect of the present disclosure, there is provided a text processing device, including: an acquisition module configured to acquire a reference unlabeled text, and input the reference unlabeled text into a pre-trained reference text processing model to obtain feature vectors of the reference unlabeled texts, wherein the reference text processing model includes a plurality of encoding layers; a determination module configured to acquire probability values output by at least one of the encoding layers, and determine a plurality of target unlabeled texts in the reference unlabeled text according to the probability values and the feature vectors; an update module configured to determine true labels corresponding to the target unlabeled texts, and update the reference text processing model by using the target unlabeled texts and the true labels until the reference text processing model meets a preset condition; and a processing module configured to process a text to be processed by using the reference text processing model that meets the preset condition to obtain a processing result.

[0008] According to a third aspect of the present disclosure, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, the above method is implemented.

[0009] According to a fourth aspect of the present disclosure, there is provided an electronic device, characterized by including: one or more processors; and a memory configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the above method.

[0010] The text processing method provided by an embodiment of the present disclosure includes: acquiring a reference unlabeled text, and inputting the reference unlabeled text into a pre-trained reference text processing model to obtain feature vectors of the reference unlabeled texts, wherein the reference text processing model includes a plurality of encoding layers; acquiring probability values output by at least one of the encoding layers, and determining a plurality of target unlabeled texts in the reference unlabeled text according to the probability values and the feature vectors; determining true labels corresponding to the target unlabeled texts, and updating the reference text processing model by using the target unlabeled texts and the true labels until the reference text processing model meets a preset condition; and processing a text to be processed by using the reference text processing model that meets the preset condition to obtain a processing result. Compared with the prior art, the probability values output by at least one of the encoding layers are used to determine the target unlabeled texts, reducing the number of manual annotations and the computational amount. Further, the true labels corresponding to the target unlabeled texts are determined, and the reference text processing model is updated by using the target unlabeled texts and the true labels, fully utilizing the information in each encoding layer of the reference text processing model, and at the same time updating the reference text processing model by using the target unlabeled texts, reducing the complexity of model training and the computational amount.

[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0012] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:

[0013] Figure 1 A schematic diagram showing an exemplary system architecture to which the embodiments of the present disclosure can be applied;

[0014] Figure 2 A schematic data flow diagram of an active learning model for text processing in the related art;

[0015] Figure 3 A schematic flowchart of a text processing method in an exemplary embodiment of the present disclosure;

[0016] Figure 4 A schematic data flow diagram of a text processing method in an exemplary embodiment of the present disclosure;

[0017] Figure 5 A schematic flowchart of obtaining target unlabeled text in an exemplary embodiment of the present disclosure;

[0018] Figure 6 A schematic data flow diagram of obtaining target unlabeled text in an exemplary embodiment of the present disclosure;

[0019] Figure 7 A schematic diagram showing the composition of a text processing device in an exemplary embodiment of the present disclosure;

[0020] Figure 8 A schematic diagram of an electronic device to which the embodiments of the present disclosure can be applied. Detailed Embodiments

[0021] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.

[0022] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0023] Figure 1 A schematic diagram of a system architecture is shown. The system architecture 100 may include a terminal 110 and a server 120. Among them, the terminal 110 may be a terminal device such as a smart phone, a tablet computer, a desktop computer, a laptop computer, etc., and the server 120 generally refers to a background system that provides services related to text processing in this exemplary embodiment, and may be a single server or a cluster formed by multiple servers. A connection may be formed between the terminal 110 and the server 120 through a wired or wireless communication link for data interaction.

[0024] In one embodiment, the above text processing method may be executed by the terminal 110. For example, a user uses the terminal 110 to obtain a pre-trained reference text processing model and the text to be processed, and then uses the method of the present disclosure to update the above reference text processing model and complete the processing of the text to be processed to obtain an output result.

[0025] In one embodiment, the above text processing method may be executed by the server 120. For example, a user uses the terminal 110 to obtain a pre-trained reference text processing model and the text to be processed, and the terminal 110 uploads the reference text processing model and the text to be processed to the server 120. The server 120 updates the reference text processing model and performs text processing on the text to be processed, and returns the processing result to the terminal 110.

[0026] As can be seen from the above, the execution subject of the text processing method in this exemplary embodiment may be the above terminal 110 or server 120, and the present disclosure does not limit this.

[0027] The exemplary embodiment of the present disclosure further provides an electronic device for executing the above text processing method. The electronic device may be the above terminal 110 or server 120. Generally, the electronic device may include a processor and a memory. The memory is used to store executable instructions of the processor, and the processor is configured to execute the above text processing method by executing the executable instructions.

[0028] In related technologies, current active learning query strategies are basically designed around uncertainty, diversity, or representativeness. After AI performs deep learning, active learning has gradually been combined with deep learning. However, current active learning query strategies based on deep learning often calculate using the output of the last layer of the neural network or design the query function using the Monte Carlo approach. In scenarios related to active learning of text, in existing solutions, the pre-trained language model is simply regarded as a network model for calculation, and the output of the encoder of the last layer is used to calculate metrics such as uncertainty, or simply through the Monte Carlo approach, multiple network structures are generated through Monte Carlo random control of the fully connected layer for metric calculation.

[0029] In the artificial intelligence annotation scenario in related technologies, as shown in Figure 2 It is often necessary to perform all manual annotations on the collected data and then select a suitable model for training. In the scenario of artificial intelligence deep learning, a large amount of labeled data is often required to start model training. Therefore, a large amount of manpower is first required to label the data. Using active learning can save the annotation cost. For example, in an annotation scenario, first, we have a small amount of labeled text 210 and unlabeled text 240. Based on the labeled text 210, a machine learning model 220 is trained, and then this model is used to predict the labels of the unlabeled text. The predicted label results are sorted through the designed query 230 strategy, and the data that is predicted to be the most inaccurate and the most difficult for the model to judge among the unlabeled text is selected for manual annotation 250. The newly manually annotated data is added back to the labeled text, and then a new round of model training begins, and so on. Since in each round, the data that is predicted to be the most inaccurate and the most difficult for the model to judge is selected from the unlabeled text for manual annotation, after multiple rounds of cycling, the model will well learn the characteristics of the entire dataset, and there is no need to perform manual annotation on the remaining large amount of data in the unlabeled text, and the model can well complete the work of automatic annotation. Therefore, the annotation cost can be greatly saved and the annotation efficiency can be improved. However, in the existing active learning technology, less resources are utilized, and a relatively large number of unlabeled texts are selected, resulting in the training process of the model still being very complex and the computational amount being relatively large.

[0030] To address the above drawbacks, the present disclosure provides a new text processing method. The following describes Figure 3 the text processing method in this exemplary embodiment. Figure 3 The exemplary process of the image quality evaluation method is shown and may include:

[0031] Step S310: Obtain the reference unlabeled text, and input the reference unlabeled text into a pre-trained reference text processing model to obtain the feature vectors of each reference unlabeled text. Among them, the reference text processing model includes multiple encoding layers;

[0032] Step S320: Obtain the probability values output by at least one of the encoding layers, and determine multiple target unlabeled texts in the reference unlabeled text according to the probability values and the feature vectors;

[0033] Step S330: Determine the true labels corresponding to the target unlabeled texts, and update the reference text processing model using the target unlabeled texts and the true labels until the reference text processing model meets the preset conditions;

[0034] Step S340: Use the reference text processing model that meets the preset conditions to process the text to be processed to obtain a processing result.

[0035] Based on the above method, compared with the prior art, the probability values output by at least one encoding layer are used to determine the target unlabeled texts, reducing the number of manual annotations and the computational amount. Further, the true labels corresponding to the target unlabeled texts are determined, and the reference text processing model is updated using the target unlabeled texts and the true labels, making full use of the information in each encoding layer of the reference text processing model. At the same time, the reference text processing model is updated using the target unlabeled texts, reducing the complexity of model training and the computational amount.

[0036] Next, Figure 3 each step in

[0037] is specifically described. Figure 3 Referring to

[0038] In an exemplary embodiment of the present disclosure, the reference labeled text and the pre-trained reference text processing model can be obtained from a database. When obtaining the pre-trained reference text processing model, at least one initial model can be obtained first. The initial model can be BERT, GPT, XLNet, ERNIE, etc., and can also be customized according to user needs, which is not specifically limited in this exemplary embodiment.

[0039] In the present exemplary embodiment, the labeled text and the corresponding true labels of the labeled text can be obtained from the database 413, and then the initial model can be trained with the above-mentioned labeled text and the corresponding true labels of the labeled text to obtain a pre-trained reference text model.

[0040] In the present exemplary embodiment, referring to Figure 4 As shown, the overall structure of the pre-trained text processing model generally includes multiple encoding layers. For the input text, an embedding 410 operation is first performed, and then the final feature vector 470 is obtained through the cascaded 420 method of multiple encoding layers. Taking the pre-trained text processing model bert-base as an example for illustration, this model includes 12 encoding layers, and the encoding layers from the bottom to the top can often extract phrase-level, syntactic-level, and semantic-level information respectively.

[0041] In the present exemplary embodiment, the above-mentioned reference unlabeled text 4131 can be input into the above-mentioned reference text processing model to obtain the feature vectors 470 of each unlabeled text.

[0042] In step S320, the probability values output by at least one of the encoding layers are obtained, and multiple target unlabeled texts are determined in the reference unlabeled text according to the probability values and the feature vectors.

[0043] In an exemplary embodiment of the present disclosure, the probability values of at least one of the above-mentioned encoding layers can be obtained at a preset interval. Among them, the encoding layers in the above-mentioned reference text processing model can be 12 layers, 24 layers, etc., or can be customized according to user needs, and are not specifically limited in the present exemplary embodiment. Among them, the above-mentioned preset interval can be two, or four, or can be customized according to user needs, and is not specifically limited in the present exemplary embodiment.

[0044] In the present exemplary embodiment, referring to Figure 4 As shown, taking the above-mentioned reference text processing model bert-base as an example for illustration, among them, the number of encoding layers 420 can be 12 layers, and the preset interval can be 2 layers. Among them, the probability values 440 of the outputs of the third, sixth, ninth, and twelfth encoding layers can be obtained.

[0045] In the present exemplary embodiment, when obtaining the probability values output by the encoding layers, a fully connected layer and a normalization loss function 430 can be connected to the above-mentioned encoding layers to obtain the probability values of the above-mentioned respective encoding layers 430.

[0046] In the present exemplary embodiment, after obtaining the above-mentioned probability values, multiple target unlabeled texts 411 can be determined in the reference unlabeled text according to the probability values and the feature vectors. Specifically, referring to Figure 5As shown, it may include step S510 and step S520.

[0047] In step S510, mutual information 450 and voting entropy 460 of each of the reference unlabeled texts are calculated according to the respective probability values of the reference unlabeled texts.

[0048] In the present exemplary embodiment, with reference to Figure 4 and Figure 6 As shown, the mutual information of each of the above-mentioned reference text annotations 4131 can be calculated first:

[0049]

[0050] where represents the probability value that the output of the nth coding layer is the ith classification result, and N represents the number of coding layers. For example, the probability values of the outputs of the coding layers of the third, sixth, ninth, and twelfth layers are obtained. At this time, N is 4, and the value of n can be 1, 2, 3, 4. If the above classification results include 10, then the value of i can be a positive integer greater than or equal to 1 and less than or equal to 10. Among them, the classification results can include news, advertisements, etc., and i is the serial number order of the above classification results. Among them, the serial number order can be customized according to user needs., x MI represents the above-mentioned mutual information.

[0051] Then, the voting entropy 460 of each of the above-mentioned reference unlabeled texts can be calculated, and specifically, it can be calculated through the following formula:

[0052]

[0053] where V(ci) represents the number of coding layers that vote for c i , N represents the number of coding layers, and x VE represents the above-mentioned voting entropy, where c i can represent classification results, such as news, advertisements, etc., and is not specifically limited in the present exemplary embodiment.

[0054] In step S520, a plurality of target unlabeled texts are determined according to the mutual information, the voting entropy, and the respective reference unlabeled text feature vectors of each of the reference unlabeled texts.

[0055] In the present exemplary embodiment, a preset number of intermediate unlabeled texts 610 can be determined first according to the mutual information and voting entropy of the above-mentioned reference unlabeled texts; the intermediate unlabeled texts 610 are clustered according to the feature vectors to determine a plurality of target unlabeled texts 411.

[0056] Specifically, the priority order of each reference unlabeled text can be determined first according to the above-mentioned mutual information and voting entropy, and a preset number of intermediate unlabeled texts 610 can be determined from the reference unlabeled texts according to the priority order.

[0057] In the present exemplary embodiment, the above-mentioned mutual information x MI and the voting entropy can be fused by an operation 490 x VE and a priority score can be obtained x ID , x ID The calculation method is as follows.

[0058] x ID = x MI + x VE

[0059] For each of the above-mentioned reference unlabeled texts 4131, the total score is calculated according to x ID and sorted to determine the priority order, and then a preset number of intermediate unlabeled texts 610 can be selected according to the above-mentioned priority order, that is, M intermediate unlabeled texts 610 are selected, where M represents the above-mentioned preset number.

[0060] It should be noted that the value of the above-mentioned preset number is less than the number of the above-mentioned reference labeled texts. For example, if the number of the above-mentioned reference labeled texts is 10,000, the above-mentioned preset number can be 100, 200, etc. The specific preset number can be customized according to user needs and is not specifically limited in the present exemplary embodiment.

[0061] In the present exemplary embodiment, after determining the above-mentioned preset number of intermediate unlabeled texts 610, the feature vectors obtained after the above-mentioned intermediate unlabeled samples pass through the above-mentioned reference text processing model can be obtained, and then, the above-mentioned intermediate unlabeled samples are clustered according to the above-mentioned feature vectors. Among them, the k-means clustering 480 algorithm can be used to cluster the above-mentioned intermediate unlabeled texts 610 to obtain K clusters, and the central sample point of each cluster is selected as the above-mentioned target unlabeled text.

[0062] In step S330, the true label corresponding to the target unlabeled text is determined, and the reference text processing model is updated by using the target unlabeled text and the true label until the reference text processing model meets the preset conditions.

[0063] In an exemplary embodiment of the present disclosure, with reference to Figure 4, after obtaining the above-mentioned target unlabeled text 411, the true labels corresponding to the above-mentioned unlabeled texts can be determined. The acquisition of the true labels can be manually labeled 412 or obtained through other means. For example, a trained text processing model can be used to obtain them. In this exemplary embodiment, it is not specifically limited, that is, after labeling, the labeled text 4132 corresponding to the target unlabeled text is obtained.

[0064] In this exemplary embodiment, after obtaining the above-mentioned target unlabeled text and the true labels corresponding to the above-mentioned target unlabeled text, the above-mentioned reference text processing model is updated by using the target unlabeled text and the true labels corresponding to the above-mentioned target unlabeled text.

[0065] Then, steps S310 to step 330 can be executed again until the above-mentioned reference text processing model meets the preset conditions, where the above-mentioned preset conditions can be the accuracy, recall rate, and F1 score of the above-mentioned reference text processing model meeting the preset conditions. For example, the accuracy is greater than or equal to 90% and the recall rate is greater than or equal to 90%, and at the same time the F1 score is greater than or equal to 0.8; or the accuracy is greater than or equal to 90% and the recall rate is greater than or equal to 90%, and at the same time the F1 score is greater than or equal to 0.9; the above-mentioned preset conditions can also be customized according to user needs and are not specifically limited in this exemplary embodiment.

[0066] In step S340, the reference text processing model that meets the preset conditions is used to process the text to be processed to obtain a processing result.

[0067] In an exemplary embodiment of the present disclosure, after obtaining the above-mentioned reference text processing model that meets the preset conditions, the above-mentioned text to be processed can be obtained, and then the above-mentioned text to be processed is input into the above-mentioned reference text processing model that meets the preset conditions to obtain the feature vector corresponding to the above-mentioned text to be processed. Then, the processing result of the above-mentioned text brought out can be obtained by using a fully connected layer and a softmax function.

[0068] Among them, the text to be processed model can be a text classification model or other models, such as a text sentence segmentation model, which is not specifically limited in this exemplary embodiment.

[0069] In summary, in this exemplary embodiment, compared with the prior art, using the probability values output by at least one encoding layer to determine the target unlabeled text reduces the number of manual annotations and the computational amount. Further, determining the true label corresponding to the target unlabeled text and using the target unlabeled text and the true label to update the reference text processing model makes full use of the information in each encoding layer of the reference text processing model. At the same time, using the target unlabeled text to update the reference text processing model reduces the complexity of model training and the computational amount. Further, the mutual information and the voting entropy are calculated using the probability values output by each encoding layer to increase the accuracy of selecting the target unlabeled text. At the same time, the k-means clustering algorithm is used to reduce the number of target unlabeled texts, improve the representativeness of the target unlabeled texts, improve the speed of model training, and reduce the computational amount of the training model.

[0070] It should be noted that the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0071] Furthermore, referring to Figure 7 as shown, in the embodiment of this example, a text processing device 700 is further provided, including an acquisition module 710, a determination module 720, an update module 730, and a processing module 740. Among them:

[0072] The acquisition module 710 can be used to acquire reference unlabeled texts and input the reference unlabeled texts into a pre-trained reference text processing model to obtain the feature vectors of the reference unlabeled texts. Among them, the reference text processing model includes multiple encoding layers.

[0073] When acquiring the above reference text processing model, the acquisition module 710 can first acquire an initial model; acquire labeled texts and the true labels corresponding to the labeled texts; and perform training on the initial model according to the labeled texts and the true labels corresponding to the labeled texts to obtain a pre-trained reference text model

[0074] The determination module 720 can be used to acquire the probability values output by at least one encoding layer and determine multiple target unlabeled texts from the reference unlabeled texts according to the probability values and the feature vectors;

[0075] Specifically, the mutual information and the voting entropy of each reference unlabeled text can be calculated first according to the probability values of each reference unlabeled text, and then multiple target unlabeled texts can be determined according to the mutual information and the voting entropy of each reference unlabeled text and the feature vectors of each reference unlabeled text.

[0076] When determining multiple target unlabeled texts based on the mutual information and voting entropy of each reference unlabeled text and the feature vectors of each reference unlabeled text, the above-mentioned determination module 720 may first determine a preset number of intermediate unlabeled texts based on the mutual information and voting entropy of each reference unlabeled text, and then cluster the intermediate unlabeled texts according to the feature vectors to determine multiple target unlabeled texts.

[0077] In the present exemplary embodiment, when determining a preset number of intermediate unlabeled texts based on the mutual information and voting entropy of each reference unlabeled text, the above-mentioned determination module 720 determines the priority order of each reference unlabeled text according to the mutual information and voting entropy; determines a preset number of intermediate unlabeled texts from the reference unlabeled texts according to the priority order.

[0078] When obtaining the probability values output by at least one encoding layer, the determination module 720 may obtain the probability values output by at least one encoding layer at preset intervals. In an exemplary embodiment of the present disclosure, a fully connected layer and a normalization loss function may be used to convert the output of the encoding layer into probability values.

[0079] The update module 730 may be used to determine the true label corresponding to the target unlabeled text, and update the reference text processing model using the target unlabeled text and the true label until the reference text processing model meets the preset conditions.

[0080] The processing module 740 may be used to process the text to be processed using the reference text processing model that meets the preset conditions to obtain a processing result.

[0081] The specific details of each module in the above-mentioned device have been described in detail in the method part of the embodiment. The details not disclosed can be seen in the content of the embodiment in the method part, and thus will not be repeated.

[0082] Next, taking Figure 8 the mobile terminal 800 in Figure 8 as an example, the structure of the electronic device will be described exemplarily. Those skilled in the art should understand that, except for the components specifically used for mobile purposes,

[0083] such as Figure 8 shown, the mobile terminal 800 may specifically include: a processor 801, a memory 802, a bus 803, a mobile communication module 804, an antenna 1, a wireless communication module 805, an antenna 2, a display screen 806, a camera module 807, an audio module 808, a power module 809, and a sensor module 810.

[0084] The processor 801 may include one or more processing units. For example, the processor 801 may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit), etc. The text processing method in this exemplary embodiment may be executed by the AP, the GPU, or the DSP. When the method involves neural network-related processing, it may be executed by the NPU.

[0085] The encoder may encode (i.e., compress) an image or a video. For example, it may encode a target image into a specific format to reduce the data size for easy storage or transmission. The decoder may decode (i.e., decompress) the encoded data of the image or the video to restore the image or video data. For example, it may read the encoded data of the target image, decode it through the decoder to restore the data of the target image, and then perform relevant processing for text processing on this data. The mobile terminal 800 may support one or more encoders and decoders. In this way, the mobile terminal 800 can process images or videos in multiple encoding formats, such as: image formats like JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap), etc., and video formats like MPEG (Moving Picture Experts Group) 1, MPEG2, H.263, H.264, HEVC (High Efficiency Video Coding), etc.

[0086] The processor 801 may form a connection with the memory 802 or other components through the bus 803.

[0087] The memory 802 may be used to store computer-executable program code, and the executable program code includes instructions. The processor 801 executes various functional applications and data processing of the mobile terminal 800 by running the instructions stored in the memory 802. The memory 802 may also store application data, such as storing images, videos, and other files.

[0088] The communication function of the mobile terminal 800 can be implemented by the mobile communication module 804, antenna 1, wireless communication module 805, antenna 2, modulation and demodulation processor, baseband processor, etc. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module 204 can provide mobile communication solutions such as 2G, 3G, 4G, 5G, etc. applied to the mobile terminal 800. The wireless communication module 805 can provide wireless communication solutions such as wireless local area network, Bluetooth, near field communication, etc. applied to the mobile terminal 200.

[0089] The display screen 806 is used to implement the display function, such as displaying user interfaces, images, videos, etc. The camera module 807 is used to implement the shooting function, such as shooting images, videos, etc. The audio module 808 is used to implement the audio function, such as playing audio, collecting voices, etc. The power module 809 is used to implement the power management function, such as charging the battery, powering the device, monitoring the battery status, etc. The sensor module 810 may include a depth sensor 8101, a pressure sensor 8102, a gyroscope sensor 8103, a barometric pressure sensor 8104, etc. to implement corresponding sensing and detection functions.

[0090] Those skilled in the art of the present technology can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.

[0091] The exemplary embodiments of the present disclosure also provide a computer-readable storage medium, on which a program product capable of implementing the above methods of this specification is stored. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Methods" section of this specification.

[0092] It should be noted that the computer-readable medium shown in this disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0093] In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0094] In addition, the program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0095] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0096] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A text processing method, characterized in that, Including: Obtain reference unlabeled text, and input the reference unlabeled text into a pre-trained reference text processing model to obtain feature vectors of each reference unlabeled text, where the reference text processing model includes multiple encoding layers; Obtain probability values output by at least one of the encoding layers, and determine multiple target unlabeled texts in the reference unlabeled text according to the probability values and the feature vectors; Determine the true labels corresponding to the target unlabeled texts, and update the reference text processing model by using the target unlabeled texts and the true labels until the reference text processing model meets a preset condition; Process the text to be processed by using the reference text processing model that meets the preset condition to obtain a processing result; The determining multiple target unlabeled texts in the reference unlabeled text according to the probability values and the feature vectors includes: Calculate the mutual information and voting entropy of each reference unlabeled text according to the probability values of each reference unlabeled text; Determine a preset number of intermediate unlabeled texts according to the mutual information and the voting entropy of each reference unlabeled text; Cluster the intermediate unlabeled texts according to the feature vectors to determine multiple target unlabeled texts; Wherein, the mutual information of each reference unlabeled text is calculated by using the following formula Among them represents the probability value that the output of the nth coding layer is the ith classification result, and N represents the number of coding layers; x MI represents the above mutual information; The voting entropy of each reference unlabeled text is calculated by using the following formula Among them, V(ci) represents the number of coding layers that vote for c i , N represents the number of coding layers, and x VE represents the above-mentioned voting entropy, where c i can represent the classification result.

2. The method according to claim 1, wherein The determining a preset number of intermediate unlabeled texts according to the mutual information and the voting entropy of each reference unlabeled text includes: Determine the priority order of each reference unlabeled text according to the mutual information and the voting entropy; Determine a preset number of intermediate unlabeled texts in the reference unlabeled text according to the priority order.

3. The method according to claim 1, wherein The obtaining probability values output by at least one of the encoding layers includes: Obtain probability values output by at least one of the encoding layers at a preset interval.

4. The method according to claim 1, wherein The obtaining probability values output by at least one of the encoding layers includes: Convert the output of the encoding layer into the probability values by using a fully connected layer and a normalization loss function.

5. The method according to claim 1, characterized in that The method further includes: Obtain an initial model; Obtain labeled text and the true labels corresponding to the labeled text; Perform pre-training on the initial model according to the labeled text and the true labels corresponding to the labeled text to obtain a pre-trained reference text model.

6. A text processing device, characterized in that, Including: An obtaining module, configured to obtain reference unlabeled text, and input the reference unlabeled text into a pre-trained reference text processing model to obtain feature vectors of each reference unlabeled text, where the reference text processing model includes multiple encoding layers; A determining module, configured to obtain probability values output by at least one of the encoding layers, and determine multiple target unlabeled texts in the reference unlabeled text according to the probability values and the feature vectors; An updating module, configured to determine the true labels corresponding to the target unlabeled texts, and update the reference text processing model by using the target unlabeled texts and the true labels until the reference text processing model meets a preset condition; A processing module, configured to process a text to be processed by using the reference text processing model that meets a preset condition to obtain a processing result; Wherein, the determining a plurality of target unannotated texts from the reference unannotated texts according to the probability value and the feature vector includes: calculating the mutual information and the voting entropy of each of the reference unannotated texts according to the probability values of each of the reference unannotated texts; determining a preset number of intermediate unannotated texts according to the mutual information and the voting entropy of each of the reference unannotated texts; clustering the intermediate unannotated texts according to the feature vector to determine the plurality of target unannotated texts; Wherein, the mutual information of each reference unannotated text is calculated by using the following formula Among them represents the probability value that the output of the nth coding layer is the ith classification result, and N represents the number of coding layers; x MI represents the above mutual information; The voting entropy of each reference unannotated text is calculated by using the following formula Among them, V(ci) represents the number of encoding layers that vote for c i , N represents the number of encoding layers, and x VE represents the above-mentioned voting entropy, where c i can represent the classification result.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the text processing method according to any one of claims 1 to 5.

8. An electronic device, characterized in that, Comprising: One or more processors; And A memory, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the text processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Text classification method and device, equipment and storage medium

    CN113064973A

  • Financial field event extraction method and device based on lifelong learning

    CN113850064A