Data Processing Method, Apparatus, Electronic Device, and Computer-Readable Storage Medium

The method improves text classification accuracy by using a pre-defined text library to identify and correct misclassifications in online data, enhancing precision through a depth learning model.

CN114328816BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111398929.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-19
Publication Date
2025-07-15
Estimated Expiration
2041-11-19

AI Technical Summary

Technical Problem

In the prior art, the effect of identifying online data based on the trained model is not ideal enough to meet practical needs, and there are misjudgment problems.

Method used

Detect the text to be detected by the preset text library to determine whether it belongs to the predetermined misjudgment type. If so, it will be directly determined as this type. Otherwise, use the deep learning model to classify to avoid repeated recognition and realize instant online hot repair.

Benefits of technology

It improves the classification recognition accuracy of the text to be detected, meets the text recognition needs in various implementation scenarios, and provides an instant misjudgment and repair mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114328816B_ABST
    Figure CN114328816B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a data processing method, apparatus, electronic device, and computer-readable storage medium, which relate to the fields of artificial intelligence, natural language processing, and cloud technology. The method includes obtaining a text to be detected, detecting the text to be detected based on a preset text library, if it is determined that the text to be detected is a text belonging to a misjudgment type, determining the predetermined misjudgment type as the classification result of the text to be detected; if it is determined that the text to be detected does not belong to a text of the misjudgment type, determining the classification result of the text to be detected based on a deep learning model. Since the preset text library includes text data corresponding to texts marked as belonging to a predetermined misjudgment type, and texts belonging to the predetermined misjudgment type are texts for which there may be a misjudgment in the text classification result determined by the deep learning model, therefore, the solution provided by the embodiments of the present application can avoid misjudging texts to be detected belonging to the predetermined misjudgment type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence, natural language processing, and cloud technology. Specifically, this application relates to a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] With the development of science and technology, online learning has gradually become a popular research field in deep learning.

[0003] In related technologies, the online data is usually directly recognized based on a trained model to obtain the classification result of the online data. Although this method can achieve the classification processing of the online data to a certain extent, the current recognition effect is still not ideal and cannot meet the practical requirements. Summary of the Invention

[0004] Embodiments of this application provide a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can better recognize the text to be detected, avoid misjudging the text to be detected, and better meet the practical requirements.

[0005] According to one aspect of the embodiments of this application, a data processing method is provided. The method includes:

[0006] Obtain the text to be detected;

[0007] Detect the text to be detected based on a preset text library, where the preset text library includes text data corresponding to texts marked as belonging to a predetermined misjudgment type;

[0008] Determine that the text to be detected belongs to a text of the predetermined misjudgment type, and determine the predetermined misjudgment type as the classification result of the text to be detected;

[0009] Determine that the text to be detected does not belong to a text of the predetermined misjudgment type, and determine the classification result of the text to be detected based on a deep learning model;

[0010] Among them, a text belonging to the predetermined misjudgment type is a text for which there may be a misjudgment in the text classification result determined by the deep learning model.

[0011] According to another aspect of the embodiments of this application, a data processing apparatus is provided. The apparatus includes:

[0012] A text acquisition module, configured to obtain the text to be detected;

[0013] A text detection module, configured to detect the text to be detected based on a preset text library, where the preset text library includes text data corresponding to texts marked as belonging to a predetermined misjudgment type;

[0014] The classification determination module is used to determine that the text to be detected belongs to the text of a predetermined misjudgment type, and determine the predetermined misjudgment type as the classification result of the text to be detected; and,

[0015] Determine that the text to be detected does not belong to the text of the predetermined misjudgment type, and determine the classification result of the text to be detected based on the deep learning model;

[0016] Among them, the text belonging to the predetermined misjudgment type is the text whose text classification result determined by the deep learning model may be misjudged.

[0017] Optionally, the text data includes at least one of a target keyword library or a target text library. When the text detection module detects the text to be detected based on the preset text library, it is specifically used for:

[0018] Detect the text to be detected based on the text data;

[0019] Determine that the text to be detected meets at least one of the following:

[0020] Any target keyword in the target keyword library is included in the text to be detected;

[0021] The text to be detected matches any target text in the target text library;

[0022] Determine that the text to be detected belongs to the text of the predetermined misjudgment type.

[0023] Optionally, the target text library includes text vectors of at least one target text. When the text detection module determines that the text to be detected matches any target text in the target text library, it is specifically used for:

[0024] Generate a text vector corresponding to the text to be detected according to the text to be detected;

[0025] Based on the target text library, perform a similarity match of the text vectors on the text vector corresponding to the text to be detected;

[0026] Determine that the similarity between the text vector corresponding to the text to be detected and the text vector corresponding to any target text in the target text library is greater than or equal to the predetermined similarity threshold, then determine that the text to be detected matches any target text in the target text library.

[0027] Optionally, the device further includes a corpus update module and a model update module. If it is determined that the text to be detected is the text of the predetermined misjudgment type,

[0028] The corpus update module is used to store the text to be detected and the classification result of the text to be detected into the corpus to obtain an updated corpus;

[0029] A model update module, configured to update a trained deep learning model based on the updated corpus to obtain an updated deep learning model.

[0030] Optionally, the apparatus further includes a text library update module, configured to:

[0031] If feedback information indicating that the classification result of the text to be detected is a misjudgment is received, parse the text to be detected to determine keywords in the text to be detected;

[0032] Store the keywords in the text to be detected into a target keyword library to update the target keyword library;

[0033] Store the text to be detected into a target text library to update the target text library.

[0034] Optionally, the apparatus may further include a scenario information acquisition module and a fusion module,

[0035] The scenario information acquisition module is configured to acquire text scenario information of the text to be detected;

[0036] The fusion module is configured to obtain a fusion result based on the classification result of the text to be detected and the text scenario information of the text to be detected, and perform corresponding processing on the text to be detected according to the fusion result.

[0037] Optionally, the deep learning model is trained through the following method;

[0038] Obtain training samples, where the training samples include at least one sample data and the true classification result of each sample data;

[0039] Based on the training samples, perform iterative training on a first deep learning model until a preset training end condition is met, to obtain the above-mentioned deep learning model.

[0040] Optionally, performing iterative training on the first deep learning model based on the training samples until a preset training end condition is met includes:

[0041] Divide the training samples into a training set and an evaluation set according to a first preset ratio;

[0042] Train the first deep learning model according to the training set until a first training condition is met, to obtain a second deep learning model;

[0043] Evaluate the second deep learning model based on the evaluation set according to a preset evaluation metric, and when the metric evaluation result meets a second training condition, determine the second deep learning model as the deep learning model;

[0044] When the index evaluation result does not meet the second training condition, adjust the model parameters of the second deep learning model and continue to train the adjusted model based on the training set;

[0045] The training end condition includes the first training condition and the second training condition.

[0046] Optionally, the evaluation set includes at least one of the validation set or the test set, and the preset evaluation index includes at least one of the validation evaluation index or the test evaluation index;

[0047] Evaluate the second deep learning model based on the evaluation set according to the preset evaluation index, including at least one of the following:

[0048] Evaluate the second deep learning model based on the validation set according to the validation evaluation index to obtain the first evaluation result;

[0049] Evaluate the second deep learning model based on the test set according to the test evaluation index to obtain the second evaluation result;

[0050] Among them, the second training condition includes at least one of the first evaluation condition or the second evaluation condition, and the index evaluation result meeting the second training condition includes: at least one of the first evaluation result meeting the first evaluation condition or the second evaluation result meeting the second evaluation condition.

[0051] Optionally, when the model update module updates the training deep learning model based on the updated corpus to obtain the updated deep learning model, it is specifically used for:

[0052] Obtain updated training samples based on the updated corpus, where the updated training samples include at least one text to be detected and the classification results of each text to be detected;

[0053] Update the training deep learning model according to the updated training samples to obtain the updated deep learning model.

[0054] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above method.

[0055] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and the computer program implements the steps of the above method when executed by a processor.

[0056] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program, and the computer program implements the steps of the above method when executed by a processor.

[0057] The beneficial effects brought by the technical solution provided in the embodiments of the present application are as follows:

[0058] In the technical solution provided in the embodiments of the present application, since the preset text library includes text data corresponding to texts marked as a predetermined misjudgment type, and a text belonging to the predetermined misjudgment type is a text for which there may be a classification misjudgment in the text classification result determined by the deep learning model. Therefore, when it is determined based on the detection result of detecting the text to be detected against the preset text library that the text to be detected belongs to the text of the predetermined misjudgment type, the predetermined misjudgment type is directly determined as the classification result of the text to be detected; when it is determined that the text to be detected does not belong to the text of the predetermined misjudgment type, the classification result of the text to be detected is determined based on the deep learning model. Instead of, when it is determined that the text to be detected belongs to the text of the predetermined misjudgment type, classifying and identifying the text to be detected again based on the deep learning model, an instant online hotfix is realized, the recognition accuracy of classifying the text to be detected is improved, and a premise guarantee is provided for better meeting the text recognition requirements of the product in various implementation scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments of the present application.

[0060] Figure 1 The structural schematic diagram of an optional data processing system in this scenario is shown;

[0061] Figure 2 The flowchart of the data processing method in this application scenario is shown;

[0062] Figure 3 The flowchart of the data processing method provided in the embodiments of the present application is shown;

[0063] Figure 4 The schematic diagram showing the training process of obtaining the deep learning model in the embodiments of the present application is shown;

[0064] Figures 5a to 5e The schematic diagram of a specific application scenario of the present application is shown;

[0065] Figure 6 The schematic diagram of the data processing device provided in the embodiments of the present application is shown;

[0066] Figure 7 The structural schematic diagram of the electronic device provided in this optional embodiment is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions of the embodiments of the present application.

[0068] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components and / or their combinations supported by the art of the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" indicates being implemented as "A", or being implemented as "A", or being implemented as "A and B".

[0069] To make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0070] Bad case: A term in the field of algorithms, used to represent a result different from the expected result generated during the inference stage. For example, in a text classification task, text A should be classified as a positive class, but the algorithm classifies it as a negative class, and this is a bad case. In the embodiments of the present application, a bad case means that if the label of a certain piece of text does not conform to the expected label, then the text is a bad case.

[0071] In the related art, usually, the online data is directly identified according to the trained model to obtain the classification result of the online data. However, the effect of directly identifying the online data in this way is still not ideal enough to meet the practical requirements.

[0072] In view of the above problems, the present application provides a data processing method. Since the preset text library includes text data corresponding to texts marked as a predetermined misjudgment type, and texts belonging to the predetermined misjudgment type are texts for which there may be a misclassification in the text classification result determined by a deep learning model, first, the obtained text to be detected is detected based on the preset text library to determine whether the text to be detected belongs to the text of the predetermined misjudgment type (i.e., bad case). Then, based on the detection result, when it is determined that the text to be detected belongs to the text of the predetermined misjudgment type, the predetermined misjudgment type is determined as the classification result of the text to be measured. When the text to be detected does not belong to the text of the predetermined misjudgment type, the classification result of the text to be detected is determined based on the deep learning model. Instead of, when it is determined that the text to be detected belongs to the text of the predetermined misjudgment type, classifying and identifying the text to be detected again through the deep learning model, real-time online hotfix is realized, the recognition accuracy of classifying and identifying the text to be detected is improved, and a prerequisite guarantee is provided for better meeting the text recognition requirements of products in various implementation scenarios.

[0073] Optionally, the data processing method provided by the embodiments of the present application can be implemented based on artificial intelligence (AI) technology. AI is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. With the research and progress of artificial intelligence technology, artificial intelligence technology has been widely studied and applied in multiple fields. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0074] Optionally, the data processing method provided by the embodiments of the present application can be implemented based on natural language processing (NLP) technology. NLP is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language used by people in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0075] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0076] Optionally, the data processing method provided in the embodiments of the present application can be implemented based on cloud technology. For example, the data calculation involved in the process of updating and training a deep learning model can adopt cloud computing. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, which can form a resource pool, be used as needed, and be flexible and convenient. Cloud computing technology will become an important support. Cloud computing refers to the delivery and use mode of IT infrastructure, which means obtaining the required resources in a on-demand and easily expandable manner through the network; in a broad sense, cloud computing refers to the delivery and use mode of services, which means obtaining the required services in a on-demand and easily expandable manner through the network. Such services can be related to IT and software, the Internet, or other services. With the development of the Internet, real-time data streams, diverse connected devices, and the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, cloud computing has developed rapidly. Different from the previous parallel distributed computing, the emergence of cloud computing will drive a revolutionary change in the entire Internet model and enterprise management model conceptually.

[0077] To facilitate the understanding of the application value of the data processing method provided in the embodiments of the present application, the data processing method will be described below in combination with a specific application scenario example.

[0078] Figure 1 shows a schematic structural diagram of an optional data processing system in this scenario, as Figure 1As shown in the figure, the system includes the user's terminal device 11, a network 12, an application server 13, and a model training server 14. The terminal device 11 communicates with the application server 13 through the network 12, and interaction can be achieved between the application server 13 and the model training server 14. For example, the application server 13 can receive the deep learning model / updated deep learning model sent by the model training server 14. Among them, an application program for data processing can be installed in the terminal device 11, or a plug-in for data processing can be set in a certain application program in the terminal device 11. By opening the application program for data processing or the application program with the above-mentioned plug-in for data processing, the terminal is started to perform the above data processing method. Among them, the terminal device 11 can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, or a wearable device, etc. The data processing method can also be implemented by a processor calling computer-readable instructions stored in a memory.

[0079] Among them, the model training server 14 can be used to train the deep learning model based on the training text to obtain the deep learning model, and update and train the deep learning model based on the updated training text to obtain the updated deep learning model, and send the deep learning model and the updated deep learning model to the application server 13. After receiving the deep learning model and the updated deep learning model, the application server 13 can deploy the deep learning model and the updated deep learning model to execute the data processing method provided in the embodiments of the present application, detect the text to be detected based on a preset text library, and directly determine the classification result of the text to be detected based on the detection result, or detect the text to be detected according to the deep learning model (or the updated deep learning model), determine the classification result of the text to be detected, and perform subsequent processing according to the classification result of the text to be detected. Among them, the preset text library includes texts marked as belonging to a predetermined misjudgment type. For example, the system can be applied to mobile instant messaging to detect the text published or about to be published in the shared area of instant messaging. After obtaining the classification result of the text to be detected, operations such as restricting sending, stopping sending, issuing a warning, blocking an account, and canceling an account are performed according to the classification result of the text to be detected.

[0080] The following combines Figure 1 the data processing system shown in the figure to illustrate the data processing method in this application scenario. Figure 2 The flowchart of the data processing method in this application scenario is shown. As Figure 2As shown, the method may include the following steps S21 to S26.

[0081] Step S21: Obtain the text to be detected.

[0082] Step S22: Detect the text to be detected according to a preset text library, where the preset text library includes text data corresponding to texts marked as belonging to a predetermined misjudgment type, and the text data includes a target keyword library (which can also be called keyword detection when detecting the text to be detected according to the keyword library) and a target text library (which can also be called text similarity detection when performing text vector similarity matching on the text vector corresponding to the text to be detected according to the text library).

[0083] Step S23: If it is determined according to the preset text library that the text to be detected is a text belonging to the predetermined misjudgment type, then determine the predetermined misjudgment type as the classification result of the text to be detected;

[0084] If it is determined according to the preset text library that the text to be detected is not a text belonging to the predetermined misjudgment type, then determine the classification result of the text to be detected based on a deep learning model.

[0085] Step S24: When it is determined that the text to be detected is a text belonging to the predetermined misjudgment type, store the text to be detected and the classification result of the text to be detected in a corpus (that is, Figure 2 the corpus management system in Figure 2 ), and perform a first mixing operation on the text to be detected and the classification result of the text to be detected (that is, Figure 2 the new corpus and the cumulative corpus in Figure 2 ), as well as at least one preset text and the classification result of each preset text (that is, Figure 2 the historical corpus in

[0086] Among them, during the process of updating and training a deep learning model, the updated training samples can be split into an updated training set, an updated validation set, and an updated test set, and the deep learning model can be trained based on the updated training set to obtain an updated second deep learning model; based on the updated second deep learning model and the updated validation set, the hyperparameters of the updated second deep learning model are adjusted and selected to obtain a second deep learning model that meets the validation metrics; based on the deep learning model that meets the validation metrics and the updated test set, the generalization ability of the deep learning model that meets the validation metrics is evaluated, and the deep learning model that meets the evaluation metrics is determined as the updated deep learning model.

[0087] Step S25: Obtain other features of the text scene information including the text to be detected.

[0088] Specifically, obtain at least one of the following text scene information of the text to be detected:

[0089] Sharing information; Discussion information; Interaction information; Statement information; Explanation information; Interactive information.

[0090] Step S26: Based on the classification result of the text to be detected and the text scene information of the text to be detected, obtain a fusion result, so as to perform corresponding processing on the text to be detected according to the fusion result.

[0091] Specifically, based on the classification result of the text to be detected and the text scene information of the text to be detected, information fusion is performed through a predetermined fusion method to obtain a fusion result, so as to perform corresponding business processing matching the fusion result on the text to be detected according to the fusion result.

[0092] Among them, the optional implementation methods of the predetermined fusion method include: extracting the classification features of the classification result of the text to be detected and the scene features of the text scene information of the text to be detected; based on the classification features and the scene features, generating fusion features, and based on the fusion features, determining the fusion result.

[0093] Optionally, a new classification result can be determined based on the fusion features, and this new classification result can be used as the above-mentioned fusion result, that is, the final classification result of the text to be detected.

[0094] As another optional method, it can be to obtain the classification result of the text to be detected through a deep learning model, and determine another classification result based on the scene features of the text scene information of the text to be detected, and determine the final classification result of the text to be detected according to the classification result of the text to be detected and this another classification result.

[0095] Among them, the text scenario information includes, but is not limited to, any sharing information (which can be information published on a sharing platform, information published in the shared space of an instant messaging tool, information published on a PGC (professionally generated content) platform, information published on a UGC (user-generated content) platform, etc.), discussion information (which can be comments on the information published in PGC, comments on the information published in the shared platform of an instant messaging tool, discussions on a certain topic on a UGC platform, etc.), interactive information (which can be information during the interaction process of an instant messaging tool, such as text, expressions, animated pictures, voice, etc.), statement or explanation information (which can be a published novel or article), etc., and interactive information for any sharing information, discussion information, interactive information, statement or explanation information, etc. Among them, the interactive information can include all interactive information for the above text scenario information, such as the publishing frequency, the number of likes received, the number of comments received, etc. Optionally, the text scenario information can include the historical text scenario information corresponding to the text to be detected.

[0096] Optionally, the processing of the downstream service according to the fusion result may include: performing corresponding processing on the text to be detected (which may include, but is not limited to, at least one of the following operations: prompt information such as restricting sending, stopping sending, canceling sending, etc.), and performing corresponding processing on the account that sends the text to be detected (which may include, but is not limited to, at least one of the following operations: issuing a warning, blocking the account, canceling the account, etc.). For example, if it is determined that the fusion result obtained by fusing the classification result of the text to be detected with other features further determines that the current text to be detected is an unsendable text, a prompt message of "stop sending" can be sent to the terminal device that publishes the text to be detected through the corresponding background server.

[0097] Among them, step S21, step S22, step S23, step S25 and step S26 can form the deployment end of the data processing system, and step S24 can form the training end of the data processing system. Step S22, step S23 and step S26 can be implemented by application server 14, step S21 and step S25 can be implemented through terminal device 11, and step S24 can be implemented through model training server 13.

[0098] In the data processing method provided by the embodiments of the present application, since the preset text library includes text data corresponding to texts marked as a predetermined misjudgment type, and texts belonging to the predetermined misjudgment type are texts with a possible classification misjudgment in text classification determined by a deep learning model. Therefore, based on the preset text library, the text to be detected is detected to determine whether the text to be detected belongs to the text of the predetermined misjudgment type. Based on the result of detecting the text to be detected using the preset text library, when it is determined that the text to be detected belongs to the text of the predetermined misjudgment type, the predetermined misjudgment type is directly determined as the classification result of the text to be measured. When the text to be detected does not belong to the text of the predetermined misjudgment type, based on the deep learning model, the classification result of the text to be detected is determined. Instead of having to identify the text to be detected again based on the deep learning model when it is determined that the text to be detected belongs to the text of the predetermined misjudgment type, instant online hotfix is achieved, improving the recognition accuracy of classifying the text to be detected, and providing a prerequisite guarantee for better meeting the text recognition requirements of the product in various implementation scenarios.

[0099] By obtaining a fusion result based on the classification result of the text to be detected and the text scenario information of the text to be detected, and performing corresponding processing on the text to be detected according to the fusion result, some behaviors related to bad information can be stopped, and the language environment can be maintained.

[0100] Moreover, by splitting the updated training samples, the deep learning model can be updated and iteratively trained more precisely, improving the efficiency of updating and iterating the deep learning model, continuously optimizing the deep learning model, as well as the coverage rate and precision rate of evaluating the deep learning model. And by continuously updating and iterating the deep learning model, the closed-loop processing of the system can be completed. When it is determined that the text to be detected does not belong to the text of the predetermined misjudgment type, the classification result of the text to be detected can be determined more precisely according to the updated deep learning model.

[0101] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below by describing several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.

[0102] Figure 3The flowchart of the data processing method provided by the embodiments of the present application is shown. The execution subject of the data processing method may be a data processing device. Optionally, the data processing device may include, but is not limited to, a terminal device or a server. Optionally, the server may be a cloud server. Among them, the terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, or a wearable device, etc. The data processing method may also be implemented by a processor calling computer-readable instructions stored in a memory.

[0103] Optionally, the method may be executed by a user terminal. For example, the user terminal includes, but is not limited to, a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, a wearable electronic device, an AR / VR device, etc.

[0104] As Figure 3 shown, the present application provides a data processing method, and the method includes the following steps S31 to S33.

[0105] Step S31: Obtain the text to be detected.

[0106] Optionally, the text to be detected may be any text, and the present application does not limit this. For example, the text to be detected may be an online text (such as real-time comments during a live broadcast, a circle of friends / Weibo / log to be sent, etc.), and the text to be detected may also be an offline text (such as the text in any document or picture). Based on this, the data processing method proposed by the present application can be applied to scenarios with strong real-time performance such as video live broadcast (for example, filtering live comments in real time during a live broadcast), and can also be applied to other scenarios.

[0107] Among them, the present application does not limit the language form of the text to be detected. Among them, the language form may include different languages, different character combinations, etc. For example, the text to be detected may be a text combined with at least one language such as Chinese, English, Spanish, etc., and the text to be detected may also be a text combined with at least one character combination such as Chinese characters, letters, numbers, etc.

[0108] Step S32: Detect the text to be detected based on a preset text library, where the preset text library includes text data corresponding to texts marked as belonging to a predetermined misjudgment type.

[0109] Optionally, the present application also places no restrictions on the language form of the text marked as belonging to a predetermined misjudgment type in the preset text library. For example, the text marked as belonging to a predetermined misjudgment type in the preset text library can be text in at least one language combination such as Chinese, English, Spanish, etc., and the text in the preset text library can also be text in at least one character combination such as Chinese characters, letters, numbers, etc.

[0110] Among them, when detecting the text to be detected based on the preset text library, it can be determined whether the text to be detected is a text of a predetermined misjudgment type according to whether the text to be detected includes the text data corresponding to the text marked as belonging to a predetermined misjudgment type, and whether the similarity between the text to be detected and the text data corresponding to the text marked as a predetermined misjudgment type exceeds a certain threshold.

[0111] Step S33: Determine that the text to be detected is a text of a predetermined misjudgment type, and determine this predetermined misjudgment type as the classification result of the text to be detected;

[0112] Determine that the text to be detected is not a text of a predetermined misjudgment type, and determine the classification result of the text to be detected based on the deep learning model;

[0113] Among them, the text of a predetermined misjudgment type is a text for which there may be a misjudgment in the text classification result determined by the deep learning model.

[0114] Optionally, the predetermined misjudgment type includes at least one type, which can be determined according to the actual situation, and the present application places no restrictions on this. For example, the predetermined misjudgment type can include the type to which the text that is misjudged by the deep learning model as an abnormal text from a normal text belongs, or can also include the type to which the text that is misjudged by the deep learning model as a normal text from an abnormal text belongs.

[0115] Optionally, the classification result of the text to be detected, that is, the text type of the text to be detected, the type to which the text to be detected belongs, can be represented by a label. Among them, the present application places no restrictions on the form of the label. For example, it can be represented by letters, numbers, combinations of letters and numbers, etc. For example, it can be represented by 0 for brushing orders text, 1 for weather text, 2 for mood text, etc., and the present application places no restrictions on this.

[0116] Optionally, the deep learning model may be obtained by training a first deep learning model based on a training data set containing a large number of training samples. Herein, the specific network structure of the first deep learning model is not limited in the embodiments of the present application and may be configured according to actual requirements. Optionally, the first deep learning model may be a model based on a convolutional neural network, and may include, but is not limited to, a neural network model based on model structures such as CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), and S-ANN (Self-Attention Neural Network). Wherein, the input of the first deep learning model may be a piece of text or text data obtained by vectorizing a piece of text, and the output may be a text type and / or the confidence level corresponding to each text type, completing the mapping of text data - text classification. That is, by inputting the text to be detected into the first deep learning model, the text type of the text to be detected or the confidence level corresponding to each text type can be obtained. Optionally, the text type with the highest confidence level or the text with a confidence level exceeding the confidence level threshold among the confidence levels corresponding to each text type may be determined as the text type corresponding to the text to be detected. Among them, the higher the confidence level, the more accurate the obtained text type. Optionally, the confidence level threshold may be configured according to actual requirements (such as an empirical value or an experimental value), and the present application does not limit this. For example, the confidence level threshold may be set to 0.75.

[0117] When calling the deep learning model to classify text, the situations where classification misjudgment may occur include, but are not limited to, the following situations:

[0118] Situation 1: For some texts to be detected that are sensitive texts in special fields (such as fields with a relatively high confidentiality level), when classifying and processing them by calling the above deep learning model, the text to be detected may be identified as a normal text, resulting in misjudgment of the text to be detected.

[0119] Situation 2: With the development of science and technology, during the dissemination of text, the forms of some texts may be processed multiple times, resulting in a large difference between the text to be detected and the training texts used to train the first deep learning model to obtain the deep learning model. Therefore, when classifying and processing by calling the above deep learning model, the text to be detected may not be correctly identified, resulting in misjudgment of the classification result of the text to be detected. For example, when the training sample is "Please find me for Taobao order brushing", and after being processed multiple times during the dissemination of the text, the obtained text to be detected may be " "Please find me for Taobao order brushing", then by calling the deep learning model trained according to the training text "Please find me for Taobao order brushing", it may not be possible to determine the classification result of the text to be detected " "Please find me for Taobao order brushing", resulting in misjudgment of the text to be detected.

[0120] Case 3: Due to the presence of some special characters in the text to be detected (for example, "mouth" as a unit of measurement), when classifying through the above deep learning model, the text to be detected may be recognized as an abnormal text, resulting in misjudgment of the text to be detected. For example, when calling the above deep learning model to classify the text to be detected "I want a 24-port switch", the text to be detected may be recognized as pornographic text, resulting in misjudgment of the text to be detected.

[0121] Based on the above situation, and because it takes a certain amount of time to re-update the deep learning model when the deep learning model is relatively complex, therefore, by parsing the text that may have misclassification during text classification, the text that may have misclassification and the type of the text that may have misclassification can be determined, and the text that may have misclassification and the type of the text that may have misclassification are stored in a preset text library, so as to detect the text to be detected according to the preset text library to determine whether the text to be detected belongs to the text of the predetermined misjudgment type. When it is determined that the text to be detected belongs to the text of the predetermined misjudgment type, the predetermined misjudgment type is determined as the classification result of the text to be detected, rather than determining the classification result of the text to be detected through the above deep learning model, thereby avoiding misjudgment of the text to be detected by the above deep learning model.

[0122] Among them, the preset text library can be classified according to the type of text misjudged by the deep learning model. For example, the preset text library can be divided into a white text library and a black text library. Among them, the white text library can represent a text library formed by texts that were originally normal texts but were misjudged as abnormal texts by the deep learning model (for example, the text in Case 3 above), and the black text library can represent a text library formed by texts that were originally abnormal texts but were misjudged as normal texts by the deep learning model (for example, the text in Case 1 above).

[0123] Optionally, it is possible to determine the recognition of the text that may have misclassification based on artificial intelligence technology, parse the text that may have misclassification, obtain the type of the text that may have misclassification, and determine the type of the text that may have misclassification as the predetermined misjudgment type. It is also possible to determine the predetermined misjudgment type by relevant technical personnel. The present application does not limit the method for determining the predetermined misjudgment type.

[0124] In addition, with the development of science and technology and the diversification of communication channels, the forms of texts have gradually become diversified. Based on this, texts in the preset text library can also be added / deleted according to actual needs.

[0125] The technical solution provided by the embodiment of the present application obtains the text to be detected, detects the text to be detected based on the preset text library, and determines whether the text to be detected belongs to the text of a predetermined misjudgment type. Since the preset text library includes the text data corresponding to the text marked as the predetermined misjudgment type, and the text belonging to the predetermined misjudgment type is the text whose text classification result determined by the deep learning model may have a classification misjudgment, therefore, when it is determined that the text to be detected belongs to the text of the predetermined misjudgment type based on the detection result of the text to be detected based on the preset text library, the predetermined misjudgment type is directly determined as the classification result of the text to be measured. When the text to be detected does not belong to the text of the predetermined misjudgment type, the classification result of the text to be detected is determined based on the deep learning model. Instead of, when it is determined that the text to be detected belongs to the text of the predetermined misjudgment type, classifying and identifying the text to be detected again based on the deep learning model, which avoids misjudgment when the text to be detected belonging to the predetermined misjudgment type is classified and identified by the deep learning model, realizes instant online hotfix, improves the recognition accuracy of classifying the text to be detected, and provides a prerequisite guarantee for better meeting the text recognition requirements of the product in various implementation scenarios.

[0126] In a possible implementation manner, the above deep learning model is trained in the following way;

[0127] Obtain training samples, where the training samples include at least one sample data and the true classification result of each sample data;

[0128] Based on the training samples, perform iterative training on the first deep learning model until the preset training end condition is met, and obtain the above deep learning model.

[0129] Optionally, the sample data may include at least one of text or keywords, that is, the form of the sample data is not restricted, so that the diversification of the training samples can be realized. Furthermore, the obtained deep learning model can also be widely applied.

[0130] Among them, the true classification result of each sample data may include one of multiple preset classification results. By inputting the sample data into the first deep learning model, the output result may be the predicted classification result corresponding to the sample data, or the confidence levels of the sample data corresponding to each preset classification result. Among them, in the case where the output result is the confidence levels of the sample data corresponding to each preset classification result, the preset classification result with the highest confidence level or the preset classification results with confidence levels exceeding the confidence level threshold among the confidence levels of each preset classification result may be determined as the predicted classification result corresponding to the text to be detected. Among them, the higher the confidence level, the more accurate the obtained predicted classification result. Optionally, the confidence level threshold can be configured according to actual needs (such as an empirical value or an experimental value), and this application does not limit this. For example, the confidence level threshold can be set to 0.75.

[0131] It can be understood that multiple sample data and the true classification results of each sample data can be stored in a corpus in advance, and the above training samples are determined based on this corpus. The training samples may include at least one sample data in the corpus and the true classification results of each sample data, and specifically, the number of sample data in the training samples can be determined according to actual needs.

[0132] Optionally, different training samples can be determined according to different training tasks to obtain a more accurate deep learning model. For example, training texts identical to the text scenario information can be selected according to different text scenario information, and the first deep learning model is trained to obtain the above deep learning model.

[0133] Optionally, the preset training end condition can be configured according to requirements, and may include but is not limited to the convergence of the loss function, the value of the loss function being less than a set value, or the number of training times reaching a set number. Among them, the smaller the set value, the higher the accuracy of the obtained deep learning model.

[0134] Through the above training method, the obtained deep learning model can accurately identify the classification result of the text to be detected.

[0135] Optionally, based on the training samples, the first deep learning model is iteratively trained until the preset training end condition is met, including:

[0136] The training samples are split into a training set and an evaluation set according to a first preset ratio;

[0137] The first deep learning model is trained according to the training set until the first training condition is met to obtain a second deep learning model;

[0138] Evaluate the second deep learning model based on the evaluation dataset according to the preset evaluation metrics. When the index evaluation result meets the second training condition, determine the second deep learning model as the deep learning model;

[0139] When the index evaluation result does not meet the second training condition, adjust the model parameters of the second deep learning model and continue to train the adjusted model based on the training set;

[0140] The training end conditions include the first training condition and the second training condition.

[0141] In this implementation, the first preset ratio can be configured according to actual needs (such as an empirical value or an experimental value), and this application does not limit it. For example, the first preset ratio can be set to 6:4, that is, after splitting, the ratio of the number of training samples in the training set to the number of training samples in the evaluation set is 6:4.

[0142] Optionally, the first training end condition can be configured according to requirements, and can include but are not limited to the convergence of the loss function, the value of the loss function being less than a set value, or the number of training times reaching a set number. Among them, the smaller the set value, the higher the accuracy of the obtained deep learning model.

[0143] Among them, the preset evaluation metrics can include but are not limited to the generalization ability, hyperparameters, etc. of the second deep learning model. The corresponding second training condition can include that the generalization ability of the second deep learning model reaches a certain value, and this value can be an evaluation metric configured according to actual needs. For example, the second training condition can be set such that the generalization ability of the second deep learning model can reach 87%, that is, the second deep learning model can be applied to 87% of the classification and recognition tasks in the same field.

[0144] Optionally, the model parameters of the second deep learning model can include the hyperparameters of the second deep learning model. For example, the hyperparameters include but are not limited to the learning rate, the number of iterations, the number of layers of each network in the model, etc.

[0145] By splitting the training samples, after training the first deep learning model based on the training set to obtain a second deep learning model that meets the first training condition, then evaluate the second deep learning model based on the evaluation dataset according to the preset evaluation metrics, and determine the second deep learning model whose index evaluation result meets the second training condition as the deep learning model, an accurate deep learning model with good generalization ability can be obtained, improving the coverage rate and accuracy rate of the deep learning model.

[0146] Optionally, the evaluation set includes at least one of the validation set and the test set, and the preset evaluation metrics include at least one of the validation evaluation metrics and the test evaluation metrics;

[0147] Evaluating the second deep learning model based on the evaluation dataset according to the preset evaluation metrics includes at least one of the following:

[0148] Evaluating the second deep learning model based on the validation set according to the validation evaluation metrics to obtain a first evaluation result;

[0149] Evaluating the second deep learning model based on the test set according to the test evaluation metrics to obtain a second evaluation result;

[0150] Among them, the second training condition includes at least one of the first evaluation condition or the second evaluation condition, and the index evaluation result satisfying the second training condition includes: the first evaluation result satisfies the first evaluation condition or the second evaluation result satisfies the second evaluation condition at least one of.

[0151] In this implementation, when the evaluation set includes the validation set and the test set, the evaluation set can be divided into the validation set and the test set according to the second preset ratio, where the second preset ratio can be configured according to actual needs (such as an empirical value or an experimental value), and this application does not limit this. For example, the second preset ratio can be set to 3:1, that is, after splitting, the ratio of the number of training samples in the validation set to the number of training samples in the test set is 3:1.

[0152] In this implementation, the validation evaluation metrics can include but are not limited to the hyperparameters of the second deep learning model. Among them, the first evaluation condition can be configured according to actual needs (such as an empirical value or an experimental value), and this application does not limit this. For example, the first evaluation condition can be that the learning rate of the second deep learning model reaches 90%. Among them, when the second deep learning model does not meet the first evaluation condition, the hyperparameters of the second deep learning model can be adjusted, the second deep learning model can be optimized, and the second deep learning model can be continuously trained according to the training set and the validation set until the second deep learning model meets the first evaluation condition.

[0153] Optionally, the test evaluation metrics can include but are not limited to the generalization ability of the second deep learning model. Among them, the second evaluation condition can be configured according to actual needs (such as an empirical value or an experimental value), and this application does not limit this. For example, the second evaluation condition can be that the generalization ability of the second deep learning model reaches 90%, that is, the generalization ability of the second deep learning model that meets the second evaluation condition is greater than or equal to 90%. When the second evaluation result does not meet the second evaluation condition, continue to train the second deep learning model according to the training set and the validation set until a second deep learning model that meets the second evaluation condition is obtained, so that the final performance of the obtained deep learning model is good.

[0154] It can be understood that the final performance of the deep learning model obtained when the second training condition includes the first evaluation condition and the second evaluation condition is better than the final performance of the deep learning model obtained when the second training condition only includes the first evaluation condition or only includes the second evaluation condition.

[0155] Through the above method, the accuracy of the deep learning model can be better, the generalization ability can be stronger, and it has a higher coverage rate and precision rate.

[0156] Figure 4 A schematic diagram showing the training process of obtaining the deep learning model in the embodiment of the present application. As Figure 4 shown, in this implementation, the first preset ratio is 6:4, and the second preset ratio is 3:1. That is, after splitting the training samples, the ratio of the number of training samples in the training set, the number of training samples in the validation set, and the number of training samples in the test set is 6:3:1. The specific training process is as follows:

[0157] Train the first deep learning model according to the training set until the first training condition is met, and obtain the second deep learning model;

[0158] Evaluate the second deep learning model based on the validation set according to the validation evaluation index, and obtain the first evaluation result;

[0159] When the first evaluation result does not meet the first evaluation condition, adjust the model parameters of the second deep learning model, and continue to train the adjusted model based on the training set until the first evaluation result meets the first evaluation condition;

[0160] When the first evaluation result meets the first evaluation condition, evaluate the second deep learning model based on the test set according to the test evaluation index, and obtain the second evaluation result;

[0161] When the second evaluation result meets the second evaluation condition, determine that the index evaluation result meets the second training condition, and determine the second deep learning model as the deep learning model;

[0162] When the second evaluation result does not meet the first evaluation condition, adjust the model parameters of the second deep learning model, and continue to train the adjusted model based on the training set until the second evaluation result meets the second training condition.

[0163] Optionally, the method further includes:

[0164] Obtain the text scene information of the text to be detected;

[0165] Based on the classification result of the text to be detected and the text scene information of the text to be detected, obtain a fusion result, so as to perform corresponding processing on the text to be detected according to the fusion result.

[0166] As an example, when the business requirement is to identify and manage malicious text for brushing orders posted in the Moments, the tags corresponding to this business requirement can be set to two tags: normal and brushing orders. Different classification results can be represented by numbers, letters, or combinations of numbers and letters. For example, 0 can be used to represent normal and 1 can be used to represent brushing orders. Set the target text as "□ Women's clothing order, as long as the female account is not downgraded, with 3". Then, the above method can be used to determine whether the text to be detected is malicious text for brushing orders. When it is determined that the text to be detected is malicious text for brushing orders, and the Moments posts corresponding to the text to be detected are frequent (when the number of Moments posts exceeds the threshold, it can be determined that the Moments posts are frequent), but the number of likes and comments for each Moments post is very small, then the sending of the text to be detected can be restricted.

[0167] By obtaining a fusion result based on the classification result of the text to be detected and the text scenario information of the text to be detected, and performing corresponding processing on the text to be detected according to the fusion result, some behaviors related to bad information can be stopped and the language environment can be maintained.

[0168] Optionally, the text data can also be one of including a target keyword library or a target text library. Detecting the text to be detected based on a preset text library includes:

[0169] Detecting the text to be detected based on the text data;

[0170] Determine that the text to be detected meets at least one of the following:

[0171] The text to be detected includes any target keyword in the target keyword library;

[0172] The text to be detected matches any target text in the target text library;

[0173] Determine the text to be detected as text belonging to a predetermined misjudgment type.

[0174] Optionally, when detecting keywords and matching texts for the text to be detected based on the target keyword library and the target text library, and when it is determined that the text to be detected belongs to text of a predetermined misjudgment type, the predetermined misjudgment type corresponding to the target keyword included in the text to be detected, and / or, the predetermined misjudgment type corresponding to the target text whose similarity to the target text corresponding to the text to be detected is greater than or equal to the predetermined similarity threshold can be determined as the classification result of the text to be detected. Among them, when the predetermined misjudgment type corresponding to the target keyword is consistent with the predetermined misjudgment type corresponding to the target text, directly determine the predetermined misjudgment type corresponding to the target keyword or the predetermined misjudgment type corresponding to the target text as the classification result of the text to be detected.

[0175] In the case where the predetermined misjudgment type corresponding to the target keyword is inconsistent with the predetermined misjudgment type corresponding to the target text, the predetermined misjudgment type corresponding to the target keyword or the predetermined misjudgment type corresponding to the target text can be determined as the classification result of the text to be detected according to the preset priority between the two. As several examples, it may include: ① The predetermined misjudgment type corresponding to the target keyword and the predetermined misjudgment type corresponding to the target text can both be determined as the classification result of the text to be detected; ② When the number of target keywords included in the text to be detected is greater than or equal to a first value, the predetermined misjudgment type corresponding to the target keyword is determined as the classification result of the text to be detected; ③ When the similarity between the target text and the text to be detected is greater than or equal to a second value, the predetermined misjudgment type corresponding to the target text is directly determined as the classification result of the text to be detected. Optionally, both the first value and the second value can be configured according to actual needs (such as empirical values or experimental values), and the present application does not limit this. For example, the first value can be set to 5 and the second value can be set to 0.8.

[0176] Optionally, the preset text library may further include other databases in addition to the above-mentioned target keyword library and target text library, as long as through this database, the text to be detected can be correspondingly detected to determine whether the text to be detected belongs to the text of the predetermined misjudgment type.

[0177] By determining that the text to be detected belongs to the text of the predetermined misjudgment type in any case of determining that the text to be detected includes any target keyword and the text to be detected matches any target text in the target text library, the text to be detected can be detected more precisely, so that the predetermined misjudgment type corresponding to the target keyword and / or target text is determined as the classification result of the text to be detected, rather than determining the classification result of the text to be detected through a deep learning model, which can avoid misjudgment when the text to be detected belonging to the predetermined misjudgment type is classified and recognized by the deep learning model again, and realize instant online hotfix.

[0178] Optionally, when the text data includes a target keyword library, detecting the text to be detected based on the preset text library includes:

[0179] Detecting keywords of the text to be detected based on the target keyword library;

[0180] Determine that the text to be detected includes any target keyword, and determine the text to be detected as the text of the predetermined misjudgment type, where the target keyword is a keyword belonging to the predetermined misjudgment type.

[0181] Optionally, the present application does not limit the language form of the target keyword either, which can be determined according to the actual situation. Among them, the target keyword can be determined according to different languages and different characters. For example, the target keyword can be a word composed of at least one language combination such as Chinese, English, Spanish, etc., and the target keyword can also be a word composed of at least one character combination such as Chinese characters, letters, numbers, etc. As an example, when the predetermined misjudgment type includes "brush order", a target keyword can be set as "brush dan". Then, when detecting the keyword of the text to be detected based on the keyword library, the text " Please find me for brush dan" belonging to the text of the predetermined misjudgment type can be accurately detected.

[0182] Optionally, the present application does not limit the number of target keywords either, and the number of target keywords can be determined according to the actual situation. When it is determined that any one of the target keywords is included in the text to be detected, it can be determined that the text to be detected belongs to the text of the predetermined misjudgment type.

[0183] Based on the above introduction of the preset text library, the target keyword library can also be divided into a white target keyword library and a black target keyword library, and the target keyword library can also be updated according to the actual situation, etc., which will not be elaborated here.

[0184] Among them, when detecting the keyword of the text to be detected based on the target keyword library and determining that the text to be detected belongs to the text of the predetermined misjudgment type, the predetermined misjudgment type corresponding to the target keyword included in the text to be detected can be determined as the classification result of the text to be detected. Among them, when at least two target keywords are included in the text to be detected, the predetermined misjudgment types corresponding to each target keyword included in the text to be detected can be combined, and the combined predetermined misjudgment type can be determined as the classification result of the text to be detected.

[0185] By setting up the target keyword library, detecting the text to be detected based on the target keyword library, determining whether the text to be detected belongs to the text of the predetermined misjudgment type, and determining that the text to be detected belongs to the text of the predetermined misjudgment type when it is determined that any one of the target keywords is included in the text to be detected, the text to be detected can be detected more precisely, so that the predetermined misjudgment type corresponding to the target keyword can be determined as the classification result of the text to be detected, without having to determine the classification result of the text to be detected through the deep learning model again, avoiding misjudgment when the text to be detected belonging to the predetermined misjudgment type is classified and recognized by the deep learning model, and realizing instant online hotfix.

[0186] Optionally, the target text library includes text vectors of at least one target text. Determining that the text to be detected matches any one of the target texts in the target text library includes:

[0187] Generate a text vector corresponding to the text to be detected according to the text to be detected;

[0188] Based on the target text library, perform a similarity matching of the text vectors for the text vector corresponding to the text to be detected;

[0189] If it is determined that the similarity between the text vector corresponding to the text to be detected and the text vector corresponding to any target text in the target text library is greater than or equal to a predetermined similarity threshold, it is determined that the text to be detected matches any target text in the target text library.

[0190] Among them, the target text library can store only each target text, or only store the text vectors corresponding to each target text, or store each target text and the text vectors corresponding to each target text. This application does not limit this. Among them, when the target text library includes the text vectors corresponding to the target texts, the similarity matching of the text vectors for the text vector corresponding to the text to be detected can be performed more quickly.

[0191] Optionally, an operation of determining whether the text to be detected matches any target text in the target text library based on the target text library can be performed based on a text detection module. Among them, the text detection module can be constructed according to an unsupervised deep learning model and is used for generating text vectors and comparing similarities. That is, through this text detection module, the text to be detected can be parsed to obtain the text vector corresponding to the text to be detected, and the text detection module has the ability to recognize shallow semantic information and can compare the similarity between the text vector corresponding to the text to be detected and any target text vector in the text library, determine the similarity between the text vector corresponding to the text to be detected and the text vector corresponding to any target text in the text library, and the magnitude relationship between this similarity and the similarity threshold. When this similarity is greater than or equal to the similarity threshold, it is determined that the text to be detected matches any target text in the target text library, that is, it is determined that the text to be detected is a text belonging to a predetermined misjudgment type; when this similarity is less than the similarity threshold, it is determined that the text to be detected does not match all the target texts in the target text library, that is, it is determined that the text to be detected does not belong to the predetermined misjudgment type of text. Optionally, the similarity threshold can be configured according to actual needs (such as it can be an empirical value or an experimental value), and this application does not limit this. For example, the similarity threshold can be set to 0.7.

[0192] As an example, when the predetermined misjudgment type includes "shopping", a target text can be set as "I want 24 switches". Then, when performing similarity matching of text vectors on the text to be detected based on the text library, it can accurately detect that the text "I want 24 port switches" that may be misjudged by the deep learning model in the above-mentioned case three belongs to the text of the predetermined misjudgment type.

[0193] Optionally, the present application does not limit the number of target texts either, and the number of target texts can be determined according to the actual situation. In the case where the similarity between the text vector corresponding to the text to be detected and the text vector corresponding to any one of the target texts in the text library is greater than or equal to the predetermined similarity threshold, it can be determined that the text to be detected belongs to the text of the predetermined misjudgment type.

[0194] Based on the above introduction of the preset text library, the target text library can also be divided into a white target text library and a black target text library, and the target text library can also be updated according to the actual situation, etc., which will not be elaborated here.

[0195] Among them, in the case where it is determined whether the text to be detected matches any one of the target texts in the target text library and it is determined that the text to be detected belongs to the text of the predetermined misjudgment type, the type of the target text corresponding to the target text vector whose similarity to the text vector corresponding to the text to be detected is greater than or equal to the predetermined similarity threshold can be determined as the classification result of the text to be detected. Among them, when there are multiple target text vectors whose similarity to the text vector corresponding to the text to be detected is greater than or equal to the predetermined similarity threshold, the type of the target text corresponding to the target text vector with the highest similarity to the text vector corresponding to the text to be detected can be determined as the classification result of the detected text.

[0196] By setting up the target text library and detecting the text to be detected based on the target text library to determine whether the text to be detected matches any one of the target texts in the target text library, that is, to determine whether the text to be detected belongs to the text of the predetermined misjudgment type. In the case where the similarity between the text vector corresponding to the text to be detected and the text vector of any one of the target texts in the text library is greater than or equal to the predetermined similarity threshold, it can be determined that the text to be detected belongs to the text of the predetermined misjudgment type, and the text to be detected can be detected more precisely. Thus, the type of the target text corresponding to the target text vector whose similarity to the text vector corresponding to the text to be detected is greater than or equal to the similarity threshold is determined as the classification result of the text to be detected, without having to determine the classification result of the text to be detected through the deep learning model again, avoiding misjudgment when the text to be detected belonging to the predetermined misjudgment type is classified and recognized by the deep learning model again, and realizing instant online hotfix.

[0197] Optionally, the method further includes:

[0198] If feedback information indicating that the classification result of the text to be detected is a misjudgment is received, then parse the text to be detected to determine the keywords in the text to be detected;

[0199] Store the keywords in the text to be detected in the target keyword library and update the target keyword library;

[0200] Store the text to be detected in the target text library and update the target text library.

[0201] Among them, the feedback information indicating that the classification result of the text to be detected is a misjudgment can be realized through manual operations, and this application does not limit this.

[0202] Optionally, if feedback information indicating that the classification result of the text to be detected is a misjudgment is received, the text to be detected can also be parsed to determine the text vector corresponding to the text to be detected, and store the text vector corresponding to the text to be detected in the target text library and update the target text library.

[0203] By parsing the text to be detected when receiving feedback information indicating that the classification result of the text to be detected is a misjudgment, and updating the keyword library and the text library according to the parsing result, the update of the keyword library and the text library can be automatically completed, and the text to be detected can be better detected.

[0204] Optionally, if it is determined that the text to be detected is a text belonging to a predetermined misjudgment type, the method further includes:

[0205] Store the text to be detected and its classification result in the corpus to obtain an updated corpus;

[0206] Based on the updated corpus, update the trained deep learning model to obtain an updated deep learning model. Thus, when it is determined that the text to be detected is not a text belonging to the predetermined misjudgment type, the classification result of the text to be detected can be determined according to the updated deep learning model.

[0207] Among them, storing the text to be detected and its classification result in the corpus includes: parsing the text to be detected to obtain, including but not limited to, the keywords in the text to be detected and the text vector corresponding to the text to be detected, so as to store the text to be detected, the keywords in the text to be detected, the text vector corresponding to the text to be detected, etc. in the corpus.

[0208] Specifically, the text to be detected and its classification result can be stored in the corpus according to at least one of the following methods, including but not limited to: the text scene information of the text to be detected, the storage time of the text to be detected in the library, and the classification result of the text to be detected.

[0209] By detecting the text to be detected based on a preset text library, when the text to be detected belongs to a predetermined misjudgment type, storing the text to be detected and the classification result of the text to be detected in a corpus to obtain an updated corpus, the generalization ability of the preset text library can be utilized to quickly update the corpus and improve the data collection ability of the corpus. Moreover, by updating and training a deep learning model based on the updated corpus to obtain an updated trained learning model, the reliability of the text in the corpus can be improved without manual labeling. Furthermore, through the automatic update of the corpus, the efficiency of the update and iteration of the deep learning model is enhanced, and the coverage rate and accuracy rate of the deep learning model and the judgment model are continuously optimized.

[0210] Optionally, the above-mentioned updating and training of the deep learning model based on the updated corpus to obtain an updated deep learning model includes:

[0211] Based on the updated corpus, obtain updated training samples, where the updated training samples include at least one text to be detected and the classification result of each text to be detected;

[0212] According to the updated training samples, update and train the deep learning model to obtain an updated deep learning model.

[0213] After the corpus is updated, then update and train the deep learning model according to the training samples obtained by mixing the text to be detected and the classification result of the text to be detected, at least one preset text and the classification result of each preset text to obtain an updated deep learning model.

[0214] Optionally, in practical applications, the above-mentioned updated training samples may also include updated first training samples and updated second training samples. Among them, the updated first training samples can be the training samples obtained by performing a first mixing operation on the keywords in the text to be detected and the classification result of the text to be detected, at least one preset keyword and the classification result of each keyword when it is determined that the text to be detected belongs to a text of a predetermined misjudgment type. The updated second training samples can be the training samples obtained by performing a second mixing operation on the text to be detected and the classification result of the text to be detected, at least one preset text and the classification result of each preset text when it is determined that the text to be detected belongs to a text of a predetermined misjudgment type.

[0215] It can be understood that, in order to reduce the amount of data and improve the data processing efficiency, the above first mixing operation and second mixing operation can also be performed separately according to a certain ratio to obtain the updated first training sample and the updated second training sample respectively. For example, taking the second mixing operation as an example, the ratio between the number of texts to be detected (that is, the number of classification results of the texts to be detected) and at least one preset text (that is, the number of classification results of each preset text) can be 1:6, and the present application does not limit this.

[0216] It can be understood that the updated deep learning model can also be obtained by cutting the updated training samples according to the above cutting method and then performing updated training on the deep learning model. Further, the updated deep learning model can also include an updated first deep learning model and an updated second deep learning model. Among them, the updated first deep learning model can be obtained by cutting the updated first training sample according to the above cutting method and then performing updated training on the deep learning model. The updated second deep learning model can be obtained by cutting the updated second training sample according to the above cutting method and then performing updated training on the deep learning model.

[0217] Optionally, different updated training samples can still be determined according to different training tasks to more accurately update the deep learning model. For example, updated training texts that are the same as the text scene information can be selected according to different text scene information to update the deep learning model.

[0218] By updating the corpus in the case of determining that the text to be detected belongs to the text of a predetermined misjudgment type, thereby obtaining an updated training sample including at least one text to be detected that has been determined to belong to the predetermined misjudgment type according to the updated corpus, and updating and training the above deep learning model according to the updated training sample, the efficiency of the update and iteration of the deep learning model can be improved, and the coverage rate and accuracy rate of the deep learning model and the judgment model can be continuously optimized. Moreover, according to the method of continuously updating the deep learning model, the closed-loop of the method can be more systematically realized.

[0219] The following will detail the data processing method in the embodiments of the present application with an example in a specific application scenario. Refer to Figures 5a to 5e , Figures 5a to 5e FIG. shows a schematic diagram of a specific application scenario of the present application. In this application scenario, the above method can be implemented through an application program of a terminal device or a plug-in in the application program. Taking the text to be detected as the text to be published in the instant messaging sharing area, and taking the above method as an example implemented through a plug-in in an instant messaging application program of the terminal, the above method will be further described. Specifically, the method is implemented through the following steps A1 to A5:

[0220] Step A1: Figure 5a As shown, the friend circle interface of user 1 displays the text to be published and the corresponding picture, and the text is "Urgent! Send an order, please +1". When user 1 clicks the "Publish" button, the instant messaging server ( Figure 1 The application server in the process obtains the text to form the above-mentioned text to be detected.

[0221] Step A2: By detecting the text to be detected using the above data processing method, it can be determined that the text "Urgent! shua order, yong+1" belongs to a predetermined misjudgment type of text, and the classification result of the text is a shua order text.

[0222] Step A3: Obtain the text scene information of the text, that is Figure 5b The friend circle of user 1 shown in the figure shows that in the friend circle of the user's history, the specific content of the posting time 1 "today 8:59" (such as Figure 5c As shown in the figure, the posting text 1 corresponding to posting time 1, "brushing orders, yong3", is also a text of brushing orders, and the number of likes and comments in the circle of friends corresponding to posting text 1 is 0. For the specific content of posting time 2, "Today 8:23" (such as Figure 5d As shown in the figure, the published text 2 "shua single, yong5" corresponding to the published time 2 is also a text of brushing orders, and the number of likes and comments in the circle of friends corresponding to the published text 2 is 0. It can be known that in the circle of friends of this user, the time interval between the published time 1 "today 8:59" and the published time 2 "today 8:23" is 36 minutes, that is, within 1 hour, the user published two circle of friends about "brushing orders".

[0223] Step A4: fuse the text "Urgent! Shua orders, yong+1" that the user is about to publish with the text scene information corresponding to the text "In the user's circle of friends, within 1 hour, the user posted two circles of friends about 'shua orders'". It can be determined that the fusion result is "the user posts to the circle of friends frequently, and most of the posts are about shua orders".

[0224] Step A5: On the instant messaging server ( Figure 1 After the application server in the example obtains the fusion result, a prompt "sending failed, sending is restricted within 30 days" may be issued to the user (e.g. Figure 5e As shown), in order to stop the behavior similar to User 1 using the circle of friends to conduct abnormal "brushing orders".

[0225] It is understandable that in the specific embodiments of the present application, data related to user information (such as the user's circle of friends) etc. are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0226] Figure 6 The schematic diagram of the data processing device provided by the embodiment of the present application is shown. As Figure 6 shown, the data processing device 600 includes a text acquisition module 610, a text detection module 620, and a classification determination module 630.

[0227] The text acquisition module 610 is used to acquire the text to be detected;

[0228] The text detection module 620 is used to detect the text to be detected based on a preset text library, where the preset text library includes text data corresponding to texts marked as belonging to a predetermined misjudgment type;

[0229] The classification determination module 630 is used to determine that the text to be detected belongs to the text of the predetermined misjudgment type, and determine the predetermined misjudgment type as the classification result of the text to be detected; and,

[0230] Determine that the text to be detected does not belong to the text of the predetermined misjudgment type, and determine the classification result of the text to be detected based on a deep learning model;

[0231] Among them, the text belonging to the predetermined misjudgment type is the text for which there may be a misjudgment in the text classification result determined by the deep learning model.

[0232] Optionally, the text data includes at least one of a target keyword library or a target text library. When the text detection module 620 detects the text to be detected based on the preset text library, it is specifically used for:

[0233] Based on the text data, perform keyword detection on the text to be detected;

[0234] Determine that the text to be detected satisfies at least one of the following:

[0235] Any target keyword in the target keyword library is included in the text to be detected;

[0236] The text to be detected matches any target text in the target text library;

[0237] Determine the text to be detected as the text belonging to the predetermined misjudgment type.

[0238] Optionally, the target text library includes text vectors of at least one target text. When the text detection module 620 determines that the text to be detected matches any target text in the target text library, it is specifically configured to:

[0239] Generate a text vector corresponding to the text to be detected according to the text to be detected;

[0240] Based on the target text library, perform a similarity match on the text vector corresponding to the text to be detected;

[0241] Determine that the similarity between the text vector corresponding to the text to be detected and the text vector corresponding to any target text in the target text library is greater than or equal to a predetermined similarity threshold, then determine that the text to be detected matches any target text in the target text library.

[0242] Optionally, the device further includes a corpus update module and a model update module. If it is determined that the text to be detected belongs to a text of a predetermined misjudgment type,

[0243] The corpus update module is configured to store the text to be detected and the classification result of the text to be detected into the corpus to obtain an updated corpus;

[0244] The model update module is configured to update the trained deep learning model based on the updated corpus to obtain an updated deep learning model.

[0245] Optionally, the device further includes a text library update module, and the text library update module is configured to:

[0246] If receiving feedback information indicating that the classification result of the text to be detected is a misjudgment, parse the text to be detected to determine the keywords in the text to be detected;

[0247] Store the keywords in the text to be detected into the target keyword library to update the target keyword library;

[0248] Store the text to be detected into the target text library to update the target text library.

[0249] Optionally, the device may further include a scenario information acquisition module and a fusion module,

[0250] The scenario information acquisition module is configured to acquire the text scenario information of the text to be detected;

[0251] The fusion module is configured to obtain a fusion result based on the classification result of the text to be detected and the text scenario information of the text to be detected, so as to perform corresponding processing on the text to be detected according to the fusion result.

[0252] Optionally, the deep learning model is trained in the following manner;

[0253] Obtain training samples, where the training samples include at least one sample data and the true classification result of each sample data;

[0254] Based on the training samples, iteratively train the first deep learning model until a preset training end condition is satisfied, and obtain the above deep learning model.

[0255] Optionally, based on the training samples, iteratively train the first deep learning model until a preset training end condition is satisfied, including:

[0256] Divide the training samples into a training set and an evaluation set according to a first preset ratio;

[0257] Train the first deep learning model according to the training set until a first training condition is satisfied, and obtain a second deep learning model;

[0258] Evaluate the second deep learning model based on the evaluation set according to a preset evaluation metric. When the metric evaluation result satisfies the second training condition, determine the second deep learning model as the deep learning model;

[0259] When the metric evaluation result does not satisfy the second training condition, adjust the model parameters of the second deep learning model, and continue to train the adjusted model based on the training set;

[0260] The training end condition includes a first training condition and a second training condition.

[0261] Optionally, the evaluation set includes at least one of a validation set or a test set, and the preset evaluation metric includes at least one of a validation evaluation metric or a test evaluation metric;

[0262] Evaluating the second deep learning model based on the evaluation set according to a preset evaluation metric includes at least one of the following:

[0263] Evaluate the second deep learning model based on the validation set according to the validation evaluation metric to obtain a first evaluation result;

[0264] Evaluate the second deep learning model based on the test set according to the test evaluation metric to obtain a second evaluation result;

[0265] Among them, the second training condition includes at least one of a first evaluation condition or a second evaluation condition, and the metric evaluation result satisfying the second training condition includes: at least one of the first evaluation result satisfying the first evaluation condition or the second evaluation result satisfying the second evaluation condition.

[0266] Optionally, when the model update module updates the training deep learning model based on the updated corpus to obtain an updated deep learning model, it is specifically used for:

[0267] Based on the updated corpus, obtain updated training samples, where the updated training samples include at least one text to be detected and the classification result of each text to be detected;

[0268] Update the training of the deep learning model according to the updated training samples to obtain an updated deep learning model.

[0269] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the method according to the embodiments of the present application. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, and details will not be repeated here.

[0270] According to another aspect of the embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above method.

[0271] In an alternative embodiment, an electronic device is provided, Figure 7 which shows a schematic structural diagram of the electronic device provided by the alternative embodiment. As Figure 7 shown, Figure 7 the electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiment of the present application.

[0272] The processor 4001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in combination with the disclosure of the present application. The processor 4001 may also be a combination for implementing computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0273] The bus 4002 may include a path for transmitting information among the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.

[0274] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0275] The memory 4003 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.

[0276] Based on the same principle as the method provided in the embodiments of the present application, the embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in any optional embodiment of the present application described above.

[0277] The embodiments of the present application provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps and corresponding contents shown in the foregoing method embodiments can be implemented.

[0278] The embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding content of the foregoing method embodiment when executed by a processor.

[0279] It should be understood that although the flowchart of the embodiment of the present application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiment of the present application does not limit this.

[0280] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, using other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiment of the present application.

Claims

1. A data processing method, characterized in that Including: Obtain the text to be detected; Detect the text to be detected based on a preset text library, where the preset text library includes text data corresponding to texts marked as belonging to a predetermined misjudgment type; Determine that the text to be detected belongs to a text of a predetermined misjudgment type, and determine the predetermined misjudgment type as the classification result of the text to be detected; Determine that the text to be detected does not belong to a text of a predetermined misjudgment type, and based on a deep learning model, determine the classification result of the text to be detected; Among them, the text belonging to the predetermined misjudgment type is a text for which there may be a misjudgment in the text classification result determined by the deep learning model; Obtain the text scenario information of the text to be detected; the text scenario information includes at least one of the following: sharing information, discussion information, interaction information, explanation information, interaction information; Based on the classification result of the text to be detected and the text scenario information of the text to be detected, obtain a fusion result, so as to perform corresponding processing on the text to be detected according to the fusion result.

2. The method according to claim 1, wherein The text data includes at least one of a target keyword library or a target text library. The detecting the text to be detected based on the preset text library includes: Detect the text to be detected based on the text data; Determine that the text to be detected satisfies at least one of the following: The text to be detected includes any target keyword in the target keyword library; The text to be detected matches any target text in the target text library; Determine that the text to be detected belongs to a text of a predetermined misjudgment type.

3. The method according to claim 2, characterized in that, The target text library includes text vectors of at least one target text. Determining that the text to be detected matches any target text in the target text library includes: Generate a text vector corresponding to the text to be detected according to the text to be detected; Based on the target text library, perform a similarity match of the text vector on the text vector corresponding to the text to be detected; Determine that the similarity between the text vector corresponding to the text to be detected and the text vector corresponding to any target text in the target text library is greater than or equal to a predetermined similarity threshold, then determine that the text to be detected matches any target text in the target text library.

4. The method according to claim 1, wherein If it is determined that the text to be detected belongs to a text of a predetermined misjudgment type, the method further includes: Store the text to be detected and the classification result of the text to be detected in a corpus to obtain an updated corpus; Based on the updated corpus, update and train the deep learning model to obtain an updated deep learning model.

5. The method according to claim 2 or 3, characterized in that, The method further includes: If feedback information indicating that the classification result of the text to be detected is a misjudgment is received, then parse the text to be detected to determine the keywords in the text to be detected; Store the keywords in the text to be detected in the target keyword library to update the target keyword library; Store the text to be detected in the target text library to update the target text library.

6. The method according to claim 1, wherein The deep learning model is trained in the following manner; Obtain training samples, where the training samples include at least one sample data and the true classification result of each sample data; Based on the training samples, iteratively train the first deep learning model until a preset training end condition is met, and obtain the deep learning model.

7. The method according to claim 6, characterized in that, The iteratively training the first deep learning model based on the training samples until a preset training end condition is met includes: Divide the training samples into a training set and an evaluation set according to a first preset ratio; Train the first deep learning model according to the training set until a first training condition is met, and obtain a second deep learning model; Evaluate the second deep learning model based on the evaluation set according to a preset evaluation metric. When the metric evaluation result meets the second training condition, determine the second deep learning model as the deep learning model; When the metric evaluation result does not meet the second training condition, adjust the model parameters of the second deep learning model, and continue to train the adjusted model based on the training set; The training end condition includes the first training condition and the second training condition.

8. The method according to claim 7, wherein The evaluation set includes at least one of a validation set or a test set, and the preset evaluation metric includes at least one of a validation evaluation metric or a test evaluation metric; The evaluating the second deep learning model based on the evaluation set according to a preset evaluation metric includes at least one of the following: Evaluate the second deep learning model based on the validation set according to the validation evaluation metric, and obtain a first evaluation result; Evaluate the second deep learning model based on the test set according to the test evaluation metric, and obtain a second evaluation result; Wherein, the second training condition includes at least one of the first evaluation condition or the second evaluation condition, and the metric evaluation result meeting the second training condition includes: at least one of the first evaluation result meeting the first evaluation condition or the second evaluation result meeting the second evaluation condition.

9. The method according to claim 4, characterized in that, The updating and training the deep learning model based on the updated corpus to obtain an updated deep learning model includes: Based on the updated corpus, obtain updated training samples, where the updated training samples include at least one of the to-be-detected texts and the classification results of each to-be-detected text; According to the updated training samples, update and train to obtain an updated deep learning model.

10. A data processing device, characterized in that, Includes: A text acquisition module, configured to acquire a to-be-detected text; A text detection module, configured to detect the to-be-detected text based on a preset text library, where the preset text library includes text data corresponding to texts marked as belonging to a predetermined misjudgment type; A classification determination module is configured to determine that the to-be-detected text belongs to a text of a predetermined misjudgment type, and determine the predetermined misjudgment type as the classification result of the to-be-detected text; and, Determine that the to-be-detected text does not belong to a text of a predetermined misjudgment type, and based on a deep learning model, determine the classification result of the to-be-detected text; Among them, the text belonging to the predetermined misjudgment type is the text for which there is a possibility of misjudgment in the text classification result determined by the deep learning model; The device further includes a scenario information acquisition module and a fusion module; The scenario information acquisition module is configured to acquire the text scenario information of the text to be detected; the text scenario information includes at least one of the following: sharing information, discussion information, interaction information, description information, interaction information; The fusion module is configured to obtain a fusion result based on the classification result of the text to be detected and the text scenario information of the text to be detected, so as to perform corresponding processing on the text to be detected according to the fusion result.

11. The device according to claim 10, characterized in that, The text data includes at least one of a target keyword library or a target text library. When the text detection module detects the text to be detected based on a preset text library, it is configured to: Detect the text to be detected based on the text data; Determine that the text to be detected meets at least one of the following: Any target keyword in the target keyword library is included in the text to be detected; The text to be detected matches any target text in the target text library; Determine the text to be detected as a text belonging to the predetermined misjudgment type.

12. The device according to claim 11, wherein The target text library includes text vectors of at least one target text. When the text detection module determines that the text to be detected matches any target text in the target text library, it is configured to: Generate a text vector corresponding to the text to be detected according to the text to be detected; Perform a similarity match of the text vector on the text vector corresponding to the text to be detected based on the target text library; Determine that the similarity between the text vector corresponding to the text to be detected and the text vector corresponding to any target text in the target text library is greater than or equal to a predetermined similarity threshold, then determine that the text to be detected matches any target text in the target text library.

13. The device according to claim 10, characterized in that, The device further includes a corpus update module and a model update module. If the classification determination module determines that the text to be detected is a text belonging to the predetermined misjudgment type, The corpus update module is configured to store the text to be detected and the classification result of the text to be detected in a corpus to obtain an updated corpus; The model update module is configured to update and train the deep learning model based on the updated corpus to obtain an updated deep learning model.

14. The device according to claim 11 or 12, characterized in that, The device further includes a text library update module, and the text library update module is configured to: If feedback information indicating that the classification result of the text to be detected is a misjudgment is received, then parse the text to be detected to determine the keywords in the text to be detected; Store the keywords in the text to be detected in the target keyword library to update the target keyword library; Store the text to be detected in the target text library to update the target text library.

15. The device according to claim 10, characterized in that, The deep learning model is trained in the following manner; Obtain training samples, where the training samples include at least one sample data and the true classification result of each sample data; Based on the training samples, iteratively train the first deep learning model until a preset training end condition is satisfied, and obtain the deep learning model.

16. The device according to claim 15, characterized in that, The iteratively training the first deep learning model based on the training samples until a preset training end condition is satisfied includes: Dividing the training samples into a training set and an evaluation set according to a first preset ratio; Training the first deep learning model according to the training set until a first training condition is satisfied, and obtaining a second deep learning model; Evaluating the second deep learning model based on the evaluation set according to a preset evaluation metric. When the metric evaluation result satisfies a second training condition, determine the second deep learning model as the deep learning model; When the metric evaluation result does not satisfy the second training condition, adjust the model parameters of the second deep learning model, and continue to train the adjusted model based on the training set; The training end condition includes the first training condition and the second training condition.

17. The device according to claim 16, wherein The evaluation set includes at least one of a validation set or a test set, and the preset evaluation metric includes at least one of a validation evaluation metric or a test evaluation metric; The evaluating the second deep learning model based on the evaluation set according to a preset evaluation metric includes at least one of the following: Evaluating the second deep learning model based on the validation set according to the validation evaluation metric, and obtaining a first evaluation result; Evaluating the second deep learning model based on the test set according to the test evaluation metric, and obtaining a second evaluation result; Wherein, the second training condition includes at least one of a first evaluation condition or a second evaluation condition, and the metric evaluation result satisfying the second training condition includes: at least one of the first evaluation result satisfying the first evaluation condition or the second evaluation result satisfying the second evaluation condition.

18. The device according to claim 17, characterized in that, When the model update module updates and trains the deep learning model based on the updated corpus to obtain an updated deep learning model, it is used for: Based on the updated corpus, obtain updated training samples, where the updated training samples include at least one of the texts to be detected and the classification results of each text to be detected; According to the updated training samples, update and train to obtain an updated deep learning model.

19. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-9.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-9.

21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Satellite internet text sensitive information detection method and device based on deep learning

    CN111078879A