A system and method for updating a speech text classification model

By introducing the method of log module filtering data, replacing annotation and clustering to generate training data in the speech text classification model, the problem of model update lag is solved, and faster, accurate and objective model updates are achieved.

CN115719593BActive Publication Date: 2025-06-06CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211352212.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2025-06-06
Estimated Expiration
2042-10-31

AI Technical Summary

Technical Problem

There is a lag problem in the update process of the existing speech text classification model, which is difficult to cover the needs of real scenes in a timely manner, and the test results are subjective and difficult to guarantee the accuracy.

Method used

It provides an update system and method for the speech text classification model. It obtains log data that is inconsistent with the classification results as filter data through the log module, replaces and annotates the entity name, clusters and generates training data, and updates the vocabulary, sentence structure and statement classification module.

Benefits of technology

It improves the timeliness and accuracy of model updates, ensures that the model can adapt to changes in scenario requirements more quickly, and improves the objectivity and reliability of model updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115719593B_ABST
    Figure CN115719593B_ABST
Patent Text Reader

Abstract

The present application relates to a system and method for updating a speech-text classification model, and the system for updating the speech-text classification model includes: a model device, a log device, and a data device. The log device is used to perform semantic recognition on user speech, obtain log data, and determine the log data inconsistent with the classification result as screening data; the data device replaces and annotates the entity names in the screening data when the screening data is greater than or equal to the data volume threshold, the data device obtains training data, and the data device clusters the training data according to the length of the training data and the number of entity names, and obtains first data for updating a vocabulary classification module, second data for updating a sentence classification module, and third data for updating a sentence classification module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and in particular to a system and method for updating a speech-text classification model. Background Art

[0002] With the improvement of neural network technology and computer computing power, the artificial intelligence industry has made great progress. The classification of speech and text can be completed by deploying classification models, which have been applied to various fields, such as vehicle-machine interaction, intelligent customer service, information distribution, etc. In order to ensure the reliable performance of the classification model, it needs to be continuously updated after it goes online to meet the needs of the scene. In this process, developers, operators, and testers are required to provide feedback based on the test results and update the model based on the feedback results, which causes the lag of model updates. It is not only difficult to cover the real scene requirements and ensure timeliness, but also the test results are highly subjective and difficult to guarantee accuracy. Summary of the invention

[0003] Based on this, a system and method for updating a speech-text classification model are provided to improve the problem of delayed model updates.

[0004] On the one hand, a system for updating a speech-text classification model is provided, comprising:

[0005] The model device includes: a vocabulary classification module, a sentence classification module and a sentence classification module;

[0006] A vocabulary classification module, wherein the vocabulary classification module includes a dictionary for classification, and the speech text information to be processed is classified and processed by the dictionary to obtain a first classification result and a first output result, wherein the first output end of the vocabulary classification module is used to output the first classification result, and the second output end of the vocabulary classification module is used to output the first output result;

[0007] A sentence classification module, wherein the sentence classification module includes a vector space unit for calculating vector similarity, the vector space unit performs classification processing on the first output result to obtain a second classification result and a second output result, the first output end of the sentence classification module is used to output the second classification result, and the second output end of the sentence classification module is used to output the second output result;

[0008] A sentence classification module, the sentence classification module comprising a neural network unit for sentence classification, the neural network unit classifies the second output result to obtain a third classification result and outputs it from an output end of the sentence classification module;

[0009] A log device, used to perform semantic recognition on the user's voice to obtain log data, and determine the log data inconsistent with the classification result as screening data, wherein the classification result includes a first classification result, a second classification result and a third classification result;

[0010] A data device, when the filtered data is greater than or equal to a data volume threshold, the data device replaces and labels the entity names in the filtered data, the data device obtains training data, and the data device clusters the training data according to the length of the training data and the number of entity names, obtains first data for updating a vocabulary classification module and outputs it through a first output end of the data device, second data for updating a sentence classification module and outputs it through a second output end of the data device, and third data for updating a sentence classification module and outputs it through a third output end of the data device.

[0011] Optionally, the sentence classification module also includes a database interface, which is used to obtain a remote dictionary service, and the remote dictionary service is used to determine whether the second classification result is greater than or equal to a similarity threshold. If the second classification result is greater than or equal to the similarity threshold, the second classification result is output through the first output end of the sentence classification module; if the second classification result is less than the similarity threshold, the second output result is output through the second output end of the sentence classification module.

[0012] Optionally, the sentence classification module further includes a preprocessing unit, and the preprocessing unit is used to vectorize the second output result;

[0013] The neural network unit includes an input layer, a fully connected layer and an output layer;

[0014] Among them, the input end of the preprocessing unit is connected to the second output end of the sentence classification module, and the output end of the preprocessing unit is connected to the input layer.

[0015] The present invention provides a method for updating a speech-text classification model, which updates the model device, and the method comprises:

[0016] Performing semantic recognition on the user voice to obtain log data, and determining the log data inconsistent with the classification result as screening data, wherein the classification result includes a first classification result, a second classification result, and a third classification result;

[0017] When the filtered data is greater than or equal to the data volume threshold, the entity names in the filtered data are replaced and annotated to obtain training data, and the training data are clustered according to the length of the training data and the number of entity names to obtain first data for updating a vocabulary classification module, second data for updating a sentence classification module, and third data for updating a sentence classification module;

[0018] updating the dictionary according to the first data to obtain an updated vocabulary classification module;

[0019] Update the vector space unit according to the second data to obtain an updated sentence classification module;

[0020] Vectorizing and labeling the third data to obtain a sentence vector and a corresponding sentence label;

[0021] Input the sentence vector and the corresponding sentence label into the initial neural network unit for classification processing to obtain a sample result;

[0022] Iteratively training the initial neural network unit according to the matching degree between the sample result and the sentence label to obtain a trained neural network unit;

[0023] The trained neural network unit is configured into the sentence classification module to obtain an updated sentence classification module.

[0024] Optionally, updating the vector space unit according to the second data includes:

[0025] The vector space unit is updated according to the second data, and the updated vector space unit is transmitted through a database interface so as to be stored by a remote dictionary service.

[0026] Optionally, clustering the training data according to the length of the training data and the number of entity names to obtain first data for updating the vocabulary classification module, second data for updating the sentence classification module, and third data for updating the sentence classification module includes:

[0027] Acquire training data whose data length is less than or equal to a length threshold, and determine the data as the first data;

[0028] Acquire training data whose data length is greater than the length threshold and whose number of entity names is greater than or equal to the number threshold, and determine the data as the second data;

[0029] The training data whose data length is greater than the length threshold and whose number of entity names is less than the number threshold is obtained and determined as the third data.

[0030] Optionally, iteratively training the initial neural network unit according to the matching degree between the sample result and the sentence label to obtain a trained neural network unit includes:

[0031] Training the initial neural network unit based on a cross entropy loss function to reduce the loss between the sample result and the sentence label to increase the matching degree between the sample result and the sentence label;

[0032] The neural network unit is iteratively trained, and the weight parameters of the neuron nodes in the neural network unit are updated to obtain a trained neural network unit.

[0033] Optionally, also include:

[0034] Connecting the input end of the updated vocabulary classification module to the text module, wherein the first output end of the vocabulary classification module is used to output a first classification result, and the second output end of the vocabulary classification module is used to output a first output result, wherein the text module is used to sample user speech and convert it into speech text information;

[0035] Connecting the input end of the updated sentence classification module to the second output end of the vocabulary classification module, wherein the first output end of the sentence classification module is used to output the second classification result, and the second output end of the sentence classification module is used to output the second output result;

[0036] The updated input end of the sentence classification module is connected to the second output end of the sentence classification module, and the output end of the sentence classification module is used to output a third classification result.

[0037] The present invention provides a computer device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods when executing the computer program.

[0038] The present invention provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the methods described above are implemented.

[0039] The above-mentioned system and method for updating the speech-text classification model obtains log data inconsistent with the classification results as screening data through the log module, obtains training data by replacing and marking the entity names in the screening data, and clusters the training data based on the length of the training data and the number of entity names, respectively obtaining the first data, second data and third data for updating different models, and updates the vocabulary classification module, sentence classification module and sentence classification module in the model device, thereby improving the problem of lagging model updates. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 An application environment diagram of a method for updating a speech-text classification model in one embodiment;

[0041] Figure 2 A schematic diagram of the structure of a system for updating a speech-text classification model in one embodiment;

[0042] Figure 3 A flowchart of the steps of a method for updating a speech-text classification model in one embodiment;

[0043] Figure 4 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0045] The updating method of the speech-text classification model provided in this application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 through a network. The terminal 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server 104 may be implemented as an independent server or a server cluster consisting of multiple servers.

[0046] like Figure 2 As shown, a system for updating a speech-text classification model is provided, comprising:

[0047] The model device includes: a vocabulary classification module, a sentence classification module and a sentence classification module;

[0048] A vocabulary classification module, wherein the vocabulary classification module includes a dictionary for classification, and the dictionary is used to classify the speech text information to be processed to obtain a first classification result and a first output result. The first output end of the vocabulary classification module is used to output the first classification result, and the second output end of the vocabulary classification module is used to output the first output result.

[0049] By way of example, human voice information can be collected through a microphone or a handset, and voice recognition can be performed after preprocessing to obtain voice text information carrying voice semantics. The dictionary can classify some voice text information and obtain a first classification result. Some voice text information cannot be classified by the dictionary, and a first output result is obtained.

[0050] A sentence classification module, wherein the sentence classification module includes a vector space unit for calculating vector similarity, the vector space unit classifies the first output result to obtain a second classification result and a second output result, the first output end of the sentence classification module is used to output the second classification result, and the second output end of the sentence classification module is used to output the second output result.

[0051] Exemplarily, the vector space unit can calculate the vector similarity of the first output result based on the stored vectors, and when the vector similarity is large, output the second classification result corresponding to the stored vector, and when the vector similarity is small, output the second output result.

[0052] A sentence classification module, wherein the sentence classification module includes a neural network unit for sentence classification, and the neural network unit classifies the second output result to obtain a third classification result and outputs it from the output end of the sentence classification module.

[0053] Exemplarily, the input object of the neural network unit is the speech text information that cannot be processed well by the dictionary and vector space unit, which reduces the discreteness and processing quantity of the input amount of the neural network unit. Accordingly, it saves the resources for training the neural network unit and the complexity of building the model. The speech text information at the vocabulary level can be processed by the vocabulary classification model, and the speech text information at the sentence level can be processed by the sentence classification model, and finally the speech text information at the sentence level can be processed by the sentence classification model. That is, the clear memory layer in the human brain memory curve can be simulated by the vocabulary classification module, the fuzzy memory layer in the human brain memory curve can be simulated by the sentence classification module, and the forgetting layer in the human brain memory curve can be simulated by the sentence classification module, avoiding the reliance on machine learning or deep learning models to build complex neural networks, and avoiding the need for large computing power support and training data samples when training neural networks.

[0054] The log device is used to perform semantic recognition on user speech to obtain log data, and determine the log data inconsistent with the classification result as screening data, wherein the classification result includes a first classification result, a second classification result and a third classification result.

[0055] Exemplarily, the log device can provide a redundant speech-text recognition system to extract user speech for semantic recognition and obtain log data. The log data stores and records the data and is also used to compare the differences between the log data and the classification results. When there are a large number of log data that are inconsistent with the classification results, the system can be updated in response.

[0056] A data device, when the filtered data is greater than or equal to a data volume threshold, the data device replaces and labels the entity names in the filtered data, the data device obtains training data, and the data device clusters the training data according to the length of the training data and the number of entity names, obtains first data for updating a vocabulary classification module and outputs it through a first output end of the data device, second data for updating a sentence classification module and outputs it through a second output end of the data device, and third data for updating a sentence classification module and outputs it through a third output end of the data device.

[0057] Exemplarily, the data device user replaces and labels the entity names in the screening data, and gives examples according to the length of the training data and the number of entity names to transfer the first data, the second data and the third data to different modules for updating.

[0058] In some embodiments, the sentence classification module also includes a database interface, which is used to obtain a remote dictionary service, and the remote dictionary service is used to determine whether the second classification result is greater than or equal to a similarity threshold. If the second classification result is greater than or equal to the similarity threshold, the second classification result is output through the first output end of the sentence classification module; if the second classification result is less than the similarity threshold, the second output result is output through the second output end of the sentence classification module.

[0059] In some embodiments, the sentence classification module further includes a preprocessing unit, and the preprocessing unit is used to vectorize the second output result;

[0060] The neural network unit includes an input layer, a fully connected layer and an output layer;

[0061] Among them, the input end of the preprocessing unit is connected to the second output end of the sentence classification module, and the output end of the preprocessing unit is connected to the input layer.

[0062] like Figure 3 As shown, the present invention provides a method for updating a speech-text classification model, and updates the model device, the method comprising:

[0063] S1: performing semantic recognition on the user voice to obtain log data, and determining the log data inconsistent with the classification result as screening data, wherein the classification result includes a first classification result, a second classification result and a third classification result;

[0064] S2: When the filtered data is greater than or equal to the data volume threshold, the entity names in the filtered data are replaced and annotated to obtain training data, and the training data are clustered according to the length of the training data and the number of entity names to obtain first data for updating a vocabulary classification module, second data for updating a sentence classification module, and third data for updating a sentence classification module;

[0065] S3: updating the dictionary according to the first data to obtain an updated vocabulary classification module; updating the vector space unit according to the second data to obtain an updated sentence classification module; vectorizing and labeling the third data to obtain a sentence vector and a corresponding sentence label;

[0066] S4: Input the sentence vector and the corresponding sentence label into the initial neural network unit for classification processing to obtain a sample result; iteratively train the initial neural network unit according to the matching degree between the sample result and the sentence label to obtain a trained neural network unit; configure the trained neural network unit into the sentence classification module to obtain an updated sentence classification module.

[0067] Log data that is inconsistent with the classification results is obtained as screening data through the log module, and training data is obtained by replacing and labeling the entity names in the screening data. Clustering is performed based on the length of the training data and the number of entity names, and the first data, second data and third data for updating different models are obtained respectively. The vocabulary classification module, sentence classification module and sentence classification module in the model device are updated to improve the problem of lagging model updates.

[0068] In some embodiments, updating the vector space unit according to the second data includes:

[0069] The vector space unit is updated according to the second data, and the updated vector space unit is transmitted through a database interface so as to be stored by a remote dictionary service.

[0070] In some embodiments, clustering the training data according to the length of the training data and the number of entity names to obtain first data for updating the vocabulary classification module, second data for updating the sentence classification module, and third data for updating the sentence classification module includes:

[0071] Acquire training data whose data length is less than or equal to a length threshold, and determine the data as the first data;

[0072] Acquire training data whose data length is greater than the length threshold and whose number of entity names is greater than or equal to the number threshold, and determine the data as the second data;

[0073] The training data whose data length is greater than the length threshold and whose number of entity names is less than the number threshold is obtained and determined as the third data.

[0074] In some embodiments, iteratively training the initial neural network unit according to the matching degree between the sample result and the sentence label to obtain a trained neural network unit includes:

[0075] Training the initial neural network unit based on a cross entropy loss function to reduce the loss between the sample result and the sentence label to increase the matching degree between the sample result and the sentence label;

[0076] The neural network unit is iteratively trained, and the weight parameters of the neuron nodes in the neural network unit are updated to obtain a trained neural network unit.

[0077] In some embodiments, it also includes:

[0078] Connecting the input end of the updated vocabulary classification module to the text module, wherein the first output end of the vocabulary classification module is used to output a first classification result, and the second output end of the vocabulary classification module is used to output a first output result, wherein the text module is used to sample user speech and convert it into speech text information;

[0079] Connecting the input end of the updated sentence classification module to the second output end of the vocabulary classification module, wherein the first output end of the sentence classification module is used to output the second classification result, and the second output end of the sentence classification module is used to output the second output result;

[0080] The updated input end of the sentence classification module is connected to the second output end of the sentence classification module, and the output end of the sentence classification module is used to output a third classification result.

[0081] It should be understood that although Figure 3-4 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 3-4 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0082] The specific definition of the updating device of the speech-text classification model can be found in the definition of the updating method of the speech-text classification model in the above text, which will not be repeated here. Each module in the updating device of the speech-text classification model can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0083] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store update data of the speech text classification model. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for updating the speech text classification model is implemented.

[0084] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0085] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0086] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0087] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A system for updating speech-text classification models, It is characterized in that include: The model device includes: a vocabulary classification module, a sentence classification module and a sentence classification module; A vocabulary classification module, wherein the vocabulary classification module includes a dictionary for classification, and the speech text information to be processed is classified by the dictionary to obtain a first classification result and a first output result, wherein the speech text information corresponding to the first output result cannot be classified by the dictionary, and the first output end of the vocabulary classification module is used to output the first classification result, and the second output end of the vocabulary classification module is used to output the first output result; A sentence classification module, the sentence classification module comprising a vector space unit for calculating vector similarity, the vector space unit classifies the first output result to obtain a second classification result and a second output result, wherein the vector similarity of the second output result is lower than a similarity threshold, the first output end of the sentence classification module is used to output the second classification result, and the second output end of the sentence classification module is used to output the second output result; A sentence classification module, the sentence classification module comprising a neural network unit for sentence classification, the neural network unit classifies the second output result to obtain a third classification result and outputs it from an output end of the sentence classification module; A log device, used to perform semantic recognition on the user's voice to obtain log data, and determine the log data inconsistent with the classification result as screening data, wherein the classification result includes a first classification result, a second classification result and a third classification result; A data device, when the filtered data is greater than or equal to a data volume threshold, the data device replaces and labels the entity names in the filtered data, the data device obtains training data, and the data device clusters the training data according to the length of the training data and the number of entity names, obtains first data for updating a vocabulary classification module and outputs it through a first output end of the data device, second data for updating a sentence classification module and outputs it through a second output end of the data device, and third data for updating a sentence classification module and outputs it through a third output end of the data device.

2. The updating system of the speech-text classification model according to claim 1, It is characterized in that The sentence classification module also includes a database interface, which is used to obtain a remote dictionary service, and the remote dictionary service is used to determine whether the second classification result is greater than or equal to the similarity threshold. If the second classification result is greater than or equal to the similarity threshold, the second classification result is output through the first output end of the sentence classification module; if the second classification result is less than the similarity threshold, the second output result is output through the second output end of the sentence classification module.

3. The updating system of the speech-text classification model according to claim 1, It is characterized in that The sentence classification module further includes a preprocessing unit, and the preprocessing unit is used to vectorize the second output result; The neural network unit includes an input layer, a fully connected layer and an output layer; Among them, the input end of the preprocessing unit is connected to the second output end of the sentence classification module, and the output end of the preprocessing unit is connected to the input layer.

4. A method for updating a speech-text classification model, It is characterized in that Updating the model device according to any one of claims 1 to 3, the method comprising: Performing semantic recognition on the user voice to obtain log data, and determining the log data inconsistent with the classification result as screening data, wherein the classification result includes a first classification result, a second classification result, and a third classification result; When the filtered data is greater than or equal to the data volume threshold, the entity names in the filtered data are replaced and annotated to obtain training data, and the training data are clustered according to the length of the training data and the number of entity names to obtain first data for updating a vocabulary classification module, second data for updating a sentence classification module, and third data for updating a sentence classification module; updating the dictionary according to the first data to obtain an updated vocabulary classification module; Update the vector space unit according to the second data to obtain an updated sentence classification module; Vectorizing and labeling the third data to obtain a sentence vector and a corresponding sentence label; Input the sentence vector and the corresponding sentence label into the initial neural network unit for classification processing to obtain a sample result; Iteratively training the initial neural network unit according to the matching degree between the sample result and the sentence label to obtain a trained neural network unit; The trained neural network unit is configured into the sentence classification module to obtain an updated sentence classification module.

5. The updating method of the speech-text classification model according to claim 4, It is characterized in that Updating the vector space unit according to the second data includes: The vector space unit is updated according to the second data, and the updated vector space unit is transmitted through a database interface so as to be stored by a remote dictionary service.

6. The updating method of the speech-text classification model according to claim 4, It is characterized in that Clustering the training data according to the length of the training data and the number of entity names to obtain first data for updating the vocabulary classification module, second data for updating the sentence classification module, and third data for updating the sentence classification module, including: Acquire training data whose data length is less than or equal to a length threshold, and determine the data as the first data; Acquire training data whose data length is greater than the length threshold and whose number of entity names is greater than or equal to the number threshold, and determine the data as the second data; The training data whose data length is greater than the length threshold and whose number of entity names is less than the number threshold is obtained and determined as the third data.

7. The updating method of the speech-text classification model according to claim 4, It is characterized in that According to the matching degree between the sample result and the sentence label, iteratively training the initial neural network unit to obtain a trained neural network unit includes: Training the initial neural network unit based on a cross entropy loss function to reduce the loss between the sample result and the sentence label to increase the matching degree between the sample result and the sentence label; The neural network unit is iteratively trained, and the weight parameters of the neuron nodes in the neural network unit are updated to obtain a trained neural network unit.

8. The updating method of the speech-text classification model according to claim 4, It is characterized in that Also includes: Connecting the input end of the updated vocabulary classification module to the text module, wherein the first output end of the vocabulary classification module is used to output a first classification result, and the second output end of the vocabulary classification module is used to output a first output result, wherein the text module is used to sample user speech and convert it into speech text information; Connecting the input end of the updated sentence classification module to the second output end of the vocabulary classification module, wherein the first output end of the sentence classification module is used to output the second classification result, and the second output end of the sentence classification module is used to output the second output result; The updated input end of the sentence classification module is connected to the second output end of the sentence classification module, and the output end of the sentence classification module is used to output a third classification result.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 4 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 4 to 8 are implemented.

Citation Information

Patent Citations

  • Updating training method, device and device for text classification model

    CN109241288A

  • Text classification model updating method and system, electronic equipment and storage medium

    CN111737472A