Text data classification method and device, vehicle, storage medium and program product

By combining text data classification models with zero-sample training and label sample training, the vehicle feedback data is classified, which solves the problems of high difficulty and low accuracy of manual classification, and achieves more efficient and accurate text data classification.

CN120104803APending Publication Date: 2025-06-06XIAOMI EV TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510215366.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

As the number of vehicles increases, the number of problems and labels reported by users increases, and manual classification methods are difficult and have low accuracy.

Method used

A text data classification method is adopted to classify the text data to be classified through the first classification model (zero sample training) and the second classification model (label sample training), and the first candidate label and the second candidate label are obtained respectively, combining the results of both to obtain a comprehensive and accurate classification label.

Benefits of technology

It improves the accuracy of text data classification, can effectively process text data in new categories and existing categories, and reduces the difficulty and error rate of manual classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104803A_ABST
    Figure CN120104803A_ABST
Patent Text Reader

Abstract

The invention provides a text data classification method and device, a vehicle, a storage medium and a program product, and relates to the technical field of vehicles, and the method comprises the steps: classifying to-be-classified text data through a first classification model, and obtaining a first candidate tag; the text data are classified through the first classification model to obtain the first candidate labels, the text data are classified through the second classification model to obtain the second candidate labels, the first classification model is obtained through zero sample training, so that the first classification model can classify the text data of a new category, and the second classification model is obtained through label sample training and is high in classification accuracy; according to the first candidate label and the second candidate label, the classification label for the text data is obtained, the comprehensive and accurate classification label can be obtained by combining the classification results of the first classification model and the second classification model, and the classification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of vehicle technology, and in particular to a text data classification method, device, vehicle, storage medium and program product. Background Art

[0002] During the use of the vehicle, users can report vehicle-related issues through the user feedback system. Backstage staff count these issues and label each issue so that each issue has a corresponding label, which is convenient for subsequent personnel to handle. As the number of vehicles increases, the number of issues reported by users continues to rise, and the number of labels has also become huge. Manual classification is difficult and has low accuracy. Summary of the invention

[0003] To overcome the problems existing in the related art, the present disclosure provides a text data classification method, device, vehicle, storage medium and program product.

[0004] According to a first aspect of an embodiment of the present disclosure, a text data classification method is provided, the method comprising: acquiring text data to be classified; classifying the text data through a first classification model to obtain a first candidate label; classifying the text data through a second classification model to obtain a second candidate label, wherein the first classification model is obtained through zero-sample training, and the second classification model is obtained through labeled sample training; obtaining a classification label for the text data based on the first candidate label and the second candidate label.

[0005] Optionally, the first classification model includes a recall model and a rearrangement model, and classifying the text data through the first classification model to obtain a first candidate label includes: classifying the text data through the recall model to obtain an intermediate label; and obtaining the first candidate label based on the text data and the intermediate label through the rearrangement model.

[0006] Optionally, the recall model obtains the intermediate labels in the following manner: vectorize the text data to obtain a text vector, and vectorize M text labels to obtain M label vectors, where M is a positive integer; obtain a first similarity measure value between the text vector and each label vector in the M label vectors to obtain M first similarity measure values; determine N intermediate labels from the M text labels based on the M first similarity measure values, where N is a positive integer less than M.

[0007] Optionally, determining N intermediate labels from M text labels based on the M first similarity measurement values ​​includes: determining N first target similarity measurement values ​​from the M first similarity measurement values, wherein the N first target similarity measurement values ​​are all greater than the first similarity measurement values ​​of the M first similarity measurement values ​​except the N first target similarity values; determining text labels corresponding to the N first target similarity values ​​from the M text labels as intermediate labels, to obtain the N intermediate labels.

[0008] Optionally, the rearrangement model obtains the first candidate label in the following manner: obtaining a second similarity measurement value between the text data and each of the N intermediate labels to obtain N second similarity measurement values; and determining X first candidate labels from the N intermediate labels based on the N second similarity measurement values, where X is a positive integer less than or equal to N.

[0009] Optionally, determining X first candidate tags from N intermediate tags according to the N second similarity measurement values ​​includes: determining X second target similarity measurement values ​​from the N second similarity measurement values, wherein the X second target similarity measurement values ​​are all greater than second similarity measurement values ​​among the N second similarity measurement values ​​except the X second target similarity measurement values; and determining, from the N intermediate tags, intermediate tags corresponding to the X second target similarity measurement values ​​as first candidate tags, to obtain the X first candidate tags.

[0010] Optionally, obtaining a classification label for the text data according to the first candidate label and the second candidate label includes: if the second candidate label is different from all X first candidate labels, replacing the first candidate label with the smallest second target similarity measure value among the X first candidate labels with the second candidate label, to obtain X classification labels.

[0011] Optionally, the method also includes: sorting the X-1 first candidate tags in the classification tags in descending order according to the second target similarity measure value, and arranging the second candidate tags before the first candidate tag with the largest second target similarity measure value, and displaying the classification tags according to the tag sorting.

[0012] Optionally, obtaining a classification label for the text data according to the first candidate label and the second candidate label includes: if the second candidate label is the same as one of the X first candidate labels, using the X first candidate labels as classification labels to obtain X classification labels.

[0013] Optionally, the method also includes: sorting the X-1 first candidate tags in the classification tags that are different from the second candidate tag in order from large to small according to the second target similarity measure value, and arranging the first candidate tag that is the same as the second candidate tag before the first candidate tag with the largest second target similarity measure value, and displaying the classification tags according to the tag sorting.

[0014] Optionally, the text data includes data obtained by a feedback system corresponding to the vehicle.

[0015] According to a second aspect of an embodiment of the present disclosure, a text data classification device is provided, the device comprising: an acquisition module, configured to acquire text data to be classified; a first acquisition module, configured to classify the text data through a first classification model to obtain a first candidate label; a second acquisition module, configured to classify the text data through a second classification model to obtain a second candidate label, wherein the first classification model is obtained through zero-sample training, and the second classification model is obtained through labeled sample training; a classification module, configured to obtain a classification label for the text data based on the first candidate label and the second candidate label.

[0016] According to a third aspect of an embodiment of the present disclosure, a vehicle is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the steps of the method described in the first aspect when executing the instructions.

[0017] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the method provided in the first aspect of the present disclosure are implemented.

[0018] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which implements the steps of the method provided in the first aspect when executed by a processor.

[0019] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: classifying the text data to be classified through the first classification model to obtain the first candidate label; and classifying the text data through the second classification model to obtain the second candidate label. The first classification model is obtained through zero-sample training, so the first classification model can classify text data of a new category. The second classification model is obtained through label sample training, and its classification accuracy is high; then, based on the first candidate label and the second candidate label, a classification label for the text data is obtained. Combining the classification results of the first classification model and the second classification model, a comprehensive and accurate classification label can be obtained, thereby improving the classification accuracy.

[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0022] Figure 1 The figure is a flowchart of a text data classification method according to an exemplary embodiment.

[0023] Figure 2 yes Figure 1 Flowchart of step S120 in FIG.

[0024] Figure 3 The figure is a flowchart of a text data classification method according to an exemplary embodiment.

[0025] Figure 4 The figure is a block diagram of a text data classification device according to an exemplary embodiment.

[0026] Figure 5 is a block diagram of a vehicle according to an exemplary embodiment. DETAILED DESCRIPTION

[0027] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0028] During the use of the vehicle, users can report vehicle-related issues through the user feedback system. Backstage staff count these issues and label each issue so that each issue has a corresponding label, which is convenient for subsequent personnel to handle. As the number of vehicles increases, the number of issues reported by users continues to rise, and the number of labels has also become huge. Manual classification is difficult and has low accuracy.

[0029] In order to improve classification efficiency, related questions can be classified through a single classification model, and a classification label can be output to achieve automatic classification.

[0030] To solve the above problems, the present disclosure provides a text data classification method, see Figure 1 , the text data classification method can be applied to Figure 4 The text data classification device 200 shown, Figure 5 The vehicle 600, computer program product, and computer readable storage medium shown in FIG. Figure 1 The process shown is described in detail, and the text data classification method may include the following steps: Step S110: Obtain text data to be classified.

[0031] When the present disclosure is applied to a vehicle scenario, the text data includes data obtained by a feedback system corresponding to the vehicle. When a user encounters a vehicle-related problem during the use of the vehicle, the problem can be fed back through the feedback system corresponding to the vehicle, and the feedback system thereby obtains text data. Exemplarily, a user-side feedback system is configured on the vehicle, and an icon corresponding to the feedback system is displayed on the central control screen of the vehicle. The user enters the feedback interface by clicking on the feedback system, and feeds back relevant issues on the feedback interface, and the feedback system thereby obtains text data. For example, relevant issues are fed back in the form of text input on the feedback interface. For another example, after the user clicks the voice input button on the feedback interface, voice input is performed, and the feedback system converts the input voice into text, thereby feeding back relevant issues.

[0032] Optionally, the text input by the user is incorrect, or there is an error when the feedback system converts speech into text. This part of the text data cannot accurately reflect the problem reported by the user. The text data with the above problems can be cleared and the remaining text data can be classified.

[0033] Since the same user may encounter multiple problems and the number of vehicles is large, the amount of text data obtained is large. In order to facilitate the subsequent classification of these problems, the text data can be classified by assigning labels, as follows.

[0034] Step S120: classify the text data using a first classification model to obtain a first candidate label.

[0035] Input the text data into the first classification model to obtain the first candidate label output by the first classification model. The number of first candidate labels can be 1 or more. The first classification model is obtained through zero-sample training, that is, through zero-shot learning. The first classification model can use text data for classification without relying on a large amount of sample data annotated with labels, which means that the first classification model can classify text data of categories that have not appeared and has strong generalization ability. Exemplarily, the first classification model may include a recall model and a rearrangement model.

[0036] For example, if the text data is "On x month x day, I was driving on x road and collided with a large red truck. The airbag did not pop out, which was outrageous", the first candidate labels output by the first classification model may be "date", "x road", "truck", "airbag" and "damaged".

[0037] Step S130: classify the text data through a second classification model to obtain a second candidate label, wherein the first classification model is obtained through zero-sample training and the second classification model is obtained through labeled sample training and supervised learning.

[0038] The text data is input into the second classification model to obtain the second candidate label output by the second classification model, and the second candidate label can be 1. The second classification model is obtained by training with label samples. The second classification model obtained by training with labels has a higher accuracy of outputting the second candidate label. For example, the second classification model may include a single label model.

[0039] Combining the above example, the second candidate tag may be "airbag".

[0040] Step S140: Obtain a classification label for the text data according to the first candidate label and the second candidate label.

[0041] The first candidate label output by the first classification model and the second candidate label output by the second classification model are combined to obtain a classification label for the text data. This step combines the advantages of the two models, and can not only process text data of new categories, but also achieve high-precision classification of text data of existing categories.

[0042] Optionally, the number of classification labels can be 1 or more.

[0043] The text data classification method provided in this embodiment classifies the text data to be classified through a first classification model to obtain a first candidate label; and classifies the text data through a second classification model to obtain a second candidate label. The first classification model is obtained through zero-sample training, so the first classification model can classify text data of a new category. The second classification model is obtained through label sample training, and its classification accuracy is high. Then, based on the first candidate label and the second candidate label, a classification label for the text data is obtained. By combining the classification results of the first classification model and the second classification model, a comprehensive and accurate classification label can be obtained, thereby improving the classification accuracy.

[0044] In one embodiment, the first classification model may include a retrieval ranking model, for example, the first classification model includes a recall model and a re-ranking model. Figure 2, step S120 may include the following steps: Step S121: classify the text data using the recall model to obtain intermediate labels.

[0045] The recall model is obtained through zero-sample training. The recall model can classify text data according to the similarity between the text data and the text labels in the label library to obtain intermediate labels.

[0046] As a method, the recall model obtains the intermediate label in the following manner: vectorize the text data to obtain a text vector. And vectorize M text labels to obtain M label vectors, where M is a positive integer, and the M text labels can be all text labels in the label library. M can be 1000, 2000, etc. according to actual conditions. Then obtain the first similarity measurement value between the text vector and each label vector in the M label vectors to obtain M first similarity measurement values. The first similarity measurement value is used to characterize the similarity between the text vector and each label vector, and the first similarity measurement value is positively correlated with the similarity. For example, the first similarity measurement value can include cosine similarity distance, edit distance, etc. Then, based on the M first similarity measurement values, determine N intermediate labels from the M text labels, where N is a positive integer less than M, for example, N can be 20, 30, 35. Exemplarily, N first target similarity measurement values ​​are determined from the M first similarity measurement values, wherein the N first target similarity measurement values ​​are all greater than the first similarity measurement values ​​among the M first similarity measurement values ​​except the N first target similarity measurement values. It can be understood that the top N first similarity measurement values ​​in the sorting are selected as the first target similarity measurement values; and the text labels corresponding to the N first target similarity values ​​are determined from the M text labels as intermediate labels to obtain the N intermediate labels.

[0047] The first similarity metric value is the cosine similarity distance, and the first similarity metric value can be calculated by the following formula:

[0048] In the above formula, is the first similarity measure, A1 is the text vector, B1 is the label vector, a i is the element in the text vector A1, b i is the element in the label vector B1, and n is the number of elements.

[0049] Step S122: Obtain the first candidate label according to the text data and the intermediate label through the rearrangement model.

[0050] The rearrangement model is obtained through zero-sample training. The rearrangement model can classify the text data according to the similarity between the text data and the intermediate label to obtain the first candidate label.

[0051] As a method, the rearrangement model obtains the first candidate label in the following manner: obtain the second similarity measurement value between the text data and each of the N intermediate labels, and obtain N second similarity measurement values, the second similarity measurement value is used to characterize the similarity between the text data and the intermediate label, and the second similarity measurement value is positively correlated with the similarity. For example, the first similarity measurement value may include cosine similarity distance, edit distance, etc. According to the N second similarity measurement values, X first candidate labels are determined from the N intermediate labels, where X is a positive integer less than or equal to N, and X can be 4, 5, 6, etc. Exemplarily, X second target similarity measurement values ​​are determined from the N second similarity measurement values, where the X second target similarity measurement values ​​are all greater than the second similarity measurement values ​​of the N second similarity measurement values ​​except the X second target similarity measurement values. It can be understood that the first X second similarity measurement values ​​in the sorting are selected as the second similarity measurement values. The intermediate tags corresponding to the X second target similarity measurement values ​​are determined from the N intermediate tags as first candidate tags to obtain the X first candidate tags.

[0052] The second similarity metric value is the cosine similarity distance, and the second similarity metric value can be calculated by the following formula:

[0053] In the above formula, is the second similarity measure, A2 is the text data, B2 is the text label, a i2 is the element in text data A2, b i2 is the element in text label B2, and n is the number of elements.

[0054] In this embodiment, the first classification model includes a recall model and a rearrangement model. The recall model is used to classify text data to obtain intermediate labels, and then the classification results of the recall model are verified by the rearrangement model. The rearrangement model obtains the first candidate label based on the intermediate model and the text data, thereby improving the accuracy of the first classification model.

[0055] In one embodiment, when the number of first candidate labels and the number of classification labels are the same, step S140 includes: if the second candidate label is different from all X first candidate labels, the second candidate label replaces the first candidate label with the smallest second target similarity metric value among the X first candidate labels to obtain X classification labels. Since the second classification model is obtained by training with label samples, which is a supervised training method, the second candidate label classified by the second classification model has a high confidence level, and the second candidate label can be preferentially selected as the classification label, and then the label with the lowest second target similarity metric value among the X first candidate labels is discarded, and the remaining X-1 first candidate labels are used as classification labels to obtain X classification labels.

[0056] Optionally, the method also includes: sorting the X-1 first candidate tags in the classification tags in descending order according to the second target similarity measure value, and arranging the second candidate tags before the first candidate tag with the largest second target similarity measure value, and displaying the classification tags according to the tag sorting.

[0057] The user terminal of the labeler is installed with a feedback system of the backend, and the classification labels are displayed on the feedback system according to the label sorting, so that the labeler can select the classification label of the text data according to the displayed X classification labels, and use the classification label selected by the labeler as the final classification label of the text data. In the scheme of a single classification model, only one classification label is obtained. If the classification label does not match the actual text data, the labeler needs to find a label that matches the text data from a label library containing a large number of labels. This embodiment outputs X labels. If the labeler believes that the classification label with the highest ranking does not match the text data, he can also choose from other classification labels without having to find labels from the label library, thereby improving the efficiency of text data classification.

[0058] In another embodiment, step S140 includes: if the second candidate tag is the same as one of the X first candidate tags, taking the X first candidate tags as classification tags to obtain X classification tags.

[0059] Optionally, X-1 first candidate tags in the classification tags that are different from the second candidate tag are sorted in descending order according to the second target similarity measure value, and the first candidate tag that is the same as the second candidate tag is arranged before the first candidate tag with the largest second target similarity measure value, and the classification tags are displayed according to the tag sorting.

[0060] The user terminal of the labeler is installed with a backend feedback system. The classification labels are displayed in the feedback system in the order of labels, so that the labeler can select the classification label of the text data according to the displayed X classification labels, and the classification label selected by the labeler is used as the final classification label of the text data.

[0061] In another embodiment, the first classification model can predict text data of a newly appeared category, and the second classification model has a higher classification accuracy. Based on this, step S140 includes: if the second candidate label is different from the X first candidate labels, the text data to be classified may be text of a new category, therefore, the second candidate label is deleted, and the X first candidate labels obtained by the first classification model are used as classification labels. If the second candidate label is the same as one of the X first candidate labels, it means that the text to be classified may not contain text data of a new category, and the second candidate label obtained by the second classification model is used as the classification label.

[0062] It should be noted that the above text data is not limited to vehicle-related data, but can also be determined according to specific application scenarios. For example, when applied to home appliance scenarios, the text data includes data corresponding to the home appliance, and the text data can be data generated during the operation of the home appliance. The text data is classified to obtain classification labels, so that the home appliance manufacturer can assign corresponding maintenance personnel to repair the home appliance according to the classification labels.

[0063] For another example, when applied to a smartphone scenario, the text data may include data during the use of the smartphone.

[0064] For another example, it can also be applied to the scenario of classifying dynamic text. Then, the text data may include dynamic text, and the dynamic text may be news or social media content. When the text data is dynamic text on a short video software, after classifying the dynamic text on the short video software to obtain a classification label, relevant videos can be pushed based on the classification label.

[0065] The present disclosure also provides a text data classification method, see Figure 3 , the method comprises the following steps: Collect data. Collect historical user feedback data, including user feedback text and its corresponding labels.

[0066] Clean the data. Clean the data to remove low-quality data.

[0067] Model training: training recall models, re-ranking models, and single-label classification models.

[0068] (1) Retrieval model: The pre-trained model bge-large-zh-v1.5 based on the deep neural network dual-encoder architecture is trained to fit the dataset and aims to quickly recall relevant tags.

[0069] (2) Reranking model: The pre-trained model bge-reranker-large based on the deep neural network cross-encoder architecture is trained to fit the dataset and aims to rerank the recalled labels with high precision.

[0070] (3) Single-label classification model: Based on the pre-trained chinese-roberta-wwm-ext-large model, the dataset is fitted and trained to generate a single-label classification model to improve the accuracy of the top1 label.

[0071] Model deployment: Deploy the successfully trained model on the labeler’s device.

[0072] Label vectorization: Use the recall model to vectorize the labels in the label library to obtain label vectors.

[0073] Vectorization of text data: The text data is vectorized through the recall model to obtain text vectors.

[0074] Preliminary recall. The recall model is recalled in the following manner: vectorize the text data to obtain a text vector. And vectorize M text labels to obtain M label vectors, where M is a positive integer, and the M text labels can be all text labels in the label library. M can be 1000, 2000, etc. according to actual conditions. Then obtain the first similarity measure between the text vector and each label vector in the M label vectors to obtain M first similarity measures. The first similarity measure is used to characterize the similarity between the text vector and each label vector, and the first similarity measure is positively correlated with the similarity. For example, the first similarity measure may include cosine similarity distance, edit distance, etc. Then, based on the M first similarity measures, determine N intermediate labels from the M text labels.

[0075] Rearrangement. That is, the rearrangement model obtains the second similarity measurement value between the text data and each of the N intermediate labels, and obtains N second similarity measurement values, the second similarity measurement value is used to characterize the similarity between the text data and the intermediate labels, and the second similarity measurement value is positively correlated with the similarity. For example, the first similarity measurement value may include cosine similarity distance, edit distance, etc. According to the N second similarity measurement values, X first candidate labels are determined from the N intermediate labels.

[0076] Single-label classification. Classify text data using a single-label model.

[0077] Adjust the label order. Sort the X-1 first candidate labels in the classification label in descending order according to the second target similarity metric value, and arrange the second candidate label before the first candidate label with the largest second target similarity metric value. Or sort the X-1 first candidate labels in the classification label that are different from the second candidate label in descending order according to the second target similarity metric value, and arrange the first candidate label that is the same as the second candidate label before the first candidate label with the largest second target similarity metric value.

[0078] Labels are applied and displayed on the labeler's device in the order in which the labels are applied.

[0079] Based on the same inventive concept, the present disclosure also provides a text data classification device, see Figure 4 , the text data classification device 200 comprises: An acquisition module 210 is configured to acquire text data to be classified; A first obtaining module 220 is configured to classify the text data through a first classification model to obtain a first candidate label; A second obtaining module 230 is configured to classify the text data through a second classification model to obtain a second candidate label, wherein the first classification model is obtained through zero-sample training, and the second classification model is obtained through label sample training; The classification module 240 is configured to obtain a classification label for the text data according to the first candidate label and the second candidate label.

[0080] Optionally, the first classification model includes a recall model and a rearrangement model, and the first obtaining module 220 includes: An intermediate label obtaining module is configured to classify the text data through the recall model to obtain an intermediate label; The first candidate tag acquisition module is configured to obtain the first candidate tag according to the text data and the intermediate tag through the rearrangement model.

[0081] Optionally, the intermediate label obtaining module includes: A vectorization module is configured to vectorize the text data to obtain a text vector, and vectorize M text labels to obtain M label vectors, where M is a positive integer; A first similarity measurement value acquisition module is configured to acquire a first similarity measurement value between the text vector and each of the M label vectors to obtain M first similarity measurement values; The intermediate label selection module is configured to determine N intermediate labels from the M text labels according to the M first similarity measurement values, wherein N is a positive integer less than M.

[0082] Optionally, the intermediate label selection module includes: A first target similarity metric value determination module is configured to determine N first target similarity metric values ​​from the M first similarity metric values, wherein the N first target similarity metric values ​​are all greater than first similarity metric values ​​among the M first similarity metric values ​​except the N first target similarity metric values; The intermediate label obtaining module is configured to determine the text labels corresponding to the N first target similarity values ​​from the M text labels as intermediate labels, and obtain the N intermediate labels.

[0083] Optionally, the first candidate tag acquisition module includes: A second similarity measurement value acquisition module is configured to acquire a second similarity measurement value between the text data and each of the N intermediate tags to obtain N second similarity measurement values; The first candidate tag determination module is configured to determine X first candidate tags from the N intermediate tags according to the N second similarity measurement values, where X is a positive integer less than or equal to N.

[0084] Optionally, the first candidate tag determination module includes: A second target similarity measure value determination module is configured to determine X second target similarity measure values ​​from the N second similarity measure values, wherein the X second target similarity measure values ​​are all greater than the second similarity measure values ​​other than the X second target similarity measure values ​​among the N second similarity measure values; The first candidate tag obtaining module is configured to determine the intermediate tags corresponding to the X second target similarity measurement values ​​from the N intermediate tags as the first candidate tags, and obtain the X first candidate tags.

[0085] Optionally, the classification module 240 includes: The first classification module is configured to replace the first candidate label with the smallest second target similarity measure value among the X first candidate labels with the second candidate label if the second candidate label is different from all the X first candidate labels, so as to obtain X classification labels.

[0086] Optionally, the text data classification device 200 further includes: The first display module is configured to sort the X-1 first candidate tags in the classification tags in descending order according to the second target similarity measurement value, and arrange the second candidate tags before the first candidate tags with the largest second target similarity measurement value, and display the classification tags according to the tag sorting.

[0087] Optionally, the classification module 240 includes: The second classification module is configured to use the X first candidate tags as classification tags if the second candidate tag is the same as one of the X first candidate tags, so as to obtain X classification tags.

[0088] Optionally, the text data classification device 200 further includes: The second display module is configured to sort the X-1 first candidate tags in the classification tags that are different from the second candidate tag in order from large to small according to the second target similarity measurement value, and arrange the first candidate tag that is the same as the second candidate tag before the first candidate tag with the largest second target similarity measurement value, and display the classification tags according to the tag sorting.

[0089] Optionally, the text data includes data obtained by a feedback system corresponding to the vehicle.

[0090] Regarding the text data classification device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0091] The present disclosure also provides a computer-readable storage medium on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the text data classification method provided by the present disclosure are implemented.

[0092] Figure 5 6 is a block diagram of a vehicle 600 according to an exemplary embodiment. For example, the vehicle 600 may be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 600 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.

[0093] Reference Figure 5 , the vehicle 600 may include various subsystems, for example, an infotainment system 610, a perception system 620, a decision control system 630, a drive system 640, and a computing platform 650. The vehicle 600 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 600 may be interconnected by wire or wireless means.

[0094] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, and a navigation system, among others.

[0095] The perception system 620 may include several sensors for sensing information about the environment around the vehicle 600. For example, the perception system 620 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system or other positioning systems), an inertial measurement unit (IMU), a laser radar, a millimeter wave radar, an ultrasonic radar, and a camera.

[0096] The decision control system 630 may include a computing system, a vehicle controller, a steering system, a throttle, and a braking system.

[0097] The drive system 640 may include components that provide powered motion for the vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting energy provided by the energy source into mechanical energy.

[0098] Some or all functions of the vehicle 600 are controlled by a computing platform 650. The computing platform 650 may include at least one processor 651 and a memory 652, and the processor 651 may execute instructions 653 stored in the memory 652.

[0099] The processor 651 may be any conventional processor, such as a commercially available CPU. The processor may also include a graphics processor (Graphic Process Unit, GPU), a field programmable gate array (Field Programmable Gate Array, FPGA), a system on chip (System on Chip, SOC), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC) or a combination thereof.

[0100] The memory 652 may be implemented by any type of volatile or nonvolatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0101] In addition to the instructions 653 , the memory 652 may also store data, such as road maps, route information, and data such as the location, direction, and speed of the vehicle. The data stored in the memory 652 may be used by the computing platform 650 .

[0102] In the embodiment of the present disclosure, the processor 651 may execute the instruction 653 to complete all or part of the steps of the above-mentioned text data classification method.

[0103] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above-mentioned text data classification method when executed by the programmable device.

[0104] In addition, the word "exemplary" is used herein to indicate serving as an example, instance, or diagram. Any aspect or design described as "exemplary" in this article is not necessarily understood to be advantageous compared to other aspects or designs. On the contrary, the use of the word exemplary is intended to present concepts in a specific way. As used herein, the term "or" is intended to represent an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X applies A or B" is intended to represent any one of the natural inclusive arrangements. That is, if X applies A; X applies B; or X applies both A and B, "X applies A or B" is satisfied under any of the aforementioned examples. In addition, unless otherwise specified or clearly pointed to a singular form from the context, the articles "one" and "an" as used in this application and the appended claims are generally understood to mean "one or more".

[0105] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art after reading and understanding the specification and drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components (e.g., elements, resources, etc.) described above, unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific functions of the described components, even if the structure is not equivalent to the disclosed structure. In addition, although specific features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and beneficial to any given or specific application. In addition, with respect to "including", "having", "having", "having", or variations thereof used in a specific embodiment or claim, such terms are intended to be inclusive in a manner similar to the term "comprising".

[0106] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.

[0107] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A text data classification method, characterized in that: The method comprises: Get the text data to be classified; Classify the text data using a first classification model to obtain a first candidate label; Classify the text data by using a second classification model to obtain a second candidate label, wherein the first classification model is obtained by zero-sample training, and the second classification model is obtained by label sample training; A classification label for the text data is obtained according to the first candidate label and the second candidate label.

2. The method according to claim 1, characterized in that The first classification model includes a recall model and a rearrangement model, and the text data is classified by the first classification model to obtain a first candidate label, including: Classifying the text data by using the recall model to obtain an intermediate label; The first candidate label is obtained according to the text data and the intermediate label by the rearrangement model.

3. The method according to claim 2, characterized in that The recall model obtains the intermediate labels in the following way: Vectorizing the text data to obtain a text vector, and vectorizing the M text labels to obtain M label vectors, where M is a positive integer; Obtaining a first similarity measurement value between the text vector and each of the M label vectors to obtain M first similarity measurement values; According to the M first similarity measurement values, N intermediate labels are determined from the M text labels, where N is a positive integer less than M.

4. The method according to claim 3, characterized in that Determining N intermediate labels from the M text labels according to the M first similarity measurement values ​​includes: Determine N first target similarity measurement values ​​from the M first similarity measurement values, wherein the N first target similarity measurement values ​​are all greater than first similarity measurement values ​​other than the N first target similarity measurement values ​​among the M first similarity measurement values; The text labels corresponding to the N first target similarity values ​​are determined from the M text labels as intermediate labels to obtain the N intermediate labels.

5. The method according to claim 3, characterized in that: The rearrangement model obtains the first candidate label in the following manner: Obtaining a second similarity measurement value between the text data and each of the N intermediate tags to obtain N second similarity measurement values; According to the N second similarity measurement values, X first candidate tags are determined from the N intermediate tags, where X is a positive integer less than or equal to N.

6. The method according to claim 5, characterized in that Determining X first candidate tags from the N intermediate tags according to the N second similarity measurement values ​​includes: Determine X second target similarity measurement values ​​from the N second similarity measurement values, wherein the X second target similarity measurement values ​​are all greater than the second similarity measurement values ​​other than the X second target similarity measurement values ​​among the N second similarity measurement values; The intermediate tags corresponding to the X second target similarity measurement values ​​are determined from the N intermediate tags as first candidate tags to obtain the X first candidate tags.

7. The method according to claim 6, characterized in that The step of obtaining a classification label for the text data according to the first candidate label and the second candidate label includes: If the second candidate label is different from all the X first candidate labels, the first candidate label with the smallest second target similarity measure value among the X first candidate labels is replaced by the second candidate label to obtain X classification labels.

8. The method according to claim 7, characterized in that The method further comprises: Sort the X-1 first candidate tags in the classification tags in descending order according to the second target similarity metric value, arrange the second candidate tags before the first candidate tags with the largest second target similarity metric value, and display the classification tags according to the tag sorting.

9. The method according to claim 6, characterized in that The step of obtaining a classification label for the text data according to the first candidate label and the second candidate label includes: If the second candidate tag is the same as one of the X first candidate tags, the X first candidate tags are used as classification tags to obtain X classification tags.

10. The method according to claim 9, characterized in that The method further comprises: Sort the X-1 first candidate tags in the classification tags that are different from the second candidate tag in order from large to small according to the second target similarity measure value, and arrange the first candidate tag that is the same as the second candidate tag before the first candidate tag with the largest second target similarity measure value, and display the classification tags according to the tag sorting.

11. The method according to any one of claims 1 to 10, characterized in that: The text data includes data obtained by a feedback system corresponding to the vehicle.

12. A text data classification device, characterized in that: The device comprises: An acquisition module, configured to acquire text data to be classified; A first obtaining module is configured to classify the text data through a first classification model to obtain a first candidate label; A second obtaining module is configured to classify the text data through a second classification model to obtain a second candidate label, wherein the first classification model is obtained through zero-sample training, and the second classification model is obtained through labeled sample training; The classification module is configured to obtain a classification label for the text data according to the first candidate label and the second candidate label.

13. A vehicle, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the steps of any one of the methods described in claims 1 to 11 when executing the instructions.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 11 are implemented.

15. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Text classification method and device, electronic equipment and computer readable storage medium

    CN113239204A

  • Zero sample text classification method and system based on CLIP and medium

    CN116701637A

  • Text content multi-label classification method and device

    CN117150026A

  • Zero sample classification method and device, electronic equipment and medium

    CN119312130A