Data classification method and apparatus

By determining the target coarse category and historical data of the data to be classified, adjusting the granularity of the classification standard knowledge base, and using a large language model to identify customer complaint types, the problem of insufficient accuracy of large model identification is solved, and high-precision and high-accuracy customer complaint type identification is achieved.

CN119691592BActive Publication Date: 2026-02-13BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411775345.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2026-02-13
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

Existing customer complaint type identification solutions based on large models require a large amount of high-precision data resources and it is difficult to guarantee the accuracy of the identification results in the long term.

Method used

By determining the knowledge data under the target coarse category to which the data to be classified belongs, identifying target historical data that is similar to the knowledge data, and adjusting the granularity of the classification standard knowledge base based on the classification accuracy of the historical data, the classification results of unclassified data are determined using a large language model, and the granularity of the classification standard is improved over time.

Benefits of technology

By utilizing knowledge data and historical data from coarse classification, high-precision and high-accuracy identification of customer complaint types was achieved, and the accuracy and precision of the classification results continued to improve over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691592B_ABST
    Figure CN119691592B_ABST
Patent Text Reader

Abstract

The application discloses a data classification method and device. One embodiment of the method comprises: determining knowledge data under a target coarse classification to which to-be-classified data belongs; determining target historical data similar to the knowledge data; in response to determining that the target historical data includes unclassified data, determining a target fineness of a classification standard to be adopted by a classification standard knowledge base according to an accuracy rate of a classification result of historical to-be-classified data; determining and storing a classification result of the unclassified data by a large language model according to the classification standard knowledge base adopting the classification standard of the target fineness, and the fineness of the classification standard of the classification standard knowledge base is improved in stages; and determining a classification result of the to-be-classified data according to the classification result of the target historical data. The application can obtain a classification result with high precision and high accuracy by using low-precision and easily-obtained data resources such as knowledge data under coarse classification, historical data and a classification standard knowledge base, and the fineness of the classification standard of the classification standard knowledge base is improved in stages, so that the precision and accuracy of the classification result of the to-be-classified data are continuously improved over time.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of computer, in particular to the technical field of large model and natural language understanding, and more particularly to a data classification method and device, computer readable medium, electronic equipment and computer program product. BACKGROUND

[0002] In the retail advertising business, there are a large number of customer complaint conversations (i.e., VOB, voice of business) every year. In order to avoid the recurrence of customer complaints at the root and improve the customer experience of merchants, identifying the complaint type of customer complaint conversations is the primary task. In the existing scheme for identifying complaint types based on large models, a large model with high accuracy can be obtained only based on learning of a large amount of high-precision data resources, but a large amount of high-precision data resources is difficult to obtain, and the large model cannot guarantee the accuracy of the identification results in the long term. SUMMARY

[0003] Embodiments of the present application provide a data classification method, device, computer readable medium and electronic equipment.

[0004] In a first aspect, the embodiments of the present application provide a data classification method, comprising: determining knowledge data to which target coarse classification to which the to-be-classified data belongs; determining target historical data similar to the knowledge data; in response to determining that the target historical data includes unclassified data, determining a target fineness of a classification standard to be used by a classification standard knowledge base according to an accuracy rate of a classification result of historical to-be-classified data, wherein the classification result of the historical to-be-classified data is determined based on a classification standard knowledge base using a current fineness of the classification standard; determining and storing a classification result of the unclassified data by a large language model according to a classification standard knowledge base using a classification standard of the target fineness; and determining a classification result of the to-be-classified data according to the classification result of the target historical data.

[0005] In some examples, the determining of the classification result of the to-be-classified data according to the classification result of the target historical data comprises: determining a weight of the classification result of the target historical data according to a determination manner of the classification result of the target historical data, wherein the determination manner comprises: determining the classification result by querying classified data in the target historical data, or determining the classification result of unclassified data according to the classification standard knowledge base; and determining the classification result of the to-be-classified data according to the respective classification result and weight of the target historical data.

[0006] In some examples, the manner of determining the classification result of the target historical data comprises: in response to the classification result of the target historical data being determined based on the query classified data, determining the weight of the classification result of the classified data according to a strategy that the weight increases step by step with the lapse of the time period; and in response to the classification result of the target historical data being determined according to the classification standard knowledge base, determining the weight of the classification result of the unclassified data according to a strategy that the weight decreases step by step with the lapse of the time period.

[0007] In some examples, the manner of determining the knowledge data under the target coarse classification to which the data to be classified belongs comprises: determining a plurality of approximate knowledge data similar to the data to be classified from the coarse classification knowledge base; determining the aggregation degree of the plurality of approximate knowledge data based on the proportion of the approximate knowledge data belonging to the same coarse classification in the plurality of approximate knowledge data; and determining the target coarse classification and the knowledge data under the target coarse classification according to the aggregation degree of the plurality of approximate knowledge data.

[0008] In some examples, the manner of determining the target coarse classification according to the aggregation degree of the plurality of approximate knowledge data comprises: in response to determining that the aggregation degree is less than a first preset threshold, determining new approximate knowledge data similar to the data to be classified from the coarse classification knowledge base until it is determined that the aggregation degree of all the approximate knowledge data obtained at present under one coarse classification in the coarse classification knowledge base is greater than or equal to the first preset threshold, or all the knowledge data under one coarse classification in the coarse classification knowledge base is included in all the approximate knowledge data obtained at present; and determining the coarse classification as the target coarse classification.

[0009] In some examples, before the target historical data similar to the knowledge data is determined, the method further comprises: determining whether the historical data includes to-be-changed data requiring to change the classification result according to the fineness of the classification result corresponding to the historical data; and in response to determining that the historical data includes the to-be-changed data, deleting the classification result corresponding to the to-be-changed data.

[0010] In some examples, the manner of determining whether the historical data includes to-be-changed data requiring to change the classification result according to the fineness of the classification result corresponding to the historical data comprises: determining whether the historical data includes the to-be-changed data according to the fineness of the classification result corresponding to the historical data and the fineness of the classification result of the data to be classified that can be obtained at present.

[0011] In some examples, the manner of determining the target historical data similar to the knowledge data comprises: determining the number of the target historical data required to be recalled, wherein the number increases step by step with the lapse of the time period; and determining the number of target historical data similar to the knowledge data from the historical data.

[0012] In some examples, the determining, according to the classification results of the target historical data respectively corresponding to the weights, the classification result of the to-be-classified data comprises: in response to the aggregation degree of the classification results of the plurality of target historical data respectively corresponding to the weights in one classification being less than the second preset threshold, determining new target historical data similar to the knowledge data from the historical data, until the aggregation degree of the classification results of all the target historical data respectively corresponding to the weights in one classification is greater than or equal to the second preset threshold according to the weights of all the target historical data respectively corresponding to the weights, and determining the classification with the aggregation degree greater than or equal to the second preset threshold as the classification result of the to-be-classified data.

[0013] In some examples, the determining, according to the accuracy of the classification results of the historical to-be-classified data, the target fineness of the classification standard to be adopted by the classification standard knowledge base comprises: in response to the accuracy of the classification results of the historical to-be-classified data exceeding a preset accuracy threshold, increasing the current fineness of the classification standard adopted by the classification standard knowledge base, and determining the target fineness.

[0014] In some examples, the method further comprises: storing the to-be-classified data and the classification result of the to-be-classified data as historical data.

[0015] In a second aspect, an embodiment of the present application provides a data classification device, comprising: a first determining unit configured to determine knowledge data in a target coarse classification to which to-be-classified data belongs; a second determining unit configured to determine target historical data similar to the knowledge data; a third determining unit configured to, in response to determining that the target historical data includes unclassified data, determine a target fineness of a classification standard to be adopted by a classification standard knowledge base according to an accuracy of classification results of historical to-be-classified data, wherein the classification results of the historical to-be-classified data are determined based on the classification standard knowledge base adopting a classification standard with a current fineness; a fourth determining unit configured to determine and store a classification result of the unclassified data by a large language model according to the classification standard knowledge base adopting the classification standard with the target fineness; and a fifth determining unit configured to determine a classification result of the to-be-classified data according to the classification results of the target historical data.

[0016] In some examples, the fifth determining unit is further configured to: determine weights of the classification results of the target historical data according to a determination manner of the classification results of the target historical data, wherein the determination manner comprises: determining the classification results by querying classified data in the target historical data, or determining the classification results of unclassified data according to the classification standard knowledge base; and determining the classification result of the to-be-classified data according to the classification results of the target historical data respectively corresponding to the weights.

[0017] In some examples, the fifth determining unit is further configured to: in response to the classification result of the target historical data being determined based on the queried classified data, determine the weight of the classification result of the classified data according to a strategy that the weight increases step by step over time; and in response to the classification result of the target historical data being determined according to the classification standard knowledge base, determine the weight of the classification result of the unclassified data according to a strategy that the weight decreases step by step over time.

[0018] In some examples, the first determining unit is further configured to: determine a plurality of approximate knowledge data similar to the data to be classified from the coarse classification knowledge base; determine the aggregation degree of the plurality of approximate knowledge data based on the proportion of the approximate knowledge data belonging to the same coarse classification in the plurality of approximate knowledge data; and determine the target coarse classification and the knowledge data under the target coarse classification according to the aggregation degree of the plurality of approximate knowledge data.

[0019] In some examples, the first determining unit is further configured to: in response to determining that the aggregation degree is less than a first preset threshold, determine new approximate knowledge data similar to the data to be classified from the coarse classification knowledge base until it is determined that the aggregation degree of all the approximate knowledge data obtained at present under one coarse classification in the coarse classification knowledge base is greater than or equal to the first preset threshold, or all the approximate knowledge data obtained at present includes all the knowledge data under one coarse classification in the coarse classification knowledge base; and determine the coarse classification as the target coarse classification.

[0020] In some examples, the apparatus further includes a changing unit configured to: before the target historical data similar to the knowledge data is determined, determine whether the historical data includes to-be-changed data requiring a change in the classification result according to the fineness of the classification result corresponding to the historical data; and in response to a determination that the historical data includes the to-be-changed data, delete the classification result corresponding to the to-be-changed data.

[0021] In some examples, the changing unit is further configured to: determine whether the historical data includes to-be-changed data according to the fineness of the classification result corresponding to the historical data and the fineness of the classification result of the data to be classified that can be obtained at present.

[0022] In some examples, the second determining unit is further configured to: determine the number of target historical data required for recall, wherein the number increases step by step over time; and determine the number of target historical data similar to the knowledge data from the historical data.

[0023] In some examples, the fifth determining unit is further configured to: in response to the aggregation degree of the classification results of the number of target historical data in the one classification being less than the second preset threshold according to the weights respectively corresponding to the number of target historical data, determine new target historical data similar to the knowledge data from the historical data until the aggregation degree of the classification results of all the target historical data in the one classification is greater than or equal to the second preset threshold according to the weights respectively corresponding to all the target historical data obtained at present, and determine the classification with the aggregation degree greater than or equal to the second preset threshold as the classification result of the data to be classified.

[0024] In some examples, the third determining unit is further configured to: in response to the accuracy of the classification result of the historical data to be classified exceeding a preset accuracy threshold, improve the current fineness of the classification standard adopted by the classification standard knowledge base, and determine the target fineness.

[0025] In some examples, the apparatus further includes a storage unit configured to store the data to be classified as historical data, and store the data to be classified and the classification result of the data to be classified.

[0026] In a third aspect, an embodiment of the present application provides a computer readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect.

[0027] In a fourth aspect, an embodiment of the present application provides an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the implementations of the first aspect.

[0028] In a fifth aspect, a computer program product is provided, including: a computer program, wherein the computer program, when executed by a processor, implements the method described in any of the implementations of the first aspect.

[0029] The data classification method and apparatus provided in this application provide the following steps: First, determine the knowledge data under the target coarse category to which the data to be classified belongs. Second, determine the target historical data that is similar to the knowledge data. Third, in response to determining that the target historical data includes unclassified data, determine the target granularity of the classification standard to be adopted by the classification standard knowledge base based on the accuracy of the classification results of the historical data to be classified. The classification results of the historical data to be classified are determined based on the classification standard knowledge base using the current granularity classification standard. Fourth, through a large language model, determine and store the classification results of the unclassified data based on the classification standard knowledge base using the target granularity classification standard. Fifth, determine the classification result of the data to be classified based on the classification results of the target historical data. This allows for the acquisition of classification results with higher precision and accuracy by utilizing readily available data resources such as knowledge data under the coarse category, historical data, and the classification standard knowledge base. Furthermore, the granularity of the classification standard in the classification standard knowledge base increases progressively over time, resulting in a continuous improvement in the precision and accuracy of the classification results of the data to be classified over time. Attached Figure Description

[0030] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0031] Figure 1 This is an exemplary system architecture diagram in which one embodiment of this application can be applied;

[0032] Figure 2 This is a flowchart of an embodiment of the data classification method according to this application;

[0033] Figure 3 This is a schematic diagram illustrating an application scenario of the data classification method according to this embodiment;

[0034] Figure 4 This is a flowchart of yet another embodiment of the data classification method according to this application;

[0035] Figure 5 This is a flowchart of an embodiment of the VOB data classification method according to this application;

[0036] Figure 6 A structural diagram of an embodiment of the data classification apparatus according to this application;

[0037] Figure 7 A schematic diagram of the structure of a computer system suitable for implementing the embodiments of this application. Detailed Implementation

[0038] The application will be described in further detail below with reference to the drawings and embodiments. It is to be understood that the specific embodiments described herein are intended to be illustrative only and not limiting of the application. Additionally, it is to be understood that the drawings are diagrammatic and schematic and that therefore their dimensions, positions and the like can not bear a strict relation to how they would appear on a real device.

[0039] It should be noted that the embodiments and features of the embodiments in the present application can be combined with each other without conflict. The present application will be described in detail below with reference to the drawings and embodiments.

[0040] It should be noted that in the technical solutions of the present application, the collection, collection, updating, analysis, processing, use, transmission, storage and the like of user personal information comply with relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data, and to maintain user personal information security, network security and national security.

[0041] Figure 1 An exemplary architecture 100 to which the data classification method and apparatus of the present application can be applied is shown.

[0042] As shown in Figure 1 The system architecture 100 can include terminal devices 101, 102, 103, a network 104 and a server 105. The terminal devices 101, 102, 103 are communicatively connected to form a topological network, and the network 104 is used as a medium to provide communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0043] The terminal devices 101, 102, 103 can interact with the server 105 through the network 104 to receive or send data, etc. The terminal devices 101, 102, 103 can be hardware devices or software that support network connection to interact and process data. When the terminal devices 101, 102, 103 are hardware, they can be various electronic devices that support network connection, information acquisition, interaction, display, processing, etc., including but not limited to smartphones, vehicle-mounted computers, tablet computers, e-book readers, laptop computers and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they can be installed in the above-mentioned electronic devices. They can be implemented as multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made herein.

[0044] Server 105 can be a server that provides various services. For example, it can be a background processing server that determines the classification result of the data to be classified provided by terminal devices 101, 102, and 103 using low-precision, easily accessible data resources such as knowledge data under coarse classification, historical data, and classification standard knowledge base. As an example, server 105 can be a cloud server.

[0045] It should be noted that a server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software programs or software modules (such as software programs or software modules used to provide distributed services), or as a single software program or software module. No specific limitations are made here.

[0046] It should also be noted that the data classification method provided in the embodiments of this application can be executed by a server, by a terminal device, or by a combination of both. Accordingly, the various parts (e.g., units) of the data classification device can be all located in the server, all located in the terminal device, or located separately in the server and the terminal device.

[0047] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. When the electronic devices on which the data classification method runs do not need to transmit data with other electronic devices, the system architecture may only include the electronic devices on which the data classification method runs (e.g., servers or terminal devices).

[0048] Continue to refer to Figure 2 The flowchart 200, illustrating an embodiment of a data classification method, includes the following steps:

[0049] Step 201: Determine the knowledge data under the target coarse category to which the data to be classified belongs.

[0050] In this embodiment, the entity executing the data classification method (e.g.) Figure 1 The terminal device or server in the process can obtain the data to be classified from a remote location or from a local location through a wired network connection or a wireless network connection, and determine the knowledge data under the target coarse category to which the data to be classified belongs.

[0051] The to-be-classified data can be various data represented in forms such as text, voice, image, etc. For example, the to-be-classified data is text data, specifically customer complaint dialogue data, and it is needed to determine the complaint type to which the customer complaint dialogue data belongs. For another example, the to-be-classified data is image data, specifically image data uploaded by a user through a social platform or a short video platform, and it is needed to determine the type to which the image data belongs, so as to perform targeted data pushing on other users of interest.

[0052] As an example, first, for different application fields, different coarse classification knowledge bases are constructed in a targeted manner, and the coarse classification knowledge base includes multiple coarse classifications, and each coarse classification includes corresponding knowledge data. The knowledge data under the coarse classification represents the data types included under the coarse classification.

[0053] Among them, the classification standard for performing coarse classification on data in an application field includes: 1. Significantly lower than the classification granularity required by the classification business in the application field. For example: the classification business requires 200 classifications, and the coarse classification is suggested to be about 20. 2. Each coarse classification has or exceeds a preset number threshold (for example, 3) of knowledge.

[0054] Continuing to take the customer complaint dialogue scene as an example, the knowledge data under the coarse classification is the complaint type that may appear under the classification.

[0055] For example, the coarse classification is: advertisement - delivery end - CPS business (commissioned advertising) - plan setting link - commission setting

[0056] Knowledge 1: The commission setting threshold does not meet the merchant's expectation

[0057] Knowledge 2: Error occurs when setting commission

[0058] Knowledge 3: Special products hope to specify fixed commission

[0059] Based on the above 3 knowledge, the VOB under the business classification of “advertisement - delivery end - CPS business (commissioned advertising) - plan setting link - commission setting” is divided into the above 3 categories.

[0060] Then, the feature distance between the to-be-classified data and the coarse classification in the coarse classification knowledge base under the application field to which the to-be-classified data belongs is determined. Specifically, the feature extraction can be performed on the to-be-classified data and the knowledge data under the coarse classification in the coarse classification set, respectively, to obtain a classification data vector and a knowledge data vector, and then the feature distance between the classification data vector and the knowledge data vector under each coarse classification is calculated.

[0061] Finally, the coarse classification with the closest distance is determined as the target coarse classification, and then the knowledge data under the target coarse classification is determined.

[0062] In some optional implementations of the embodiment, the execution subject can execute the step 201 in the following manner:

[0063] First, the execution subject determines a plurality of approximate knowledge data from the coarse classification knowledge base that are approximate to the data to be classified.

[0064] As an example, the execution subject can determine the distance between the knowledge data in each coarse classification in the coarse classification knowledge base and the data to be classified in the feature vector, and sort the knowledge data in the coarse classification knowledge base in the order of distance from small to large, and then determine the plurality of knowledge data in the front of the order as the plurality of approximate knowledge data.

[0065] Second, the execution subject determines the aggregation degree of the plurality of approximate knowledge data based on the proportion of the approximate knowledge data belonging to the same coarse classification in the plurality of approximate knowledge data.

[0066] As an example, the execution subject can determine the coarse classification to which each of the plurality of approximate knowledge data belongs, and determine the number and proportion of the approximate knowledge data in each coarse classification; and take the maximum proportion as the aggregation degree of the plurality of approximate knowledge data. The proportion represents the proportion of the data amount of the approximate knowledge data in each coarse classification in the total number corresponding to the plurality of approximate knowledge data.

[0067] Third, the execution subject determines the target coarse classification and the knowledge data in the target coarse classification according to the aggregation degree of the plurality of approximate knowledge data.

[0068] In the embodiment, in response to the aggregation degree exceeding a first preset threshold, the coarse classification corresponding to the aggregation degree is determined as the target coarse classification, and then the knowledge data in the target coarse classification in the coarse classification knowledge base is determined.

[0069] The first preset threshold can be set according to actual conditions, for example, 0.5 or 0.75.

[0070] In the implementation, a specific manner of determining the target coarse classification to which the data to be classified belongs and the knowledge data in the target coarse classification is provided, and the accuracy of the determined knowledge data is improved based on the aggregation degree of the plurality of approximate knowledge data approximate to the data to be classified.

[0071] In some optional implementations of the embodiment, the execution subject can execute the third step in the following manner:

[0072] Firstly, in response to determining that the aggregation degree is less than the first preset threshold, new approximate knowledge data similar to the data to be classified is determined from the coarse classification knowledge base until it is determined that the aggregation degree of all the approximate knowledge data obtained at present under one coarse classification in the coarse classification knowledge base is greater than or equal to the first preset threshold, or all the knowledge data under one coarse classification in the coarse classification knowledge base is included in all the approximate knowledge data obtained at present.

[0073] Then, the coarse classification is determined as the target coarse classification.

[0074] For example, in response to determining that the aggregation degree is less than the first preset threshold, a specified number of new approximate knowledge data similar to the data to be classified is determined from the coarse classification knowledge base, and the aggregation degree of all the approximate knowledge data obtained at present (including the approximate knowledge data in the third step and the specified number of new approximate knowledge data) is determined; in response to the aggregation degree of all the approximate knowledge data obtained at present under one coarse classification in the coarse classification knowledge base being greater than or equal to the first preset threshold, or all the knowledge data under one coarse classification in the coarse classification knowledge base being included in all the approximate knowledge data obtained at present, the coarse classification is determined as the target coarse classification.

[0075] In response to determining that the aggregation degree of all the approximate knowledge data obtained at present under one coarse classification in the coarse classification knowledge base is still less than the first preset threshold, and all the knowledge data under one coarse classification in the coarse classification knowledge base is not included in all the approximate knowledge data obtained at present, a specified number of new approximate knowledge data similar to the data to be classified is continuously determined from the coarse classification knowledge base until the aggregation degree of all the approximate knowledge data obtained at present under one coarse classification in the coarse classification knowledge base is greater than or equal to the first preset threshold, or all the knowledge data under one coarse classification in the coarse classification knowledge base is included in all the approximate knowledge data obtained at present.

[0076] In the present implementation, the number of cycles is not limited, and only one of the two end conditions, i.e., the aggregation degree under one coarse classification in the coarse classification knowledge base being greater than or equal to the first preset threshold and all the knowledge data under one coarse classification in the coarse classification knowledge base being included in all the approximate knowledge data obtained at present, needs to be met to end the cycle.

[0077] In the present implementation, a determination method of the target coarse classification when the aggregation degree is less than the first preset threshold is provided, and the flexibility of the determination process is improved; since the data to be classified needs to be classified by the knowledge data under the target coarse classification, and the accuracy and precision are improved based on the stage, the flexible determination method in the present implementation is helpful to improve the accuracy and precision of the classification result.

[0078] Step 202, determining target historical data approximate to the knowledge data.

[0079] In this embodiment, the execution subject can determine the target historical data approximate to the knowledge data.

[0080] The execution subject can determine the target historical data approximate to the knowledge data from the historical data corresponding to the data to be classified.

[0081] As an example, for each knowledge data under the target coarse classification, the execution subject can determine at least one target historical data according to the distance between the historical data corresponding to the data to be classified and the knowledge data in the feature vector; and combine the target historical data corresponding to each of the multiple knowledge data under the target coarse classification to obtain the final multiple target historical data.

[0082] As another example, the execution subject can determine the feature vector distance between each knowledge data under the target coarse classification and each historical data; sort the historical data in the order of the feature vector distance from small to large, and select the multiple historical data in the front of the order as the target historical data.

[0083] In some optional implementations of this embodiment, the execution subject can execute the step 202 in the following manner:

[0084] First, determine the number of target historical data required to be recalled.

[0085] The number increases in steps with the passage of time period.

[0086] As an example, the execution subject can adopt the strategy that the number increases in steps with the passage of time period to determine the number of target historical data required to be recalled.

[0087] Second, determine the number of target historical data approximate to the knowledge data from the historical data.

[0088] In this implementation, the execution subject can determine the number of target historical data approximate to the knowledge data from the historical data according to the above two examples.

[0089] It can be understood that, with the passage of stages, the classification result of the data to be classified determined by using more number of target historical data helps to further improve the accuracy of the classification result.

[0090] Step 203, in response to determining that the target historical data includes unclassified data, determine the target fineness of the classification standard to be adopted by the classification standard knowledge base according to the accuracy of the classification result of the historical data to be classified.

[0091] In this embodiment, the execution subject can determine the target fineness of the classification standard to be used by the classification standard knowledge base according to the accuracy of the classification result of the historical data to be classified, in response to determining that the target historical data includes unclassified data. The classification result of the historical data to be classified is determined based on the classification standard knowledge base using the classification standard of the current fineness.

[0092] In this embodiment, the fineness of the classification standard of the classification standard knowledge base is improved in stages. The fineness of the classification standard of the classification standard knowledge base is different in different stages, and is improved as the stages progress. As an example, there are two stages, an early stage and a later stage. In the early stage, the classification granularity of the classification standard used by the classification standard knowledge base is relatively large. When entering the later stage, the fineness of the classification standard used by the classification standard knowledge base is improved, so that the classification granularity is reduced.

[0093] More specifically, the key feature of the classification standard knowledge base in the early stage is that the classification standard is coarse enough, so that the accuracy of the classification result of the unclassified data obtained in step 203 is not less than a first threshold (for example, 50%), and the accuracy of the classification result of the data to be classified obtained in subsequent step 204 is not less than a second threshold (for example, 80%). If the first threshold and the second threshold are not reached, the classification standards in the classification standard knowledge base need to be further combined. The classification standard in the classification standard knowledge base can be as low as binary classification identification of data. Binary classification identification refers to the judgment of data under the most coarse classification granularity, such as whether it is real customer complaint data.

[0094] If only binary classification identification of the unclassified data in the target historical data is performed, the classification result of the unclassified data obtained in step 203 selects one of the two classifications, so that the accuracy of the classification result of the unclassified data obtained in step 203 is not less than the first threshold, and the accuracy of the classification result of the data to be classified obtained in subsequent step 204 is not less than the second threshold.

[0095] The key feature of the classification standard knowledge base in the later stage is that the classification standard is fine enough, so that the classification result can directly guide the classification business optimization. "Directly guide the classification business optimization" is a relatively vague concept, which depends on the application of the classification business system of the present application. The judgment standard can be defined as: the fineness of the classification standard of the classification standard knowledge base in the early stage.

[0096] It should be noted that the above example is not limited to setting two stages, and multiple stages can be set according to the needs of implementation.

[0097] The historical data includes unclassified data and classified data, wherein the unclassified data does not include a classification result, and the classified data includes a classification result. It can be understood that at the initial stage, the historical data is mostly unclassified data, but as time goes by, more and more unclassified data is determined to have a classification result, so that the historical data is mostly classified data.

[0098] As an example, according to the accuracy of the classification result of the historical data to be classified, it can be determined whether the classification standard knowledge base of the large language model mounted with the classification standard of the current fineness can meet the classification task of the current stage. For example, if the accuracy of the classification result is high, it indicates that the classification standard knowledge base of the large language model mounted with the classification standard of the current fineness can meet the classification task of the current stage, and the fineness corresponding to the next stage can be taken as the target fineness; otherwise, the current fineness is continued to be taken as the target fineness.

[0099] In some optional implementations of the embodiment, the execution subject can execute the step 203 by increasing the current fineness of the classification standard adopted by the classification standard knowledge base and determining the target fineness in response to the accuracy of the classification result of the historical data to be classified exceeding a preset accuracy threshold.

[0100] As an example, the accuracy of the classification result of the historical data to be classified determined by the classification standard knowledge base adopting the current classification standard can be periodically counted, and in response to determining that the accuracy exceeds a preset accuracy threshold, the fineness of the classification standard of the classification standard knowledge base is increased in a preset manner.

[0101] For example, the fineness of the classification standard is divided into different levels, and in response to the accuracy of the classification result of the historical data to be classified determined by the classification standard knowledge base adopting the classification standard of the current level exceeding a preset accuracy threshold, the level of the fineness of the classification standard of the classification standard knowledge base is increased from the current level to the next level higher than the current level.

[0102] In the implementation, a specific way of improving the fineness of the classification standard of the classification standard knowledge base is provided, which helps to ensure continuous improvement of the fineness and accuracy of the classification result.

[0103] Step 204, determining and storing the classification result of the unclassified data by the large language model according to the classification standard knowledge base adopting the classification standard of the target fineness.

[0104] In the embodiment, the execution subject can determine and store the classification result of the unclassified data by the large language model according to the classification standard knowledge base adopting the classification standard of the target fineness.

[0105] The core implementation of the embodiment is the natural language understanding capability based on a large (language) model. Compared with the traditional classification scheme based on a large model, the large model in the embodiment does not require high accuracy, and the large model disclosed by the industry can meet the requirements of the embodiment after loading the classification standard knowledge base. The large model with the lowest classification effect can also achieve the effect of the present scheme, which is the core value of the embodiment. Therefore, the present application does not need to optimize the reasoning ability of the large model, and can be used in a conventional manner.

[0106] In step 205, the classification result of the to-be-classified data is determined according to the classification result of the target historical data.

[0107] In the embodiment, the execution subject can determine the classification result of the to-be-classified data according to the classification result of the target historical data.

[0108] As an example, when the number of target historical data is one, the classification result of the target historical data can be directly determined as the classification result of the to-be-classified data.

[0109] As another example, when the number of target historical data is multiple, the classification result obtained by election according to the classification results of the multiple target historical data can be determined as the classification result of the to-be-classified data. The election method includes but is not limited to more than half winning method and more than 2 / 3 winning method.

[0110] In some optional implementation modes of the embodiment, the execution subject can execute the above step 205 in the following manner:

[0111] First, the weight of the classification result of the target historical data is determined according to the determination method of the classification result of the target historical data.

[0112] The determination method includes querying the classified data in the target historical data to determine the classification result, or determining the classification result of the unclassified data according to the classification standard knowledge base.

[0113] For the classified data in the target historical data, the determination method of the classification result is to query the classified data in the target historical data to determine the classification result; for the classified data in the target historical data, the determination method of the classification result is to determine the classification result of the unclassified data according to the classification standard knowledge base.

[0114] In the implementation mode, different weights can be set for different determination methods, and then the weight of the classification result of the target historical data can be determined according to the determination method of the classification result of the target historical data.

[0115] Second, the classification result of the to-be-classified data is determined according to the classification result and the weight corresponding to each target historical data.

[0116] The execution subject can select the classification result of the to-be-classified data from the classification results of the target historical data according to the weights.

[0117] As an example, the execution subject can combine the same classification results of the classification results of the target historical data, determine the weights of the combined classification results, and determine the classification result corresponding to the maximum weight as the classification result of the to-be-classified data.

[0118] In the present implementation, different determination manners of the classification results of the target historical data correspond to different weights, and the weights determined based on the determination manners improve the accuracy of the classification result of the to-be-classified data.

[0119] In some optional implementations of the present embodiment, the execution subject can perform the first step in the following manner:

[0120] In response to the classification result of the target historical data being determined based on the query of the classified data, the weight of the classification result of the classified data is determined according to a strategy that the weight increases step by step with the passage of time.

[0121] Specifically, in response to the classification result of the target historical data being determined based on the query of the classified data, the weight of the classification result increases with the passage of stages. For example, in the early stage, the weight of the classification result of the classified data is small, and in the later stage, the weight of the classification result of the classified data is large.

[0122] In response to the classification result of the target historical data being determined according to the classification standard knowledge base, the weight of the classification result of the unclassified data is determined according to a strategy that the weight decreases step by step with the passage of time.

[0123] Specifically, in response to the classification result of the target historical data being determined according to the classification standard knowledge base, the weight of the classification result decreases with the passage of stages. For example, in the early stage, the weight of the classification result of the unclassified data is large, and in the later stage, the weight of the classification result of the unclassified data is small.

[0124] That is, in the early stage, the classification result of the target historical data determined according to the classification standard knowledge base is preferentially referred to, and in the later stage, the classification result determined based on the query of the classified data in the target historical data is preferentially referred to. Preferential reference refers to a weight ratio of more than 50%.

[0125] It can be understood that in the early stage of the classification business development, most of the historical data do not have classification results, at this time the inference ability of the large model mounted with the classification standard knowledge base is more important, and in the weight setting, the classification result of the target historical data determined according to the classification standard knowledge base should be given priority to reference. In the later stage of the classification business development, most of the historical data have classification results, at this time the existing information is more important, and in the weight setting, the classification result determined according to the classified data in the query target historical data should be given priority to reference.

[0126] In the present implementation, the weights of the classification results of the target historical data corresponding to different determination manners change step by step with the passage of time, and the change trend of the weights corresponding to different determination manners is different, which can make the reference focus different in the early stage and the late stage, and realize the automatic upgrading of the classification precision and accuracy of the classification business.

[0127] In some optional implementations of the present embodiment, the above-mentioned execution subject can execute the above-mentioned second step in the following manner:

[0128] In response to determining that the aggregation degree of the classification results of the above-mentioned number of target historical data in one classification is less than the second preset threshold value according to the weights of the above-mentioned number of target historical data, new target historical data similar to the knowledge data is determined from the historical data until the aggregation degree of the classification results of all target historical data in one classification is greater than or equal to the second preset threshold value according to the weights of all target historical data obtained at present, and the classification with the aggregation degree greater than or equal to the second preset threshold value is determined as the classification result of the data to be classified.

[0129] As an example, in response to determining that the aggregation degree of the classification results of the above-mentioned number of target historical data in one classification is less than the second preset threshold value according to the weights of the above-mentioned number of target historical data, a specified number of new target historical data similar to the knowledge data is determined from the historical data, and whether the aggregation degree of the classification results of all target historical data in one classification is greater than or equal to the second preset threshold value is determined according to the weights of all target historical data obtained at present; in response to determining that the classification with the aggregation degree greater than or equal to the second preset threshold value is determined as the classification result of the data to be classified.

[0130] In response to determining that the classification with the aggregation degree greater than or equal to the second preset threshold value is determined as the classification result of the data to be classified.

[0131] In the present embodiment, a specific determination manner of the classification result of the data to be classified is provided, the flexibility of the determination process is improved, the determination of the classification result of the data to be classified is ensured, and the accuracy of the classification result is improved.

[0132] With reference to the foregoing Figure 3 , Figure 3 is a schematic diagram 300 of an application scenario of the data classification method according to the present embodiment. In the Figure 3 application scenario, a user 301 sends data to be classified through a terminal device 302 to a server 303. After receiving the data to be classified, the server first determines knowledge data under a target coarse classification to which the data to be classified belongs from a coarse classification knowledge base; then determines target historical data similar to the knowledge data from a historical database; then, in response to determining that the target historical data includes unclassified data, determines and stores a classification result of the unclassified data according to a classification standard knowledge base, wherein the fineness of the classification standard of the classification standard knowledge base increases in steps as time elapses; and finally, determines a classification result of the data to be classified according to the classification result of the target historical data.

[0133] The method provided by the foregoing embodiments of the present application determines knowledge data under a target coarse classification to which data to be classified belongs; determines target historical data similar to the knowledge data; in response to determining that the target historical data includes unclassified data, determines a target fineness of a classification standard to be adopted by a classification standard knowledge base according to the accuracy of a classification result of historical data to be classified, wherein the classification result of the historical data to be classified is determined based on the classification standard knowledge base adopting a current fineness of the classification standard; determines and stores a classification result of the unclassified data according to the classification standard knowledge base adopting the classification standard of the target fineness through a large language model; and determines a classification result of the data to be classified according to the classification result of the target historical data, so that a classification result with high precision and high accuracy can be obtained by using data resources with low precision and easy to obtain, such as knowledge data under coarse classification, historical data, and a classification standard knowledge base, and the fineness of the classification standard of the classification standard knowledge base increases in steps as time elapses, so that the precision and accuracy of the classification result of the data to be classified continuously improve as time elapses.

[0134] In some optional implementations of the present embodiment, before step 202 is performed, the above-mentioned subject of execution can perform the following operations:

[0135] First, determine whether the historical data includes data to be changed that needs to change the classification result according to the fineness of the classification result corresponding to the historical data.

[0136] It can be understood that one of the purposes of the present application is to continuously improve the accuracy and precision of the classification results. In order to achieve the above purpose, it is necessary to improve the accuracy and precision of the classification results corresponding to the historical data as the stage progresses.

[0137] In the present implementation, the historical data to be changed can be determined from the historical data corresponding to the classification results according to the fineness of the classification results at intervals of a period of time.

[0138] As an example, the classification criteria are divided into first, second and third levels, and the fineness of the third classification criteria gradually increases. At the current stage, the large model can implement second-level classification on the unclassified data in the historical data, but the historical data still includes classification results using first-level classification criteria. At this time, the historical data using the classification results of the first-level classification criteria can be used as the historical data to be changed.

[0139] Second, in response to the determination including, deleting the classification results corresponding to the data to be changed.

[0140] The classification results corresponding to the data to be changed are deleted. That is, the data to be changed can be changed from classified data to unclassified data.

[0141] It can be understood that in the subsequent information processing process, in the execution process as shown in step 203, the data to be changed becomes unclassified data, which needs to be classified according to the classification criteria knowledge base to obtain classification results with higher fineness.

[0142] In the present implementation, a scheme for continuously improving the fineness of the classification results of the historical data is provided, which helps to quickly improve the accuracy and precision of the classification business.

[0143] In some optional implementations of the present embodiment, the above-mentioned execution subject can execute the above-mentioned first step in the following manner: according to the fineness of the classification results corresponding to the historical data and the fineness of the classification results of the data to be classified that can be obtained at present, determine whether the historical data includes data to be changed.

[0144] As an example, the classification criteria are divided into first, second and third levels, and the fineness of the third classification criteria gradually increases. At the current stage, the large model can implement second-level classification on the unclassified data in the historical data, but the historical data still includes classification results using first-level classification criteria. At this time, the historical data using the classification results of the first-level classification criteria can be used as the historical data to be changed.

[0145] It can be understood that the target historical data determined from the historical data is finally used to determine the classification result of the to-be-classified data, and therefore, the classification result of the historical data directly affects the classification result of the to-be-classified data. According to the fineness of the classification result corresponding to the historical data and the fineness of the classification result of the to-be-classified data that can be obtained at present, whether the to-be-changed data is included in the historical data is determined, and the determination accuracy of the to-be-changed data is further improved.

[0146] In some optional implementations of the embodiment, the execution subject can further perform the following operation: storing the to-be-classified data and the classification result of the to-be-classified data as historical data.

[0147] In the implementation, after obtaining the classification result of the to-be-classified data, the to-be-classified data and the classification result of the to-be-classified data are stored as historical data, so that the historical data is supplemented, the richness of the historical data is improved, and the accuracy of the classification result of the subsequent to-be-classified data is improved.

[0148] With reference to Figure 4 , a schematic flow 400 of yet another embodiment of the data classification method according to the present application is shown, including the following steps:

[0149] Step 401: determining, from the coarse classification knowledge base, a plurality of approximate knowledge data similar to the to-be-classified data.

[0150] Step 402: determining the aggregation degree of the plurality of approximate knowledge data based on the proportion of the approximate knowledge data belonging to the same coarse classification in the plurality of approximate knowledge data.

[0151] The aggregation degree of the plurality of knowledge data is used to represent the proportion information of the knowledge data belonging to the same coarse classification in the plurality of knowledge data.

[0152] Step 403: in response to determining that the aggregation degree is less than a first preset threshold, determining new approximate knowledge data similar to the to-be-classified data from the coarse classification knowledge base until it is determined that the aggregation degree of all the obtained approximate knowledge data in one coarse classification of the coarse classification knowledge base is greater than or equal to the first preset threshold, or all the obtained approximate knowledge data include all the knowledge data in one coarse classification of the coarse classification knowledge base.

[0153] Step 404: determining the coarse classification as a target coarse classification, and determining the knowledge data under the target coarse classification.

[0154] Step 405: determining whether the to-be-changed data that needs to change the classification result is included in the historical data according to the fineness of the classification result corresponding to the historical data.

[0155] Step 406, in response to the determination including, deleting the classification result corresponding to the data to be changed.

[0156] Step 407, determining the number of target historical data required to be recalled.

[0157] Wherein, the number increases stepwise with the passage of time period.

[0158] Step 408, determining the number of target historical data similar to the knowledge data from the historical data.

[0159] Step 409, in response to the determination that the target historical data includes unclassified data, and the accuracy of the classification result of the historical data to be classified exceeds the preset accuracy threshold, increasing the current fineness of the classification standard adopted by the classification standard knowledge base, and determining the target fineness.

[0160] Wherein, the classification result of the historical data to be classified is determined based on the classification standard knowledge base adopting the classification standard of the current fineness.

[0161] Step 410, determining and storing the classification result of the unclassified data according to the classification standard knowledge base adopting the classification standard of the target fineness through the large language model.

[0162] Step 411, in response to the classification result of the target historical data being determined based on the classified data, determining the weight of the classification result of the classified data according to the strategy that the weight increases stepwise with the passage of time period; in response to the classification result of the target historical data being determined according to the classification standard knowledge base, determining the weight of the classification result of the unclassified data according to the strategy that the weight decreases stepwise with the passage of time period.

[0163] Step 412, in response to determining that the aggregation degree of the classification result of each of the number of target historical data under one classification is less than the second preset threshold according to the weight corresponding to each of the number of target historical data, determining new target historical data similar to the knowledge data from the historical data, until determining that the aggregation degree of the classification result of each of all the target historical data under one classification is greater than or equal to the second preset threshold according to the weight corresponding to each of all the target historical data currently obtained, and determining the classification with the aggregation degree greater than or equal to the second preset threshold as the classification result of the unclassified data.

[0164] Step 413, storing the unclassified data and the classification result of the unclassified data as historical data.

[0165] As can be seen from the present embodiment, compared with Figure 2Compared with the corresponding embodiments, the flow 400 of the data classification method in the embodiment specifically illustrates the determination process of the knowledge data and the target historical data under the target coarse classification, the refinement process of the classification standard of the classification standard knowledge base, and the classification result of the to-be-classified data, so that the classification result with high precision and high accuracy can be obtained by using the knowledge data, the historical data, the classification standard knowledge base and other low-precision and easily-obtained data resources under coarse classification, so that the precision and accuracy of the classification result of the to-be-classified data are continuously improved over time.

[0166] With reference to the foregoing Figure 5 Taking the to-be-classified data as the to-be-classified complaint data, the execution process of the data classification method is specifically described as follows:

[0167] Step 1: Historical VOB data storage and vectorization

[0168] A historical VOB database is established, and all merchant and customer service dialogues are stored in the database. Each historical VOB data in the VOB database is vectorized, and the historical VOB vector is obtained and stored to establish a historical VOB vector library.

[0169] Step 1.1:

[0170] A database is built, and a VOB storage structure is defined to store historical VOB data.

[0171] The storage structure of the historical VOB data necessarily includes at least three fields: a unique VOB data identifier, historical VOB data, and a classification result. If the historical VOB data does not include the classification result, the classification result field is empty.

[0172] The database necessarily includes a function of persisting all historical VOB data, and optionally includes a function of web interface operation. The input link of the historical VOB data can be a data stream, a file or other formats, and the storage link of the historical VOB data can be a storage product such as mysql or redis.

[0173] As an example, first, a mysql database is built, and a database is created by using a create database command, and a table is created by using a create table command; then, the historical VOB data is saved in a txt file; finally, the file is imported into the table created by using the create table command by using the mysql import capability.

[0174] The database also needs to include the maintenance capability of data, which must include the capabilities of "query", "add", "modify", etc., and optionally includes the "delete" capability. The data maintenance capability necessarily contains the function of engineering execution operation, and the reason for containing the engineering operation is that the automatic upgrade capability of the present application relies on cross-system data modification and query operation.

[0175] The engineering execution operation refers to the capability of operating stored data through computer system server technology, and common methods include jdbc, etc. As an example, a jdbc connection is provided for a mysql database, which provides 4 operations of addition, deletion, modification, and query; a java service is set up to connect the jdbc connection, which can realize data maintenance in the form of an interface; when an external modification instruction requests the interface, data modification is realized; when an external query instruction requests the interface, data query is realized.

[0176] Step 1.2: Vectorization of historical VOB data

[0177] 1.2.1: Select a large model. The present application can select a large model with low parameters and low precision, as long as it can guarantee a classification accuracy of 50% (lower than the industry standard accuracy of 80%), which is the key feature that distinguishes the present application from traditional high-precision operations.

[0178] 1.2.2: Generate historical VOB vectors based on the vectorization capability of the large model.

[0179] As an example, a large model with low parameters and low precision is selected, and its input and output capabilities are enabled; each piece of historical VOB data in the historical VOB database in step 1.1 is input into the large language model, and the historical VOB vector corresponding to the historical VOB data is output.

[0180] 1.3: Store historical VOB vectors

[0181] 1.3.1: Select an open-source vector library in the industry, which contains necessary vector addition, deletion, modification, and query capabilities.

[0182] 1.3.2: Use the saving capability provided by the vector library to store the historical VOB vector.

[0183] The historical VOB vector in the historical VOB vector library necessarily contains 2 fields: VOB data identifier and historical VOB vector.

[0184] The historical VOB vector library provides the capability of searching historical VOB data based on historical VOB vectors. As an example, the implementation steps of searching historical VOB data based on historical VOB vectors include:

[0185] Based on the vector retrieval operation, at least one historical VOB vector is confirmed; the VOB data identifier corresponding to each of the at least one historical VOB vector is extracted; and the historical VOB data corresponding to the VOB data identifier is determined based on the query capability of the historical VOB database according to the VOB data identifier.

[0186] Step 2: Domain knowledge identification

[0187] The customized domain coarse classification knowledge base is vectorized, the knowledge election capability is provided, the election threshold setting capability is provided, and the weight automatic adjustment capability based on the election result and the threshold judgment is provided, and finally the knowledge vector of the knowledge data under the target coarse classification to which the VOB data to be classified belongs is output, which is used to assist the subsequent step 3 to identify the approximate VOB data.

[0188] Step 2.1: Create coarse classification knowledge base and vectorize knowledge data under coarse classification

[0189] 2.1.1: Domain experts manually perform knowledge coarse classification.

[0190] The standards for coarse classification include: 1. Significantly lower than the classification granularity required by the classification business. For example: the business requires 200 VOB classifications, and the coarse classification is recommended to be about 20. 2. Each classification has or exceeds 3 pieces of knowledge (necessary number of judgment aggregation degree and approximate vector recall).

[0191] 2.1.2: Vectorize knowledge data under each coarse classification by a large model and store. The storage structure includes at least three fields: knowledge data, coarse classification to which the knowledge data belongs, and knowledge vector.

[0192] Step 2.2: Approximate knowledge retrieval

[0193] Retrieve approximate knowledge vectors similar to the VOB vector of the VOB data to be classified from the coarse classification knowledge base.

[0194] The number of approximate knowledge vectors obtained is equal to the number of knowledge data under each coarse classification in the coarse classification knowledge base.

[0195] Step 2.3: Automatic adjustment

[0196] 2.3.1: Set the aggregation rate threshold in advance. For 3 approximate knowledge vectors, the aggregation rate threshold can be set to 2 / 3.

[0197] 2.3.2: Judge the aggregation degree of the approximate knowledge vectors retrieved in step 2.2, and compare it with the aggregation rate threshold in step 2.3.1.

[0198] 2.3.3: If the aggregation degree exceeds the aggregation rate threshold, go to step 2.4.

[0199] 2.3.4: If the degree of aggregation is lower than the aggregation rate threshold, it means that the large model cannot identify the coarse classification corresponding to the VOB data this time, and it is even less likely to identify the approximate knowledge vector. Therefore, flexible adjustment is needed.

[0200] The adjustment process is as follows: increase the number of recalled approximate knowledge vectors by a predetermined proportion or quantity, for example, from 3 to 4; perform the first type of judgment: judge the degree of aggregation of the approximate knowledge vector after increasing the number of recalls. If it meets the first preset threshold, the knowledge vector under the coarse classification corresponding to the first preset threshold is taken as the result and transferred to step 2.4; if it does not meet the first preset threshold, the second type of judgment is performed: whether the approximate knowledge vector after increasing the number of recalls contains all the knowledge vectors under a certain coarse classification. If it contains all the knowledge vectors under a certain coarse classification, the coarse classification is taken as the result and transferred to step 2.4; if it does not contain all the knowledge vectors under any coarse classification, the second type of judgment fails.

[0201] The above weight process is repeatedly performed until the degree of aggregation of the approximate knowledge vector after increasing the number of recalls meets the first preset threshold or contains all the knowledge vectors under a certain coarse classification.

[0202] Step 2.4: Output the knowledge vector under the determined target coarse classification as the domain knowledge vector.

[0203] Step 3: Approximate VOB vector retrieval

[0204] Set a threshold N, and select the N closest historical VOB vectors to the domain knowledge vector determined in step 2.4 from the historical VOB vectors generated in step 1.2 as the approximate VOB vectors.

[0205] 1) Calculate the vector distance between the stored historical VOB vectors and the domain knowledge vector;

[0206] 2) Sort the historical VOB vectors in ascending order of vector distance;

[0207] 3) Take the TopN of the sorting result;

[0208] 4) Output the TopN historical VOB vectors to the approximate vector list.

[0209] This step must include threshold setting capabilities, and optionally, threshold adjustment capabilities. The threshold can be set to a fixed value, such as n=5 (a more optimal value in business scenarios). If a fixed value is chosen, the number of approximate VOB vectors retrieved in step 3 and the number of approximate VOB vectors involved in the subsequent election process in step 5 are both fixed. Similarly, if threshold adjustment capabilities are included, the number of approximate VOB vectors involved in both the retrieval process in step 3 and the election process in step 5 are adjusted.

[0210] Step 4: Retrieve classification results or model inference classification results

[0211] Determine if a classification result already exists for the approximate VOB vector. If it does, retrieve the classification result. If it does not exist, infer the classification result using the larger model.

[0212] 4.1: Retrieval and Classification Results

[0213] 4.1.1: VOB data identifiers extracted from the list of approximate vectors generated in step 3.

[0214] 401.2: Based on the query function provided in step 1.1, use VOB data identifiers to query historical VOB data.

[0215] 401.3: Extract the classification result field from the historical VOB data and form an array.

[0216] The array example is as follows: [

[0218] {vobId:1, Recognition result: "Provides financial recharge inquiry capability"}

[0219] {vobId:2, Recognition result: "Provides express train deployment query capability"}

[0220] {vobId:3, Recognition Result: "Provides the ability to display plan queries"}

[0221] {vobId:4,Recognition result:“”}, ]

[0223] 4.1.4: Determine the existence of the classification result.

[0224] If the classification result field is empty, then there is no classification result in the original historical VOB data that was entered into the database. The classification result needs to be obtained through reasoning in the subsequent step 4.2.

[0225] 4.2: Model Inference Classification Results

[0226] The core implementation of the process of reasoning VOB data based on a large model is the natural language understanding capability of the large model. The application does not require the large model to have high accuracy compared to traditional solutions, so the large model in the industry can meet the requirements of the application after loading the classification standard knowledge base. The minimum effect of the large model can also achieve the effect of the application, which is the core value of the application. Therefore, the application does not need to optimize the reasoning capability of the large model, and can use a conventional way.

[0227] As an example, a large language model is selected; a classification standard knowledge base is set for the large language model; the large model is used to reason the candidate classification results of the historical VOB data based on the classification standard knowledge base and recall; and the recall result of the large model is used as the classification result of the historical VOB data.

[0228] The classification standard knowledge base used by the large model needs to be optimized to adapt to the VOB classification scene. Specifically, the optimization method of the classification standard knowledge base is to divide the fineness of the classification standard into at least two levels, at least one level of which can achieve relatively accurate classification, and at least one level of which can achieve high-precision classification.

[0229] The key feature of the classification standard that can achieve relatively accurate classification is that the classification granularity is coarse enough, so that the accuracy rate of the classification result obtained by the large model (reasoning accuracy) is not less than 50%, and the accuracy rate of the classification result after step 5 election (election accuracy) is not less than 80%. If this standard is not met, the classification granularity must be further combined, and at least only two classifications of historical VOB data are identified.

[0230] If only two classifications of historical VOB data are identified, the classification standard knowledge base only has two records: system optimization required and system optimization not required. When the large model reasons the results, it will choose one of the two classifications, which will inevitably achieve a reasoning accuracy of not less than 50% and an election accuracy of not less than 80%.

[0231] The key feature of the classification standard that can achieve high-precision classification is that the classification standard is fine enough, so that its classification result can directly guide business optimization. "Directly guide business optimization" is a relatively vague concept, which depends on the requirements of the business system applying the application. Its judgment standard can be defined as: more than two classifications (i.e., higher than the minimum value of the first classification).

[0232] It should be noted that the at least two levels of classification standard of the classification standard knowledge base do not take effect at the same time. For example, the first level of classification standard is selected in the early stage of execution, and the second level of classification standard is selected in the later stage of execution.

[0233] The indicator of the transition between the initial and later stages of implementation is that the election accuracy using the first-level classification standard exceeds the preset accuracy threshold (e.g., 92%).

[0234] 4.3: Generate a classification result array

[0235] After obtaining the classification result by reasoning through the historical VOB data excluding the classification result using a large model, an array of classification results corresponding to the approximate VOB vector is generated. The data structure used in the classification result array must contain three fields: whether it is the VOB data to be identified this time, the classification result, and the method of determining the classification result.

[0236] The value range for "Whether it is a VOB to be identified this time" is: true, false (or other fields that can represent true / false, such as 0 / 1). For each approximate VOB vector in the approximate vector list, if the classification result corresponding to the approximate VOB vector is obtained through the current inference of the large model, then this field is true; otherwise, it is false.

[0237] The value range for "Identification Method" is: model inference, result query (or other methods that can represent the binary classification, such as search / reasoning).

[0238] If the "Whether it is a VOB to be identified this time" field is true, then the value of "Identification method" must be model inference.

[0239] The classification result array is shown below as an example: [

[0241] {isOrigin:true, Identification Result: "Provides financial recharge query capability", source: "Result query"},

[0242] {isOrigin:false, Identification result: "Provides financial invoice query capability", source: "Result query"},

[0243] {isOrigin:false, Identification result: "Provides the ability to query express car consumption data", source: "Result query"},

[0244] {isOrigin:false, Identification Result: "Provides plan to demonstrate query capabilities", source: "Model Inference"} ]

[0246] Step 5: Select the classification results of the VOB data to be classified and store them in the database.

[0247] Set election rules and provide election capabilities, determine the classification results of the VOB data to be classified based on the election process, and store the VOB data to be classified and its classification results in the historical VOB database.

[0248] 5.1: Election

[0249] 5.1.1: Weight Setting

[0250] This step is used to weight and differentiate the classification results of the approximate VOB vectors determined in step 4, thereby generating more accurate election results, i.e., the classification results of the VOB data to be classified.

[0251] Based on the business development stage assessment described in step 4 above, different weighting models need to be selected: priority is given to classification results recorded in the historical VOB database, and priority is given to classification results obtained through model inference. Here, "priority" means that the weight percentage exceeds 50%.

[0252] Regarding the setting that prioritizes classification results obtained from model inference, in the early stages of business development, most historical VOB data do not have classification results. At this time, the ability of the large model to infer classification results is more important, and the classification results obtained from model inference should be given priority in the weight setting.

[0253] The setting that prioritizes classification results recorded in the historical VOB database is problematic in later stages of business development, when most historical VOB data already has classification results. At this point, the existing information in the historical VOB database becomes more important, and the weighting should prioritize the classification results recorded in the historical VOB database. This is also the core logic behind how the classification results of the VOB data to be classified in this embodiment can be automatically upgraded to truly high precision.

[0254] Here is an example of data with weights set: [

[0256] {isOrigin:true, Identification result: "Provides financial recharge inquiry capability", weight: "5"},

[0257] {isOrigin:false,Recognition result:“Provides financial invoice query capability”,weight:“3”},

[0258] {isOrigin:false, Identification result: "Provides the ability to query express car consumption data", weight: "2"},

[0259] {isOrigin:false, Identification Result: "Provides plan display query capabilities", weight: "2"} ]

[0261] 5.1.2: Election process setting

[0262] According to the set weight, the weight score of each record in the classification result array obtained in step 4.3 can be obtained, and the score is counted in the election; the aggregated dimension is the "identification result" field. As an example, the election method is, for example, a more than half winning method, a more than 2 / 3 winning method.

[0263] Optionally, the election failure mode can be selected when the aggregation degree does not exceed the second preset threshold. If this mode is selected, the threshold expansion and re-search capabilities need to be set in step 3 to determine new approximate historical vectors from the historical VOB database until the aggregation degree of the classification result of each approximate historical vector in a classification is greater than or equal to the second preset threshold according to the weight corresponding to each of the obtained approximate historical vectors.

[0264] 5.1.3: Election result

[0265] Regardless of the election method selected, the final intelligent result is to determine the winning side or not to determine the winning side. If the winning side is determined, the classification result of the winning side is determined as the classification result of the VOB data to be classified.

[0266] If the winning side is not determined, the following two strategies can be selected: 1. Discard the model reasoning result of this time, mark the model reasoning process of this time as "not accurate", and change it to manual identification. 2. The classification result obtained by the large model reasoning on the historical VOB data is the final result. This step must generate a final result.

[0267] 5.2: Save the classification result of the VOB data to be classified

[0268] The "modification" capability provided in step 1.1 is called to retrieve the corresponding historical VOB data in the historical VOB database according to the VOB data identifier, and modify the "classification result" field to the classification result determined by the election in step 5.1.3.

[0269] 5.3: Tuning scheme

[0270] According to the judgment of the business development stage described in step 4 above, there are two different tuning schemes as follows:

[0271] In the early stage of business development, the classification results saved in step 5.2 are periodically counted, the accuracy of the manual judgment result is manually judged, and the classification results of the data whose election fails in step 5.1.3 are revised to the correct value.

[0272] In the later stage of business development, as more historical VOB data has accurate classification results, the weight setting of step 5.1.1 tends to prioritize the classification results recorded in the historical VOB database. By adjusting the threshold of step 3 and increasing the threshold, more historical VOB data with accurate classification results can be recalled, thereby realizing the automatic upgrade of classification results in progress and accuracy.

[0273] With reference to the above figures Figure 6 , as an implementation of the method shown in the above figures, the present application provides an embodiment of a data classification device, which corresponds to the method embodiment shown in Figure 2 . The device can be applied to various electronic devices.

[0274] As shown in Figure 6 , a data classification device includes: a first determination unit 601 configured to determine knowledge data under a target coarse classification to which the data to be classified belongs; a second determination unit 602 configured to determine target historical data similar to the knowledge data; a third determination unit 603 configured to, in response to determining that the target historical data includes unclassified data, determine a target granularity of a classification standard to be adopted by a classification standard knowledge base according to an accuracy rate of classification results of historical data to be classified, wherein the classification results of the historical data to be classified are determined based on a classification standard knowledge base adopting a current granularity of the classification standard; a fourth determination unit 604 configured to determine and store classification results of the unclassified data according to a classification standard knowledge base adopting a target granularity of the classification standard through a large language model; and a fifth determination unit 605 configured to determine classification results of the data to be classified according to the classification results of the target historical data.

[0275] In some examples, the fifth determination unit 605 is further configured to: determine weights of the classification results of the target historical data according to a determination manner of the classification results of the target historical data, wherein the determination manner includes: determining the classification results by querying classified data in the target historical data, or determining the classification results of unclassified data according to the classification standard knowledge base; and determining the classification results of the data to be classified according to the respective classification results and weights of the target historical data.

[0276] In some examples, the fifth determination unit 605 is further configured to: in response to the classification results of the target historical data being determined based on the querying of the classified data, determine weights of the classification results of the classified data according to a strategy that the weights increase step by step as time periods elapse; and in response to the classification results of the target historical data being determined according to the classification standard knowledge base, determine weights of the classification results of the unclassified data according to a strategy that the weights decrease step by step as time periods elapse.

[0277] In some examples, the first determining unit 601 is further configured to: determine, from the coarse classification knowledge base, a plurality of approximate knowledge data similar to the data to be classified; determine an aggregation degree of the plurality of approximate knowledge data based on a proportion of approximate knowledge data belonging to the same coarse classification in the plurality of approximate knowledge data; and determine the target coarse classification and the knowledge data under the target coarse classification according to the aggregation degree of the plurality of approximate knowledge data.

[0278] In some examples, the first determining unit 601 is further configured to: in response to determining that the aggregation degree is less than a first preset threshold, determine new approximate knowledge data similar to the data to be classified from the coarse classification knowledge base until it is determined that all the approximate knowledge data obtained at present has an aggregation degree greater than or equal to the first preset threshold under a coarse classification in the coarse classification knowledge base, or all the approximate knowledge data obtained at present includes all the knowledge data under the coarse classification in the coarse classification knowledge base; and determine the coarse classification as the target coarse classification.

[0279] In some examples, the apparatus further includes a changing unit (not shown in the figure) configured to: before determining the target historical data similar to the knowledge data, determine whether the historical data includes to-be-changed data requiring a changed classification result according to a fineness of a classification result corresponding to the historical data; and in response to determining that the historical data includes the to-be-changed data, delete the classification result corresponding to the to-be-changed data.

[0280] In some examples, the changing unit (not shown in the figure) is further configured to: determine whether the historical data includes the to-be-changed data according to the fineness of the classification result corresponding to the historical data and a fineness of a classification result of the data to be classified at present.

[0281] In some examples, the second determining unit 602 is further configured to: determine a number of target historical data required to be recalled, wherein the number increases step by step as a time period elapses; and determine the number of target historical data similar to the knowledge data from the historical data.

[0282] In some examples, the fifth determining unit 605 is further configured to: in response to determining that an aggregation degree of a classification result corresponding to each of the number of target historical data under a classification is less than a second preset threshold according to a weight corresponding to each of the number of target historical data, determine new target historical data similar to the knowledge data from the historical data until it is determined that an aggregation degree of a classification result corresponding to each of all the target historical data obtained at present under the classification is greater than or equal to the second preset threshold according to the weight corresponding to each of all the target historical data, and determine the classification with the aggregation degree greater than or equal to the second preset threshold as the classification result of the data to be classified.

[0283] In some examples, the third determining unit 603 is further configured to: in response to the accuracy rate of the classification result of the historical data to be classified exceeding a preset accuracy rate threshold, increase the current fineness of the classification standard adopted by the classification standard knowledge base to determine the target fineness.

[0284] In some examples, the apparatus further includes a storage unit (not shown in the figure) configured to: store the data to be classified and the classification result of the data to be classified as historical data.

[0285] In this embodiment, the first determining unit in the data classification apparatus determines the knowledge data under the target coarse classification to which the data to be classified belongs; the second determining unit determines the target historical data similar to the knowledge data; the third determining unit determines the target fineness of the classification standard to be adopted by the classification standard knowledge base according to the accuracy rate of the classification result of the historical data to be classified in response to the determination that the target historical data includes unclassified data, wherein the classification result of the historical data to be classified is determined based on the classification standard knowledge base adopting the classification standard of the current fineness; the fourth determining unit determines and stores the classification result of the unclassified data by the large language model according to the classification standard knowledge base adopting the classification standard of the target fineness; and the fifth determining unit determines the classification result of the data to be classified according to the classification result of the target historical data, so that the classification result with high precision and high accuracy can be obtained by using the knowledge data under the coarse classification, the historical data, the classification standard knowledge base and other data resources with low precision and easy to obtain, and the fineness of the classification standard of the classification standard knowledge base is improved in steps with the passage of time, so that the precision and accuracy of the classification result of the data to be classified are continuously improved with the passage of time.

[0286] Reference will now be made to the following description Figure 7 which shows a computer system 700 suitable for implementing the devices (for example Figure 1 devices 101, 102, 103, 105) of embodiments of the present application. Figure 7 The device shown is merely an example and should not impose any limitations on the functions and scope of use of embodiments of the present application.

[0287] As shown in Figure 7 the computer system 700 includes a processor (for example, a CPU, central processing unit) 701 which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 702 or programs loaded from a storage section 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0288] The following components are connected to the I / O interface 705: an input part 706 including a keyboard, a mouse, etc.; an output part 707 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 708 including a hard disk, etc.; and a communication part 709 including a network interface card such as a LAN card, a modem, etc. The communication part 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as necessary. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 710 as necessary, so that a computer program read out therefrom is installed in the storage part 708 as necessary.

[0289] In particular, according to embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above-described functions defined in the methods of the present application are performed.

[0290] It is to be appreciated that the computer readable medium of the present application can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a computer readable storage medium can include, but are not limited to, the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. In the present application, a computer readable signal medium can include a computer readable program code, propagated by any suitable medium of transmission, including, but not limited to, wireless, wire line, optical fiber cable, R.F, etc., or any suitable combination of the foregoing. The computer readable medium of the present application can be a computer readable signal medium or a computer readable storage medium or any combination thereof.

[0291] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0292] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or can sometimes be executed in reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowcharts, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.

[0293] The units described in the embodiments of the present application can be implemented by software, or can be implemented by hardware. The described units can also be arranged in a processor, for example, can be described as: a processor comprising a first determination unit, a second determination unit, a third determination unit, a fourth determination unit, and a fifth determination unit. In some cases, the names of these units do not constitute a limitation on the units themselves, for example, the third determination unit can also be described as: a unit for determining the target fineness of the classification standard to be used by the classification standard knowledge base according to the accuracy rate of the classification result of the historical to-be-classified data, in response to determining that the target historical data includes unclassified data.

[0294] As another aspect, the present application also provides a computer readable medium, which can be included in the device described in the above embodiments, or can exist separately without being assembled into the device. The above computer readable medium carries one or more programs, which, when executed by the device, cause the computer device to: determine knowledge data of a target coarse classification to which the to-be-classified data belongs; determine target historical data similar to the knowledge data; in response to determining that the target historical data includes unclassified data, determine a target fineness of a classification standard to be used by a classification standard knowledge base according to an accuracy rate of a classification result of historical to-be-classified data, wherein the classification result of the historical to-be-classified data is determined based on the classification standard knowledge base using the current fineness of the classification standard; determine and store the classification result of the unclassified data by a large language model according to the classification standard knowledge base using the target fineness of the classification standard; and determine the classification result of the to-be-classified data according to the classification result of the target historical data.

[0295] The above description is only the preferred embodiment of the present application and the explanation of the technical principles. It should be understood by those skilled in the art that the scope of the protection of the present application is not limited to the technical solutions formed by the specific combinations of the above technical features. It should also cover other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the concept of the present application. For example, the technical solutions formed by the mutual replacement of the above features and the technical features with similar functions disclosed (but not limited to) in the present application.

Claims

1. A data classification method, comprising: The target coarse category to which the data to be classified belongs is determined based on the feature distance between the data to be classified and the coarse category in the coarse category knowledge base of the application domain to which the data to be classified belongs. The knowledge data under the target coarse classification is determined from the coarse classification knowledge base; Identify target historical data that approximates the knowledge data; In response to determining that the target historical data includes unclassified data, the target granularity of the classification standard to be adopted by the classification standard knowledge base is determined based on the accuracy of the classification results of the historical unclassified data, wherein the classification results of the historical unclassified data are determined based on the classification standard knowledge base adopting the current granularity of the classification standard. Using a large language model, the classification results of the unclassified data are determined and stored based on a classification standard knowledge base that adopts the classification standard with the target level of refinement. Based on the classification results of the target historical data, the classification result of the data to be classified is determined.

2. The method according to claim 1, wherein, The step of determining the classification result of the data to be classified based on the classification result of the target historical data includes: The weight of the classification result of the target historical data is determined according to the method of determining the classification result of the target historical data, wherein the method of determining the classification result includes: querying the classified data in the target historical data to determine the classification result, or determining the classification result of the unclassified data according to the classification standard knowledge base; The classification result of the data to be classified is determined based on the classification results and weights corresponding to the target historical data.

3. The method according to claim 2, wherein, The method for determining the classification results of the target historical data, including determining the weights of the classification results of the target historical data, includes: The classification result of the target historical data is determined based on querying the classified data, and the weight of the classification result of the classified data is determined according to the strategy of increasing the weight stepwise with the time period. In response to the classification result of the target historical data being determined based on the classification standard knowledge base, the weight of the classification result of the unclassified data is determined according to a strategy in which the weight decreases stepwise over time.

4. The method according to claim 1, wherein, The step of determining the knowledge data under the target coarse classification from the coarse classification knowledge base includes: Multiple approximate knowledge data that are similar to the data to be classified are identified from the coarse classification knowledge base; The degree of clustering of the multiple approximate knowledge data is determined based on the proportion of approximate knowledge data belonging to the same coarse category among the multiple approximate knowledge data. Based on the degree of aggregation of the multiple approximate knowledge data, the target coarse classification and the knowledge data under the target coarse classification are determined.

5. The method according to claim 4, wherein, The step of determining the target coarse classification based on the aggregation degree of the multiple approximate knowledge data includes: In response to determining that the degree of aggregation is less than a first preset threshold, new approximate knowledge data that is similar to the data to be classified is determined from the coarse classification knowledge base until it is determined that the degree of aggregation of all currently obtained approximate knowledge data under a coarse classification in the coarse classification knowledge base is greater than or equal to the first preset threshold, or that all currently obtained approximate knowledge data includes all knowledge data under a coarse classification in the coarse classification knowledge base. The coarse classification is determined as the target coarse classification.

6. The method according to claim 1, wherein, Before determining the target historical data that approximates the knowledge data, the method further includes: Based on the granularity of the classification results corresponding to the historical data, determine whether the historical data includes data that needs to be modified and whose classification results need to be changed; In response to the determination that the classification result corresponding to the data to be changed is deleted.

7. The method according to claim 6, wherein, The step of determining whether the historical data includes data whose classification results need to be changed based on the level of detail of the historical data classification results includes: Based on the granularity of the classification results corresponding to the historical data and the granularity of the classification results of the data to be classified that can be obtained now, it is determined whether the historical data includes the data to be changed.

8. The method according to claim 2, wherein, The determination of target historical data that approximates the knowledge data includes: Determine the number of target historical data to be recalled, wherein the number increases stepwise over time. The number of target historical data that are similar to the knowledge data are determined from the historical data.

9. The method according to claim 8, wherein, The step of determining the classification result of the data to be classified based on the classification results and weights corresponding to the target historical data includes: In response to determining that the degree of clustering of the classification results of the target historical data under a category is less than a second preset threshold based on the weights corresponding to the target historical data, new target historical data that is similar to the knowledge data is determined from the historical data, until the degree of clustering of the classification results of the target historical data under a category is greater than or equal to the second preset threshold based on the weights corresponding to all the target historical data obtained so far, and the category with a degree of clustering greater than or equal to the second preset threshold is determined as the classification result of the data to be classified.

10. The method according to claim 1, wherein, The determination of the target precision of the classification standard to be adopted by the classification standard knowledge base based on the accuracy of the classification results of historical data to be classified includes: In response to the fact that the accuracy of the classification results of the historical data to be classified exceeds a preset accuracy threshold, the current precision of the classification standard adopted by the classification standard knowledge base is increased, and the target precision is determined.

11. The method according to claim 1, wherein, Also includes: The data to be classified is used as historical data, and the data to be classified and the classification results of the data to be classified are stored.

12. A method for classifying customer complaint data, comprising: The first determining unit is configured to determine the target coarse classification to which the data to be classified belongs based on the feature distance between the data to be classified and the coarse classification in the coarse classification knowledge base of the application domain to which the data to be classified belongs. The knowledge data under the target coarse classification is determined from the coarse classification knowledge base; The second determining unit is configured to determine target historical data that is similar to the knowledge data; The third determining unit is configured to, in response to determining that the target historical data includes unclassified data, determine the target granularity of the classification standard to be adopted by the classification standard knowledge base based on the accuracy of the classification results of the historical unclassified data, wherein the classification results of the historical unclassified data are determined based on the classification standard knowledge base adopting the current granularity of the classification standard. The fourth determining unit is configured to determine and store the classification result of the unclassified data by using a large language model and a classification standard knowledge base that adopts the classification standard of the target fineness. The fifth determining unit is configured to determine the classification result of the data to be classified based on the classification result of the target historical data.

13. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-11.

14. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-11.

15. A computer program product comprising: A computer program that, when executed by a processor, implements the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Data classification method and device, model training method and device and electronic equipment

    CN111783861A

  • Article category identification method and device, electronic equipment and storage medium

    CN115017384A