Data classification method and device, medium and product
By training and deploying the inference model in a network-free or low-computing-power environment, the problem of traditional large models being unable to work is solved, and the function of data classification is realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-10
AI Technical Summary
In environments without network access or with low computing power, traditional large-scale pre-trained models cannot function properly, making it impossible to meet data classification requirements.
The inference model is trained using the classified data samples output by the data classification model, and the interface of the inference model is deployed locally. Data classification is then performed in a network-free environment using the Hypertext Transfer Protocol.
It enables data classification in environments without network or with low computing power, and simulates the classification function of large models. It is suitable for intranet environments without network terminals and where large models cannot be deployed.
Smart Images

Figure CN121637178A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the fields of finance and artificial intelligence technology, and in particular to a data classification method, device, medium and product. Background Technology
[0002] In today's technological landscape, data classification technology has made significant progress. With the continuous development of artificial intelligence and machine learning, classifiers can now be analyzed and processed using large pre-trained models (large models). These large models typically possess powerful data processing capabilities and high-precision classification performance, enabling them to accurately classify large amounts of data in a short time. The application of this technology greatly reduces the workload of manual judgment, lowers labor costs, and improves work efficiency.
[0003] However, despite the excellent performance of large-scale models in data classification, their deployment and application face several practical challenges. First, running large-scale models requires powerful computing capabilities, necessitating high-performance hardware and substantial computing resources. Second, the training and inference processes of large-scale models typically require a stable network environment to access and process large volumes of data. In certain scenarios, environments without network access or computing power may arise. In such cases, traditional classification methods relying on large-scale models cannot function properly, yet the need for data classification remains. Therefore, achieving efficient and accurate data classification in these constrained environments has become a pressing issue. Summary of the Invention
[0004] This invention provides a data classification method, device, medium, and product to enable the establishment of an inference model based on classified data samples output by a data classification model. Based on a large model with interface-based calls, the classifier can be deployed in a network-free environment with limited computing power to meet data classification requirements.
[0005] According to one aspect of the present invention, a data classification method is provided, comprising:
[0006] Obtain the data to be classified;
[0007] The interface of the inference model is invoked to classify the data to be classified based on the inference model; both the data to be classified and the inference model are deployed locally; the inference model is trained from the classification data samples output by the data classification model.
[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0009] At least one processor; and
[0010] A memory communicatively connected to the at least one processor; wherein,
[0011] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data classification method according to any embodiment of the present invention.
[0012] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data classification method according to any embodiment of the present invention.
[0013] According to another aspect of the present invention, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the data classification method described in any embodiment of the present invention.
[0014] This invention trains an inference model using categorized data samples output by a data classification model, acquires locally deployed data to be classified, and calls the interface of the locally deployed inference model to classify the data based on the inference model. Through this invention's technical solution, an inference model is established based on categorized data samples output by the data classification model. Using a large model with interface-based calls as a foundation, a classifier can be deployed in a network-free environment with limited computing power to meet data classification needs. In the financial field, this invention is applicable to data classification functions in scenarios such as intranet environments without network terminals or where large models cannot be deployed.
[0015] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart of a data classification method according to an embodiment of the present invention;
[0018] Figure 2 This is a flowchart of an exemplary data classification method in an embodiment of the present invention;
[0019] Figure 3This is a schematic diagram of the structure of a data classification device according to an embodiment of the present invention;
[0020] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the data classification method of this invention. Detailed Implementation
[0021] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and their derivatives, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0023] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0024] Example 1
[0025] Figure 1 This is a flowchart of a data classification method according to an embodiment of the present invention. This embodiment is applicable to data classification situations. The method can be executed by the data classification device of the present invention, which can be implemented in software and / or hardware, such as... Figure 1 As shown, the method specifically includes the following steps:
[0026] S101. Obtain the data to be classified.
[0027] In this embodiment, the data to be classified can be financial text data to be classified. Preferably, the data to be classified in this embodiment can be locally deployed data.
[0028] For example, in the financial sector, the data classification method of this embodiment can be used for data classification in scenarios such as intranet environments without network terminals or where large models cannot be deployed.
[0029] Specifically, users can input data to be classified, and the data classification method in this embodiment can identify the specific type of the data to be classified, such as suggestion, question, or answer.
[0030] S102. Call the interface of the inference model to classify the data to be classified based on the inference model.
[0031] Both the data to be classified and the inference model are deployed locally, and the inference model is trained from the classified data samples output by the data classification model.
[0032] In this embodiment, the inference model can be a locally deployed model used for category judgment inference on the data to be classified. Preferably, the inference model can be, for example, a trained deep learning model. In specific implementation, the inference model can be trained using categorized data samples output by the data classification model. In this embodiment, the data classification model can be an existing, known model capable of data classification, and this embodiment does not limit the model type or specific parameter details of the data classification model. Preferably, the data classification model can be, for example, a large language model.
[0033] Specifically, an inference model is trained using the categorized data samples output by the data classification model, and then deployed locally. In low-computing-power, network-free environments, when data to be classified is detected, the interface of the locally deployed inference model can be called to classify the data based on the inference model and return the data classification result.
[0034] In practice, a simple network service can be deployed to expose the inference capabilities of the deep learning model. The inference model deployment can consist of the model output from the deep learning trainer and a simple script program. It can be deployed independently in various system environments, and even in low-configuration, offline environments, a simple interface can be deployed to enable local calls to the simplified classifier functionality.
[0035] This invention trains an inference model using categorized data samples output by a data classification model, acquires locally deployed data to be classified, and calls the interface of the locally deployed inference model to classify the data based on the inference model. Through this invention's technical solution, an inference model is established based on categorized data samples output by the data classification model. Using a large model with interface-based calls as a foundation, a classifier can be deployed in a network-free environment with limited computing power to meet data classification needs. In the financial field, this invention is applicable to data classification functions in scenarios such as intranet environments without network terminals or where large models cannot be deployed.
[0036] Optionally, before calling the inference model's interface, the following may also be included:
[0037] Deploy the application interface based on the Hypertext Transfer Protocol corresponding to the inference model locally.
[0038] In practice, a simple network service can be deployed locally, such as an application interface based on the Hypertext Transfer Protocol, to expose the inference function of the deep learning model.
[0039] The interface for calling the inference model includes:
[0040] If the current environment is detected to be without network access, the application interface based on the Hypertext Transfer Protocol corresponding to the inference model deployed locally will be invoked.
[0041] It should be noted that the current environment refers to the network / computing power environment at the moment the data to be classified is acquired. A "no network environment" can be understood as an environment without network or computing power.
[0042] Specifically, by deploying the application interface based on the Hypertext Transfer Protocol (HTTP) corresponding to the inference model locally, large-scale model classification can be achieved even in low-computing-power, network-free environments. When data to be classified is detected, the application interface based on the HTTP corresponding to the locally deployed inference model is called to classify the data based on the inference model, thus achieving data classification.
[0043] The technical solution of this embodiment can simulate the function of large language model classification in a network-free and low-computing-power environment by establishing an inference model and deploying the corresponding application interface based on the hypertext transfer protocol on the local machine. In the form of interface-based calls, the corresponding data classification requirements can be realized in a network-free environment with limited computing power.
[0044] Optionally, before calling the inference model's interface, the following may also be included:
[0045] Build a deep learning model.
[0046] The deep learning model can be an initial, untrained model.
[0047] Obtain the target data sample set.
[0048] In this embodiment, the target data sample set may be a collection of financial text data samples used to train a deep learning model.
[0049] The target data sample set includes training data samples and the corresponding data classification result samples.
[0050] It should be noted that training data samples can be input samples used when training a deep learning model, specifically, they can be several collected financial text data samples. Data classification result samples can be samples used as a control group during the training of the deep learning model, specifically, they can be the result samples after classifying the collected financial text data samples. For example, training data samples could be financial data 1, and its corresponding data classification result sample could be category 1; training data samples could be financial data 2, and its corresponding data classification result sample could be category 2; training data samples could be financial data 3, and its corresponding data classification result sample could be category 1; training data samples could be financial data 4, and its corresponding data classification result sample could be category 2.
[0051] Specifically, a number of financial text data were collected and classified as training samples.
[0052] The inference model is obtained by iteratively training a deep learning model using the target data sample set.
[0053] Specifically, the target data sample set is obtained, the deep learning model is iteratively trained, and the inference model is obtained.
[0054] The technical solution of this embodiment establishes a deep learning model, obtains a target data sample set for the deep learning model, and trains the deep learning model based on the target data sample set to obtain an inference model. The inference model can be deployed as a classifier in a network-free environment with limited computing power to meet data classification requirements, thus realizing the function of simulating large model classification in a network-free and low-computing-power environment.
[0055] Optionally, a deep learning model is iteratively trained using the target data sample set to obtain an inference model, including:
[0056] The training data samples in the target data sample set are input into the deep learning model to obtain the predicted classification results.
[0057] It should be explained that the predicted classification result can be the classification result corresponding to the training data sample output by the deep learning model after making a prediction based on the input training data sample.
[0058] Specifically, the training data samples from the target data sample set are input into the deep learning model to predict the classification results.
[0059] The parameters of the deep learning model are trained based on the objective function formed by the predicted classification results and the corresponding data classification results samples of the training data samples.
[0060] Specifically, the weights and other parameters of the deep learning model are iteratively trained based on the objective function formed by the predicted classification results and the corresponding data classification results samples of the training data samples.
[0061] Return to the previous step and input the training data samples from the target data sample set into the deep learning model to obtain the predicted classification results, until the inference model is obtained.
[0062] Specifically, the process returns to input the training data samples from the target data sample set into the deep learning model to obtain the predicted classification results, until the maximum number of iterations is reached (the maximum number of iterations can be preset by the user according to the actual situation, and this embodiment does not limit this), or the convergence condition is reached (the convergence condition can be set by the user according to the actual situation, and this embodiment does not limit this), and the inference model is obtained.
[0063] The technical solution of this embodiment uses a target data sample set to iteratively train a deep learning model, optimize the parameters of the deep learning model, and obtain an inference model, which makes the inference model more accurate when classifying data.
[0064] Optionally, the training data samples include categorized text data and uncategorized text data.
[0065] It should be noted that categorized text data can be financial text data whose categories have been determined in advance, while uncategorized text data can be financial text data whose categories have not yet been determined.
[0066] Obtain the target data sample set, including:
[0067] Obtain the initial text dataset.
[0068] In this embodiment, the initial text dataset can be a dataset composed of several financial text data collected in advance. The data in the initial text dataset can be unclassified data.
[0069] Specifically, a sufficient amount of financial text data can be collected in advance as training samples for training deep learning models to form an initial text dataset.
[0070] Input the first dataset from the initial text dataset into the data classification model to obtain the classified text data corresponding to the first dataset.
[0071] It should be noted that the first dataset can be a dataset composed of several financial text data arbitrarily extracted from the initial text dataset.
[0072] Specifically, a portion of financial text data is selected from the initial text dataset as the first dataset. This data is then input into a data classification model to classify the financial text data in the first dataset, resulting in a classification result for each piece of financial text data in the first dataset. In this way, the financial text data in the first dataset can be considered as classified text data.
[0073] In actual operation, this embodiment configures the storage path of local text (i.e., the initial text dataset) and the interface path of the large model (i.e., the data classification model), automatically accesses the large model (i.e., the data classification model) to classify the text data (i.e., the first dataset in the initial text dataset), and then organizes and stores it in the database.
[0074] All text data in the initial text dataset other than the first dataset are treated as unclassified text data.
[0075] Specifically, all financial text data in the initial text dataset other than the first dataset are directly treated as unclassified text data. This embodiment does not limit the specific ratio of classified to unclassified text data; users can set it according to their actual needs or experience.
[0076] Classified and unclassified text data were used as training data samples.
[0077] Specifically, training data samples, consisting of categorized and uncategorized text data, are used as input to iteratively train the deep learning model.
[0078] The classification results of categorized and uncategorized text data are labeled to obtain the data classification result samples corresponding to the training data samples.
[0079] Specifically, before training a deep learning model, the classification results of categorized and uncategorized text data need to be labeled. That is, the classification results of categorized text data are labeled, and uncategorized text data can be directly labeled as uncategorized results.
[0080] The target data sample set consists of training data samples and the corresponding data classification result samples.
[0081] Specifically, the target data sample set is composed of training data samples and the corresponding data classification result samples, and the deep learning model is iteratively trained.
[0082] Specifically, uncategorized text calls the interface of an existing large model (i.e., a data classification model) to obtain the returned classification, generating categorized text; the categorized text is then formatted uniformly and stored in the database.
[0083] The technical solution of this embodiment classifies unclassified text data into classified text data by using a data classification model, thereby generating a target data sample set for iterative training of a deep learning model. This makes the training samples of the deep learning model more closely resemble the actual situation, thus making the inference model more reasonable and accurate, and making the classification judgment of the data to be classified more realistic and reliable.
[0084] Optionally, at least two deep learning models may exist.
[0085] In practice, multiple deep learning models may have different effects due to different parameters. This invention constructs model parameters in batches and trains them in parallel by means of random / incremental / traversal methods, and selects the best parameter combination as the model output.
[0086] By iteratively training a deep learning model using a target data sample set, an inference model is obtained, including:
[0087] Each deep learning model is trained iteratively using the target data sample set to obtain the target model corresponding to each deep learning model.
[0088] It should be noted that the target model can be a trained model obtained after training a deep learning model, and each deep learning model corresponds to one target model.
[0089] Specifically, multiple deep learning models can be iteratively trained based on the target data sample set, resulting in multiple target models after training.
[0090] Obtain the test dataset from the target data sample set.
[0091] The test dataset can be a dataset used to test the target model and verify its accuracy. In practice, this embodiment does not limit the ratio of the test dataset to the training data samples used to train the model; users can set it according to their actual needs or experience.
[0092] Specifically, the target data sample set can be divided into a training set and a validation set (i.e., a test dataset). The validation set does not participate in the training process but participates in model evaluation to assess the accuracy of the model.
[0093] Each target model is tested using the test dataset to obtain the classification accuracy for each target model.
[0094] It should be noted that classification accuracy can be the accuracy of the target model in classifying the data in the input test dataset. This value can be calculated based on the total number of test datasets and the number of data points that are correctly classified.
[0095] Specifically, the data in the test dataset is input into each target model for testing, and the classification accuracy of each target model is obtained.
[0096] The target model with the highest classification accuracy is used as the inference model.
[0097] Specifically, the parameter combination of the target model with the highest classification accuracy can be selected as the parameters of the final inference model to perform classification operations on the data to be classified.
[0098] The technical solution of this embodiment trains the parameters of multiple deep learning models and then selects the model with the highest classification accuracy as the inference model, which ensures the accuracy of the inference model when classifying data and improves the efficiency of data classification.
[0099] Optional, also includes:
[0100] Deploy the network service interface corresponding to the data classification model.
[0101] In practice, while deploying the application interface based on the Hypertext Transfer Protocol (HTTP) for the inference model, the network service interface for the data classification model can also be deployed. This allows users to freely choose the data classification model for data classification when the network recovers. Of course, users can also continue to use the inference model, choosing freely according to their own needs.
[0102] If a network environment is detected in the current environment, the network service interface corresponding to the data classification model will be invoked.
[0103] Specifically, when network recovery is detected, users can choose to call the network service interface corresponding to the data classification model to classify the data.
[0104] The technical solution in this embodiment, by deploying interfaces for two models, can comprehensively cover the data classification scenario requirements in both network and offline environments, thereby improving the user experience.
[0105] Optionally, the data to be classified is classified based on an inference model, including:
[0106] The data to be classified is input into the inference model for classification, and the data classification result is obtained.
[0107] The data classification result can be the result obtained by the inference model after classifying the input data to be classified.
[0108] The data classification results include at least one of the following: suggestions, questions, and answers.
[0109] Specifically, the data to be classified is input into the inference model for classification, and the classification result corresponding to the data to be classified is obtained. For example, the data to be classified may be question-type data.
[0110] The technical solution of this embodiment classifies the data to be classified through an inference model in a low-computing-power, network-free environment, thereby determining the specific category of the data to be classified. This realizes the data classification function in scenarios such as intranet environments in the financial field where there are no network terminals and large models cannot be deployed.
[0111] The technical solution of this invention generates a classifier interface that is easy to deploy in a lightweight manner by generating training data. This enables simple data classification in environments without network or computing power, and can be used as appropriate in application scenarios that require data classification.
[0112] Example 2
[0113] Figure 2 This is a flowchart of an exemplary data classification method in an embodiment of the present invention. Based on the above embodiments, this embodiment provides an exemplary data classification method flowchart, such as... Figure 2 As shown, the specific process of this method can be described as follows:
[0114] Obtain an initial text dataset; input the first dataset from the initial text dataset into a data classification model to obtain categorized text data corresponding to the first dataset; treat the other text data in the initial text dataset besides the first dataset as uncategorized text data; use the categorized and uncategorized text data as data samples, and label the categorized and uncategorized text data according to the classification results to obtain data classification result samples corresponding to the training data samples; the training data samples and the data classification result samples corresponding to the training data samples constitute the target data sample set; establish a deep learning model, and iteratively train the deep learning model using the target data sample set to obtain an inference model; obtain the data to be classified, and classify the data to be classified based on the inference model to obtain the data classification result.
[0115] The main function of the technical solution in this embodiment is to deploy a simple classifier in an environment where the hardware configuration cannot support the deployment of large models, so as to realize the function of targeted large models. In the financial field, it can be applied to scenarios such as intranet environments without network terminals and where large models cannot be deployed, so as to realize the function of data classification.
[0116] Example 3
[0117] Figure 3 This is a schematic diagram of a data classification device according to an embodiment of the present invention. This embodiment is applicable to data classification scenarios. The device can be implemented using software and / or hardware, and can be integrated into any device that provides data classification functionality, such as… Figure 3 As shown, the data classification device specifically includes: an acquisition module 201 and a classification module 202.
[0118] Among them, the acquisition module 201 is used to acquire the data to be classified;
[0119] The classification module 202 is used to call the interface of the inference model to classify the data to be classified based on the inference model; both the data to be classified and the inference model are deployed locally; the inference model is trained from the classification data samples output by the data classification model.
[0120] Optionally, the device further includes:
[0121] The first deployment unit is used to deploy the application interface based on the Hypertext Transfer Protocol corresponding to the inference model locally.
[0122] The classification module 202 is specifically used for:
[0123] If the current environment is detected to be without network access, the application interface based on the Hypertext Transfer Protocol corresponding to the inference model deployed locally will be invoked.
[0124] Optionally, the device further includes:
[0125] Establishment units are used to build deep learning models;
[0126] An acquisition unit is used to acquire a target data sample set; the target data sample set includes training data samples and data classification result samples corresponding to the training data samples.
[0127] The training unit is used to iteratively train the deep learning model using the target data sample set to obtain the inference model.
[0128] Optionally, the training unit is specifically used for:
[0129] The training data samples in the target data sample set are input into the deep learning model to obtain the predicted classification result;
[0130] The parameters of the deep learning model are trained based on the objective function formed by the predicted classification results and the corresponding data classification result samples of the training data samples;
[0131] Return to the operation of inputting training data samples from the target data sample set into the deep learning model to obtain the predicted classification result, until the inference model is obtained.
[0132] Optionally, the training data samples include categorized text data and uncategorized text data;
[0133] The acquisition unit is specifically used for:
[0134] Obtain the initial text dataset;
[0135] Input the first dataset from the initial text dataset into the data classification model to obtain the classified text data corresponding to the first dataset;
[0136] All text data in the initial text dataset other than the first dataset are treated as unclassified text data.
[0137] The categorized text data and the uncategorized text data are used as the training data samples;
[0138] The classified text data and the unclassified text data are labeled with classification results to obtain the data classification result samples corresponding to the training data samples;
[0139] The target data sample set is composed of the training data samples and the data classification result samples corresponding to the training data samples.
[0140] Optionally, there may be at least two deep learning models;
[0141] The training unit is specifically used for:
[0142] Each deep learning model is iteratively trained using the target data sample set to obtain the target model corresponding to each deep learning model.
[0143] Obtain a data classification dataset from the target data sample set;
[0144] Based on the data classification dataset, each target model is classified to obtain the classification accuracy corresponding to each target model;
[0145] The target model with the highest classification accuracy is used as the inference model.
[0146] Optionally, the device further includes:
[0147] The second deployment unit is used to deploy the network service interface corresponding to the data classification model;
[0148] The calling unit is used to call the network service interface corresponding to the data classification model if the current environment is detected to be a network environment.
[0149] The above-mentioned products can execute the data classification method provided in any embodiment of the present invention, and have the corresponding functional modules and beneficial effects of executing the method.
[0150] Example 4
[0151] Figure 4A schematic diagram of an electronic device 30 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0152] like Figure 4 As shown, the electronic device 30 includes at least one processor 31 and a memory, such as a read-only memory 32 or a random access memory 33, communicatively connected to the at least one processor 31. The memory stores computer programs executable by the at least one processor. The processor 31 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 32 or loaded from storage unit 38 into the random access memory 33. The random access memory 33 can also store various programs and data required for the operation of the electronic device 30. The processor 31, read-only memory 32, and random access memory 33 are interconnected via a bus 34. An input / output interface 35 is also connected to the bus 34.
[0153] Multiple components in electronic device 30 are connected to input / output interface 35, including: input unit 36, such as keyboard, mouse, etc.; output unit 37, such as various types of monitors, speakers, etc.; storage unit 38, such as disk, optical disk, etc.; and communication unit 39, such as network card, modem, wireless transceiver, etc. Communication unit 39 allows electronic device 30 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0154] Processor 31 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 31 include, but are not limited to, central processing units, graphics processing units, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 31 performs the various methods and processes described above, such as data classification methods:
[0155] Obtain the data to be classified;
[0156] The interface of the inference model is invoked to classify the data to be classified based on the inference model; both the data to be classified and the inference model are deployed locally; the inference model is trained from the classification data samples output by the data classification model.
[0157] In some embodiments, the data classification method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 38. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 30 via read-only memory 32 and / or communication unit 39. When the computer program is loaded into random access memory 33 and executed by processor 31, one or more steps of the data classification method described above may be performed. Alternatively, in other embodiments, processor 31 may be configured to perform the data classification method by any other suitable means (e.g., by means of firmware).
[0158] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays, application-specific integrated circuits (ASICs), application-specific standard products (ASICs), systems-on-a-chip (SoCs), payload programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0159] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0160] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (flash memory), optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0161] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0162] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0163] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to address the shortcomings of traditional physical hosts and virtual private servers, such as high management difficulty and weak business scalability.
[0164] In one embodiment, the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the data classification method of any embodiment of the present invention.
[0165] In the implementation of a computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0166] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0167] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method of data classification, characterized by, The method comprises: obtaining to-be-classified data; calling an interface of an inference model to classify the to-be-classified data based on the inference model; the to-be-classified data and the inference model are both locally deployed; the inference model is obtained by training a classification data sample output by a data classification model.
2. The method of claim 1, wherein, Before calling the interface of the inference model, the method further comprises: locally deploying a hypertext transfer protocol (HTTP) application interface corresponding to the inference model; calling the interface of the inference model comprises: if it is detected that the current environment is a network-free environment, calling the HTTP application interface corresponding to the locally deployed inference model.
3. The method of claim 1, wherein, Before calling the interface of the inference model, the method further comprises: establishing a deep learning model; obtaining a target data sample set; the target data sample set comprises training data samples and data classification result samples corresponding to the training data samples; iteratively training the deep learning model by using the target data sample set to obtain an inference model.
4. The method of claim 3, wherein, The iteratively training the deep learning model by using the target data sample set to obtain an inference model comprises: inputting the training data samples in the target data sample set into the deep learning model to obtain predicted classification results; training parameters of the deep learning model according to a target function formed by the predicted classification results and the data classification result samples corresponding to the training data samples; returning to perform the operation of inputting the training data samples in the target data sample set into the deep learning model to obtain predicted classification results until the inference model is obtained.
5. The method of claim 3, wherein, The training data samples comprise classified text data and unclassified text data; The obtaining the target data sample set comprises: obtaining an initial text data set; inputting a first data set in the initial text data set into a data classification model to obtain classified text data corresponding to the first data set; regarding other text data in the initial text data set except the first data set as unclassified text data; regarding the classified text data and the unclassified text data as the training data samples; performing classification result marking on the classified text data and the unclassified text data to obtain data classification result samples corresponding to the training data samples; composing the target data sample set by using the training data samples and the data classification result samples corresponding to the training data samples.
6. The method of claim 3, wherein, There are at least two deep learning models; The iteratively training the deep learning model by using the target data sample set to obtain an inference model comprises: iteratively training each of the deep learning models by using the target data sample set to obtain a target model corresponding to each of the deep learning models; obtaining a test data set from the target data sample set; testing each of the target models according to the test data set to obtain a classification accuracy rate corresponding to each of the target models; regarding a target model with the highest classification accuracy rate as the inference model.
7. The method of claim 1, wherein, The method further comprises: deploying a network service interface corresponding to the data classification model; if it is detected that the current environment is a network environment, calling the network service interface corresponding to the data classification model.
8. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected with the at least one processor in communication; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the data classification method in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing a processor to implement the data classification method in any one of claims 1-7 when executed.
10. A computer program product comprising a computer program which, when executed by a processor, implements the data classification method according to any one of claims 1-7.