Data processing method and device, equipment, computer storage medium and program product

By calling the target large language model to perform keyword description processing on the rule keywords, a data classification rule with higher accuracy is generated, which solves the problem of insufficient accuracy of data classification rules and improves the accuracy and efficiency of data classification business.

CN120611046APending Publication Date: 2025-09-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410237158.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-01
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In the existing technology, the accuracy of data classification rules is affected by the accuracy of data classification rules. How to generate data classification rules with higher accuracy has become an important research topic in the field of information security.

Method used

By calling the target large language model to perform keyword description processing on the rule keywords in the rule reasoning text under the semantic dimension and/or data specification dimension, accurate and detailed data classification rules are generated.

Benefits of technology

It improves the classification accuracy of data classification business, reduces manual workload, and improves the efficiency and accuracy of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611046A_ABST
    Figure CN120611046A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, equipment, a computer storage medium and a program product, and is suitable for the fields of artificial intelligence, Internet of Vehicles and the like. According to the method, after a rule reasoning request is received, a rule reasoning text is obtained from the rule reasoning request, the rule reasoning text comprises rule keywords needed to be used when a data classification rule is generated, and then a classification target of the data classification rule is reasoned based on the rule reasoning text; selecting a target description dimension matched with the classification target from preset description dimensions, wherein the target description dimension comprises at least one of a semantic dimension and a data specification dimension; and finally, calling the target large language model to perform keyword description processing on the rule keyword under the target description dimension to obtain keyword description information of the rule keyword, and generating an accurate data classification rule based on the keyword description information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computer security technology and artificial intelligence technology, and in particular to a data processing method, apparatus, device, computer storage medium, and program product. Background Art

[0002] To maintain the information security of various internet businesses and their business objects, it is often necessary to classify relevant business data according to certain data classification rules. Security management of internet businesses is then conducted based on the classification results. This makes the effectiveness of business security management closely related to the accuracy of the classification results. Because the accuracy of classification results is affected by the accuracy of the data classification rules, generating highly accurate data classification rules has become an important research topic in the field of information security. Summary of the Invention

[0003] The embodiments of the present application provide a data processing method, apparatus, device, computer storage medium, and program product, which can assist in generating data classification rules with higher accuracy by generating descriptive information with higher reference value for rule keywords.

[0004] In one aspect, an embodiment of the present application provides a data processing method, comprising:

[0005] receiving a rule inference request and obtaining a rule inference text from the rule inference request, wherein the rule inference text includes rule keywords to be used when generating data classification rules;

[0006] Inferring a classification target of the data classification rule based on the rule inference text, and selecting a target description dimension that matches the classification target from preset description dimensions, wherein the target description dimension includes at least one of a semantic dimension and a data specification dimension;

[0007] The target large language model is called to perform keyword description processing on the rule keywords under the target description dimension to obtain keyword description information of the rule keywords, and the keyword description information is used to generate the data classification rule.

[0008] In another aspect, an embodiment of the present application provides a data processing device, comprising:

[0009] A receiving module, configured to receive a rule inference request and obtain a rule inference text from the rule inference request, wherein the rule inference text includes rule keywords to be used when generating data classification rules;

[0010] an inference module, configured to infer a classification target of the data classification rule based on the rule inference text, and select a target description dimension that matches the classification target from preset description dimensions, wherein the description dimension includes at least one of a semantic dimension and a data specification dimension;

[0011] The model calling module is used to call the target large language model to perform keyword description processing on the rule keywords under the target description dimension to obtain keyword description information of the rule keywords, and the keyword description information is used to generate the data classification rules.

[0012] In another aspect, an embodiment of the present application further provides a data processing device, comprising:

[0013] a processor adapted to execute one or more computer programs;

[0014] A computer storage medium storing one or more computer programs, wherein the one or more computer programs are suitable for being loaded by the processor and executing the data processing method according to the first aspect.

[0015] On the other hand, an embodiment of the present application further proposes a computer storage medium, which stores one or more computer programs, and the one or more computer programs are suitable for being loaded by a processor and executing the data processing method of the first aspect.

[0016] On the other hand, an embodiment of the present application further proposes a program product, which includes a computer program, and the computer program is suitable for being loaded by a processor and executing the data processing method as in the first aspect.

[0017] The embodiment of the present application calls the target large language model to perform keyword description processing on the rule keywords in the rule reasoning text under the semantic dimension and / or data specification dimension, and can obtain keyword description information for indicating the specific meaning of the rule keywords or the corresponding data usage specifications. Therefore, by referring to the keyword description information, the rule keywords can be accurately and comprehensively understood. In this case, the data classification rules generated based on this understanding can also have a higher accuracy, which is conducive to improving the classification accuracy of the data classification business to which the data classification rules are applied. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 This is a schematic diagram of an implementation process of rule reasoning provided by an embodiment of the present application;

[0020] Figure 2 This is a structural diagram of a data processing system provided in an embodiment of the present application;

[0021] Figure 3 is a schematic flow chart of a data processing method provided in an embodiment of the present application;

[0022] Figure 4 is a schematic flow chart of another data processing method provided in an embodiment of the present application;

[0023] Figure 5 This is a structural diagram of a training device provided in an embodiment of the present application;

[0024] Figure 6 This is a schematic diagram of a model training process provided in an embodiment of the present application;

[0025] Figure 7 This is a schematic diagram of a training data collection process provided by an embodiment of the present application;

[0026] Figure 8 This is a schematic diagram of the structure and structural functions of an application device provided in an embodiment of the present application;

[0027] Figure 9 is a structural diagram of a data processing device provided in an embodiment of the present application;

[0028] Figure 10 It is a structural diagram of a data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] It should be noted in advance that, in order to enable those skilled in the art to better understand the technical solutions proposed in the embodiments of this application, the embodiments of this application will be combined with one or more figures to clearly and completely describe the implementation of the technical solutions proposed in the embodiments of this application. In addition, the various figures shown in the embodiments of this application are only exemplary. For example, the execution order of the various steps in the figures can be adaptively adjusted according to the actual application scenario.

[0030] In addition, in the embodiments of the present application, the block diagrams, modules and units shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities, and each module or unit can be part of an overall module or unit that includes the function of the module or unit. That is, the term "module" or "unit" mentioned in the embodiments of the present application refers to a computer program or a part of a computer program with a predetermined function, which can work together with other related parts to achieve a predetermined goal, and can also be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof, or implemented in different networks and / or processor devices and / or microcontroller devices. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units.

[0031] The embodiment of the present application mainly proposes a data processing solution suitable for rule generation scenarios. The solution performs keyword description processing on the rule keywords to be used when generating data classification rules to obtain keyword description information with strong reference value, thereby assisting in generating data classification rules with higher accuracy.

[0032] Specifically, the principle of this solution is: after receiving a rule inference request, based on the rule inference text carried in the rule inference request, the classification target corresponding to the data classification rule to be generated is inferred, and the target description dimension that matches the classification target is selected from the preset description dimension, and the target description dimension includes at least one of the semantic dimension and the data specification dimension, and then the target large language model constructed based on the artificial neural network is called to perform keyword description processing on the rule keywords contained in the rule inference text under the target description dimension to obtain keyword description information, so that the keyword description information can be used to indicate the specific meaning of the rule keyword or the corresponding data usage specification. Then, by referring to the keyword description information in the process of generating data classification rules to achieve the classification target, a data classification rule with higher accuracy can be generated, thereby effectively improving the classification accuracy of the data classification business.

[0033] Artificial neural networks can be used to mirror the behavior of the human brain, allowing computer programs to learn to recognize patterns based on training data, thereby solving common problems in the current fields of artificial intelligence, machine learning, and deep learning. In other words, artificial neural networks are a technology within the field of machine learning (ML) / deep learning. In practical applications, machine learning (or deep learning) can also include belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning techniques.

[0034] In general, machine learning (or deep learning) is a multidisciplinary interdisciplinary field used to study how to use computers to simulate or implement human learning behavior. It can specifically involve probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines, so that computers can continuously acquire new knowledge or skills based on machine learning technology, and can reorganize existing knowledge structures so that computers can continuously improve their own performance, thereby achieving better intelligent processing effects.

[0035] By using deep learning technology to construct a neural network model (such as a target large language model) for data processing, the manual workload can be reduced, thereby improving the efficiency and accuracy of data processing. Among them, the target large language model can be deployed by multiple computer devices, and one or more data processing layers of the target large language model can be deployed in each computer device to reduce the data processing volume of each computer device, thereby improving the model processing efficiency. In addition, the target large language model can be obtained by model training (or model optimization) of the constructed large language model, and the training data used in the model training process can be collected from the business data generated by one or more data processing businesses within the scope of permission.

[0036] That is to say, in an exemplary implementation, the present application may include a large language model training phase and a target large language model application phase when implementing rule reasoning. Among them, rule reasoning can be understood as the reasoning of classification targets and the generation of keyword description information. The main functions and collaboration methods of each stage involved in implementing rule reasoning can be exemplarily referred to Figure 1 The implementation flow of rule reasoning is shown.

[0037] like Figure 1 As shown, the goal of the large language model training phase is to train the large language model to obtain the target large language model. This phase specifically includes three main processes: data processing, model training, and model performance evaluation. The goal of the target large language model application phase is to respond to rule inference requests and output corresponding keyword description information. This phase specifically includes three main processes: instruction parsing, model inference, and result processing.

[0038] During the training phase of the large language model, the data processing process mainly involves cleaning the collected raw data (i.e., business data) to obtain training data that meets the model training requirements. The model training process mainly involves optimizing the model parameters of the pre-built large language model to obtain an optimized large language model that meets the optimization goals. The model performance evaluation process mainly involves performance evaluation of the optimized large language model to determine whether the optimized large language model can be used as the target large language model to complete the actual rule reasoning business, and when the current optimized large language model cannot be used as the target large language model, the model training process is triggered to retrain the optimized large language model (or pre-built large language model) until the target large language model is obtained.

[0039] During the application phase of the target large language model, the instruction parsing process mainly parses user instructions (or rule inference requests) to obtain the various data required for reference in the process of generating keyword description information, such as rule keywords and rule inference text. The model inference process mainly obtains the classification target based on the rule inference text, and calls the target large language model to perform keyword description processing on the rule keywords under the target description dimension that matches the classification target to obtain keyword description information. The result processing process mainly feeds back keyword description information to the user in a user-friendly manner, where the user-friendly manner refers to a method that facilitates the user to use the keyword description information in the process of generating data classification rules, which can usually be achieved by formatting the keyword description information.

[0040] In addition, in a specific implementation, the data processing solution proposed in this application can be implemented independently by a single computer device, or can be implemented collaboratively by multiple computer devices. As an example, in the case of collaborative implementation by multiple computer devices, each computer device can be constructed as follows: Figure 2 The data processing system shown is used.

[0041] like Figure 2As shown, the data processing system applicable to the present application may include a data providing module marked by 201, a rule reasoning module marked by 202, and a business service module marked by 203. Among them, the data providing module may be composed of n (n is a positive integer) data providers, and each data provider is associated with at least one data processing business. The rule reasoning module may be a distributed system composed of multiple computer devices, and there is a direct or indirect communication connection between each computer device to collaboratively realize the deployment of the target large language model. The business service module may be composed of m clients, and each client can at least be used to provide users with data processing services implemented based on this application, so that users can initiate rule reasoning requests to the rule reasoning module through human-computer interaction with the client, and then obtain keyword description information fed back by the rule reasoning module in response to the request.

[0042] It should be noted that in actual applications, Figure 2 The modules in the present invention are divided into more refined modules or combined into larger modules for use. Moreover, the modules are not limited to realizing the functions mentioned above, but can also realize other functions, which are not limited here.

[0043] It is also worth mentioning that the aforementioned computer device can be a terminal device or a server. When the computer device is a terminal device, it can specifically include but is not limited to: smartphones, tablet computers, laptop computers, desktop computers, vehicle-mounted terminals, smart TVs, smart wearable devices (such as smart bracelets, smart watches, etc.), game consoles, etc. In addition, the computer device can run the application used to implement the present application, as well as a variety of other applications such as image processing programs, multimedia playback programs, and navigation programs.

[0044] When the computer device is a server, the server may specifically include but is not limited to physical servers, and one or more cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), big data services and artificial intelligence platforms, etc. The embodiments of the present application do not impose specific restrictions on this.

[0045] For ease of description, the following is an example of using a data processing device to implement the various implementation methods proposed in this application. Unless otherwise specified, the data processing device may refer to the above-mentioned computer device for independently implementing this application, or it may refer to the above-mentioned data processing system for collaboratively implementing this application.

[0046] Based on the core principle of the aforementioned data processing solution, the present application embodiment specifically proposes a data processing method, which is executed by the aforementioned data processing device. Figure 3 , Figure 3 A schematic flow chart of the data processing method is shown in FIG. Figure 3 As shown, the method may include steps S301-S303:

[0047] S301: Receive a rule inference request, and obtain a rule inference text from the rule inference request, where the rule inference text includes rule keywords to be used when generating data classification rules.

[0048] In a specific embodiment, a rule inference request can be generated by a target client based on a data reference requirement when a user needs to obtain reference information to understand a keyword through the reference information. Keywords can be composed of one or more symbols or text, and the rule inference request can carry or include rule inference text, which can be used to describe the user's data reference requirement.

[0049] As an example, data reference requirements can be used to indicate the data that needs to be referenced (such as rule keywords) and / or the dimensions that need to be referenced. Among them, the dimensions that need to be referenced may include but are not limited to one or more of definition dimensions, example dimensions, and abbreviation dimensions. Among them, the definition dimension is mainly applicable to situations where it is expected to obtain the expression meaning of the keyword, such as when it is necessary to know what a logistics order number is; the example dimension is mainly applicable to situations where it is expected to obtain usage examples of the keyword, such as when it is necessary to know examples or application scenarios of logistics order numbers; the abbreviation dimension is mainly applicable to situations where it is expected to obtain abbreviated expressions of keywords, such as when it is necessary to know the English abbreviation of the logistics order number.

[0050] It's easy to understand that, in the context of generating data processing rules, rule inference text can be used to describe a user's questions about one or more rule keywords, where rule keywords can be the characters or words needed to generate data processing rules. In this case, the rule inference text can include question information generated based on the rule keywords.

[0051] For example, if a user is writing a rule for identifying a logistics tracking number and wants to obtain reference information for the "logistics tracking number," they can enter "Please explain in detail what a logistics tracking number is" in the target client, triggering a rule inference request. In this example, the rule inference text carried in the rule inference request could be "Please explain in detail what a logistics tracking number is," and the rule keyword must contain at least "logistics tracking number."

[0052] S302: Reasoning a classification target of a data classification rule based on rule reasoning text, and selecting a target description dimension that matches the classification target from preset description dimensions, where the target description dimension includes at least one of a semantic dimension and a data specification dimension.

[0053] In specific embodiments, a data classification rule refers to the classification criteria used to achieve a specific classification objective. A classification objective can also be understood as a data processing intent, primarily used to indicate the purpose of the data classification rule. For example, a classification objective could be "identify a logistics order number," "identify a valid logistics order number," or "identify the shipping company to which a valid logistics order number belongs," etc.

[0054] In a feasible implementation method, the classification target can be determined by the data processing device through intent recognition of the rule reasoning text, or it can be predicted based on semantic understanding of the rule reasoning text, combining semantic information with client characteristics of the target client (such as the type of business provided by the client, the data processing rules generated or queried by the client in the historical time period). There is no restriction here.

[0055] In addition, the preset description dimension refers to the angle that can be selected when describing the rule keywords before responding to the rule inference request. In an embodiment of the present application, the preset description dimension may include the aforementioned definition dimension, example dimension, abbreviation dimension and other one or more dimensions. Furthermore, if after describing the rule keywords under a certain description dimension, it can assist the target client (or user) to quickly construct a data classification rule that meets its expected classification target based on the obtained keyword description information, then the description dimension can be used as the target description dimension.

[0056] To facilitate a clear and detailed understanding of rule keywords, the target description dimension can include at least one of a semantic dimension and a data specification dimension. For example, the semantic dimension can include a definition dimension, primarily used to explain the concept of the rule keyword. The data specification dimension can include one or more of an example dimension, an abbreviation dimension, an application scenario dimension, a parameter value range dimension, and a data format dimension.

[0057] For example, if the rule inference text is "Please explain in detail what a logistics tracking number is," the rule keyword used to generate the data classification rule could be "logistics tracking number," and the classification target derived from inferring the rule inference text could be "identify the logistics tracking number." In this case, to facilitate the user's rapid generation of data classification rules capable of identifying physical tracking numbers, both the "semantic dimension" and the "data specification dimension" can be used as target description dimensions. This allows the user to quickly understand the definition of a logistics tracking number based on the keyword description information obtained from the semantic dimension, and to clarify the data format of the logistics tracking number based on the keyword description information from the data specification dimension, thereby making the generated data classification rule more accurate.

[0058] S303: Call the target large language model to perform keyword description processing on the rule keywords under the target description dimension to obtain keyword description information of the rule keywords. The keyword description information is used to generate data classification rules.

[0059] In a specific embodiment, the target large language model can be a large language model obtained by training the large model using a large number of language samples. The large model refers to a neural network model based on artificial intelligence and deep learning technologies and having a large number of processing parameters. The large model is trained using a large number of language samples so that the trained large model can be used to process natural language. Therefore, the trained large model can also be referred to as a large language model.

[0060] Artificial intelligence (AI) is a comprehensive field within computer science and technology, primarily used to study the design principles and implementation methods of various intelligent machines. Applying AI to machines can empower them with perception, reasoning, and decision-making capabilities. Specifically, AI can utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, enabling them to perceive their environment and acquire knowledge. In other words, AI encompasses theories, methods, techniques, and application systems that enable digital computers or related machines to use their learned knowledge to achieve optimal results.

[0061] In practical applications, artificial intelligence technology covers a wide range of fields, including both hardware-level and software-level technologies. Among them, hardware-level technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, and other technologies. Software-level technologies generally include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning technology. This application mainly utilizes deep learning technology and natural language processing (NLP) technology in artificial intelligence technology.

[0062] Natural language processing (NLP) is a discipline that integrates linguistics, computer science, and mathematics. It primarily studies theories and methods for enabling effective communication between humans and computers using natural language (i.e., the language we use daily). The embodiments of this application utilize NLP primarily through the construction of a large language model (or target large language model).

[0063] In a feasible implementation method, in order to enable the target large language model to have data processing capabilities that are more in line with actual business needs and thus obtain more reference-oriented keyword description information, the data processing device can perform secondary training on the large language model based on the business data generated in one or more similar businesses, and use the trained large language model that can meet actual business needs as the target large language model.

[0064] Among them, business requirements can be used to indicate professional fields that need assistance in generating data classification rules, and / or performance indicators that the model needs to achieve. Similar businesses can refer to data processing businesses whose business requirements are similar to the business requirements of the currently developed data processing business (referred to as the current business). And as an example, the professional fields (or knowledge fields) involved in each similar business should cover the professional fields involved in the current business requirements, so that the target large language model can perform excellent processing effects under the current business requirements.

[0065] In addition, when keyword description information is used to generate data classification rules, the data processing device can first convert the keyword description information into a certain (or multiple) data feedback format, and the data feedback format can include the data format commonly used when using the rule keyword for rule writing, so as to facilitate the rapid and efficient use of keyword description information when generating data classification rules.

[0066] Specifically, when determining the data feedback format, the data processing device may first use a knowledge matching algorithm to calculate the knowledge matching degree between the keyword description information and each knowledge domain, and then use the common data format associated with the knowledge domain with the highest knowledge matching degree as the data feedback format. Furthermore, as an example, the principle of the knowledge matching algorithm may be: after calculating the feature similarity between the semantic characteristics of the keyword description information and the expression features of each knowledge domain, based on the principle that feature similarity and knowledge matching degree are positively correlated, the knowledge matching degree between the keyword description information and each knowledge domain is determined.

[0067] After receiving a rule inference request, the embodiment of the present application obtains the classification target corresponding to the data classification rule to be generated by inferring the rule inference text carried in the inference rule inference request, and selects the target description dimension that matches the classification target from the preset description dimension, so as to call the target large language model to perform keyword description processing on the rule keywords in the rule inference text under the target description dimension to obtain keyword description information. Since the target description dimension in the embodiment of the present application includes at least one of the semantic dimension and the data specification dimension, the keyword description information can clearly express the meaning or usage specification of the rule keywords, and thus referring to the keyword description information can also accurately and thoroughly understand the rule keywords, so that the data classification rules generated based on this understanding can accurately achieve the classification target, which is ultimately beneficial to the improvement of the classification accuracy of the data classification business.

[0068] based on Figure 3 The method shown in the embodiment of the present application also proposes another data processing method, which can still be executed by the data processing device mentioned above. Figure 4 , Figure 4 A schematic flow chart of the data processing method is shown in FIG. Figure 4 As shown, the method may include steps S401-S406:

[0069] S401: Receive a rule inference request, and obtain a rule inference text from the rule inference request, where the rule inference text includes rule keywords to be used when generating data classification rules.

[0070] In a specific embodiment, the partial implementation principle of step S401 can refer to the relevant embodiment of the aforementioned step S301 and will not be repeated here. However, it should be noted that in actual applications, the data processing device can respond to rule inference requests in batches to improve the overall data processing efficiency.

[0071] Specifically, the data processing device can continuously receive rule inference requests and detect the number of rule inference requests currently to be responded to, and when the number of requests is greater than or equal to the target number, select a target number of rule inference requests from each rule request currently to be responded to, and then obtain the rule inference text from each rule inference request in parallel for the target number of rule inference requests.

[0072] It is not difficult to understand that compared with responding to N rule inference requests one by one, responding to N rule inference requests in parallel can effectively shorten the overall response time, thereby improving the overall efficiency of data processing, where N is an integer greater than 1.

[0073] S402: Inferring a classification target of a data classification rule based on rule reasoning text, and selecting a target description dimension that matches the classification target from preset description dimensions, where the target description dimension includes at least one of a semantic dimension and a data specification dimension.

[0074] In a specific embodiment, a data processing device can infer a classification target by obtaining the question intent of a rule inference text. Specifically, the data processing device can first perform knowledge question and answer analysis on the rule inference text to obtain one or more knowledge question information, and then predict the classification target of the data classification rule by identifying the question intent of each knowledge question information. The knowledge question information can be used to initiate a question for a rule keyword from a dimension, and the classification target can be used to indicate the function that can be achieved based on the answer information expected to be obtained based on the question intent.

[0075] For example, if the rule-based inference text is "Please describe the logistics order number in detail," the knowledge question information obtained from knowledge question answering analysis of the rule-based inference text may include "What is a logistics order number" and "What is the data composition of a logistics order number?" In this case, after performing intent recognition on each knowledge question information, the question intents obtained may be "Get the meaning of a logistics order number" and "Get an example of a logistics order number." It is easy to understand that the answer information expected based on these two question intents can be used to identify the logistics order number, and thus the predicted classification target may be "Identify the logistics order number."

[0076] S403: Call the target large language model to perform keyword description processing on the rule keywords under the target description dimension to obtain keyword description information of the rule keywords. The keyword description information is used to generate data classification rules.

[0077] In a specific embodiment, the keyword description information obtained under the semantic dimension can be used to describe the definition of the rule keywords under one or more knowledge fields; the keyword description information obtained under the data specification dimension can be used to describe at least one of the usage specifications and classification scenarios of the rule keywords when used to generate data classification rules.

[0078] The data classification rules can be generated by a data processing device. For example, when the keyword description information includes description information under the data specification dimension, the data processing device can first select a target rule template that matches the classification target from one or more preset rule templates. The target rule template includes one or more information to be configured, so that the data processing device can parameterize the rule keywords according to the description information under the data specification dimension, obtain the configuration parameters required for the information to be configured, and obtain the data classification rules by filling the configuration parameters into the target rule template.

[0079] For example, suppose the data classification rule currently being constructed is a rule for identifying logistics order numbers, with the logistics order number being the key word in the rule. The target rule template can then be a template that can be used to construct a rule for identifying a certain type of data item, and the to-be-configured information in the target rule template can include the data item to be identified and the value range of the data item. In this case, parameterizing the rule key word can include generating a parameter name (e.g., X) for indicating the logistics order number and a parameter value type (e.g., a numeric type or a string type) corresponding to the parameter name.

[0080] S404: Obtain the client type of the target client, where the target client refers to the client that sends the rule inference request.

[0081] In one implementation, the client type may include a web page type. A web page type client refers to a client capable of displaying a target service page to a user, such as various search engines. The target service refers to a data processing service that requires utilizing one or more of the embodiments proposed in this application to obtain keyword description information.

[0082] For web-type clients, the user of the client (or user) can trigger the generation of rule inference requests by performing human-computer interaction operations on the service page, and after receiving the keyword description information, the keyword description information can be visually displayed in a certain layout to facilitate reading by the user.

[0083] In another implementation, the client type may also include a program development type. A program development type client refers to a client used to write program development code. Such a client typically initiates rule inference requests by calling an Application Programming Interface (API). Therefore, a program development type client can also be understood as a client that needs to initiate rule inference requests by calling an application programming interface.

[0084] S405 : Formatting the keyword description information according to the information feedback format associated with the client type to obtain formatted keyword description information.

[0085] In one implementation, when the client type is a web page type, the data processing device, when performing formatting processing, may include a process of highlighting some information fragments in the keyword description information so that these information fragments can be highlighted in the target client. The highlighting may include at least one of magnified display and highlighted display.

[0086] Generally speaking, the information fragments that are subjected to the highlighted arrangement processing are the information fragments that have high reference value for the generation process of data classification rules. Highlighting such information fragments can enable related objects to locate important reference information more quickly and clearly, and ultimately can improve the efficiency and accuracy of data classification rules to a certain extent.

[0087] In a specific implementation, the data processing device may first perform knowledge domain identification on the keyword description information to determine information segments (hereinafter referred to as knowledge information segments) whose knowledge relevance to the preset knowledge domain is greater than or equal to a relevance threshold from the keyword description information, and then identify each knowledge information segment as an information segment requiring prominent arrangement processing. The preset knowledge domain may include the knowledge domain involved in the aforementioned classification target, or may include one or more highly specialized knowledge domains pre-configured in the data processing device, without limitation.

[0088] In another implementation, when the client type is a program development type, the data processing device may include converting the keyword description information into a certain data format during formatting. Specifically, the data processing device may first determine the program interface used by the target client when sending the rule inference request, and then convert the keyword description information into the data format indicated by the interface protocol of the program interface to obtain the formatted keyword description information.

[0089] The data format indicated by the interface protocol may include a data format recognizable by the program interface (such as the JSON format), or may include a data format that can be quickly put into use by the target client based on its own data processing needs (such as the code writing format of the C language program code), and this is not limited here. However, it is understood that by converting the keyword description information according to the data format indicated by the interface protocol, the target client can efficiently generate data classification rules based on the formatted keyword description information.

[0090] S406: Feedback the edited keyword description information to the target client, so that the target client generates data classification rules based on the keyword description information and rule keywords.

[0091] In a specific embodiment, the keyword description information is used to describe the rule keywords to deepen or enlighten the target client's understanding of the rule keywords, so that the target client can generate more accurate data classification rules. However, it should be noted that in one implementation, the target client can generate data classification rules based on program instructions input by the object, and the program instructions can be used to indicate the construction method of the data classification rules. In another implementation, the target client can also generate data classification rules in accordance with the method of generating data classification rules based on the target rule template described in the aforementioned step S403, which will not be described in detail.

[0092] The embodiment of the present application performs keyword description processing on the rule keywords in the rule reasoning text under the semantic dimension and / or data specification dimension by calling the target large language model, and can obtain keyword description information for accurately describing the rule keywords. Therefore, the data classification rules generated with reference to the keyword description information can also have a higher accuracy, which is conducive to improving the classification accuracy of the data classification business that applies the data classification rules. In addition, after generating the keyword description information, the embodiment of the present application can also use the client type of the target client of the keyword description information as needed to format the keyword description information, so that the keyword description information can be fed back in a data format that is convenient for the target client to use, thereby improving the generation efficiency of the data classification rules to a certain extent.

[0093] Based on the above embodiments, the present application also proposes a method for training a large language model to obtain a target large language model. The computer device used to perform model training can be the aforementioned data processing device, that is, it can be another device different from the aforementioned data processing device (hereinafter referred to as a training device for ease of description). Figure 5 The structure of the training device shown, and Figure 6 The model training process shown in Figure 1 details a specific implementation method for obtaining the target large language model. Figure 6 , the model training process may include steps S601-S604:

[0094] S601: Collect training data from one or more question-answering services.

[0095] In a specific embodiment, the training device may call Figure 5 The data acquisition module in the embodiment realizes the collection of training data. And as an example, when collecting training data, the data acquisition module may mainly include Figure 5 The three processes of data collection, data cleaning and data construction are shown in .

[0096] Among them, data collection mainly refers to the process of collecting public business data or business data that is allowed to be commercially used from one or more Internet intelligent question-and-answer services (hereinafter referred to as question-and-answer services) and storing the collected business data. As an example, the knowledge fields involved in different question-and-answer services can be different, so that the embodiments of the present application can collect business data from various industries. In addition, the data collection module can use Cloud Object Storage (COS) to store the collected business data.

[0097] COS is a distributed storage service that supports access via the Hypertext Transfer Protocol (HTTP) and Hypertext Transfer Protocol Secure (HTTPS). COS has no directory hierarchy or data format restrictions and can accommodate massive amounts of data. However, the model training process requires frequent reading or updating of training data. Therefore, using COS to store collected business data facilitates data reading and management during subsequent model training, thereby improving the efficiency of model training to a certain extent.

[0098] Data cleaning mainly refers to the process of obtaining relevant data for constructing training data from the collected business data. Optionally, data cleaning can include two processing stages: data parsing and data screening. Among them, the data formats supported by the data parsing stage may include but are not limited to: Comma-Separated Values ​​(CSV) format, slide (Power Point, PPT) format, data tables (such as Excel), text documents (such as Word), etc. In addition, the data screening stage can be used to deduplicate the relevant data obtained in the data parsing stage and remove empty fields, so as to obtain more streamlined data for generating training data, thereby facilitating the generation of training data with more reference value.

[0099] Data construction mainly refers to the process of constructing training data with a unified format based on the relevant data obtained from data cleaning. Since the parameters required for large language models to process training data in different data formats are usually different, it is difficult to achieve the expected training goals (such as training efficiency, model measurement effect, etc.) by using training data with non-uniform formats to train large language models. However, this application can effectively improve the model training efficiency and training effect by constructing training data with a unified format during the data construction process.

[0100] Taking the format of training data into question-answer pairs as an example, the data collection module can collect enough business data and then Figure 7 The training data collection process shown processes the business data in the COS to obtain training data.

[0101] like Figure 7 As shown, the data collection module can first read one or more business data from COS, and then parse the business data based on the format of each business data to obtain the text content in the business data. After removing the empty fields in the text content (i.e., fields with empty data values, such as spaces), the data is deduplicated. Finally, one or more question-answer pairs are constructed based on the text content obtained by deduplication, and each question-answer pair is used as training data. The above process is repeated until all business data collected in COS are cleaned.

[0102] For example, assuming that the business data that currently needs data cleaning is in PDF file format, the training device can parse the business data to obtain the text content in the business data, and then remove the characters indicating blank content in the text (such as space characters), and then further deduplicate the obtained text content, and finally convert the deduplicated text content into a predetermined data format (such as the format of question and answer pairs).

[0103] Because the business data in COS is collected from multiple vertical knowledge domains (i.e., different knowledge domains), the training data constructed based on this collected business data also involves multiple knowledge domains. By training the large language model multiple times with training data from different knowledge domains to obtain the target large language model, the target large language model's ability to transfer across different knowledge domains is guaranteed, allowing it to still perform well in data processing businesses related to other knowledge domains. This can reduce manual repetitive work (such as repeated model training) and lower business development costs.

[0104] S602: Using the training data, and performing model optimization processing on the large language model according to preset optimization execution parameters to obtain an optimized large language model.

[0105] In a specific embodiment, the training device may call Figure 5 The model training module in the performs model optimization (or model training) on ​​the large language model. And as an example, when training the large language model, the model training module can mainly include Figure 5 The three processes of model parameter setting, model optimization, and model saving are shown in Figure 1. Model parameter setting is mainly used to configure and initialize the various parameters required for the model training process. Model optimization is mainly to optimize the various parameters (such as weight parameters) in the large language model in the direction of enhancing the model processing capability (such as accuracy). Model saving mainly refers to saving the optimized large language model.

[0106] Regarding the model parameter setting process, in one implementation, it at least includes setting the optimization execution parameters used when optimizing the parameters of the large language model, and the optimization execution parameters can specifically include the learning rate, the number of samples taken for a single training (i.e., Batch Size), the number of training rounds (i.e., the number of epochs), the training accuracy, the length of the input data, etc.

[0107] In other implementations, the parameter categories required for setting the model parameters may also specifically include one or more of the following: training framework parameter category, training parameter category, hardware parameter category, and storage parameter category. The following is an exemplary description of the parameters included in each parameter category:

[0108] (1) Training framework parameter class

[0109] The training framework parameter class can specifically include one or more of the following: framework indicator parameters, optimization framework indicator parameters, and management node ports (also known as master ports). The framework indicator parameters indicate the deep learning framework used for model training, such as TensorFlow and PyTorch; the optimization framework indicator parameters indicate the optimizer used for model training, such as Megatron and DeepSpeed; and the master port is the communication port used by the management node of the distributed training system.

[0110] The distributed training system consists of multiple training nodes. A management node sends and receives data with each training node via a communication port, thereby scheduling the nodes to collaborate on training a large language model. In a distributed training system, model training can be implemented using either data parallelism or model parallelism.

[0111] The key principle of data parallelism is to deploy a complete large language model on each training node and divide the training dataset into multiple subsets, allowing different training nodes to process different subsets. Each training node calculates its own forward and backward gradients during processing and sends these gradients to the management node. The management node then determines the update information for each relevant parameter in the large language model based on the gradients calculated by each training node. This update is then broadcast to each training node, allowing each training node to update the parameters of the large language model based on the updated information. This completes a training iteration.

[0112] By employing data parallelism for model training, the computing resources of the distributed training system can be fully utilized, achieving high model training efficiency even with large training datasets. Furthermore, since each training node processes a portion of the training data, different training nodes can use training data with different data distributions. This allows the overall model training process to adapt to uneven training data distributions, reducing the requirements for training data collection and thus improving training efficiency to a certain extent.

[0113] The main principle of model parallelism is to divide a complete large language model into multiple sub-models and deploy them collaboratively across multiple training nodes. This means that different training nodes perform different data processing tasks during model training. Therefore, during a complete model training process, different training nodes are used to perform data processing at different stages of training. Data connectivity between stages is achieved by the management node through the Master port.

[0114] By adopting model parallelism for model training, the problem of insufficient memory of a single electronic device to train large models can be solved. In addition, since the sub-models on different training nodes can be processed in parallel, the communication overhead between training nodes is reduced, thereby improving the model training speed.

[0115] (2) Training parameter class

[0116] In addition to the aforementioned optimization execution parameters, the training parameter class may further include one or more of the optimizer (i.e., Optimizer), the Zero Redundancy Optimizer (ZeRO) stage, the performance evaluation interval, etc. By utilizing ZeRO, computing resources and memory resources can be aggregated in parallel during model training, thereby reducing the memory and computing requirements of each electronic device used for model training. Optionally, the electronic device used for model training may include one or both of a central processing unit (CPU) and a graphics processing unit (GPU), and each electronic device may be used as a training node to construct the aforementioned distributed training system.

[0117] (3) Hardware parameter class

[0118] The hardware parameter class can specifically include one or more of the following: number of GPUs, number of GPU cards per machine, GPU card model, image used in the runtime environment, etc. When the number of GPUs is greater than one, the hardware parameter class can further include: the number of samples taken for a single training run on a GPU, and / or the number of samples taken for performance evaluation of a GPU.

[0119] (4) Save parameter class

[0120] The saving parameter class may specifically include one or more of the model saving path, model saving steps, evaluation indication parameter (used to indicate whether to trigger the performance evaluation of the model), coverage indication parameter (used to indicate whether to cover the intermediate model), total model saving times, etc.

[0121] For ease of understanding, the model optimization process will be explained using question-answer pairs as training data. A question-answer pair is a data pair consisting of a question and a reference answer.

[0122] In this case, the model training module can first obtain the question intent of the question information in the training data, then select the description dimension that matches the question intent from the preset description dimensions, and call the large language model to answer the question information under the selected description dimension, thereby obtaining the model answer information. Then, according to the preset optimization execution parameters, the large language model is optimized with the optimization goal of reducing the information difference between the reference answer information and the model answer information, until a large language model that meets the optimization goal (called the optimized large language model) is obtained.

[0123] S603: Perform a performance evaluation on the optimized large language model to obtain a performance evaluation result.

[0124] In a specific embodiment, the training device may call Figure 5 The performance evaluation module in step S602 performs a performance evaluation on the optimized large language model obtained in step S602. The main principle of performance evaluation is to use the optimized large language model to process test data and evaluate whether the optimized large language model is sufficient for actual data processing business based on one or more performance indicators such as accuracy and false positive rate achieved in the processing results.

[0125] As an example, when the performance evaluation module is used to perform performance evaluation, it can mainly include Figure 5 The three processes of reading test data, model testing, and calculating performance indicators are shown in FIG. Among them, the test data refers to the data used for performance testing of the optimized large language model, and its format is the same as that of the training data.

[0126] In one implementation, the test data may be one or more pieces of training data selected after the training data are constructed in step S601. The selected test data is not involved in the model optimization process to ensure that the performance evaluation results obtained when the test data is used for performance evaluation have a high degree of credibility.

[0127] Model testing refers to the process of inputting test data into the optimized large language model to obtain model responses corresponding to the test data. Performance metric calculation refers to the process of comparing the model responses corresponding to the test data with the reference responses in the test data to calculate one or more performance metrics, such as the accuracy and / or false positive rate of the model responses. Based on these performance metrics, the optimized large language model is then determined to meet performance requirements.

[0128] S604: When the performance evaluation result indicates that the model performance of the optimized large language model meets the preset performance requirements, the optimized large language model is used as the target large language model.

[0129] In a specific embodiment, when the performance evaluation result obtained in step S603 indicates that the optimized large language model does not meet the performance requirements, the optimization execution parameters mentioned in step S602 can be updated, and the large language model (or optimized large language model) can be re-optimized according to the execution principle of step S603 based on the updated optimization execution parameters until an optimized large language model that meets the performance requirements is obtained.

[0130] Correspondingly, it can be understood that when the performance evaluation results indicate that the optimized large language model meets the performance requirements, the training device can save the optimized large language model and can further output the optimized large language model to the aforementioned data processing device so that the data processing device can apply it as the target large language model.

[0131] It is worth mentioning that when applying the target large language model, the model input data can be kept consistent with the format of the training data, so that the target large language model can fully exert the model processing performance and obtain a more accurate model processing result (such as keyword description information).

[0132] When generating a target large language model, the embodiment of the present application first uses training data to train the large language model, and then performs a performance evaluation on the optimized large language model that achieves the optimization target after training, so that the optimized large language model is only used as the target large language model for actual application when the performance evaluation results show that it meets the performance requirements. This can improve the model quality of the target large language model, and thus make the keyword description information obtained by the embodiment of the present application using the target large language model have higher reference value.

[0133] In order to facilitate the embodiment of this application Figure 3 、 Figure 4 and Figure 6 The various implementation methods proposed in the article are quickly applied. The following is combined with the data security classification scenario and Figure 8 An application method of an embodiment of the present application is described in detail.

[0134] Data security classification relies primarily on data detection rules, which are primarily used to detect sensitive data within an industry. As various industries are rapidly developing, new concepts are constantly emerging, making it more difficult for technicians to write data detection rules. However, the embodiments of this application can help technicians quickly understand industry concepts and improve the efficiency of rule writing.

[0135] Specifically, in the data security classification scenario, the format of the training data used for model training of the large language model can continue to be question-answer pairs, and the question form of the question information in the question-answer pair can be any preset form of explanation form, example form and abbreviation form.

[0136] Explanation-style questions are primarily used to instruct the data processing device to provide a detailed explanation of the definition of a rule keyword, allowing the user to understand the rule keyword based on the answer to the question. For example, an explanation-style question might be something like, "Explain the meaning of XX," where XX is a sequence of one or more text characters, typically used to indicate a proprietary concept within a professional field (or knowledge domain).

[0137] Example-based questions are primarily used to instruct data processing devices to provide usage specifications for rule keywords (data format, value range, usage scenarios, etc.), allowing users to quickly apply these rule keywords when writing data detection rules. For example, an example-based question might be, "Introduce an example of using XX."

[0138] The abbreviated question information is mainly used to instruct the data device to provide abbreviated expressions of regular keywords so that users can understand abbreviated proper nouns. As an example, the abbreviated question information can be such as "Introduce the abbreviation or proper noun of XX".

[0139] In actual applications, in order to ensure the processing effect of the target large language model in actual applications, the input data of the target large language model can be limited to the same format as the training data. However, in order to adapt to the question-answering habits of different users, the embodiment of the present application will not restrict the user's question format, but will instead perform instruction parsing on the question content input by the user to construct question information in a preset form, thus taking into account both the model processing effect and user convenience.

[0140] See Figure 8 , Figure 8 The structure and function of an application device are shown. The application device refers to a computer device that applies the target large language model to provide users with description services of rule keywords. Figure 8 , the application device may exemplarily include an instruction parsing module, a model reasoning module and a result processing module.

[0141] The instruction parsing module can be used to receive, parse and prompt instructions (i.e. Figure 8 The model inference module is mainly used to call the target large language model, which can be used to implement batch processing (i.e. Figure 8 The result processing module is mainly used to convert the answer data output by the target large language model into a user-friendly feedback format. Specifically, it can be used to implement the three main functions of output data parsing, data formatting, and result sending.

[0142] Specifically, when writing data detection rules, users can use the client to generate question instructions for the professional terms that the user wants to know. The application device can call the instruction receiving function in the instruction parsing module to receive the question instructions (or rule inference requests) sent by the user, and then call the instruction parsing function to parse the question content (or rule inference text) carried in the question instruction. The question intention, and finally call the instruction prompt function to construct the corresponding form of question information according to the question intention, and use the question information as the input data of the target large language model.

[0143] Furthermore, the Batch processing function in the model inference module is called to detect the number of question information currently to be processed, and after reaching a certain number, the model inference function is called to determine the target description dimension corresponding to the professional terms in each question information, and then the target large language model is called to perform keyword description processing under the target description dimension to obtain the description information corresponding to each professional term, and finally the data output function is called to output the description information of each professional term.

[0144] In one implementation, the descriptive information can be directly output to the client that initiated the question, or it can be output to a result processing module of an application device, which then optimizes the format of the descriptive information before feeding it back to the corresponding client. For example, the result processing module can invoke an output data parsing function to determine the client type of the receiving client corresponding to the descriptive information, determine the information feedback format associated with that client type, then invoke a data formatting function to format the descriptive information according to the information feedback format, and finally invoke a result sending module to feed the formatted descriptive information back to the receiving client.

[0145] Applying this application in a data security classification scenario can provide rule configuration personnel with highly referenceable descriptive information for various professional concepts, thereby increasing their ability to understand sensitive data and reducing the interference caused by various professional concepts on rule configuration personnel, enabling rule configuration personnel to construct accurate data detection rules, and thus enabling program products that apply the data detection rules to have stronger detection capabilities for sensitive data (such as higher accuracy and lower false alarm rate). It can be seen that this application can demonstrate strong practicality in data security classification scenarios.

[0146] Based on the above Figure 3 and Figure 4 The data processing method shown in the present application also discloses a data processing device for performing the data processing method shown in the present application. Figure 3 or Figure 4 The data processing means may be a computer program (including program code) running on a data processing device. Figure 9 , the data processing device may include at least Figure 9 The receiving module 901, the inference module 902 and the model calling module 903 in FIG.

[0147] A receiving module 901 is configured to receive a rule inference request and obtain a rule inference text from the rule inference request, wherein the rule inference text includes rule keywords to be used when generating data classification rules;

[0148] An inference module 902 is configured to infer a classification target of the data classification rule based on the rule inference text, and select a target description dimension that matches the classification target from preset description dimensions, wherein the description dimension includes at least one of a semantic dimension and a data specification dimension;

[0149] The model calling module 903 is used to call the target large language model to perform keyword description processing on the rule keywords under the target description dimension to obtain keyword description information of the rule keywords, and the keyword description information is used to generate the data classification rule.

[0150] In one embodiment, the keyword description information obtained by the target large language model under the semantic dimension is called to describe the definition of the rule keyword under one or more knowledge fields; the keyword description information obtained by the target large language model under the data specification dimension is called to describe at least one of the usage specifications and classification scenarios of the rule keyword when used to generate data classification rules.

[0151] In another embodiment, the rule inference request is sent by the target client, and the data processing device may further include an information feedback module 904, which may be configured to execute:

[0152] Obtaining the client type of the target client;

[0153] formatting the keyword description information according to the information feedback format associated with the client type to obtain formatted keyword description information;

[0154] The formatted keyword description information is fed back to the target client, so that the target client generates the data classification rule based on the keyword description information and the rule keyword.

[0155] In yet another embodiment, the client type includes a webpage type; and when the information feedback module 904 formats the keyword description information according to the information feedback format associated with the client type to obtain the formatted keyword description information, the information feedback module 904 may be specifically configured to perform:

[0156] By performing knowledge domain identification on the keyword description information, a knowledge information segment in the keyword description information is determined; wherein the knowledge information segment refers to an information segment whose knowledge relevance to a preset knowledge domain is greater than or equal to a relevance threshold;

[0157] The knowledge information segments in the keyword description information are highlighted and arranged to obtain the arranged keyword description information, so as to highlight and display the knowledge information segments during the display of the arranged keyword description information.

[0158] In yet another embodiment, the client type includes a program development type; and when the information feedback module 904 formats the keyword description information according to the information feedback format associated with the client type to obtain the formatted keyword description information, the information feedback module 904 may be specifically configured to execute:

[0159] Obtaining a program interface used by the target client to send the rule inference request;

[0160] The keyword description information is format-converted according to the data format indicated by the interface protocol of the program interface to obtain the formatted keyword description information.

[0161] In yet another embodiment, the data processing apparatus may further include a formatting module 905, which may be configured to execute:

[0162] Calculating the knowledge matching degree between the keyword description information and each knowledge field using a knowledge matching algorithm, and obtaining the commonly used data format associated with the knowledge field with the highest knowledge matching degree;

[0163] The keyword description information is format-converted according to the commonly used data format, and the keyword description information after format conversion is used to generate the data classification rule.

[0164] In another embodiment, the data processing device may further include a rule generation module 906, the keyword description information includes description information obtained under the data specification dimension, and the rule generation module 906 may be used to:

[0165] Selecting a target rule template that matches the classification target from one or more preset rule templates, wherein the target rule template includes one or more pieces of information to be configured;

[0166] Parameterize the rule keywords according to the description information under the data specification dimension to obtain the configuration parameters required for the information to be configured;

[0167] Fill the configuration parameters into the target rule template to obtain the data classification rule.

[0168] In yet another embodiment, when inferring the classification target of the data classification rule based on the rule inference text, the reasoning module 902 may be specifically configured to perform:

[0169] Performing knowledge question-answering analysis on the rule-based reasoning text to obtain one or more pieces of knowledge question information;

[0170] The question intention of each knowledge question information is identified, and the classification target of the data classification rule is predicted based on the identified question intentions.

[0171] In yet another embodiment, the data processing apparatus may further include a model training module 907, which may be configured to execute:

[0172] Collect training data from one or more question-answering services;

[0173] Using the training data, and performing model optimization processing on the large language model according to preset optimization execution parameters to obtain an optimized large language model;

[0174] Performing a performance evaluation on the optimized large language model to obtain a performance evaluation result;

[0175] When the performance evaluation result indicates that the model performance of the optimized large language model meets the preset performance requirements, the optimized large language model is used as the target large language model.

[0176] In another embodiment, the training data is a question-answer pair consisting of question information and reference answer information; when the model training module 907 uses the training data and performs model optimization processing on the large language model according to preset optimization execution parameters to obtain an optimized large language model, it can be specifically used to perform:

[0177] Obtaining the question intention of the question information in the training data;

[0178] Selecting a description dimension that matches the question intention from the preset description dimensions;

[0179] Calling the large language model to answer the question information under the selected description dimension to obtain model answer information;

[0180] According to the preset optimization execution parameters, with the optimization goal of reducing the information difference between the reference answer information and the model answer information, the large language model is subjected to model optimization processing to obtain the optimized large language model.

[0181] In yet another embodiment, the receiving module 901 may also be configured to execute:

[0182] Detect the number of rule inference requests received;

[0183] When the number of requests is greater than or equal to the target number, selecting the target number of rule inference requests from the received rule requests;

[0184] For the target number of rule inference requests, the step of obtaining rule inference texts from the rule inference requests is performed in parallel.

[0185] According to one embodiment of the present application, Figure 9 The modules in the data processing device shown are divided based on logical functions. The modules can be individually or all combined into one or more other modules to form a module, or one (or some) of the modules can be further divided into multiple modules with smaller functions to form a module, which can achieve the same operation without affecting the realization of the technical effects of the embodiment of the present application. In other embodiments of the present application, the data processing device can also include other modules. In actual applications, these functions can also be implemented with the assistance of other modules, and can be implemented with the assistance of multiple modules.

[0186] According to another embodiment of the present application, a computer program that can execute the following operations can be executed on a general-purpose computing device such as a data processing device including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM), and other processing elements and storage elements. Figure 3 and Figure 4 A computer program (including program code) for each step of the method shown is used to construct Figure 9 The data processing device shown in the figure can be used to implement the data processing method of the embodiment of the present application. The computer program can be recorded on a computer storage medium, for example, and loaded into the above-mentioned data processing device through the computer storage medium and can be run therein.

[0187] The data processing device proposed in the embodiment of the present application calls the target large language model to perform keyword description processing on the rule keywords in the rule reasoning text under the semantic dimension and / or data specification dimension, and can obtain keyword description information for indicating the specific meaning of the rule keywords or the corresponding data usage specifications. Therefore, by referring to the keyword description information, the rule keywords can be accurately and comprehensively understood. In this case, the generated data classification rules can also have a higher accuracy, which is conducive to improving the classification accuracy of the data classification business to which the data classification rules are applied.

[0188] Based on the relevant descriptions of the above method embodiments and apparatus embodiments, the present application also provides a data processing device, see Figure 10 The data processing device at least includes a processor 1001 and a computer storage medium 1002, and the processor 1001 and the computer storage medium 1002 are connected via a bus or other means.

[0189] The aforementioned computer storage medium 1002 is a memory device within a data processing device, used to store programs and data. It is understood that the computer storage medium 1002 herein may include both built-in storage media within the data processing device and, of course, extended storage media supported by the data processing device. Computer storage medium 1002 provides storage space, which stores the operating system of the data processing device. Furthermore, this storage space also stores one or more computer programs suitable for loading and execution by processor 1001. These computer programs may be one or more program codes.

[0190] It should be noted that the computer storage medium may be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage; optionally, it may be at least one storage medium located remote from the processor. Processor 1001 (or CPU (Central Processing Unit)) is the computing and control core of a data processing device, and is suitable for implementing one or more computer programs, and is specifically suitable for loading and executing one or more computer programs to implement corresponding method flows or corresponding functions.

[0191] In addition, optionally, the data processing device can also have a communication connection with one or more target output devices 1003 (such as a client, a mobile terminal device, etc.), and can trigger rule inference requests and receive keyword description information by interacting with the target output device 1003.

[0192] In one embodiment, the processor 1001 may load and execute one or more computer programs stored in the computer storage medium 1002 to implement the above-mentioned Figure 3 as well as Figure 4 Corresponding method steps in the method embodiment shown. In a specific implementation, one or more computer programs in the computer storage medium 1002 can be loaded and executed by the processor 1001:

[0193] receiving a rule inference request and obtaining a rule inference text from the rule inference request, wherein the rule inference text includes rule keywords to be used when generating data classification rules;

[0194] Inferring a classification target of the data classification rule based on the rule inference text, and selecting a target description dimension that matches the classification target from preset description dimensions, wherein the description dimension includes at least one of a semantic dimension and a data specification dimension;

[0195] The target large language model is called to perform keyword description processing on the rule keywords under the target description dimension to obtain keyword description information of the rule keywords, and the keyword description information is used to generate the data classification rule.

[0196] In one embodiment, the keyword description information obtained by the target large language model under the semantic dimension is called to describe the definition of the rule keyword under one or more knowledge fields; the keyword description information obtained by the target large language model under the data specification dimension is called to describe at least one of the usage specifications and classification scenarios of the rule keyword when used to generate data classification rules.

[0197] In yet another embodiment, the rule inference request is sent by a target client, and the processor 1001 may be configured to load and execute:

[0198] Obtaining the client type of the target client;

[0199] formatting the keyword description information according to the information feedback format associated with the client type to obtain formatted keyword description information;

[0200] The formatted keyword description information is fed back to the target client, so that the target client generates the data classification rule based on the keyword description information and the rule keyword.

[0201] In another embodiment, the client type includes a web page type; when the processor 1001 formats the keyword description information according to the information feedback format associated with the client type and obtains the formatted keyword description information, it can be specifically used to load and execute:

[0202] By performing knowledge domain identification on the keyword description information, a knowledge information segment in the keyword description information is determined; wherein the knowledge information segment refers to an information segment whose knowledge relevance to a preset knowledge domain is greater than or equal to a relevance threshold;

[0203] The knowledge information segments in the keyword description information are highlighted and arranged to obtain the arranged keyword description information, so as to highlight and display the knowledge information segments during the display of the arranged keyword description information.

[0204] In yet another embodiment, the client type includes a program development type; and when the processor 1001 formats the keyword description information according to the information feedback format associated with the client type to obtain the formatted keyword description information, the processor 1001 may be specifically configured to execute:

[0205] Obtaining a program interface used by the target client to send the rule inference request;

[0206] The keyword description information is format-converted according to the data format indicated by the interface protocol of the program interface to obtain the formatted keyword description information.

[0207] In yet another embodiment, the processor 1001 may also be configured to load and execute:

[0208] Calculating the knowledge matching degree between the keyword description information and each knowledge field using a knowledge matching algorithm, and obtaining the commonly used data format associated with the knowledge field with the highest knowledge matching degree;

[0209] The keyword description information is format-converted according to the commonly used data format, and the keyword description information after format conversion is used to generate the data classification rule.

[0210] In another embodiment, the keyword description information includes description information obtained under the data specification dimension, and the processor 1001 can also be used to load and execute:

[0211] Selecting a target rule template that matches the classification target from one or more preset rule templates, wherein the target rule template includes one or more pieces of information to be configured;

[0212] Parameterize the rule keywords according to the description information under the data specification dimension to obtain the configuration parameters required for the information to be configured;

[0213] Fill the configuration parameters into the target rule template to obtain the data classification rule.

[0214] In yet another embodiment, when the processor 1001 infers the classification target of the data classification rule based on the rule inference text, it can be specifically configured to load and execute:

[0215] Performing knowledge question-answering analysis on the rule-based reasoning text to obtain one or more pieces of knowledge question information;

[0216] The question intention of each knowledge question information is identified, and the classification target of the data classification rule is predicted based on the identified question intentions.

[0217] In yet another embodiment, the processor 1001 may also be configured to load and execute:

[0218] Collect training data from one or more question-answering services;

[0219] Using the training data, and performing model optimization processing on the large language model according to preset optimization execution parameters to obtain an optimized large language model;

[0220] Performing a performance evaluation on the optimized large language model to obtain a performance evaluation result;

[0221] When the performance evaluation result indicates that the model performance of the optimized large language model meets the preset performance requirements, the optimized large language model is used as the target large language model.

[0222] In another embodiment, the training data is a question-answer pair consisting of question information and reference answer information; when the processor 1001 uses the training data and performs model optimization processing on the large language model according to preset optimization execution parameters to obtain an optimized large language model, it can be specifically used to load and execute:

[0223] Obtaining the question intention of the question information in the training data;

[0224] Selecting a description dimension that matches the question intention from the preset description dimensions;

[0225] Calling the large language model to answer the question information under the selected description dimension to obtain model answer information;

[0226] According to the preset optimization execution parameters, with the optimization goal of reducing the information difference between the reference answer information and the model answer information, the large language model is subjected to model optimization processing to obtain the optimized large language model.

[0227] In yet another embodiment, the processor 1001 may also be configured to load and execute:

[0228] Detect the number of rule inference requests received;

[0229] When the number of requests is greater than or equal to the target number, selecting the target number of rule inference requests from the received rule requests;

[0230] For the target number of rule inference requests, the step of obtaining rule inference texts from the rule inference requests is performed in parallel.

[0231] The data processing device proposed in the embodiment of the present application calls the target large language model to perform keyword description processing on the rule keywords in the rule reasoning text under the semantic dimension and / or data specification dimension, and can obtain keyword description information for indicating the specific meaning of the rule keywords or the corresponding data usage specifications. Therefore, referring to the keyword description information can also make it possible to accurately and comprehensively understand the rule keywords. In this case, the generated data classification rules can also have a higher accuracy, which is conducive to improving the classification accuracy of the data classification business to which the data classification rules are applied.

[0232] The present application also provides a computer storage medium that stores one or more computer programs corresponding to the above-mentioned data processing method. When one or more processors load and execute the one or more computer programs, the description of the data processing method in the embodiment can be implemented, which is not repeated here. The description of the beneficial effects of adopting the same method is not repeated here. It is understood that the computer program can be deployed and executed on one or more devices that can communicate with each other.

[0233] It should be noted that, according to one aspect of the embodiments of the present application, a program product or computer program is also provided, the program product includes a computer program, the computer program is stored in a computer storage medium. The processor in the data processing device reads the computer program from the computer storage medium and then executes the computer program, thereby enabling the data processing device to perform the above Figure 3 as well as Figure 4 The data processing method shown in the embodiment is provided in various optional aspects.

[0234] Those skilled in the art will appreciate that all or part of the processes in the above-described embodiments can be implemented by instructing related hardware through a computer program. The computer program can be stored in a computer storage medium. When executed, the computer program can include the processes in the above-described embodiments of the data processing method. The computer storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0235] In addition, it can be understood that what is disclosed above is only a partial embodiment of the present application, and it is certainly not intended to limit the scope of rights of the present application. A person skilled in the art can understand that the implementation of all or part of the processes of the above embodiments and the equivalent changes made in accordance with the claims of the present application are still within the scope of the invention.

[0236] Finally, it should be emphasized that when applying the various embodiments in this application to specific products or technologies, data collection can only be carried out after obtaining the permission or consent of the relevant users, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

Claims

1. A data processing method, characterized in that: include: receiving a rule inference request and obtaining a rule inference text from the rule inference request, wherein the rule inference text includes rule keywords to be used when generating data classification rules; Inferring a classification target of the data classification rule based on the rule inference text, and selecting a target description dimension that matches the classification target from preset description dimensions, wherein the target description dimension includes at least one of a semantic dimension and a data specification dimension; The target large language model is called to perform keyword description processing on the rule keywords under the target description dimension to obtain keyword description information of the rule keywords, and the keyword description information is used to generate the data classification rule.

2. The method according to claim 1, characterized in that Calling the keyword description information obtained by the target large language model under the semantic dimension to describe the definition of the rule keyword under one or more knowledge fields; The keyword description information obtained by the target large language model under the data specification dimension is called to describe at least one of the usage specification and classification scenario of the rule keyword when used to generate the data classification rule.

3. The method according to claim 1, characterized in that The rule inference request is sent by a target client, and the method further includes: Obtaining the client type of the target client; formatting the keyword description information according to the information feedback format associated with the client type to obtain formatted keyword description information; The formatted keyword description information is fed back to the target client, so that the target client generates the data classification rule based on the keyword description information and the rule keyword.

4. The method according to claim 3, characterized in that The client type includes a web page type; the keyword description information is formatted according to the information feedback format associated with the client type to obtain formatted keyword description information, including: By performing knowledge domain identification on the keyword description information, a knowledge information segment in the keyword description information is determined; wherein the knowledge information segment refers to an information segment whose knowledge relevance to a preset knowledge domain is greater than or equal to a relevance threshold; The knowledge information segments in the keyword description information are highlighted and arranged to obtain the arranged keyword description information, so as to highlight and display the knowledge information segments during the display of the arranged keyword description information.

5. The method according to claim 3, characterized in that The client type includes a program development type; the keyword description information is formatted according to the information feedback format associated with the client type to obtain the formatted keyword description information, including: Obtaining a program interface used by the target client to send the rule inference request; The keyword description information is format-converted according to the data format indicated by the interface protocol of the program interface to obtain the formatted keyword description information.

6. The method according to claim 1, characterized in that The method further comprises: Calculating the knowledge matching degree between the keyword description information and each knowledge field using a knowledge matching algorithm, and obtaining the commonly used data format associated with the knowledge field with the highest knowledge matching degree; The keyword description information is format-converted according to the commonly used data format, and the keyword description information after format conversion is used to generate the data classification rule.

7. The method according to claim 1 or 2, characterized in that The keyword description information includes description information obtained under the data specification dimension; the method further includes: Selecting a target rule template that matches the classification target from one or more preset rule templates, wherein the target rule template includes one or more pieces of information to be configured; Parameterize the rule keywords according to the description information under the data specification dimension to obtain the configuration parameters required for the information to be configured; Fill the configuration parameters into the target rule template to obtain the data classification rule.

8. The method according to claim 1, characterized in that The classification target of the data classification rule based on the rule reasoning text includes: Performing knowledge question-answering analysis on the rule-based reasoning text to obtain one or more pieces of knowledge question information; The question intention of each knowledge question information is identified, and the classification target of the data classification rule is predicted based on the identified question intentions.

9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: Collect training data from one or more question-answering services; Using the training data, and performing model optimization processing on the large language model according to preset optimization execution parameters to obtain an optimized large language model; Performing a performance evaluation on the optimized large language model to obtain a performance evaluation result; When the performance evaluation result indicates that the model performance of the optimized large language model meets the preset performance requirements, the optimized large language model is used as the target large language model.

10. The method according to claim 9, characterized in that The training data is a question-answer pair consisting of question information and reference answer information; the method uses the training data and performs model optimization processing on the large language model according to preset optimization execution parameters to obtain an optimized large language model, including: Obtaining the question intention of the question information in the training data; Selecting a description dimension that matches the question intention from the preset description dimensions; Calling the large language model to answer the question information under the selected description dimension to obtain model answer information; According to the preset optimization execution parameters, with the optimization goal of reducing the information difference between the reference answer information and the model answer information, the large language model is subjected to model optimization processing to obtain the optimized large language model.

11. The method according to claim 1, characterized in that The method further comprises: Detect the number of rule inference requests received; When the number of requests is greater than or equal to the target number, selecting the target number of rule inference requests from the received rule requests; For the target number of rule inference requests, the step of obtaining rule inference texts from the rule inference requests is performed in parallel.

12. A data processing device, characterized in that: include: A receiving module, configured to receive a rule inference request and obtain a rule inference text from the rule inference request, wherein the rule inference text includes rule keywords to be used when generating data classification rules; an inference module, configured to infer a classification target of the data classification rule based on the rule inference text, and select a target description dimension that matches the classification target from preset description dimensions, wherein the description dimension includes at least one of a semantic dimension and a data specification dimension; The model calling module is used to call the target large language model to perform keyword description processing on the rule keywords under the target description dimension to obtain keyword description information of the rule keywords, and the keyword description information is used to generate the data classification rules.

13. A data processing device, characterized in that: include: a processor adapted to execute one or more computer programs; A computer storage medium storing one or more computer programs, wherein the one or more computer programs are suitable for being loaded by the processor and executing the data processing method according to any one of claims 1 to 11.

14. A computer storage medium, characterized in that The computer storage medium stores one or more computer programs, and the one or more computer programs are suitable for being loaded by a processor and executing the data processing method according to any one of claims 1 to 11.

15. A program product, characterized in that The program product comprises a computer program, which is suitable for being loaded by a processor and executing the data processing method according to claims 1-11.