Intelligent review method and system applied to data processing

By building a code library and a problem library for target review rules, training code review and running monitoring models, the algorithm code compliance and security problems in the data processing process are solved, real-time compliance monitoring and data security guarantees are achieved.

CN120295889APending Publication Date: 2025-07-11SHENZHEN SHANGSHU COM TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510343124.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the data processing process of prior art, the compliance audit of algorithm codes relies on static code inspection and manual review, which has high false alarm rate, low efficiency, and difficult to ensure data security and difficult to prevent data leakage.

Method used

By building a target review rule code base and problem library, generating use case data, training the target code review model and running monitoring model, monitoring and reviewing algorithm behavior in real time, and generating review reports to ensure compliance and security.

Benefits of technology

Real-time compliance monitoring and security of algorithm codes during data processing is realized, preventing data leakage, and improving audit efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295889A_ABST
    Figure CN120295889A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent review method and system applied to data processing. The method comprises the steps of obtaining historical review data; constructing a target review rule code library and a target review problem library according to the historical review data; obtaining a preset first model; training the first model based on the target review rule code library and the target review problem library to obtain a target code review model and a target operation monitoring model; reviewing a preset data processing algorithm through the target code reviewing model and the target operation monitoring model to obtain a target reviewing result; generating a target review report according to the target review result; and according to the target review report, updating contents in the target review rule code library and / or the target review problem library, and performing update training on the target code review model and the target operation monitoring model. By adopting the method and the device, the compliance of algorithm codes in the data processing and training process can be ensured, the data security is ensured, and data leakage is prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data review, and particularly to an intelligent review method and system applied to data processing. Background Art

[0002] With the development of the fields of artificial intelligence and big data, more and more artificial intelligence algorithms are applied to these two fields. However, the training of artificial intelligence algorithms often requires a large amount of high-quality private data, which brings new challenges to the data security and privacy protection of users.

[0003] Currently, data processing and training code review often rely on rule-based static code checking and manual review. Rule-based static code checking often either has a high false alarm rate and low efficiency, or results in a large number of missed reports. Moreover, when encountering new vulnerabilities or problems, it is necessary to manually analyze the vulnerable code and design new rules for prevention, which is not only time-consuming and laborious, but also highly dependent on the ability level of analysts, making it difficult to ensure the data security of users during the code usage process. Therefore, how to ensure the compliance of algorithm code during data processing and training, protect data security, and prevent data leakage has become an urgent problem to be solved. Summary of the Invention

[0004] The embodiments of this application provide an intelligent review method and system applied to data processing, which can ensure the compliance of algorithm code during data processing and training, protect data security, and prevent data leakage.

[0005] In a first aspect, the embodiments of this application provide an intelligent review method applied to data processing, which is applied to an electronic device and includes:

[0006] Obtain historical review data of a target object for a data processing algorithm;

[0007] Construct a target review rule code library and a target review question library according to the historical review data;

[0008] Obtain a preset first model;

[0009] Generate target use case data based on the target review rule code library and the target review question library;

[0010] Divide the target use case data into first use case data and second use case data; the first use case data includes problem code data of the data processing algorithm; the second use case data includes action data of the data processing algorithm during data processing;

[0011] Train the first model with the first use case data to obtain a target code review model;

[0012] Train the first model with the second use case data to obtain a target operation monitoring model;

[0013] When the target object runs a preset data processing algorithm on a preset training platform, review the preset data processing algorithm through the target code review model and the target operation monitoring model to obtain a target review result;

[0014] Generate a target review report according to the target review result; the target review report includes a prompt message; the prompt message is used to prompt whether there are any violations in the preset data processing algorithm;

[0015] Update the content in the target review rule code library and / or the target review question library according to the target review report, and update and train the target code review model and the target operation monitoring model through the updated target review rule code library and the target review question library.

[0016] In a second aspect, an embodiment of the present application provides an intelligent review system for data processing, which is applied to an electronic device. The system includes: an acquisition unit, a control unit, and a review unit, where:

[0017] The acquisition unit is configured to acquire historical review data of a target object for a data processing algorithm;

[0018] The control unit is configured to construct a target review rule code library and a target review question library according to the historical review data;

[0019] The acquisition unit is further configured to acquire a preset first model;

[0020] The control unit is further configured to generate target use case data based on the target review rule code library and the target review question library; divide the target use case data into first use case data and second use case data; the first use case data includes problem code data of the data processing algorithm; the second use case data includes action data of the data processing algorithm during the data processing process; train the first model with the first use case data to obtain a target code review model; train the first model with the second use case data to obtain a target operation monitoring model;

[0021] The review unit is configured to, when the target object runs a preset data processing algorithm on a preset training platform, review the preset data processing algorithm through the target code review model and the target operation monitoring model to obtain a target review result; generate a target review report according to the target review result; the target review report includes prompt information; the prompt information is used to prompt whether there are any violations in the preset data processing algorithm; update the content in the target review rule code library and / or the target review question library according to the target review report, and update and train the target code review model and the target operation monitoring model through the updated target review rule code library and the target review question library.

[0022] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, a memory, a communication interface, and one or more programs, wherein the above one or more programs are stored in the above memory and are configured to be executed by the above processor, and the above programs include instructions for performing the steps in the first aspect of the embodiments of the present application.

[0023] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program for electronic data exchange, and the computer program causes the computer to execute some or all of the steps described in the first aspect of the embodiments of the present application.

[0024] In a fifth aspect, an embodiment of the present application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause the computer to execute some or all of the steps described in the first aspect of the embodiments of the present application. The computer program product can be a software installation package.

[0025] Implementing the present application has the following beneficial effects:

[0026] It can be seen that the intelligent review method applied to data processing described in this application obtains the historical review data of the target object for the data processing algorithm; constructs a target review rule code library and a target review question library based on the historical review data; obtains a preset first model; generates target use case data based on the target review rule code library and the target review question library; divides the target use case data into first use case data and second use case data; trains the first model with the first use case data to obtain a target code review model; trains the first model with the second use case data to obtain a target operation monitoring model; when the target object runs a preset data processing algorithm on a preset training platform, reviews the preset data processing algorithm through the target code review model and the target operation monitoring model to obtain a target review result; generates a target review report according to the target review result; thus, by constructing a code library and a question library containing a large number of algorithm compliance review rules and questions, training two models, namely the target code review model and the target operation monitoring model, and monitoring and reviewing the behavior of the algorithm in the data processing process in real time through these two models, once a violation is found, it is immediately processed, thereby ensuring the compliance of the algorithm code in the data processing and training processes, protecting data security, and preventing data leakage. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the following will describe the drawings required to be used in the embodiments of the present application or the background art.

[0028] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0029] Figure 2 is a scenario application diagram of an electronic device provided by an embodiment of the present application;

[0030] Figure 3 is a flowchart of an intelligent review method applied to data processing provided by an embodiment of the present application;

[0031] Figure 4 is a schematic diagram of obtaining historical review data provided by an embodiment of the present application;

[0032] Figure 5 is a training flowchart of a target code review model provided by an embodiment of the present application;

[0033] Figure 6 is a working flowchart of detecting violation behaviors provided by an embodiment of the present application;

[0034] Figure 7 is a block diagram of the functional units of an intelligent review system applied to data processing provided by an embodiment of the present application;

[0035] Figure 8 It is a schematic structural diagram of another electronic device provided by an embodiment of the present application. Detailed implementation manners

[0036] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0037] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0038] It should be understood that the term "and / or" in this article is only an associative relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article indicates that the associated objects before and after are in an "or" relationship. The "multiple" mentioned in the embodiments of the present application refers to two or more.

[0039] Referring to "embodiment" in this article means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0040] The electronic devices described in the embodiments of the present application may include smart phones (such as Android phones, iOS phones, Windows Phone phones, etc.), tablet computers, personal digital assistants, laptop computers, video matrices, monitoring platforms, mobile internet devices (MID) or wearable devices, etc. The above are only examples, not an exhaustive list, including but not limited to the above-mentioned devices. Of course, the above-mentioned electronic device may also be a server, for example, a cloud server.

[0041] The relevant content, concepts, meanings, technical problems, technical solutions, beneficial effects, etc. involved in the embodiments of the present application will be described below.

[0042] First, some professional terms involved in the present application will be explained:

[0043] LLM large model: That is, the Large Language Model, which is an artificial intelligence model based on deep learning, with a huge parameter scale and powerful language understanding and generation capabilities. It can understand the semantic, syntactic, and pragmatic information of natural language and generate natural and fluent text by performing unsupervised or supervised learning on a large amount of text data, and can be applied to various natural language processing tasks such as text generation, question answering systems, and machine translation.

[0044] GRPO framework: Generalized Advantage Policy Optimization framework. It is an algorithm framework for optimizing the policy network, which is used in reinforcement learning to solve the problem of how to learn the optimal policy according to the environmental feedback. The GRPO framework estimates the advantage function to measure the advantage degree of an action relative to the average policy, thereby guiding the policy network to update in the direction of obtaining higher cumulative rewards to improve learning efficiency and stability.

[0045] RL training: That is, Reinforcement Learning training. Reinforcement learning is a machine learning paradigm in which an agent interacts with the environment and learns the optimal behavior policy according to the reward signal feedback from the environment. During the RL training process, the agent continuously tries different actions, observes the state changes of the environment and the obtained rewards, and adjusts its own policy to maximize the long-term cumulative reward. This training method is commonly used in fields such as robot control, games, and autonomous driving, enabling the agent to autonomously learn and make optimal decisions in complex environments.

[0046] RL training checkpoint: It contains the state of the model at a specific time point during the training process, usually including the model's parameters, the state of the optimizer, and other relevant information (such as the number of training steps, loss values, etc.).

[0047] Inference trajectory: It refers to a series of state and action sequences passed from the initial state to the final state when using the RL model for inference.

[0048] Inference Latency: In artificial intelligence and computer systems, inference latency refers to the time interval from the input data to the output result of the model. For intelligent systems such as deep learning models, when performing inference tasks (such as classifying new input data, generating text, etc.), the model needs to perform a series of calculations and processes, including data preprocessing, feature extraction, forward propagation of the model, etc. Inference latency is a measure of the total time required for these operations. It is an important indicator for evaluating model performance and system real-time performance. A lower inference latency means that the model can give results faster, which is suitable for application scenarios with high requirements for response speed, such as real-time speech recognition and real-time decision-making in autonomous driving.

[0049] Model Quantization: It is a model compression technology aimed at converting the parameters and calculations in the model from a higher-precision data type (such as 32-bit floating-point numbers) to a lower-precision data type (such as 8-bit integers), while trying to maintain the performance of the model as much as possible. By model quantization, the model storage space can be reduced, the computational amount and power consumption can be lowered, and the running efficiency of the model on hardware devices can be improved, especially suitable for resource-constrained devices such as mobile devices and embedded systems. Common quantization methods include uniform quantization, non-uniform quantization, symmetric quantization, asymmetric quantization, etc. During the quantization process, some technical means are needed to balance the model accuracy loss and compression effect, such as using quantization-aware training, fine-tuning and other methods to optimize the performance of the quantized model.

[0050] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device may include: a communication module, a model training module, a data review module, etc., which are not limited here. Among them:

[0051] The communication module can be composed of a hardware communication interface (such as a network interface card, a Bluetooth module, etc.) and related communication protocol software, and is responsible for data transmission and interaction between the electronic device and external devices or systems. For example, it obtains historical review data from an external data source (target object). Another example is that the communication module can also receive update information about data processing algorithms, send review results or alarm information, etc.

[0052] The model training module mainly includes computing resources (such as CPU, GPU), storage devices, and deep learning frameworks or machine learning algorithm libraries, etc., and is used to train and optimize the review model (for example, the LLM large model). Specifically, the model training module can obtain a large number of normal and abnormal data samples through the communication module, build a code library and a question library based on these sample data, and use the data in the code library and the question library to train the review model respectively to obtain a target code review model and a target operation monitoring model.

[0053] Further, the model training module can also optimize and update the target code review model and the target operation monitoring model according to new data and feedback, improving the accuracy and adaptability of the models, enabling them to better identify various complex abnormal behaviors, and thus providing a basis and standard for the data review module to judge whether the data processing algorithm is abnormal and determine the abnormal level.

[0054] The data review module is used to comprehensively review and analyze the data processing process. Specifically, when the target object runs a preset data processing algorithm on a preset training platform, the data review module can review the preset data processing algorithm through the target code review model and the target operation monitoring model to obtain a target review result, and generate a target review report based on the target review result.

[0055] Please refer to Figure 2 , Figure 2 which is a scenario application diagram of an electronic device provided by an embodiment of the present application. It can be seen that the user is the source of data. In the intelligent review scenario of data processing, the user can be the service object of the target object, and the target object can obtain the data to be processed from the user. Among them, the target object can be an institution or organization with high confidentiality requirements such as an enterprise, a hospital, or a bank.

[0056] Then, the electronic device can receive data from the target object. Next, the intelligent review method for data processing provided by the embodiment of the present application can be used to process it to determine whether there are abnormal behaviors in the process of the target data processing algorithm adopted by the target object. If there are abnormal behaviors, the electronic device generates an alarm message to prompt the target object to handle the abnormal behaviors, thereby ensuring the compliance of the algorithm code during data processing and training, protecting data security, and preventing data leakage.

[0057] Please refer to Figure 3 , Figure 3 which is a flowchart of an intelligent review method for data processing provided by an embodiment of the present application, applied to an electronic device. The method includes but is not limited to the following steps:

[0058] S301. Obtain historical review data of the target object for the data processing algorithm.

[0059] In the embodiment of the present application, the target object can be an institution or organization with high confidentiality requirements such as an enterprise, a hospital, or a bank.

[0060] In a specific embodiment, please refer to Figure 4 part a in Figure 4It is a schematic diagram for obtaining historical review data provided by an embodiment of the present application. When the target object is an institution or organization, the information management system of the target object can be determined first. Then, the electronic device can establish a communication connection or a physical connection with the information management system, and obtain the authorization information of the target object. The electronic device relies on this authorization information to access the information management system, and extracts the data that has been reviewed and the corresponding review results of the target object from the database of the information management system, thereby obtaining the historical review data.

[0061] It should be noted that the target object can also be an individual. If the target object is an individual, please refer to Figure 4 part b in. The historical review data can be manually uploaded by the target object to the electronic device.

[0062] S302. Construct a target review rule code library and a target review question library according to the historical review data.

[0063] In the embodiment of the present application, the historical review data may include at least one of the following: review process records, review standard documents, review result reports, code error records, etc., which are not limited here.

[0064] In a specific embodiment, the historical review data is carefully analyzed to extract general review rules to form a target review rule code library. For example, in file format review, it is found that all files passing the review follow specific naming specifications, layout formats, etc., and these requirements are refined into rules. The extracted rules are abstracted to make them more widely applicable and avoid being limited to specific cases. Further, problems found in various review processes can also be screened out from the historical review data, such as file content errors, unqualified project indicators, etc. These problems are classified, which can be divided according to the nature of the problems (such as compliance problems, technical problems, etc.), the severity level (minor, medium, severe) or the field to which they belong (finance, technology, operation, etc.) to construct a target review question library.

[0065] S303. Obtain a preset first model.

[0066] In the embodiment of the present application, the first model may include at least one of the following: LLM model, deep learning model, expert system model, etc., which are not limited here.

[0067] In a specific embodiment, the preset first model can be an LLM model, and this LLM model is a basic general LLM large model for compliance alignment training.

[0068] S304. Generate target use case data based on the target review rule code library and the target review question library.

[0069] In the embodiments of the present application, the preset format of the use case data can be obtained first. Then, the rules in the target review rule code library and the problem types and characteristics in the target review problem library can be deeply analyzed to clarify the applicable scope, conditions, and expected results of the rules, and to understand the common manifestation forms, causes, and impacts of the problems. According to the characteristics of the rules and problems, different review scenarios can be divided, for example, the review of the algorithm running process, the review of code quality, etc. Each scenario corresponds to a set of relevant rules and possible problems. Finally, data can be randomly selected from the target review rule code library and the target review problem library to generate use case data in the preset format, that is, the target use case data.

[0070] Optionally, in step S304, the generating the target use case data based on the target review rule code library and the target review problem library may include the following steps:

[0071] S41. Label the data in the target review rule code library and the target review problem library to obtain code-labeled data and problem-labeled data;

[0072] S42. Determine the code hint word set corresponding to the code-labeled data;

[0073] S43. Determine the problem hint word set corresponding to the problem-labeled data;

[0074] S44. Extract hint words from the code hint word set and the problem hint word set to obtain m hint words; m is an integer greater than 1;

[0075] S45. Invoke the first model based on the m hint words to generate the target use case data.

[0076] In the embodiments of the present application, the data in the target review rule code library and the target review problem library can be labeled to obtain code-labeled data and problem-labeled data. Specifically, a special labeling tool, such as BRAT (for text labeling), Label Studio (supporting the labeling of multiple data types), etc., can be used to label the data in the target review rule code library and the target review problem library, or manual labeling can also be performed to obtain code-labeled data and problem-labeled data.

[0077] Next, the code hint word set corresponding to the code annotation data can be determined. Specifically, the annotation types in the code annotation data, such as rule categories, code functions, compliance, etc., can be analyzed first. Different annotation types will correspond to different types of hint words to obtain the initial hint word set. For example, if there is a compliance annotation, information such as the reason for compliance or non-compliance and the relevant regulations or standards involved is extracted. Suppose a certain annotation data is "Compliance - Conforms to Financial Accounting Standards", then "Financial Accounting Standards" can be extracted as a hint word. Then, the generated initial hint word set can be screened to remove duplicate, similar, or redundant hint words. Thus, the code hint word set is obtained. For example, "Verify password strength" and "Test password strength" have similar meanings, and only one can be retained. Then, the problem hint word set corresponding to the problem annotation data can be determined. Specifically, the method for obtaining the problem hint word set can be the same as the method for obtaining the code hint word set.

[0078] Furthermore, hint words can be extracted from the code hint word set and the problem hint word set to obtain m hint words. Specifically, the random sampling method can be used to randomly extract m / 2 hint words from the code hint word set and the problem hint word set respectively. Thus, m hint words are obtained. Or, m hint words can be extracted from the code hint word set and the problem hint word set according to the preset extraction rules. Specifically, according to the business requirements or usage scenarios, an importance weight can be assigned to each hint word. For example, in the security review scenario, hint words related to security rules may be more important, and the importance weight can be determined through manual evaluation or statistical analysis based on historical data. The preset extraction rule can be: preferentially extract hint words with a greater importance weight.

[0079] Finally, based on the m hint words, the first model can be called to generate the target use case data. Specifically, first, according to the characteristics of the generated use case data, appropriate model parameters, such as temperature, maximum generation length, etc., can be set. Then, the preset use case template can be obtained, and the m hint words are input into the first model, and the first model is called to generate use case data in the format of the preset use case template to obtain the target use case data.

[0080] Among them, the temperature parameter of the model controls the randomness of the generated text. The higher the value, the more random the generated text; the lower the value, the more definite the generated text. The maximum generation length limits the maximum number of words in the generated text.

[0081] In this way, by annotating the data in the target review rule code library and the target review question library, the characteristics and attributes of the data can be clarified, making the code-annotated data and question-annotated data more targeted and accurate. In addition, by determining the code prompt word set and the question prompt word set and extracting the prompt words from them, the key information in the data can be further refined and summarized, providing more accurate guidance for generating the target use case data. Thus, the quality of the generated data can be improved to better meet the requirements of the actual application scenario.

[0082] S305. Divide the target use case data into first use case data and second use case data; the first use case data includes the problem code data of the data processing algorithm; the second use case data includes the action data of the data processing algorithm during the data processing process.

[0083] In the embodiments of the present application, the action data may include at least one of the following: reading, writing, downloading, copying, encrypting, decrypting, etc., which are not limited herein.

[0084] In a specific embodiment, since there are multiple use cases in the target use case data, the target use case data can be divided into first use case data and second use case data according to different contents in the use cases. For example, assuming that a certain use case contains the code data of the data processing algorithm, it can be divided into the first use case data. Conversely, assuming that a certain use case contains the action data of the data processing algorithm, it can be divided into the second use case data.

[0085] S306. Train the first model with the first use case data to obtain a target code review model.

[0086] In the embodiments of the present application, the first use case data can be divided into a training set and a test set according to a preset ratio (for example, 7:3). The first model is trained with the training set, and then the first model is tested with the test set to obtain a test result. Determine the number of correct model output results in the test result to obtain a first number. Divide the first number by the total number of test results to obtain a first accuracy rate. When the first accuracy rate is greater than a preset accuracy rate (for example, 90%), the trained first model is used as the target code review model. Conversely, if the first accuracy rate is not greater than the preset accuracy rate, new use case data can be obtained to continue training the first model until the accuracy rate of the first model is greater than the preset accuracy rate, and this model is used as the target code review model.

[0087] Optionally, in step S306, the step of training the first model with the first use case data to obtain a target code review model may include the following steps:

[0088] A1. Obtain the target review requirements of the target object;

[0089] A2. Determine the initial model parameters of the first model according to the target audit requirements;

[0090] A3. Adjust the model parameters of the first model to the initial model parameters;

[0091] A4. Perform cold start pre-training on the first model with the first use case data to obtain a second model;

[0092] A5. Develop a rule-based dual reward system according to the code prompt word set and the question prompt word set;

[0093] A6. Use the GRPO framework as the RL framework, and use the rule-based dual reward system as the optimization direction of RL training to perform RL training on the second model;

[0094] A7. During the RL training process, save the checkpoints of the second model at preset time intervals to obtain a RL training checkpoints; a is a positive integer;

[0095] A8. Perform rejection sampling on the a RL training checkpoints to obtain b inference trajectories; b is a positive integer;

[0096] A9. Fine-tune the parameters of the second model based on the b inference trajectories to obtain the target code review model.

[0097] In the embodiments of the present application, the model parameters may include at least one of the following: learning rate, batch size, number of training epochs, etc., which are not limited herein; the preset time interval may be preset in advance or by default.

[0098] In a specific embodiment, the target audit requirements of the target object may be obtained first. Specifically, the business requirement document of the target object may be obtained. The business requirement document contains the requirements and expectations for the audit. The target audit requirements may be determined according to the business requirement document. For example, the target audit requirements may be: the audit accuracy rate is required to be greater than 95%, or the target object may manually upload the target audit requirements to the electronic device; then, the initial model parameters of the first model may be determined according to the target audit requirements. Specifically, the initial model parameters may be the learning rate. The mapping relationship between the preset audit requirements and the model parameters may be pre-stored, and the initial model parameters corresponding to the target audit requirements may be determined based on the mapping relationship.

[0099] Next, the model parameters of the first model can be adjusted to the initial model parameters. For example, the adjustment can be performed using the model API. Then, the first model can be pre-trained by cold start with the first use case data to obtain the second model. Specifically, the first use case data can be pre-processed first (e.g., cleaning, deduplication, tokenization, etc.). Then, the first use case data is divided into several batches according to the preset batch size and sequentially input into the first model for training. In each batch, the first model performs forward propagation based on the input data, calculates the prediction result, and calculates the loss value through the preset loss function. Then, the backpropagation algorithm is used to calculate the gradient, and the model parameters are updated according to the gradient and the optimizer. This process is repeated until the preset number of training epochs is reached or other stopping conditions are met, such as the loss function no longer decreases or the performance of the validation set no longer improves, etc., to obtain the trained model, that is, the second model.

[0100] Furthermore, a rule-based dual reward system can be formulated according to the code prompt word set and the question prompt word set. Specifically, the rule-based dual reward system can include an extrinsic reward mechanism and an intrinsic reward mechanism. The extrinsic reward mechanism can be designed first. The extrinsic reward should be directly related to the successful completion of the task. If the code generated by the model completely conforms to the rules extracted from the code library and the question library, a positive reward should be given; if the rules are violated, a negative reward should be given. In addition, since different rules have different importance levels, corresponding weights need to be assigned to each rule. For example, the weights of security-related rules (such as avoiding SQL injection) should be relatively high, while the weights of code style-related rules (such as comment specifications) can be relatively low. Next, the intrinsic reward mechanism can be designed. The metrics of the intrinsic reward should be able to reflect the positive behaviors of the model during the learning process. For example, when the model uses a more efficient algorithm or the code structure is more modular during code generation, etc., these can all be used as metrics for the intrinsic reward. Design the corresponding reward function according to the intrinsic reward metrics. When the behavior of the model conforms to these metrics, a positive reward is given; otherwise, a negative reward is given. Further, different weights can also be assigned to the extrinsic reward and the intrinsic reward to balance their contributions to the total reward. The weight assignment can be adjusted according to the characteristics and requirements of the task. Finally, the extrinsic reward and the intrinsic reward are weighted and summed according to the set weights to obtain the total reward. The model will adjust its strategy according to the total reward to maximize the long-term cumulative reward, thereby obtaining a rule-based dual reward system.

[0101] It should be noted that the extrinsic reward is mainly related to the final goal of the task and is used to measure whether the model output conforms to the overall task requirements; the intrinsic reward focuses on the intermediate behaviors of the model during the learning process and encourages the model to explore and learn useful knowledge and skills.

[0102] Then, a large number of new use case data can be obtained. The GRPO framework is adopted as the RL framework, and the rule-based dual reward system is used as the optimization direction for RL training. The second model is trained with RL using the newly obtained large amount of use case data. Further, during the RL training process, the checkpoints of the second model are saved at preset time intervals, obtaining a RL training checkpoints. Specifically, before the RL training starts, a counter is initialized to 0, which is used to record the number of training steps or epochs. In the training loop, for each training step, the counter is incremented by 1. After each training step, it is checked whether the value of the counter has reached the preset time interval. If so, the checkpoint of the current model is saved, and the model data such as the parameters of the current model are saved to the database of the electronic device, obtaining a RL training checkpoints. In this way, when resuming training, these parameters and states can be loaded and training can continue from the last saved position.

[0103] Next, rejection sampling can be performed on the a RL training checkpoints to obtain b inference trajectories. Specifically, for each RL training checkpoint, the parameters and other relevant states of the model can be loaded according to the checkpoint file, and the model is restored to the corresponding training state. According to the environment settings of the task, the environment is initialized to the initial state. Then, the loaded model is used for at least one inference, actions are selected according to the policy of the model, and the environment state is continuously transformed. The state and action of each step are recorded to form at least one inference trajectory. Thus, d inference trajectories corresponding to the a RL training checkpoints are obtained; d is an integer greater than a and greater than b. Next, a preset rejection sampling criterion can be obtained, and the inference trajectories that do not meet the rejection conditions are selected from the d inference trajectories according to this rejection sampling criterion to obtain b inference trajectories. For example, assuming the rejection sampling criterion is: reject if the trajectory length of the inference trajectory is greater than the preset trajectory length, then the inference trajectories with a trajectory length not greater than the preset trajectory length can be selected from the d inference trajectories to obtain b inference trajectories. Finally, the second model can be fine-tuned based on the b inference trajectories to obtain the target code review model.

[0104] In this way, c training samples are obtained by augmenting the b inference trajectories through the preset data augmentation technique, increasing the diversity and quantity of the data. This helps the model learn richer features, improves the generalization ability of the model, and reduces the risk of overfitting. In addition, the first loss function is determined according to the target review requirements, and the loss value, target average loss value, target variance, and the first gradient corresponding to the target average loss value between the c output results and the c true results are calculated. These calculations can quantify the performance of the model and provide a clear direction and basis for the optimization of the model. Through the gradient information, the model can adjust its parameters along the direction that makes the loss function decrease to gradually optimize the model.

[0105] Optionally, in step A9, the step of fine-tuning the parameters of the second model based on the b inference trajectories to obtain the target code review model may include the following steps:

[0106] B1. Augment the b inference trajectories through a preset data augmentation technique to obtain c training samples; c is an integer greater than b;

[0107] B2. Input the c training samples into the second model to obtain c output results;

[0108] B3. Determine a first loss function according to the target review requirement;

[0109] B4. Determine c true results corresponding to the c output results;

[0110] B5. Calculate the loss values between the c output results and the c true results according to the first loss function to obtain c loss values;

[0111] B6. Determine the target average loss value and target variance corresponding to the c loss values;

[0112] B7. Determine the first gradient corresponding to the target average loss value;

[0113] B8. Determine the target optimization factor corresponding to the target variance;

[0114] B9. Adjust the first gradient according to the target optimization factor to obtain a second gradient;

[0115] B10. Optimize the initial model parameters according to the second gradient to obtain the first fine-tuned model parameters;

[0116] B11. Fine-tune the second model according to the first fine-tuned model parameters to obtain a third model;

[0117] B12. Determine whether the third model meets the preset conditions;

[0118] B13. If the third model meets the preset conditions, determine the target code review model according to the third model;

[0119] B14. If the third model does not meet the preset conditions, update the c training samples to obtain the updated c training samples, and perform parameter fine-tuning on the third model through the updated c training samples until the third model meets the preset conditions, then determine the target code review model according to the third model.

[0120] In the embodiments of the present application, both the preset data augmentation technique and the preset conditions can be preset in advance or default.

[0121] In a specific embodiment, b inference trajectories can be augmented through the preset data augmentation technique to obtain c training samples. Specifically, the preset data augmentation technique can be one of the following: noise injection technique, action replacement technique, trajectory truncation and splicing technique, etc., which are not limited herein; then, the c training samples can be input into the second model in sequence, and the second model analyzes them to obtain c output results.

[0122] Then, the first loss function can be determined according to the target audit requirement. Specifically, different audit requirements require different loss functions. An appropriate evaluation index can be selected according to the target audit requirement, and then an appropriate loss function can be selected according to the nature and characteristics of the evaluation index. For example, assuming that the evaluation index is the proportion of the number of non-compliant code lines in the total number of lines in the code, since this evaluation index is a ratio value and belongs to the category of regression problems, the mean square error loss function can be selected as the first loss function; then, the c true results corresponding to the c output results can be determined. Specifically, the true result corresponding to each training sample among the c training samples can be stored in the database of the electronic device in advance, and the above c true results can be directly extracted from the database of the electronic device.

[0123] Then, the loss values between the c output results and the c true results can be calculated according to the first loss function to obtain c loss values; then, the average value and variance corresponding to the c loss values can be calculated to obtain the target average loss value and the target variance; then, the first gradient corresponding to the target average loss value can be calculated using the backpropagation algorithm. The first gradient represents the change rate and change direction of the first loss function under the initial model parameters; then, the target optimization factor corresponding to the target variance can be determined. Specifically, the mapping relationship between the preset variance and the optimization factor can be stored in advance, and the target optimization factor corresponding to the target variance can be determined based on this mapping relationship. The value range of the target optimization factor can be -0.25 to 0.25; further, the first gradient can be adjusted according to the target optimization factor. The specific calculation formula is as follows:

[0124] The second gradient = the first gradient * (1 + the target optimization factor);

[0125] According to the above formula, the second gradient can be calculated; then, the initial model parameters can be optimized based on the second gradient to obtain the first fine-tuned model parameters. Specifically, according to a preset optimization algorithm (e.g., stochastic gradient descent algorithm, Newton's method), the initial model parameters can be optimized using the second gradient to obtain the first fine-tuned model parameters. Finally, the second model can be fine-tuned based on the first fine-tuned model parameters to obtain the third model. Specifically, the first fine-tuned model parameters can be applied to the second model to update the parameters of the second model, thereby obtaining the third model.

[0126] Next, it can be determined whether the third model meets the preset conditions; if the third model meets the preset conditions, the third model can be directly used as the target code review model.

[0127] If the third model does not meet the preset conditions, then the c training samples are updated to obtain the updated c training samples, and the third model is continuously fine-tuned with the updated c training samples until the third model meets the preset conditions, and then the third model is used as the target code review model.

[0128] In this way, c training samples are obtained by augmenting b inference trajectories through a preset data augmentation technique, increasing the diversity and quantity of data. This helps the model learn richer features, improve the generalization ability of the model, and reduce the risk of overfitting. Additionally, the target optimization factor corresponding to the target variance is determined, and the first gradient is adjusted according to it to obtain the second gradient. Considering the variance can further optimize the adjustment of the gradient, enabling the model to not only consider the mean of the loss but also take into account the fluctuations of the loss during the optimization process. This helps the model converge more stably and avoid instability in the optimization process or falling into local optima due to large fluctuations in the loss.

[0129] Optionally, in step B13, the determining the target code review model according to the third model may include the following steps:

[0130] C1. Obtain the first inference latency corresponding to the third model;

[0131] C2. When the first inference latency is less than or equal to the preset inference latency, determine the third model as the target code review model;

[0132] C3. When the first inference latency is greater than the preset inference latency, determine the required review accuracy requirement for the target object according to the target review requirement;

[0133] C4. Select a small model corresponding to the review accuracy requirement from a preset model library to obtain the first small model;

[0134] C5. Use the third model as the teacher model and the first small model as the student model to perform knowledge distillation training on the first small model to obtain a second small model;

[0135] C6. Determine the target code review model according to the second small model.

[0136] In the embodiments of the present application, the preset inference delay and the preset model library can both be preset in advance or by default; the preset model library stores multiple small models.

[0137] In a specific embodiment, multiple samples can be obtained first, these multiple samples are input into the third model to obtain multiple output results, record the input moments of these multiple sample data to obtain multiple input moments, then record the output moments of these multiple output results to obtain multiple output moments, use these multiple output moments to subtract the corresponding input moments in the multiple input moments to obtain multiple inference delays, and take the average value of these multiple inference delays as the first inference delay. Alternatively, only one sample can be obtained, this one sample is input into the third model to obtain an inference delay, and this inference delay is used as the first inference delay; when the first inference delay is less than or equal to the preset inference delay, the third model is determined as the target code review model.

[0138] When the first inference delay is greater than the preset inference delay, the required review accuracy requirement for the target object can be determined according to the target review requirement. Specifically, the target review requirement may contain multiple requirements, and only the requirements related to the review accuracy need to be extracted from the target review requirement to obtain the review accuracy requirement. For example, assuming the target review requirement is: ensure data security is not leaked, the review accuracy rate should be greater than 90%, etc., and the requirement related to the review accuracy among them is "the review accuracy rate should be greater than 90%", then the review accuracy requirement can be determined as: when conducting the review, the proportion of the number of items correctly identified as meeting or not meeting the review standard in the total number of review items needs to exceed 90%.

[0139] Then, a small model corresponding to the review accuracy requirement can be selected from the preset model library to obtain the first small model. Specifically, the review accuracy of each small model in the preset model library can be obtained to obtain multiple review accuracies, and a small model that meets the review accuracy requirement is selected from them to obtain the first small model. If there are multiple small models that meet the review accuracy requirement, the model with the highest review accuracy can be selected as the first small model; further, the third model can be used as the teacher model and the first small model as the student model to perform knowledge distillation training on the first small model to obtain a second small model. Specifically, since knowledge distillation training is a conventional technology, it will not be elaborated here; finally, the target code review model can be determined according to the second small model. Specifically, the second small model can be used as the target code review model.

[0140] In this way, by comparing the first inference latency of the third model with the preset inference latency, it can be ensured that the finally determined target code review model meets specific requirements in terms of inference speed. When the first inference latency is less than or equal to the preset inference latency, the third model is directly determined as the target code review model, ensuring that the model can quickly give review results in actual applications and improving the efficiency of code review. When the first inference latency is greater than the preset inference latency, a small model is selected from the preset model library for knowledge distillation training according to the review accuracy requirements. In this way, without reducing the review accuracy too much, the lightweight structure of the small model can be used to reduce the inference latency and achieve a better balance between review accuracy and inference speed.

[0141] In one embodiment, please refer to Figure 5 , Figure 5 which is a training flow chart of a target code review model provided by an embodiment of the present application. It can be seen that in Figure 5 , the process starts from the "Start" node. First, "obtain the preset first model", which is the starting model of the entire training process. Then, prompt words (Prompts) can be obtained from the target review rule code library, and the first model is called according to these prompt words to "generate samples" (i.e., target use case data). Next, the generated samples can be used to perform "cold start pre-training" on the first model to lay a foundation for subsequent in-depth training. Further, the model after cold start pre-training can enter the "reinforcement learning RL training" stage, and new prompt words can be obtained from the target review rule code library to formulate a rule-based dual reward system. The GRPO framework is used as the RL framework, and the rule-based dual reward system is used as the optimization direction of RL training to perform RL training on the model. In this stage, the model is further optimized through reinforcement learning. After the RL training is completed, it will be judged whether the "model converges". If it converges (i.e., reaches a certain performance index or training stable state), the "converged model" is obtained; if it does not converge, an "intermediate model" is generated, and "rejection sampling and SFT data generation" are performed. Among them, SFT data refers to supervised fine-tuning data, which refers to c training samples in the above embodiment. Then, the steps of "reinforcement learning RL training" are performed again, and the training is iterated one to two rounds until convergence.

[0142] Further, the converged model will perform "knowledge distillation training" with the "first small model". Knowledge distillation is a technology that transfers the knowledge of a large model to a small model, enabling the small model to have similar performance to the large model. After "knowledge distillation training", "model quantization" can be selected. Model quantization reduces the storage and computational costs of the model by reducing the data precision of the model parameters and improves the operation efficiency. Finally, the "target code review model" is obtained and the process ends.

[0143] This process comprehensively applies a variety of machine learning training techniques. Through gradual optimization and transformation, it aims to generate an efficient and accurate target code review model for code review work.

[0144] Optionally, in step C6, determining the target code review model according to the second smallest model may include the following steps:

[0145] D1. Obtain the second inference latency corresponding to the second smallest model;

[0146] D2. When the second inference latency is less than or equal to the preset inference latency, determine the second smallest model as the target code review model;

[0147] D3. When the second inference latency is greater than the preset inference latency, determine the target quantization method according to the audit accuracy requirement;

[0148] D4. Quantize the second smallest model through the target quantization method to obtain the third smallest model;

[0149] D5. Determine the target code review model according to the third smallest model.

[0150] In the embodiment of the present application, the second inference latency corresponding to the second smallest model can be obtained. Specifically, the method for obtaining the second inference latency can be the same as the method for obtaining the first inference latency; when the second inference latency is less than or equal to the preset inference latency, the second smallest model is determined as the target code review model;

[0151] When the second inference latency is greater than the preset inference latency, the target quantization method can be determined according to the audit accuracy requirement. Specifically, the mapping relationship between the preset accuracy requirement and the quantization method can be pre-stored, and the target quantization method corresponding to the audit accuracy requirement is determined based on this mapping relationship. The target quantization method can include one of the following: uniform quantization, non-uniform quantization, int8 quantization method, int4 quantization method, etc., which are not limited herein; then, the second smallest model can be quantized by the target quantization method to obtain the third smallest model; the third smallest model is determined as the target code review model.

[0152] In this way, when the second inference latency is greater than the preset inference latency, it means that the running speed of the second smallest model is not ideal enough and needs further optimization. Determining the target quantization method according to the audit accuracy requirement to quantize the second smallest model to obtain the third smallest model can reduce the storage space and calculation amount of the model through quantization without affecting the audit accuracy, thereby improving the running speed of the model and achieving a better balance between performance and efficiency.

[0153] S307. Train the first model using the second use case data to obtain a target operation monitoring model.

[0154] In the embodiments of the present application, the first model is trained using the second use case data to obtain a target operation monitoring model. Specifically, the training method of the target operation monitoring model can be the same as that of the target code review model, which will not be elaborated here.

[0155] S308. When the target object runs a preset data processing algorithm on a preset training platform, review the preset data processing algorithm using the target code review model and the target operation monitoring model to obtain a target review result.

[0156] In the embodiments of the present application, both the preset training platform and the preset data processing algorithm can be preset or default in advance.

[0157] In a specific embodiment, when the target object runs a preset data processing algorithm on a preset training platform, review the preset data processing algorithm using the target code review model and the target operation monitoring model to obtain a target review result.

[0158] Optionally, in step S308, the step of "review the preset data processing algorithm using the target code review model and the target operation monitoring model to obtain a target review result" may include the following steps:

[0159] S81. Obtain the target code data corresponding to the preset data processing algorithm;

[0160] S82. Review the target code data using the target code review model to obtain a first review result; the first review result includes the existence of a violation problem and the non - existence of a violation problem;

[0161] S83. If the first review result includes the existence of a violation problem, suspend the operation of the preset data processing algorithm, generate a first prompt message, and review the target code data through a manual review service to obtain a manual review result;

[0162] S84. Determine the target review result based on the manual review result and the first review result;

[0163] S85. If the first review result includes the non - existence of a violation problem, during the operation of the preset data processing algorithm, monitor the action data of the preset data processing algorithm in real time using the target operation monitoring model to obtain a second review result;

[0164] S86. Determine the target review result according to the first review result and the second review result, and update the content in the target review rule code library and / or the target review question library according to the target review report. Update and train the target code review model and the target operation monitoring model through the updated target review rule code library and the target review question library.

[0165] In the embodiments of the present application, the preset data processing algorithm can be preset in advance or by default.

[0166] In a specific embodiment, the target code data corresponding to the preset data processing algorithm can be obtained. Specifically, the target code data can be obtained from the information management system of the target object. The target code data is reviewed by the target code review model to obtain a first review result. Specifically, the target code data can be input into the target code review model to obtain the first review result. If the first review result includes a violation problem, the operation of the preset data processing algorithm is suspended, and a first prompt message is generated. The first prompt message can include the violation problem and its corresponding processing solution. In addition, the manual review service is started, and the staff conducts a manual review of the target code data to obtain a manual review result. Then, the target review result can be determined according to the manual review result and the first review result. Specifically, the manual review result and the first review result can be directly fused to obtain the target review result.

[0167] If the first review result includes no violation problem, during the operation of the preset data processing algorithm, the action data of the preset data processing algorithm is monitored in real time through the target operation monitoring model to obtain a second review result. Specifically, the action data of the preset data processing algorithm can be input into the target operation monitoring model to obtain the second review result. Then, the target review result can be determined according to the first review result and the second review result. Specifically, the first review result and the second review result can be directly fused to obtain the target review result.

[0168] For example, please refer to Figure 6 , Figure 6It is a workflow chart for detecting illegal behaviors provided by an embodiment of the present application. It can be seen that after the "start" process, first, "obtain the code of the preset data processing algorithm", which can be to extract the code of the preset data processing algorithm from the database of the target object as the object for subsequent review. Then, "run the target code review model" can be carried out to review the obtained code by using the trained target code review model. Determine whether the "code is compliant". According to the review result of the target code review model, determine whether the code is compliant. If the code is compliant (Y branch), enter the node of "run the code of the preset data processing algorithm"; if it is not compliant (N branch), then "start the manual review service" to introduce manual review, and the staff will further review the code. After manual review, determine again whether the "code is compliant". If it is compliant, the process can be directly ended; if it is not compliant (N branch), a review report will be generated.

[0169] Furthermore, "run the code of the preset data processing algorithm" can be carried out: when the target code review model determines that the code is compliant, start running the code of this algorithm. Then, "run the target operation monitoring model" can be carried out to monitor its action data in real time during the algorithm operation process to determine whether the operation process is normal. It should be noted that whether it is after the target code review model determines compliance or after the manual review determines compliance, the target operation monitoring model can be run.

[0170] Then, it can be determined whether the "operation process is normal": judge according to the monitoring result of the target operation monitoring model. If the operation process is normal (Y branch), continue to run the code until the algorithm ends; if it is not normal (N branch), then "start the manual review service".

[0171] Next, "generate a review report" can be carried out: when the code is still not compliant after manual review or an abnormality occurs during the operation process, generate a review report to summarize the relevant information in the code review and operation monitoring processes. Furthermore, the review report can also be added to the target review rule code library and the target review problem library.

[0172] Among them, the data in the target review rule code library and the target review problem library are used for "model training" of the first model to obtain the target code review model and the target operation monitoring model.

[0173] Finally, "output the review report for filing and end" can be carried out: output the generated review report for filing, indicating the end of the entire detection process.

[0174] Generally speaking, this process comprehensively uses automated model review, manual review, and operation monitoring to detect illegal behaviors of the code of the preset data processing algorithm in all aspects to ensure the compliance and normal operation of the algorithm.

[0175] It should be noted that Figure 5 and Figure 6 in Figure 5 and Figure 6 , Y represents "Yes" and N represents "No".

[0176] In this way, by using the target code review model to preliminarily review the target code data, it is possible to quickly determine whether there are any violations in the code. Since the model has been trained with a large amount of data, it can identify common code violation patterns, such as security vulnerabilities and non-compliance with coding specifications, and filter out the potentially problematic code at an early stage, preventing the problematic code from entering the subsequent running process, thereby effectively ensuring the compliance and security of the code.

[0177] S309. Generate a target review report according to the target review result; the target review report includes prompt information; the prompt information is used to prompt whether there are any violations in the preset data processing algorithm.

[0178] In the embodiment of the present application, a preset review report template can be obtained first. The review report template may include: report title, report overview, prompt information, detailed review result, suggestions and measures, technical appendix, etc. Fill the above review report template according to the target review result to obtain the target review report.

[0179] S310. Update the content in the target review rule code library and / or the target review question library according to the target review report, and update and train the target code review model and the target running monitoring model through the updated target review rule code library and the target review question library.

[0180] In the embodiment of the present application, update the content in the target review rule code library and / or the target review question library according to the target review report. Then, new test case data can also be generated through the updated target review rule code library and the target review question library, and the target code review model and the target running monitoring model can be updated and trained based on these test case data.

[0181] Implementing the present application has the following beneficial effects:

[0182] It can be seen that the intelligent review method applied to data processing described in this application obtains historical review data of a target object for a data processing algorithm; constructs a target review rule code library and a target review question library based on the historical review data; obtains a preset first model; generates target use case data based on the target review rule code library and the target review question library; divides the target use case data into first use case data and second use case data; trains the first model with the first use case data to obtain a target code review model; trains the first model with the second use case data to obtain a target operation monitoring model; when the target object runs a preset data processing algorithm on a preset training platform, reviews the preset data processing algorithm through the target code review model and the target operation monitoring model to obtain a target review result; generates a target review report according to the target review result; thus, by constructing a code library and a question library containing a large number of algorithm compliance review rules and questions, training two models, namely a target code review model and a target operation monitoring model, and using these two models to monitor and review the behavior of the algorithm in the data processing process in real time, once a violation is found, it is immediately processed, thereby ensuring the compliance of the algorithm code in the data processing and training processes, protecting data security, and preventing data leakage.

[0183] Please refer to Figure 7 , Figure 7 FIG. is a functional unit composition block diagram of an intelligent review system 700 applied to data processing provided by an embodiment of this application. The intelligent review system 700 applied to data processing includes: an acquisition unit 701, a control unit 702, and a review unit 703, where:

[0184] The acquisition unit 701 is configured to acquire historical review data of a target object for a data processing algorithm;

[0185] The control unit 702 is configured to construct a target review rule code library and a target review question library according to the historical review data;

[0186] The acquisition unit 701 is further configured to acquire a preset first model;

[0187] The control unit 702 is further configured to generate target use case data based on the target review rule code library and the target review question library; divide the target use case data into first use case data and second use case data; the first use case data includes problem code data of the data processing algorithm; the second use case data includes action data of the data processing algorithm during data processing; train the first model with the first use case data to obtain a target code review model; train the first model with the second use case data to obtain a target operation monitoring model;

[0188] The review unit 703 is configured to, when the target object runs a preset data processing algorithm on a preset training platform, review the preset data processing algorithm through the target code review model and the target operation monitoring model to obtain a target review result; generate a target review report according to the target review result; the target review report includes a prompt message; the prompt message is used to prompt whether there are any violations in the preset data processing algorithm; update the content in the target review rule code library and / or the target review question library according to the target review report, and update and train the target code review model and the target operation monitoring model through the updated target review rule code library and the target review question library.

[0189] Optionally, in terms of generating target use case data based on the target review rule code library and the target review question library, the control unit 702 is specifically configured to:

[0190] Annotate the data in the target review rule code library and the target review question library to obtain code annotation data and question annotation data;

[0191] Determine the code prompt word set corresponding to the code annotation data;

[0192] Determine the question prompt word set corresponding to the question annotation data;

[0193] Extract prompt words from the code prompt word set and the question prompt word set to obtain m prompt words; m is an integer greater than 1;

[0194] Call the first model based on the m prompt words to generate the target use case data.

[0195] Optionally, in terms of training the first model with the first use case data to obtain a target code review model, the control unit 702 is specifically configured to:

[0196] Obtain the target review requirements of the target object;

[0197] Determine the initial model parameters of the first model according to the target review requirements;

[0198] Adjust the model parameters of the first model to the initial model parameters;

[0199] Perform cold start pre-training on the first model through the first use case data to obtain a second model;

[0200] Formulate a rule-based dual reward system according to the code prompt word set and the question prompt word set;

[0201] Use the GRPO framework as the RL framework, take the rule-based dual reward system as the optimization direction of RL training, and perform RL training on the second model;

[0202] During the RL training process, save the checkpoints of the second model at preset time intervals to obtain a RL training checkpoints; a is a positive integer;

[0203] Perform rejection sampling on the a RL training checkpoints to obtain b inference trajectories; b is a positive integer;

[0204] Fine-tune the parameters of the second model based on the b inference trajectories to obtain the target code review model.

[0205] Optionally, in terms of fine-tuning the parameters of the second model based on the b inference trajectories to obtain the target code review model, the control unit 702 is specifically used for:

[0206] Expand the b inference trajectories through a preset data augmentation technique to obtain c training samples; c is an integer greater than b;

[0207] Input the c training samples into the second model to obtain c output results;

[0208] Determine the first loss function according to the target review requirements;

[0209] Determine the c true results corresponding to the c output results;

[0210] Calculate the loss values between the c output results and the c true results according to the first loss function to obtain c loss values;

[0211] Determine the target average loss value and target variance corresponding to the c loss values;

[0212] Determine the first gradient corresponding to the target average loss value;

[0213] Determine the target optimization factor corresponding to the target variance;

[0214] Adjust the first gradient according to the target optimization factor to obtain the second gradient;

[0215] Optimize the initial model parameters according to the second gradient to obtain the first fine-tuned model parameters;

[0216] Fine-tune the second model according to the first fine-tuned model parameters to obtain the third model;

[0217] Determine whether the third model meets the preset conditions;

[0218] If the third model meets the preset conditions, determine the target code review model according to the third model;

[0219] If the third model does not meet the preset conditions, update the c training samples to obtain the updated c training samples, and fine-tune the parameters of the third model with the updated c training samples until the third model meets the preset conditions, then determine the target code review model according to the third model.

[0220] Optionally, in terms of determining the target code review model according to the third model, the control unit 702 is specifically configured to:

[0221] Obtain the first inference latency corresponding to the third model;

[0222] When the first inference latency is less than or equal to the preset inference latency, determine the third model as the target code review model;

[0223] When the first inference latency is greater than the preset inference latency, determine the required review accuracy requirement for the target object according to the target review requirement;

[0224] Select a small model corresponding to the review accuracy requirement from the preset model library to obtain the first small model;

[0225] Use the third model as the teacher model and the first small model as the student model to perform knowledge distillation training on the first small model to obtain the second small model;

[0226] Determine the target code review model according to the second small model.

[0227] Optionally, in terms of determining the target code review model according to the second small model, the control unit 702 is specifically configured to:

[0228] Obtain the second inference latency corresponding to the second small model;

[0229] When the second inference latency is less than or equal to the preset inference latency, determine the second small model as the target code review model;

[0230] When the second inference latency is greater than the preset inference latency, determine the target quantization method according to the review accuracy requirement;

[0231] Quantize the second small model through the target quantization method to obtain the third small model;

[0232] Determine the target code review model according to the third small model.

[0233] Optionally, in terms of reviewing the preset data processing algorithm through the target code review model and the target operation monitoring model to obtain a target review result, the review unit 703 is specifically configured to:

[0234] Obtain target code data corresponding to the preset data processing algorithm;

[0235] Review the target code data through the target code review model to obtain a first review result; the first review result includes the existence of a violation problem and the non-existence of a violation problem;

[0236] If the first review result includes the existence of a violation problem, suspend the operation of the preset data processing algorithm, generate a first prompt message, and, through manual review service, review the target code data to obtain a manual review result;

[0237] Determine the target review result according to the manual review result and the first review result;

[0238] If the first review result includes the non-existence of a violation problem, during the operation of the preset data processing algorithm, the target operation monitoring model is used to monitor the action data of the preset data processing algorithm in real time to obtain a second review result;

[0239] Determine the target review result according to the first review result and the second review result.

[0240] In specific implementation, the intelligent review system 700 applied to data processing described in the embodiments of the present invention may also execute other implementation manners described in the intelligent review method applied to data processing provided in the embodiments of the present invention, which will not be elaborated here.

[0241] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of another electronic device provided in the embodiments of the present application. The electronic device may include a processor, a memory, a communication interface, and one or more programs. The processor, the memory, and the communication interface may be connected to each other through a bus; the above one or more programs are stored in the above memory and are configured to be executed by the above processor; in the embodiments of the present application, the above programs include parts or all of the steps that enable the electronic device to execute any method described in the method embodiments above.

[0242] The embodiments of the present application also provide a computer-readable storage medium, where the computer-readable storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute parts or all of the steps of any method described in the method embodiments above. The above computer includes an electronic device.

[0243] The embodiments of the present application also provide a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to execute some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer includes an electronic device.

[0244] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0245] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0246] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0247] Those of ordinary skill in the art can understand the entire or part of the process of implementing the above method embodiments. This process can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: ROM or random access memory RAM, magnetic disk, or optical disk and other media that can store program codes.

[0248] The steps of the methods or algorithms described in the embodiments of this application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a RAM, flash memory, ROM, EPROM, electrically erasable programmable read-only memory (EEPROM), register, hard disk, removable hard disk, compact disc read-only memory (CD-ROM), or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. Additionally, the ASIC can be located in a terminal device or a management device. Of course, the processor and the storage medium can also exist as discrete components in the terminal device or the management device.

[0249] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of this application can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part.

[0250] The above computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server, data center, etc. that contains one or more integrated available media.

[0251] Among them, the available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0252] Each device and product described in the above embodiments includes various modules / units, which can be software modules / units, hardware modules / units, or can be partially software modules / units and partially hardware modules / units. For example, for each device and product applied to or integrated into a chip, each module / unit it includes can be implemented in the form of hardware such as circuits. Or, at least some of the modules / units can be implemented in the form of software programs that run on the processor integrated inside the chip, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits; for each device and product applied to or integrated into a chip module, each module / unit it includes can be implemented in the form of hardware such as circuits. Different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components of the chip module. Or, at least some of the modules / units can be implemented in the form of software programs that run on the processor integrated inside the chip module, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits; for each device and product applied to or integrated into a terminal device, each module / unit it includes can be implemented in the form of hardware such as circuits. Different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components inside the terminal device. Or, at least some of the modules / units can be implemented in the form of software programs that run on the processor integrated inside the terminal device, and the remaining (if any) part of the modules / units can be implemented in the form of hardware such as circuits.

[0253] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the embodiments of the present application. It should be understood that the above is only the specific embodiments of the embodiments of the present application and is not used to limit the protection scope of the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.

Claims

1. An intelligent review method applied to data processing, characterized in that, Applied to an electronic device, including: Obtain historical review data of a target object for a data processing algorithm; Construct a target review rule code library and a target review question library according to the historical review data; Obtain a preset first model; Generate target use case data based on the target review rule code library and the target review question library; Divide the target use case data into first use case data and second use case data; the first use case data includes problem code data of the data processing algorithm; the second use case data includes action data of the data processing algorithm during the data processing process; Train the first model with the first use case data to obtain a target code review model; Train the first model with the second use case data to obtain a target operation monitoring model; When the target object runs a preset data processing algorithm on a preset training platform, review the preset data processing algorithm through the target code review model and the target operation monitoring model to obtain a target review result; Generate a target review report according to the target review result; the target review report includes prompt information; the prompt information is used to prompt whether there are any violations in the preset data processing algorithm; Update the content in the target review rule code library and / or the target review question library according to the target review report, and update and train the target code review model and the target operation monitoring model through the updated target review rule code library and the target review question library.

2. The method according to claim 1, wherein The generating of the target use case data based on the target review rule code library and the target review question library includes: Annotate the data in the target review rule code library and the target review question library to obtain code annotation data and question annotation data; Determine a code prompt word set corresponding to the code annotation data; Determine a question prompt word set corresponding to the question annotation data; Extract prompt words from the code prompt word set and the question prompt word set to obtain m prompt words; m is an integer greater than 1; Call the first model based on the m prompt words to generate the target use case data.

3. The method according to claim 2, characterized in that, The training of the first model with the first use case data to obtain a target code review model includes: Obtain the target review requirements of the target object; Determine the initial model parameters of the first model according to the target review requirements; Adjust the model parameters of the first model to the initial model parameters; Perform cold start pre-training on the first model with the first use case data to obtain a second model; Formulate a rule-based dual reward system according to the code prompt word set and the question prompt word set; Use the GRPO framework as the RL framework, and use the rule-based dual reward system as the optimization direction of RL training to perform RL training on the second model; During the RL training process, save checkpoints of the second model at preset time intervals to obtain a RL training checkpoints; a is a positive integer; Perform rejection sampling on the a RL training checkpoints to obtain b inference trajectories; b is a positive integer; Fine-tune the parameters of the second model based on the b inference trajectories to obtain the target code review model.

4. The method according to claim 3, wherein The step of fine-tuning the parameters of the second model based on the b inference trajectories to obtain the target code review model includes: Augment the b inference trajectories through a preset data augmentation technique to obtain c training samples; c is an integer greater than b; Input the c training samples into the second model to obtain c output results; Determine the first loss function according to the target review requirement; Determine the c true results corresponding to the c output results; Calculate the loss values between the c output results and the c true results according to the first loss function to obtain c loss values; Determine the target average loss value and the target variance corresponding to the c loss values; Determine the first gradient corresponding to the target average loss value; Determine the target optimization factor corresponding to the target variance; Adjust the first gradient according to the target optimization factor to obtain a second gradient; Optimize the initial model parameters according to the second gradient to obtain the first fine-tuned model parameters; Fine-tune the second model according to the first fine-tuned model parameters to obtain a third model; Determine whether the third model meets the preset conditions; If the third model meets the preset conditions, determine the target code review model according to the third model; If the third model does not meet the preset conditions, update the c training samples to obtain the updated c training samples, and fine-tune the parameters of the third model through the updated c training samples until the third model meets the preset conditions, then determine the target code review model according to the third model.

5. The method according to claim 4, wherein The step of determining the target code review model according to the third model includes: Obtain the first inference latency corresponding to the third model; When the first inference latency is less than or equal to the preset inference latency, determine the third model as the target code review model; When the first inference latency is greater than the preset inference latency, determine the required review accuracy requirement for the target object according to the target review requirement; Select a small model corresponding to the review accuracy requirement from a preset model library to obtain a first small model; Use the third model as the teacher model and the first small model as the student model to perform knowledge distillation training on the first small model to obtain a second small model; Determine the target code review model according to the second small model.

6. The method according to claim 5, characterized in that The step of determining the target code review model according to the second small model includes: Obtain the second inference latency corresponding to the second small model; When the second inference latency is less than or equal to the preset inference latency, determine the second small model as the target code review model; When the second inference latency is greater than the preset inference latency, determine the target quantization method according to the review accuracy requirement; Quantize the second small model through the target quantization method to obtain a third small model; Determine the target code review model according to the third small model.

7. The method according to any one of claims 1 to 6, characterized in that Reviewing the preset data processing algorithm through the target code review model and the target operation monitoring model to obtain a target review result, including: Obtaining target code data corresponding to the preset data processing algorithm; Reviewing the target code data through the target code review model to obtain a first review result; the first review result includes the existence of a violation problem and the non-existence of a violation problem; If the first review result includes the existence of a violation problem, suspending the operation of the preset data processing algorithm, generating a first prompt message, and reviewing the target code data through a manual review service to obtain a manual review result; Determining the target review result according to the manual review result and the first review result; If the first review result includes the non-existence of a violation problem, during the operation of the preset data processing algorithm, the action data of the preset data processing algorithm is monitored in real time through the target operation monitoring model to obtain a second review result; Determining the target review result according to the first review result and the second review result.

8. An intelligent review system applied to data processing, characterized in that, Applied to an electronic device, the system includes: an acquisition unit, a control unit, and a review unit, where: The acquisition unit is configured to acquire historical review data of a target object for a data processing algorithm; The control unit is configured to construct a target review rule code library and a target review problem library according to the historical review data; The acquisition unit is further configured to acquire a preset first model; The control unit is further configured to generate target use case data based on the target review rule code library and the target review problem library; divide the target use case data into first use case data and second use case data; the first use case data includes problem code data of the data processing algorithm; the second use case data includes action data of the data processing algorithm during the data processing process; training the first model through the first use case data to obtain a target code review model; training the first model through the second use case data to obtain a target operation monitoring model; The review unit is configured to, when the target object runs a preset data processing algorithm on a preset training platform, review the preset data processing algorithm through the target code review model and the target operation monitoring model to obtain a target review result; generating a target review report according to the target review result; the target review report includes a prompt message; the prompt message is used to prompt whether there is a violation problem in the preset data processing algorithm; updating the content in the target review rule code library and / or the target review problem library according to the target review report, and updating and training the target code review model and the target operation monitoring model through the updated target review rule code library and the target review problem library.

9. An electronic device, characterized in that, Including: A processor, a memory, a communication interface, and one or more programs; The one or more programs are stored in the memory and configured to be executed by the processor, the programs including instructions for performing the steps in the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program for electronic data interchange is stored, wherein the computer program causes a computer to execute the method according to any one of claims 1-7.

Citation Information

Cited By

  • Large language model anti-seismic report examination method and device based on domain knowledge enhancement

    CN120951987A

  • Academic paper automatic review method and device based on reinforcement learning

    CN121213010A

  • Detection and review method and device based on multi-data-source strategy

    CN121524169A