Data processing method and related device

CN122816628APending Publication Date: 2026-09-25TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610868913.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

这种数据存储方式虽然节省了存储空间并便于结构化管理,但导致了业务语境信息的丢失,会导致模型微调方向的偏移,进而影响模型训练效果

Benefits of technology

[0035]之后,响应于针对任务执行控件的触发操作,启动对预设模型的训练微调任务。由于本方案对脱水数据进行复水重构后再执行训练微调任务,因而可以使预设模型的学习目标与标注规则信息保持一致,从而减少因语境缺失造成的对齐偏差。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816628A_ABST
    Figure CN122816628A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a data processing method and related equipment; the embodiments of the present application display a task configuration page corresponding to a preset model, the task configuration page includes a task execution control and a page address of a labeling page corresponding to training data of the preset model, and the task configuration page is used for interacting with an agent; then, the agent accesses the labeling page based on the page address to display labeling rule information recognized by the agent in the labeling page on the task configuration page, and the labeling rule information is used for updating the training data; finally, in response to a trigger operation on the task execution control, a task execution page is displayed, and the task execution page includes a result of training the preset model by the updated training data; the scheme can make the learning goal of the preset model consistent with the labeling rule information, thereby reducing alignment deviation caused by context loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a data processing method and related equipment, including data processing devices, electronic devices, computer program products, and computer-readable storage media. Background Technology

[0002] With the development of artificial intelligence technology, multimodal large language models have been widely used in content moderation, intelligent customer service, and image and text understanding. To adapt general pre-trained models to specific business scenarios, supervised fine-tuning (SFT) has become a commonly adopted model customization method in the industry. In recent years, automated fine-tuning technology has gradually emerged, aiming to reduce the cost of manual intervention and improve model iteration efficiency. However, in practical industrial applications, automated fine-tuning systems face technical challenges in maintaining data semantic integrity and business context.

[0003] In typical application scenarios such as content moderation, intelligent customer service, and image / text understanding, human annotators usually complete data annotation work in a complete graphical user interface (GUI) environment. This interface environment contains rich business context information, such as prominently displayed review guidelines, warning text marked with specific colors or fonts, and image / text combinations with specific spatial layout relationships. This visual context information plays an important role in annotators' understanding of business rules and making accurate judgments. However, when the annotation results are stored in a data warehouse, the above-mentioned context information is usually stripped away, leaving only the original content to be reviewed and the corresponding annotation tags. While this data storage method saves storage space and facilitates structured management, it leads to the loss of business context information, which can cause a shift in the direction of model fine-tuning and thus affect the model training effect. Summary of the Invention

[0004] This application provides a data processing method and related equipment that can rehydrate and reconstruct dehydrated data before performing training fine-tuning tasks. This can ensure that the learning objectives of the preset model are consistent with the annotation rule information, thereby reducing alignment deviations caused by missing context.

[0005] In a first aspect, embodiments of this application provide a data processing method, including: Displays the task configuration page corresponding to the preset model. The task configuration page includes task execution controls and the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the agent. The agent is invoked to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent in the annotation page on the task configuration page. The annotation rule information is used to update the training data. In response to a trigger operation on the task execution control, a task execution page is displayed, the task execution page including the results of training the preset model with updated training data.

[0006] Accordingly, embodiments of this application provide a data processing apparatus, including: The task configuration unit is used to display the task configuration page corresponding to the preset model. The task configuration page includes task execution controls and the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the intelligent agent. The agent invocation unit is used to invoke the agent to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent in the annotation page on the task configuration page, and the annotation rule information is used to update the training data; A task execution unit is used to display a task execution page in response to a trigger operation on the task execution control. The task execution page includes the results of training the preset model with updated training data.

[0007] In some embodiments, the task configuration unit can be specifically used to display a training task creation page, the training task creation page including a task creation control and a configuration information input control; in response to an input operation on the configuration information input control, the task configuration information input through the configuration information input control is displayed on the training task creation page, the task configuration information including the page address of the annotation page corresponding to the training data of the preset model; in response to a trigger operation on the task creation control, the task configuration page corresponding to the task configuration information is displayed.

[0008] In some embodiments, the configuration information input control includes a model selection sub-control and an address input sub-control. Specifically, the task configuration unit can be used to, in response to a trigger operation on the model selection sub-control, select the model corresponding to the trigger operation from a preset model set as the preset model; and in response to an input operation on the address input sub-control, obtain the web address corresponding to the input operation and use the web address as the page address of the annotation page for the training data of the preset model.

[0009] In some embodiments, the task configuration page further includes an agent invocation control. Specifically, the agent invocation unit can be used to invoke the agent to access the annotation page based on the page address in response to a trigger operation on the invocation control, so as to identify the annotation rule information corresponding to the training data in the annotation page through the agent; and display the annotation rule information on the task configuration page.

[0010] In some embodiments, the agent invocation unit can be specifically used to invoke the agent to access the annotation page based on the page address; the agent identifies at least one text content in the annotation page and extracts features from the content hierarchy relationship between the at least one text content to obtain text structure features; the agent identifies at least one image content in the annotation page and extracts features from the at least one image content to obtain page visual features; the text structure features and the page visual features are fused to obtain annotation rule information corresponding to the training data.

[0011] In some embodiments, the agent invocation unit may be specifically used to identify the text region corresponding to the text content in the annotation page and determine the area of ​​the text region corresponding to the text region; identify the image region corresponding to the image content in the annotation page and determine the area of ​​the image region corresponding to the image region; determine the feature weights corresponding to the feature fusion based on the area of ​​the text region and the area of ​​the image region; and perform feature fusion of the text structural features and the page visual features based on the feature weights to obtain the annotation rule information corresponding to the training data.

[0012] In some embodiments, the intelligent agent invocation unit may be specifically used to determine the total area of ​​the labeled page; calculate the ratio between the text region area and the total page area to obtain the text region area ratio; calculate the ratio between the image region area and the total page area to obtain the image region area ratio; and determine the feature weights corresponding to the feature fusion based on the text region area ratio and the image region area ratio.

[0013] In some embodiments, the agent invocation unit may be specifically used to add the text region area ratio and the image region area ratio to obtain a target area ratio; calculate the ratio between the text region area ratio and the target area ratio to obtain a relative text region area ratio; determine the page type of the labeled page based on the relative text region area ratio, and determine the feature weights corresponding to the feature fusion based on the page type.

[0014] In some embodiments, the preset model includes a text encoding network and a visual encoding network. Specifically, the task execution unit can be used to determine the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training based on the relative ratio of the text region areas; and display the text learning rate and the visual learning rate on the task configuration page.

[0015] In some embodiments, the task execution unit may specifically be used to calculate the text color contrast of the text region and the image color contrast of the image region in the annotation page, and to calculate the text level depth of the text node and the image level depth of the image node in the annotation page; to identify the page center coordinates in the annotation page, and to calculate the text position centrality of the text element of the text region from the page center coordinates, and the image position centrality of the image element of the image region from the page center coordinates; to determine the text modality saliency score corresponding to the text region based on the text region area, the text color contrast, the text level depth, and the text position centrality, and to determine the visual modality saliency score corresponding to the image region based on the image region area, the image color contrast, the image level depth, and the image position centrality; and to determine the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during the training process based on the text modality saliency score and the visual modality saliency score.

[0016] In some embodiments, the task execution unit may be specifically used to add the text modality saliency score and the visual modality saliency score to obtain a target saliency score; calculate the ratio between the text modality saliency score and the target saliency score to obtain a text modality relative ratio; and determine the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during the training process based on the text modality relative ratio.

[0017] In some embodiments, the task execution unit may be specifically used to freeze the network parameters of the text encoding network during the training process when the text modality relative ratio is less than a first text modality relative ratio threshold; and to freeze the network parameters of the visual encoding network during the training process when the text modality relative ratio is greater than a second text modality relative ratio threshold.

[0018] In some embodiments, the task execution unit may be specifically used to obtain network capacity parameters of a preset model; determine the page type of the labeled page based on the text modality ratio; and allocate the network capacity parameters based on the page type to obtain the text network capacity parameters corresponding to the text encoding network and the visual network capacity parameters corresponding to the visual encoding network.

[0019] In some embodiments, the task configuration page further includes an update method selection control. Specifically, the task execution unit can be configured to, in response to a selection operation on the update method selection control, filter out the target update method corresponding to the selection operation from a preset update method set; in response to a trigger operation on the task execution control, execute the training task corresponding to the preset model to display the task execution page. The task execution page includes execution process information of the training task, which includes updating the training data using the target update method and the annotation rule information, and training the preset model using the updated training data, the text learning rate, and the visual learning rate; when the completion of the training task is detected, the task execution result of the training task is displayed on the task execution page, and the task execution result indicates the result of training the preset model.

[0020] In some embodiments, the task execution unit may be specifically used to obtain at least one annotation rule information obtained by the agent accessing the annotation page at different times, and determine the rule information generation time of the annotation rule information and the training data generation time of the training data; calculate the time difference between the rule information generation time and the training data generation time, and filter out the target annotation rule information corresponding to the training data from the annotation rule information according to the time difference; update the training data through the target update method and the target annotation rule information to obtain the updated training data.

[0021] In some embodiments, the training data includes sample data of at least one training sample. Specifically, the task execution unit can be used to convert the target annotation rule information into annotation rule prompts when the target update method is prompt word injection; and add the annotation rule prompts to the sample data to obtain the updated training data.

[0022] In some embodiments, the task execution unit may be specifically used to identify target text structural features and target page visual features in the target annotation rule information; generate text annotation rule prompts based on the target text structural features; generate page visual rule prompts based on the target page visual features; and fuse the text annotation rule prompts and the page visual rule prompts to obtain the annotation rule prompts.

[0023] In some embodiments, the training data includes sample data of at least one training sample. Specifically, the task execution unit can be used to extract features from the target annotation rule information to obtain annotation rule features when the target update method is feature concatenation, and to extract features from the sample data to obtain sample features; and to concatenate the annotation rule features and the sample features to obtain the updated training data.

[0024] In some embodiments, the task execution unit may be specifically used to perform a dot product calculation on the sample features and the annotation rule features to obtain the attention weight of the sample features for the annotation rule features; perform a weighted calculation on the annotation rule features according to the attention weight to obtain the weighted annotation rule features; and concatenate the weighted annotation rule features with the sample features to obtain the updated training data.

[0025] In some embodiments, the task execution unit may be specifically used to extract features from the training data to obtain the original feature distribution data of the training data; extract features from the updated training data to obtain the updated feature distribution data of the updated training data; calculate the feature distribution difference data between the original feature distribution data and the updated feature distribution data; adjust the updated training data according to the feature distribution difference data to obtain the adjusted training data, and use the adjusted training data as the updated training data.

[0026] Secondly, embodiments of this application also provide a data processing method, including: Obtain the page address of the annotation page corresponding to the training data of the preset model; Based on the page address, access the annotation page to identify annotation rule information in the annotation page; The training data is updated according to the annotation rule information to obtain the updated training data; Based on the updated training data, the preset model is trained to obtain the trained model.

[0027] Accordingly, embodiments of this application provide a data processing apparatus, including: The page address acquisition unit is used to obtain the page address of the labeled page corresponding to the training data of the preset model; A page address access unit is used to access the annotation page based on the page address, so as to identify annotation rule information in the annotation page; The training data update unit is used to update the training data according to the annotation rule information to obtain the updated training data. The model training unit is used to train the preset model based on the updated training data to obtain the trained model.

[0028] Furthermore, embodiments of this application also provide an electronic device, including a processor and a memory, wherein the memory stores an application program, and the processor is used to run the application program in the memory to execute the data processing method provided in embodiments of this application.

[0029] Furthermore, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the data processing method provided in embodiments of this application.

[0030] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the data processing methods provided in embodiments of this application.

[0031] This embodiment of the application displays a task configuration page corresponding to a preset model. The task configuration page includes a task execution control and the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the agent. Then, the agent is invoked to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent in the annotation page on the task configuration page. The annotation rule information is used to update the training data. Finally, in response to the trigger operation of the task execution control, the task execution page is displayed. The task execution page includes the result of training the preset model with the updated training data.

[0032] This solution adds a front-end context collection step before initiating the training task for the preset model. The task configuration page displays the page address of the annotation page corresponding to the task execution control and the training data of the preset model. Specifically, the task execution control triggers the training task, the preset model is the object to be trained, the training data is the dehydrated data of the preset model, and the annotation page is the front-end graphical user interface displayed on the terminal when human annotators annotate the data. The annotation page contains rich business context information, and the page address is a link to the annotation page.

[0033] Building upon this, an intelligent agent can autonomously access the annotation page based on the page address, allowing the agent to identify annotation rule information within the annotation page. This annotation rule information is the structured business rule information extracted by the agent through parsing the annotation page; it can also be referred to as the business prior context. The annotation rule information reflects the explicit rules and implicit visual cues used by human annotators or operators when annotating the original content, encompassing the complete business context upon which human annotators or operators base their annotations on the training data. The task configuration page can display the annotation rule information identified by the agent.

[0034] Furthermore, this solution can update the training data using the annotation rule information obtained from agent recognition. This update process involves fusing and reconstructing the annotation rule information with the training data, essentially a context rehydration process. This update process ensures that the updated training data carries the same rule guidance as when human annotators or operators make decisions.

[0035] Subsequently, in response to the trigger operation of the task execution control, the training and fine-tuning task of the preset model is initiated. Because this solution rehydrates and reconstructs the dehydrated data before executing the training and fine-tuning task, the learning objective of the preset model can be kept consistent with the annotation rule information, thereby reducing alignment deviations caused by missing context. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1A This is a schematic diagram of an application scenario of the data processing method provided in the embodiments of this application; Figure 1B This is a schematic diagram illustrating another application scenario of the data processing method provided in the embodiments of this application; Figure 2 This is a flowchart illustrating a data processing method provided in an embodiment of this application; Figure 3 This is a schematic diagram of a page change process of the data processing method provided in the embodiments of this application; Figure 4A This is a schematic diagram of a labeled page provided in an embodiment of this application; Figure 4B This is a schematic diagram of another page of the labeled page provided in the embodiments of this application; Figure 5This is a schematic diagram of the training task creation page provided in an embodiment of this application; Figure 6 This is a schematic diagram illustrating the page change process from the training task creation page to the task configuration page provided in an embodiment of this application; Figure 7 This is a schematic diagram of the invocation process of the intelligent agent provided in the embodiments of this application; Figure 8 This is a schematic diagram illustrating the process of determining annotation rule information provided in an embodiment of this application; Figure 9 This is a schematic diagram of a task configuration page provided in an embodiment of this application; Figure 10 This is another schematic diagram of the task configuration page provided in the embodiments of this application; Figure 11 This is a schematic diagram illustrating the page change process from the task configuration page to the task execution page provided in an embodiment of this application; Figure 12 This is a schematic diagram of the user operation flow corresponding to the data processing method provided in the embodiments of this application; Figure 13 This is a schematic diagram of the system architecture corresponding to the data processing method provided in the embodiments of this application; Figure 14 This is a schematic diagram of the execution logic corresponding to the data processing method provided in the embodiments of this application; Figure 15 This is a schematic diagram illustrating another application scenario of the data processing method provided in the embodiments of this application; Figure 16 This is another schematic flowchart of the data processing method provided in the embodiments of this application; Figure 17 This is a schematic diagram of the structure of the data processing apparatus provided in an embodiment of this application; Figure 18 This is another structural schematic diagram of the data processing apparatus provided in the embodiments of this application; Figure 19 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0038] The technical solutions described below, with reference to the accompanying drawings, will be clearly and completely described. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0039] The following explanations of some terms used in the embodiments of this application are provided to facilitate understanding by those skilled in the art.

[0040] A multimodal graphical user interface agent (GUI agent) is an artificial intelligence agent program with visual perception and interface interaction capabilities. The GUI agent can simulate human users accessing a graphical user interface. Through a joint anchoring technique combining DOM (Document Object Model) tree structure parsing and visual features from page screenshots, it automatically identifies and extracts text information, visual elements, and their layout relationships from the interface, transforming the semantic content of the front-end business interface into structured, machine-understandable data. In this embodiment, the GUI agent can be interacted with through a task configuration page to invoke it to access a labeled page based on a page address.

[0041] Dehydrated data refers to data that loses its original business context during the process of storing real-time operational data in an offline data warehouse due to the standardization of storage formats. Typically, it retains only the original text or image content and its labels, while losing front-end business context elements such as interface display rules, visual cues, and layout relationships. In this application embodiment, for any preset model, the dehydrated data of that preset model is the training data for that preset model.

[0042] Context rehydration refers to a data augmentation operation that extracts prior business context information from the front-end business interface and fuses it with offline, dehydrated data through methods such as global system prompt injection or feature concatenation. This operation restores complete business background information to previously dry data lacking business context. In this embodiment, after context rehydration, annotation rule information and training data can be fused and reconstructed to obtain updated training data. The updated training data can then be used to train a preset model.

[0043] Business Prior Context: This refers to the structured business rule information extracted by the GUI agent through parsing the front-end interface, including but not limited to review specification announcements, warning prompts, operation instructions, and interface visual layout features. This context information reflects the explicit rules and implicit visual cues that human operators rely on when making business decisions. In this embodiment, the business prior context is the annotation rule information corresponding to the annotation page.

[0044] Hyperparameter Dynamic Routing (HDR) refers to a mechanism that automatically adjusts hyperparameters such as the learning rate of each modal encoder in the underlying deep learning training framework based on the analysis results of the visual saliency features of the front-end interface by the GUI agent. In the embodiments of this application, the learning rate of each modal encoder in the preset model can be automatically adjusted based on the page visual saliency analysis results obtained by the agent.

[0045] Automatic Supervised Fine-Tuning (AutoSFT) refers to a technical solution that automates the process of supervising and fine-tuning models in the machine learning operation and maintenance process. It covers the entire process of training data preparation, hyperparameter configuration, training execution, effect evaluation and model deployment, aiming to reduce manual intervention and improve the efficiency of model iteration.

[0046] Visual saliency features refer to visual design elements in a graphical user interface that guide user attention, including highlighted color annotations, font size differences, element placement, and text-image area ratios. In this embodiment, the GUI agent analyzes these features to determine the modal emphasis of the business interface, thereby guiding the configuration and adjustment of the training strategy.

[0047] Cross-system control loop: In this embodiment, it refers to the end-to-end automated control link established in this embodiment, from front-end user interface visual perception → middle-layer data context reconstruction → back-end training hyperparameter intervention. This loop breaks down the information barrier between the business presentation layer and the model training layer, enabling the model fine-tuning process to respond in real time to changes in front-end business rules.

[0048] Typical solutions for automated model fine-tuning in related technologies mainly include the following categories: (I) Automated fine-tuning method based on pure data-driven approach This method directly extracts historical labeled data from the data warehouse, automatically finds a better training configuration through hyperparameter search algorithms (such as grid search, Bayesian optimization, etc.), and then performs model fine-tuning. The characteristic of this type of method is a high degree of automation throughout the process. However, due to the lack of input of business context information, the model cannot know the business background when the labeled data was generated during training. This can easily lead to a learning direction deviation when dealing with samples with ambiguous boundaries or when business rules need to be considered in conjunction with the data.

[0049] (ii) Manually assisted prompt word enhancement and fine-tuning methods This method requires algorithm engineers to manually write a business background description document or prompt word template before starting the fine-tuning task, injecting information such as business rules and review standards into the training process in text form. While this method can compensate for the lack of data semantics to some extent, it suffers from high manual costs and untimely updates. When the review rules of the front-end business system change, the underlying business description document often cannot be updated synchronously, leading to inconsistencies between the training data and the actual business rules.

[0050] (III) Context Preservation Method Based on Metadata Annotation This method requires the annotation system to record some contextual metadata at the time of annotation, such as the version number of the rule in effect at that time and the business group to which the annotator belongs. During fine-tuning, the system queries the corresponding rule document based on the metadata for correlation. The limitation of this method is that the granularity of the metadata is usually coarse, making it difficult to fully restore the visual interface state at the time of annotation, and it requires additional modifications to the data acquisition system; historical data is often not applicable.

[0051] Based on a systematic analysis of the aforementioned existing technical solutions, the following main technical shortcomings can be summarized: (1) The problem of irreversible loss of business context information during data storage In related technologies, a significant information dimensionality reduction phenomenon occurs during the process of annotated data flowing from the front-end business system to the back-end data warehouse. When manual annotators perform annotations in a graphical user interface environment, their decision-making process actually relies on the complete business context presented by the interface, including but not limited to the review specification text displayed in a prominent manner, warning information represented by specific visual styles (such as red font, bold display, flashing effects, etc.), and the spatial relationship between the content to be reviewed and the specification description. However, when the annotation results are persistently stored in a data warehouse (such as a distributed storage system based on Hive), the above-mentioned visual context information is completely stripped away, leaving only the original content to be reviewed (text or image) and the corresponding category label.

[0052] This data storage method can be formally described as follows: Let the complete labeled scene information be S. full ={C content C visual C rule C label}, where C content Represents the original content, C visual Indicating visual context, C rule Represents the business rule context, C label If the label is specified, the stored data will only contain S. stored ={C content C labelThe amount of information loss is This irreversible loss of information means that subsequent training processes cannot obtain complete information for annotation decisions.

[0053] (2) The lack of business rule awareness in the automated fine-tuning system leads to model alignment deviation. Automated fine-tuning methods based on purely data-driven approaches have fundamental technical limitations. Because the training system can only acquire structured data after dimensionality reduction, the model cannot understand the specific business rules underlying the generation of each labeled sample during the learning process. This problem is particularly pronounced when dealing with review scenarios with ambiguous boundaries. For example, in content review, the same text may produce drastically different review conclusions under different business rules, and this information about rule differences is completely missing from the dehydrated data.

[0054] From the perspective of optimization theory, let the true annotation decision function be... Where x is the original content and r is the business rule parameter. The function that the existing method attempts to learn is... Due to the lack of input for the rule parameter r, the learned function It can only be A certain marginal approximation in the rule space inevitably leads to systematic biases on rule-sensitive samples. This bias makes it difficult for the fine-tuned model to maintain consistency with the judgment intent of human annotators at complex and subtle review boundaries.

[0055] (3) The high cost and synchronization lag of manually writing business description documents To address the aforementioned lack of business context, the manual assistance approach used in related technologies requires algorithm engineers to manually write business background description documents or prompt word templates for each fine-tuning task. This approach has three significant drawbacks. First, it is labor-intensive, as engineers need to deeply understand the business rules and write documents for each business line and each fine-tuning task, consuming substantial human resources. Second, document updates suffer from a fixed lag; the review rules of the front-end business system may be hot-updated at any time (e.g., adding new violation types, adjusting judgment criteria, etc.), while the underlying business description documents usually cannot be updated synchronously, resulting in a time lag between the rule descriptions used during training and the actual effective business rules. Third, manually written description documents cannot fully cover all visual elements on the interface that affect annotation decisions; implicit information such as the relative position of elements and the degree of visual emphasis is often overlooked.

[0056] Let t be the actual update time of the business rule. update The document synchronization completion time is t. sync Then the system exists During the rule expiration window, fine-tuning tasks performed during this period will be trained based on outdated rule descriptions.

[0057] (4) The problem of coarse granularity and limited application scope of existing context preservation schemes While context preservation methods based on metadata annotation typically record some contextual information during the data acquisition phase, this approach has several technical limitations. First, the granularity of metadata recording is usually coarse, containing only discrete attributes such as rule version numbers, annotator identifiers, and timestamps, failing to recreate the complete visual state of the interface at the time of annotation. Second, this approach relies on the correlation query between metadata and rule documents, which still require manual maintenance, thus not fundamentally solving the problem of labor costs. Third, this approach requires additional modifications to the data acquisition system, necessitating the embedding of a metadata recording module into the annotation tool, which is costly for existing deployed systems. Finally, and most critically, this approach cannot be applied to historical data; a large number of annotated samples that have been collected but whose metadata has not been recorded will not receive context enhancement, limiting the utilization of data assets.

[0058] (5) There is a lack of effective connection mechanism between GUI intelligent agent technology and model fine-tuning process. Current GUI agent technology is mainly applied in automated testing and robotic process automation, characterized by its ability to perceive interface states and execute predefined interactive operations. However, the technological path of combining the visual perception capabilities of GUI agents with the fine-tuning process of deep learning models remains unexplored. Specifically, related technologies lack the following key mechanisms: first, a standardized method for converting interface information extracted by GUI agents into structured contexts that can be used for training data augmentation; second, a decision-making logic for automatically adjusting training hyperparameters based on interface visual features (such as text-to-image ratio, element saliency distribution, etc.); and third, an end-to-end automated pipeline from front-end interface perception to back-end training control. This lack of technological integration prevents the visual understanding capabilities of GUI agents from being effectively translated into guidance for model fine-tuning.

[0059] To address this, embodiments of this application provide a data processing method and related equipment. The related equipment may include a data processing device, electronic equipment, computer program products, and computer-readable storage media. In terms of product form, the data processing method and related equipment provided in this application can serve as an intelligent model training aid product for enterprise-level machine learning operations and maintenance teams. By adopting an architecture combining a web-based console and a backend service engine, users can create and manage fine-tuning tasks through a unified visual operation page, eliminating the need to directly write training scripts or manually maintain business rule documents.

[0060] The data processing method in this application uses a built-in multimodal GUI agent as its core capability unit, enabling it to autonomously access designated business front-end pages, understand interface content, and extract key business context information like a human operator. This achieves end-to-end automation from front-end business interface perception to back-end model training parameter control, completely freeing algorithm engineers from repetitive tasks such as business context analysis and training data preprocessing.

[0061] The data processing device provided in this application embodiment can be integrated into an electronic device, which can be a server or a user terminal or other similar device.

[0062] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud pre-built databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms.

[0063] The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, smart TV, in-vehicle terminal, etc., but is not limited to these. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited herein.

[0064] Figure 1A This illustration shows an application scenario diagram of the data processing method provided in an embodiment of this application. For example... Figure 1A As shown, this example illustrates a data processing device integrated into an electronic device, which serves as the terminal. Users can perform various operations on the terminal; these users could be engineers from a machine learning operations team, model training engineers, etc. The terminal can integrate a multimodal GUI agent, which can be invoked at any time to execute corresponding tasks.

[0065] Figure 1B This illustration shows another application scenario of the data processing method provided in the embodiments of this application. For example... Figure 1B As shown, this example illustrates a data processing device integrated into an electronic device, which serves as the terminal. Users can perform various operations on the terminal, which can communicate with the server. Users can be engineers from a machine learning operations team, model training engineers, etc. The server can integrate a multimodal GUI agent, allowing the terminal to invoke the agent to execute corresponding tasks.

[0066] exist Figure 1A or Figure 1BBased on this, the terminal can display a task configuration page corresponding to the preset model. The task configuration page includes task execution controls and the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the agent. The agent is invoked to access the annotation page based on the page address to display the annotation rule information identified by the agent in the annotation page on the task configuration page. The annotation rule information is used to update the training data. In response to the trigger operation of the task execution controls, the task execution page is displayed. The task execution page includes the results of training the preset model with the updated training data.

[0067] It is understood that in the specific implementation of this application, page addresses and other related data are involved. When the following embodiments of this application are applied to specific products or technologies, permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0068] The following will introduce typical application scenarios of the data processing method provided in the embodiments of this application.

[0069] Scenario 1: Continuous Optimization of Internet Content Security Review Model In the content security operation scenarios of large-scale internet platforms, review rules are frequently updated in response to changes in regulatory policies and platform governance strategies. In the traditional model, whenever the rule announcements on the front-end review workbench change, algorithm engineers need to manually read the new rules, rewrite the business background prompts, and adjust the training data configuration. The entire process is time-consuming and prone to missing key information.

[0070] After using the data processing method provided in this application, operations and maintenance personnel only need to configure the front-end page address of the target audit business (i.e., the page address of the annotation page corresponding to the training data) in the console, and the system can automatically complete the perception of rule changes and trigger the training process. The GUI intelligent system can periodically or on demand access the audit workbench. Once it detects changes in the content of the specification announcements, warning prompts, etc. on the interface, it automatically extracts the latest business prior context and starts the re-watering and fine-tuning process. This mechanism ensures that the audit model is always synchronized with the latest business rules, thereby eliminating the risk of lag caused by manual maintenance.

[0071] Scenario 2: Cross-category migration of product compliance detection models on e-commerce platforms E-commerce platforms typically have multiple product categories, each with different compliance testing standards. When it is necessary to migrate an existing compliance testing model to a new category, traditional methods require algorithm engineers to have a deep understanding of the new category's review interface layout and rule characteristics, and manually write adapted training configurations.

[0072] The data processing method provided in this application allows users to specify the URL of the review workbench for a new category (i.e., the page address of the annotation page corresponding to the training data) when creating a fine-tuning task. The GUI agent automatically accesses this interface and extracts the category-specific compliance standards and visual layout features. The system automatically adjusts the training hyperparameters based on the interface analysis results. For example, it increases the weight of the visual encoder for the image-intensive interface of the clothing category and increases the weight of the text encoder for the text-intensive interface of the book category. This adaptive mechanism reduces the manual cost of cross-category model transfer.

[0073] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0074] This embodiment will be described from the perspective of a data processing device, which can be integrated into an electronic device, such as a server or a terminal. The terminal can include tablet computers, laptops, personal computers (PCs), wearable devices, virtual reality devices, or other smart devices that can generate image files.

[0075] A data processing method includes: displaying a task configuration page corresponding to a preset model, the task configuration page including a task execution control and a page address of a labeling page corresponding to the training data of the preset model, the task configuration page being used to interact with an agent; invoking the agent to access the labeling page based on the page address, so as to display labeling rule information identified by the agent in the labeling page on the task configuration page, the labeling rule information being used to update the training data; and responding to a trigger operation on the task execution control, displaying a task execution page, the task execution page including the result of training the preset model with the updated training data.

[0076] Figure 2 A flowchart illustrating a data processing method provided in an embodiment of this application is shown. Figure 2 As shown, the specific process of this data processing method is as follows: 101. Display the task configuration page corresponding to the preset model. The task configuration page includes the task execution control and the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the agent.

[0077] Figure 3 This diagram illustrates a page change process of the data processing method provided in an embodiment of this application. Figure 3 As shown, the task configuration page corresponding to the preset model can be displayed first. The task configuration page can include task execution controls and the page address of the annotation page corresponding to the training data of the preset model.

[0078] The control corresponding to "Start Automatic Training" is the task execution control.

[0079] The preset model can be any multimodal large language model specified by the user, and the preset model is the object for subsequent model training.

[0080] The training data for a pre-defined model can be understood as the labeled data used to train that model, i.e., the dehydrated data of that model. Training data may include raw content (such as at least one of text content and image content) and its labeled tags.

[0081] Since manual data annotators typically complete data annotation work in a complete graphical user interface environment, the annotation page can be understood as the front-end graphical user interface displayed on the terminal when manual data annotators are performing data annotation. The annotation page can contain rich business context information, such as review specification announcements, warning prompts, operation instructions, and interface visual layout.

[0082] Figure 4A This illustration shows a schematic diagram of a labeled page provided in an embodiment of this application. For example... Figure 4A As shown, the annotation page can include image content and text content.

[0083] Figure 4B This illustration shows another page diagram of the labeled pages provided in an embodiment of this application. For example... Figure 4B As shown, the annotation page can also include only text content.

[0084] A page address can be understood as the link address of a labeled page (Uniform Resource Locator, or URL for short). Using a page address, you can accurately locate and access the labeled page.

[0085] In this embodiment of the application, displaying the task configuration page corresponding to the preset model may include: displaying the training task creation page, which includes a task creation control and a configuration information input control; In response to input operations on the configuration information input control, the task configuration information entered through the configuration information input control is displayed on the training task creation page. The task configuration information includes the page address of the annotation page corresponding to the training data of the preset model. In response to the trigger operation of the task creation control, the task configuration page corresponding to the task configuration information is displayed.

[0086] Figure 5 A schematic diagram of the training task creation page provided in an embodiment of this application is shown. Figure 5 As shown, the training task creation page includes task creation controls and configuration information input controls.

[0087] The control corresponding to "OK" is the task creation control.

[0088] The configuration information input control is used by the user to input a preset model and the page address of the annotation page corresponding to the training data of the preset model. The user can input a specified preset model and the page address of the annotation page corresponding to the training data of that preset model through the configuration information input control. The preset model and page address entered by the user are collectively referred to as the task configuration information.

[0089] After a user enters task configuration information, the training task creation page can display that information. Furthermore, the user can trigger the task creation control on the training task creation page; responding to this trigger will redirect to the task configuration page corresponding to the task configuration information.

[0090] Figure 6 This illustration shows a page transition process from the training task creation page to the task configuration page, as provided in an embodiment of this application. Figure 6 As shown, the configuration information input controls on the training task creation page can include a model selection sub-control and an address input sub-control.

[0091] The model selection sub-control may include a model display control ( Figure 5 and Figure 6 (Represented by "V" in Chinese) When a user clicks the model display control, the training task creation page will display a set of preset models provided by the system. This set can include multiple models for the user to choose from. The user can specify one model from the preset model set as the preset model.

[0092] For example, in Figure 6 In the example, after a user clicks the model display control, the training task creation page displays multiple models, including Model A, Model B, Model C, and so on. If the user clicks on Model B, it means that the user has selected Model B as the preset model.

[0093] The address input sub-control can be used to input the page address of the annotation page corresponding to the training data of a preset model. For example, in Figure 6 In the example, after the user selects model B as the preset model, they can input the page address of the annotation page corresponding to the training data of model B through the address input sub-control. For example, Figure 6 The page address entered by the user is XXXXXXXXXXXX.

[0094] Accordingly, in response to input operations on the configuration information input control, the task configuration information entered through the configuration information input control is displayed on the training task creation page. This may include: in response to a trigger operation on the model selection sub-control, selecting the model corresponding to the trigger operation from the preset model set as the preset model; in response to an input operation on the address input sub-control, obtaining the web address corresponding to the input operation, and using the web address as the page address of the annotation page for the training data of the preset model.

[0095] In this embodiment, when inputting task configuration information, the user can first select a model from the preset model set as the preset model using the model selection sub-control; then, the user can input the web address corresponding to the preset model using the address input sub-control. The web address input by the user is the page address of the annotation page for the training data of the preset model.

[0096] After entering the task configuration information, the user can click the task creation control on the training task creation page. In response to the trigger operation of the task creation control, the user will be redirected to the task configuration page corresponding to the task configuration information.

[0097] In this embodiment, the task configuration page is used to interact with the intelligent agent.

[0098] The intelligent agent can be a multimodal GUI intelligent agent. The intelligent agent can be integrated into a terminal or server, and the data processing device can invoke the intelligent agent to execute corresponding tasks at any time.

[0099] 102. Call the agent to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent on the annotation page on the task configuration page. The annotation rule information is used to update the training data.

[0100] In this embodiment of the application, an intelligent agent can be invoked to access the annotation page based on the page address, and then the annotation rule information identified by the intelligent agent on the annotation page can be displayed on the task configuration page.

[0101] The annotation rule information can be understood as the structured business rule information extracted by the intelligent agent through parsing the annotation page, also known as the business prior context. The annotation rule information can include feature information such as review rule announcements, warning prompts, operation instructions, and interface visual layout displayed on the annotation page. The annotation rule information reflects the explicit rules and implicit visual cues used by human annotators or operators when annotating the original content.

[0102] In this embodiment of the application, the task configuration page may also include a control for invoking the smart agent.

[0103] Figure 7A schematic diagram illustrating the invocation process of an intelligent agent provided in an embodiment of this application is shown. For example... Figure 7 As shown, the control corresponding to "Invoke Agent" is the agent invocation control. Based on this, invoking the agent accesses the annotation page based on the page address to display the annotation rule information identified by the agent on the annotation page on the task configuration page. This can include: responding to a trigger operation on the invocation control, invoking the agent to access the annotation page based on the page address, identifying the annotation rule information corresponding to the training data on the annotation page through the agent; and displaying the annotation rule information on the task configuration page.

[0104] Users can trigger the agent's call control. In response to the trigger operation of the call control, the agent can be invoked to access the annotation page based on the page address. Then, the agent can identify the annotation rule information corresponding to the training data on the annotation page and display the annotation rule information identified by the agent on the task configuration page.

[0105] In this embodiment of the application, the intelligent agent is the core technical component for automatically acquiring the front-end business context. It can identify the annotation rule information corresponding to the training data in the annotation page based on the joint anchoring mechanism of DOM tree structure parsing and page screenshot visual features.

[0106] In this embodiment, the agent accesses the currently active front-end business interface (i.e., the annotation page) in real time each time a fine-tuning task is initiated, ensuring that the extracted prior business context (annotation rule information) always reflects the latest review specifications. When the front-end system performs a hot rule update (such as adding a violation type, adjusting the judgment threshold, or modifying the warning text), the next triggered fine-tuning task will automatically detect and apply these changes, without requiring manual maintenance of the underlying rule description file. This mechanism shortens the response cycle from business rule changes to preset model updates from several days in traditional solutions to several hours, ensuring that the preset model always resonates with the latest business interface.

[0107] The work of an intelligent agent mainly includes three stages: page access control, multimodal feature extraction, and context fusion encoding.

[0108] Figure 8 A schematic diagram illustrating the process for determining annotation rule information provided in an embodiment of this application is shown. For example... Figure 8 As shown, step 102, calling the agent to access the annotation page based on the page address, so that the agent can identify the annotation rule information corresponding to the training data in the annotation page, may include: 1021. Call the intelligent agent to access the labeled page based on the page address.

[0109] Step 1021 corresponds to the page access control step of the intelligent agent. The intelligent agent can use headless browser technology to simulate the page access behavior of a real user. First, it parses the page address specified in the task configuration page, then starts the browser instance and navigates to that page address.

[0110] To ensure the complete loading of the annotation page content, a page readiness determination logic based on DOM change detection can be built into the intelligent agent. When no new nodes are detected in the page DOM structure within a set time window, the annotation page is considered to have finished loading.

[0111] 1022. The agent identifies at least one text content in the labeled page and extracts features of the content hierarchy relationship between at least one text content to obtain text structure features.

[0112] 1023. Identify at least one image content in the labeled page using an intelligent agent, and extract features from at least one image content to obtain the page's visual features.

[0113] Steps 1022 and 1023 correspond to the multimodal feature extraction stage of the agent, which is responsible for extracting structured data (i.e., text structure features) and visual features (i.e., page visual features) from the loaded labeled page.

[0114] In terms of text structure feature extraction, the agent can traverse the labeled page DOM tree to identify nodes with specific semantic tags, including bulletin board containers, warning text areas, and rule lists. The identification process can employ a dual filtering strategy based on CSS selector rules and text pattern matching. For nodes that match preset patterns, their text content and hierarchical relationships are extracted.

[0115] CSS selectors are "navigation rules" used to locate HTML (Hypertext Markup Language) elements and apply styles. Their core consists of basic selectors, combinatoric selectors, attribute selectors, pseudo-classes, and pseudo-elements.

[0116] In terms of page visual feature extraction, the agent can perform a full-screen screenshot operation on the labeled page to obtain a complete rendered image of the page. Then, a pre-trained visual encoder is used to extract features from the rendered image, obtaining fixed-dimensional page visual features. The process of feature extraction from the rendered image by the visual encoder can be represented by the following feature mapping relationship:

[0117] in, This indicates the image being rendered on the page. This represents the feature mapping function of the visual encoder. This represents the pre-trained parameters of the visual encoder. This indicates the visual characteristics of the output page.

[0118] 1024. The text structure features and page visual features are fused to obtain the annotation rule information corresponding to the training data.

[0119] Step 1024 corresponds to the context fusion encoding stage of the agent. The agent can fuse and encode the text structural features extracted in step 1022 with the page visual features extracted in step 1023 to generate the final annotation rule information.

[0120] In this embodiment of the application, text structural features and page visual features can be directly fused with a weight of 1:1 to obtain annotation rule information.

[0121] Step 1024, fusing text structural features and page visual features to obtain annotation rule information corresponding to the training data, may include: identifying text regions corresponding to text content in the annotation page and determining the area of ​​the corresponding text region; identifying image regions corresponding to image content in the annotation page and determining the area of ​​the corresponding image region; determining the feature weights corresponding to feature fusion based on the area of ​​the text region and the area of ​​the image region; and fusing text structural features and page visual features based on the feature weights to obtain annotation rule information corresponding to the training data.

[0122] In this embodiment, the process of fusing text structural features and page visual features can also employ an attention-weighted mechanism. Different weights are assigned based on the confidence level of each information source; that is, text structural features and page visual features have different feature weights. Then, feature fusion is performed based on the respective feature weights of the text structural features and page visual features to obtain the annotation rule information corresponding to the training data.

[0123] Let the set of text rules in the annotation page be... , This corresponds to a text rule fragment on the annotation page. This represents the encoding result of the set of text rules, i.e., the text structural features; the page visual features are... Then the annotation rule information ( The generation process of ) can be represented as:

[0124] in, This refers to the fusion coefficient, or feature weight, dynamically calculated based on the page type of the labeled page. Specifically, it will... As the feature weights corresponding to the text structure features, 1- As feature weights corresponding to the visual features of the page.

[0125] In practice, the text regions corresponding to the text content can be identified on the annotation page, and their areas can be determined. Simultaneously, the image regions corresponding to the image content can be identified on the annotation page, and their areas can be determined. Then, based on the text region areas and image region areas, the feature weights corresponding to the feature fusion process are calculated. By substituting the feature weights into the above formula, the annotation rule information corresponding to the training data can be obtained.

[0126] The process of determining the feature weights corresponding to feature fusion based on the text region area and the image region area can include: determining the total area of ​​the labeled page; calculating the ratio between the text region area and the total page area to obtain the text region area ratio; calculating the ratio between the image region area and the total page area to obtain the image region area ratio; and determining the feature weights corresponding to feature fusion based on the text region area ratio and the image region area ratio.

[0127] The total page area refers to the total area of ​​the entire page containing the text. The text area ratio can be understood as the proportion of the text area to the total page area, and it can be expressed as... The image area ratio can be understood as the proportion of the image area in the total page area, and it can be expressed as: .

[0128] In this embodiment of the application, the text region area ratio ( ) and the ratio of image region area ( ), determine the feature weights corresponding to feature fusion.

[0129] Specifically, determining the feature weights corresponding to feature fusion based on the ratio of text region area to image region area can include: adding the ratio of text region area to image region area to obtain the target area ratio; calculating the ratio between the text region area ratio and the target area ratio to obtain the relative ratio of text region area; determining the page type of the labeled page based on the relative ratio of text region area, and determining the feature weights corresponding to feature fusion based on the page type.

[0130] In this embodiment of the application, a page type determination index can be preset, and the page type determination index adopts... express, The calculation formula is as follows:

[0131] Among them, the denominator ( This refers to the target area ratio, a key indicator for determining page type. That is, the ratio between the text area ratio and the target area ratio (called the relative ratio of text area).

[0132] After obtaining the relative ratio of the text region area ( After that, the page type of the labeled page can be determined based on the relative ratio of the text area area. For example, when When the value is greater than 0.6, the page type of the labeled page is determined to be text-intensive; when... When the value is less than 0.4, the page type of the labeled page is determined to be image-intensive; otherwise, it is determined to be balanced.

[0133] It is understandable that the comparison thresholds for text-intensive and image-intensive types can be set or adjusted according to the actual situation. The examples of 0.6 and 0.4 mentioned above are only for illustration and do not limit the actual comparison thresholds that can be used.

[0134] Once the page type of the labeled page is determined, the feature weights corresponding to feature fusion can be determined based on the page type. Specifically, when the page type of the labeled page is text-intensive, A larger value can be taken; when the page type of the labeled page is image-intensive, A smaller value can be taken.

[0135] In this embodiment of the application, the page type of the labeled page (text-intensive, image-intensive, or balanced) can also be understood as the page visual saliency analysis result obtained by the agent analyzing the labeled page.

[0136] Based on the results of page visual saliency analysis, the training hyperparameters of the preset model can also be configured during the training process. Specifically, in this embodiment, the learning rate of each modal encoder (network) in the preset model can be automatically adjusted based on the page visual saliency analysis results obtained by the intelligent agent, thereby achieving automatic mapping from front-end interface visual features to candidate training parameter configurations.

[0137] The preset model can include a text encoding network and a visual encoding network. The text encoding network can be used to encode the text content in the input samples during training to extract the corresponding text features; the visual encoding network can be used to encode the image content in the input samples during training to obtain the corresponding image features.

[0138] Based on this, after calculating the ratio between the text region area ratio and the target area ratio to obtain the relative ratio of the text region area, it may also include: determining the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training based on the relative ratio of the text region area; and displaying the text learning rate and the visual learning rate on the task configuration page.

[0139] Let the base learning rate be... The text learning rate of the text encoding network is The visual learning rate of the visual encoding network is .

[0140] The formula for calculating the text learning rate of a text encoding network can be expressed as follows:

[0141] The formula for calculating the visual learning rate of a visual encoding network can be expressed as follows:

[0142] in, This is the routing strength coefficient, which controls the magnitude of the learning rate adjustment. The above formula ensures that when the labeled page type is text-intensive (…),… When the learning rate is greater than 0.6, the text encoding network of the preset model obtains a higher learning rate to enhance the learning of text features; conversely, the visual encoding network of the preset model obtains a higher learning rate.

[0143] In this embodiment, after calculating the ratio between the image region area and the total page area to obtain the image region area ratio, the method may further include: calculating the text color contrast of the text region and the image color contrast of the image region in the annotation page; calculating the text level depth of the text node and the image level depth of the image node in the annotation page; identifying the page center coordinates in the annotation page; calculating the text position centrality of the text element in the text region from the page center coordinates and the image position centrality of the image element in the image region from the page center coordinates; determining the text modality saliency score corresponding to the text region based on the text region area, text color contrast, text level depth, and text position centrality; determining the visual modality saliency score corresponding to the image region based on the image region area, image color contrast, image level depth, and image position centrality; and determining the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training based on the text modality saliency score and the visual modality saliency score.

[0144] In addition to the aforementioned technical solutions that rely solely on the area of ​​the text region and the area of ​​the image region to determine the learning rates of the text encoding network and the visual encoding network, embodiments of this application can also construct a saliency quantization function containing four-dimensional features to comprehensively determine the learning rates of the text encoding network and the visual encoding network.

[0145] Specifically, the annotation page can be divided into text regions and image regions. For text regions, the text modality saliency score can be calculated from four dimensions: text region area, text color contrast, text hierarchy depth, and text position centrality. For image regions, the image modality saliency score can be calculated from four dimensions: image region area, image color contrast, image hierarchy depth, and image position centrality. Then, based on the text modality saliency score and the visual modality saliency score, the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training are determined.

[0146] Text color contrast can be understood as the color contrast weight corresponding to the text in a text area. In this embodiment, the text color of the DOM text nodes in the DOM tree structure can be extracted, and the relative brightness contrast can be calculated according to the WCAG (Content Accessibility Guidelines) standard. The contrast ratio ranges from 1:1 to 21:1, and is normalized to the (0, 1] interval. The more glaring the color (such as red background with white text), the closer the text color contrast is to 1. Image color contrast can be understood as the color contrast weight corresponding to the background of an image area. Similarly, the image color of the DOM image nodes in the DOM tree structure can be extracted, and the relative brightness contrast can be calculated according to the WCAG (Content Accessibility Guidelines) standard. The contrast ratio ranges from 1:1 to 21:1, and is normalized to the (0, 1] interval. The more glaring the color, the closer the image color contrast is to 1.

[0147] Text hierarchy depth can be understood as the number of levels from the root node in the DOM tree structure for a DOM text node. Image hierarchy depth can be understood as the number of levels from the root node in the DOM tree structure for an image node. Hierarchical depth generally indicates higher visual priority.

[0148] Wherein, the coordinates of the page center identified by the annotation page are ( The center coordinates of the text elements in the text area are ( We can introduce a two-dimensional Gaussian distribution decay function to calculate the text position centrality of text elements in a text region from the center coordinates of the page, as follows:

[0149] The closer the text is to the center of the screen or the visual focal point of the first screen, the higher its centrality.

[0150] Similarly, let the coordinates of the center point of the image element in the image region be ( Then, the image centrality of the image elements in the image region from the center coordinates of the page can be expressed as follows:

[0151] The closer an image is to the center of the screen or the visual focal point of the first screen, the higher its centrality.

[0152] Based on this, the text modality saliency score can be calculated as follows:

[0153] in, Indicates text color contrast. Indicates the depth of text hierarchy, used in the formula. As a factor, it reflects the principle in front-end design logic that the shallower the nesting and the higher the visual global priority.

[0154] Similarly, the visual modality saliency score can be calculated as follows:

[0155] in, Indicates image color contrast. Representing the image layer depth, used in the formula As a factor, it reflects the principle in front-end design logic that the shallower the nesting and the higher the visual global priority.

[0156] w 1. w 2. w 3 and w 4. These four weighting coefficients reflect the proportion of different dimensions in visual saliency, and satisfy the following conditions: This application provides two methods for determining the four weighting coefficients mentioned above.

[0157] (1) Empirical Prior Configuration Method (Default): Empirical values ​​are set based on statistical priors in the fields of human-computer interaction and eye-tracking detection. For example, region area (including text region area and image region area) and location centrality (including text location center confidence and image location center confidence) usually have the greatest visual impact and can be set. =0.4, =0.3, =0.2, =0.1.

[0158] (2) Data-driven fine-tuning method: During the system cold start phase, an open-source webpage visual saliency dataset is introduced to...w 1. w 2. w 3 and w 4 is used as a learnable parameter to train a lightweight linear regressor to fit the UI design specifications of a specific enterprise.

[0159] After determining the text modality saliency score and the visual modality saliency score, the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training can be determined based on these scores. This step specifically includes: adding the text modality saliency score and the visual modality saliency score to obtain the target saliency score; calculating the ratio between the text modality saliency score and the target saliency score to obtain the text modality relative ratio; and determining the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training based on the text modality relative ratio.

[0160] For example, the relative ratio of text modalities can be expressed as follows:

[0161] Among them, the denominator ( This is the target saliency score. The relative ratio of text modalities at this point ( This can also be used as an indicator for determining page type. For example, when... When the value is greater than 0.6, the page type of the labeled page is determined to be text-intensive; when... When the value is less than 0.4, the page type of the labeled page is determined to be image-intensive; otherwise, it is determined to be balanced.

[0162] Based on this, let the base learning rate be... The text learning rate of the text encoding network is The visual learning rate of the visual encoding network is .

[0163] The formula for calculating the text learning rate of a text encoding network can be expressed as follows:

[0164] The formula for calculating the visual learning rate of a visual encoding network can be expressed as follows:

[0165] in, This is the routing strength coefficient, which controls the magnitude of the learning rate adjustment.

[0166] In this embodiment of the application, after determining the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training based on the text modality relative ratio, the method may further include: freezing the network parameters of the text encoding network during training when the text modality relative ratio is less than a first text modality relative ratio threshold; and freezing the network parameters of the visual encoding network during training when the text modality relative ratio is greater than a second text modality relative ratio threshold.

[0167] This application also proposes a modality freezing adaptive method. When extreme tendencies occur (such as excessively high or low relative ratios of text modalities), the system routing instruction will trigger a "partial freeze" mechanism, directly freezing all gradient updates of the corresponding encoding network. In this way, not only can catastrophic forgetting be prevented, but the training time of fine-tuning tasks can also be reduced to less than 60% of the original time.

[0168] For example, when the relative ratio of text modalities ( When the relative ratio of text modalities is less than the threshold of the first text modality, the network parameters of the text encoding network can be frozen during training. If the ratio of the two text modalities is greater than the threshold of the relative ratio of the second text modality, then the network parameters of the visual coding network can be frozen during training.

[0169] The relative ratio thresholds for the first and second text modalities can be set according to actual conditions, and the relative ratio threshold for the first text modality is much smaller than the relative ratio threshold for the second text modality. For example, the relative ratio threshold for the first text modality can be set to 0.2, and the relative ratio threshold for the second text modality can be set to 0.8.

[0170] In addition, after calculating the ratio between the text modality saliency score and the target saliency score to obtain the text modality relative ratio, the process may also include: obtaining the network capacity parameters of the preset model; determining the page type of the labeled page based on the text modality relative ratio; and allocating the network capacity parameters based on the page type to obtain the text network capacity parameters corresponding to the text encoding network and the visual network capacity parameters corresponding to the visual encoding network.

[0171] In this embodiment, in addition to dynamically adjusting the learning rate, the fine-tuning structural parameters (rank size of low-rank adaptation) of the underlying large language model (i.e., the preset model in this embodiment) can be dynamically routed and mapped.

[0172] The network capacity adoption number of the preset model can be understood as the rank of the network structure of the preset model based on the low-rank adaptation (LoRA) fine-tuning architecture. LoRA achieves fine-tuning by injecting a low-rank matrix, and the size of the rank directly determines how many trainable parameters (i.e., neuron capacity) can be assigned to different encoding networks (such as text encoding networks and visual encoding networks) in the preset model.

[0173] Let the network capacity parameter be... The text network capacity parameter corresponding to the text encoding network is: The visual network capacity parameter corresponding to the visual coding network is: .

[0174]

[0175]

[0176] in, This is the capacity strength coefficient, used to control the magnitude of rank adjustment.

[0177] if The interface is highly dependent on text rules (heavy text). The system dynamically expands the text network capacity parameter corresponding to the text encoding network, while shrinking the visual network capacity parameter corresponding to the visual encoding network.

[0178] In this embodiment, by adjusting the learning rate and network capacity parameters of the text encoding network and the visual encoding network respectively through the relative ratio of text modalities, a comprehensive intervention strategy can be implemented to allocate faster update speeds to network layers with larger capacities. This mechanism of tilting computing resources towards business-related modalities can improve the convergence ceiling of specific businesses under the same memory conditions.

[0179] In this embodiment of the application, during each backpropagation in the training loop, the text modality relative ratio can also be used as a basis ( Apply a gradient scaler to the multimodal fusion layer. If the business rule context of a certain batch is extremely complex (DOM nodes increase dramatically), the gradient backpropagation weight of that part of the network parameters can be temporarily increased to ensure that the complex business logic is fully absorbed by the model, forming "visual feature-guided end-to-end gradient intervention".

[0180] In this embodiment of the application, after calculating the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during the training process, the text learning rate and visual learning rate can also be displayed on the task configuration page so that the user can know the training hyperparameter configuration scheme corresponding to the preset model and confirm whether to use the text learning rate and visual learning rate to train the preset model.

[0181] The agent in this embodiment not only extracts business rules in text form but also analyzes the visual saliency features of the interface, including the area ratio, positional distribution, and visual emphasis of text and image elements. The system automatically adjusts the training hyperparameters in the deep learning framework based on these visual features. For example, when the interface presents a layout that emphasizes text over images, the learning rate of the text encoder is increased while the learning rate of the visual encoder is decreased. This dynamic routing mechanism based on business scenario characteristics improves training convergence speed by approximately 20% compared to traditional methods with fixed hyperparameter configurations or blind grid search, while avoiding modal bias problems caused by improper hyperparameter configuration.

[0182] Figure 9 This illustration shows a schematic diagram of a task configuration page provided in an embodiment of this application. For example... Figure 9 As shown, the task configuration page can simultaneously display annotation rule information, the text learning rate of the text encoding network, and the visual learning rate of the visual encoding network.

[0183] In this embodiment, annotation rule information is used to update the training data to obtain updated training data. In related technologies, training data is directly used to train a preset model. However, in this embodiment, an intelligent agent is invoked to access the annotation page corresponding to the training data of the preset model. The agent identifies annotation rule information on the annotation page. This annotation rule information represents the complete business context used by human annotators during the annotation process, including but not limited to prominently displayed review specification text, warning information displayed with specific visual styles (such as red font, bolding, flashing effects, etc.), and the spatial relationship between the content to be reviewed and the specification description. This embodiment uses annotation rule information to update the training data to obtain updated training data, and then uses the updated training data to train the preset model.

[0184] 103. In response to a trigger operation on the task execution control, display the task execution page, which includes the results of training the preset model with the updated training data.

[0185] like Figure 3 As shown, in response to a trigger action on the task execution control in the task configuration page, the task execution page can be displayed. The task execution page can be used to display the results of training a preset model using updated training data. For example, the result information could be "Model training complete," indicating that the preset model has converged.

[0186] As mentioned earlier, step 102 involves calling an intelligent agent to identify the labeled page, and the resulting labeling rule information is used to update the training data of the preset model, thus obtaining updated training data. The updated training data is then used to train the preset model.

[0187] The annotation rules are displayed on the task configuration page. In this embodiment, the update method for training data can also be displayed on the task configuration page.

[0188] Figure 10 This illustration shows another page diagram of the task configuration page provided in an embodiment of this application. For example... Figure 10 As shown, the task configuration page can also include an update method selection control, i.e. Figure 10 The control corresponding to "how to update training data".

[0189] Based on this, in response to a trigger operation on the task execution control, a task execution page is displayed, which may include: in response to a selection operation on the update method selection control, filtering the target update method corresponding to the selection operation from a preset set of update methods; in response to a trigger operation on the task execution control, executing the training task corresponding to the preset model to display the task execution page, the task execution page including the execution process information of the training task, the training task including updating the training data through the target update method and annotation rule information, and training the preset model through the updated training data, text learning rate and visual learning rate; when the completion of the training task is detected, the task execution result of the training task is displayed on the task execution page, the task execution result indicating the result of training the preset model.

[0190] Figure 11 This illustration shows a page transition process from the task configuration page to the task execution page, as provided in an embodiment of this application. Figure 11 As shown, users can select an update method from a set of preset update methods using the update method selection control. For example, when a user triggers the update method selection control ( Figure 10 When the "V" is selected, the task configuration page displays a set of preset update methods provided by the system. This set can include multiple preset update methods (such as Update Method 1, Update Method 2, etc.) for the user to choose from. If the user clicks Update Method 1, it means the user has selected Update Method 1 as the target update method.

[0191] Once the user specifies the target update method, the task execution control on the task configuration page can be triggered. In response to the triggering action of the task execution control, the training task for the preset model can begin, and the corresponding task execution page will be displayed.

[0192] In the process of starting the training task for the preset model, the training data needs to be updated according to the target update method and annotation rule information to obtain the updated training data; then the preset model is trained using the updated training data.

[0193] Furthermore, during the training of the preset model using the updated training data, it is also necessary to control the update step size of the text encoding network and the visual encoding network during the model update process based on the text learning rate and visual learning rate calculated above.

[0194] like Figure 11 As shown, after starting a training task for a preset model, the task execution page can first display the execution process information of the training task. This execution process information can be dynamically updated, reflecting the current training status of the preset model in real time. For example, the execution process information may include the training progress of the preset model, changes in the loss curve, and trends in evaluation metrics. Furthermore, this embodiment also supports setting abnormal alarm rules. When abnormal situations such as gradient explosion or metric decline occur during the training process, alarms can be triggered, prompting the user to take appropriate action.

[0195] like Figure 11 As shown, when the training task is detected to be completed (i.e., the training progress is 100%), the task execution result can be displayed on the task execution page. The task execution result indicates the result of training the preset model.

[0196] Furthermore, after training is completed, this embodiment of the application can automatically execute the evaluation process and generate an evaluation report of the model training results. If the evaluation metrics meet the preset deployment criteria, the user can trigger the model release operation with one click, thereby pushing the weights of the newly trained model to the online inference service and completing the canary deployment switch.

[0197] This application establishes a complete cross-system control link, covering three core stages: front-end UI visual extraction, offline data context rehydration, and back-end training hyperparameter intervention. The entire process eliminates the need for manual writing of business description documents or manual adjustment of training configurations. The GUI agent automatically completes interface information collection and structured encoding, and the system automatically performs data rehydration (updating training data through target update methods and annotation rules) and hyperparameter routing (dynamically adjusting the learning rates of the text encoder and visual encoder in the preset model). Compared to existing technologies that require algorithm engineers to manually write business background prompts for each fine-tuning task, this application reduces the manual intervention time for word fine-tuning tasks from an average of 4-8 hours to zero, achieving truly fully automated model evolution.

[0198] As mentioned earlier, in the process of starting the training task for the preset model, the training data needs to be updated according to the target update method and annotation rule information to obtain the updated training data.

[0199] In this embodiment of the application, updating the training data through the target update method and annotation rule information may include: obtaining at least one annotation rule information obtained by the agent accessing the annotation page at different times, and determining the rule information generation time of the annotation rule information and the training data generation time of the training data; Calculate the time difference between the rule information generation time and the training data generation time, and based on the time difference, filter out the target annotation rule information corresponding to the training data from the annotation rule information; update the training data through the target update method and the target annotation rule information to obtain the updated training data.

[0200] Simple data splicing can easily lead to the problem of "historical data using future rules." To address this, the embodiments of this application can employ a dual alignment algorithm based on timestamps and business tags, using a nearest-neighbor backtracking matching mechanism to ensure that the annotation rule information used when updating training data is strictly consistent with the interface context during manual annotation.

[0201] Specifically, during the process of invoking the agent to access the annotation page, a context version snapshot tree can be constructed. This tree can include at least one annotation rule information obtained by the agent accessing the annotation page at different times, as well as the rule information generation time corresponding to each annotation rule information. For the at least one annotation rule information obtained, its corresponding rule information generation time can be compared with the training data generation time to filter out the target annotation rule information from the at least one annotation rule information.

[0202] The method for determining the target annotation rule information can be represented as follows:

[0203] in, C represents the target annotation rule information; C represents a candidate historical context node in the context version snapshot tree (i.e., the annotation rule information captured in a certain crawl). Tree This represents a context version snapshot tree generated by the agent periodically inspecting the annotation page, which stores annotation rule information that takes effect at different time periods; The timestamp indicates when the annotation rule information officially takes effect (or is captured) in the front-end business system, which is also the time when the rule information is generated; This indicates the timestamp when the training data was generated, i.e., the time when the training data was generated.

[0204] Constraints It strictly limits the use of rules that are effective "before or at the same time as the data is generated" to prevent the introduction of future rules (i.e., data leaks or time travel issues). argmin This indicates that under the above constraints, the search is for the time difference ( The smallest annotation rule information is the version of the annotation rule information that is most recently generated from the training data and is already in effect, which is used as the target annotation rule information.

[0205] In this embodiment, the training data for the preset model can be obtained in batches from a data warehouse through a standardized data source adaptation interface. The training data may include sample data of at least one training sample. The sample data of each training sample is stored in a structured form, and each piece of sample data contains original content (text or image), annotation labels, and timestamps, etc.

[0206] Let the training data be ,in, This represents the original content of the i-th sample data. The label represents the corresponding annotation, and N represents the total number of training samples in the training samples.

[0207] In this embodiment, the process of updating training data using target annotation rule information can be understood as a process of fusing and reconstructing the target annotation rule information with the training data, i.e., a context rehydration process. During this process, the method of fusing and reconstructing the target annotation rule information with the training data varies depending on the target update method selected by the user.

[0208] In this embodiment of the application, updating the training data through the target update method and target annotation rule information to obtain updated training data may include: when the target update method is the prompt word injection method, converting the target annotation rule information into annotation rule prompt words; adding the annotation rule prompt words to the sample data to obtain updated training data.

[0209] The cue word injection method is suitable for cue word-based fine-tuning paradigms. When the target update method is cue word injection, the annotation rule information can be converted into natural language descriptions and then uniformly added as system-level cue words to the input prefix position of each sample data, thereby obtaining the updated sample data corresponding to each training sample. The updated training data includes the updated sample data corresponding to each training sample.

[0210] In the prompt injection method, the updated sample data corresponding to each training sample can be represented as:

[0211] in, The textual representation of target annotation rule information (i.e. annotation rule prompts). This indicates a sequence concatenation operation.

[0212] Based on this, the complete updated training data can be represented as:

[0213] in, This indicates the context refill operator.

[0214] Furthermore, converting target annotation rule information into annotation rule prompts may also include: identifying target text structural features and target page visual features in the target annotation rule information; generating text annotation rule prompts based on the target text structural features; generating page visual rule prompts based on the target page visual features; and fusing the text annotation rule prompts and page visual rule prompts to obtain the annotation rule prompts.

[0215] Unlike the aforementioned technical solutions that use the textual representation of target annotation rule information as annotation rule prompts, this application embodiment can also use both the textual and visual representations of target annotation rule information as annotation rule prompts. Accordingly, the updated sample data corresponding to each training sample can be represented as:

[0216] in, As a dynamic template function, it can convert visual emphasis information on the page (such as highlighting and bolding a rule) into natural language emphasis instructions (such as: "Note: The following red-highlighting rules must be strictly followed...").

[0217] This represents the spatial layout and visual structure information extracted from the labeled page (such as elements being located at the top of the first screen of the page, text highlighted in red and bold, and the enclosing relationship between images and text), which are the visual features of the target page according to the target labeling rules.

[0218] In this embodiment, when assembling annotation rule prompts, the target text structural features can be converted into text annotation rule prompts, and the target page visual features can be converted into page visual rule prompts. The text annotation rule prompts and page visual rule prompts are then fused to obtain the annotation rule prompts. The resulting annotation rule prompts not only reflect the textual rule content but also the relative importance of different rules on the annotation page, thus more closely resembling the visual judgment logic of a real person and improving the accuracy of the updated training data.

[0219] In this embodiment of the application, updating the training data through the target update method and target annotation rule information to obtain updated training data may include: when the target update method is feature concatenation, extracting features from the target annotation rule information to obtain annotation rule features; extracting features from the sample data to obtain sample features; and concatenating the annotation rule features and sample features to obtain updated training data.

[0220] Feature concatenation is suitable for feature-based fine-tuning paradigms. When the target update method is feature concatenation, feature extraction can be performed on the target annotation rule information to obtain the corresponding vector representation (i.e., annotation rule features), and feature extraction can be performed on the sample data to obtain the corresponding vector representation (i.e., sample features). Then, the annotation rule features and each sample feature are concatenated along a specified dimension to obtain the updated sample data for each training sample. The updated training data includes the updated sample data for each training sample.

[0221] The process of concatenating the labeled rule features and sample features to obtain updated training data can include: performing a dot product calculation on the sample features and the labeled rule features to obtain the attention weight of the sample features relative to the labeled rule features; performing a weighted calculation on the labeled rule features based on the attention weight to obtain weighted labeled rule features; and concatenating the weighted labeled rule features with the sample features to obtain updated training data.

[0222] In this embodiment, the labeled rule features and sample features can be concatenated using a feature-level rehydration method based on cross-modal attention to obtain updated training data. Specifically, the labeled rule features can be used as key and value vectors, and the sample features as query vectors, through a cross-attention mechanism, and the updated training data can be calculated as follows:

[0223] in, It represents the sample features, which, as a query vector, indicate what kind of annotation rules and features are needed to guide the current batch of sample data.

[0224] It represents the annotation rule features, serving as both a key vector and a value vector, and represents the structured rule library provided by the front-end annotation page.

[0225] This represents the cross-attention mechanism, which determines which part of the annotation rule features is most relevant to the sample features by calculating the clicks of the query vector and key vector, and extracts the relevant annotation rule features from the value vector to obtain the weighted annotation rule features.

[0226] After weighting, the labeled rule features and sample features can be residually connected, followed by layer normalization to obtain updated training data. Layer normalization prevents feature value explosion and ensures training stability.

[0227] In this embodiment of the application, in order to prevent the original feature distribution data of the training data from being overwhelmed ("semantic overload") due to excessive refilling during the training data update process, the semantic shift difference (SSD) distribution before and after the training data update can be calculated, and it can be determined whether the updated training data needs to be adjusted based on the semantic shift difference distribution.

[0228] Correspondingly, after updating the training data through the target update method and annotation rule information, it may also include: extracting features from the training data to obtain the original feature distribution data of the training data; Feature extraction is performed on the updated training data to obtain the updated feature distribution data of the updated training data. Calculate the feature distribution difference data between the original feature distribution data and the updated feature distribution data; adjust the updated training data according to the feature distribution difference data to obtain the adjusted training data, and use the adjusted training data as the updated training data.

[0229] Among them, the semantic offset distribution can measure whether the core semantics of the training data itself have been destroyed or buried after the (target) annotation rule information is added.

[0230] Specifically, features can be extracted from the training data to obtain the original feature distribution data of the training data (represented as follows for ease of distinction). ); it can also extract features from the updated training data to obtain the updated feature distribution data of the updated training data (for ease of distinction, it is represented as ); ).

[0231] Afterwards, the original feature distribution data can be calculated ( ) and updated feature distribution data ( The feature distribution difference data between the two is obtained by considering the semantic offset distribution before and after the training data update. The specific calculation formula can be expressed as follows:

[0232] The larger the value of the feature distribution difference data, the more drastic the change or interference of the rehydration operation on the semantics of the training data. Therefore, based on the feature distribution difference data, it can be determined whether the updated training data needs to be adjusted. When the value of the feature distribution difference data is greater than a preset threshold, the updated training data can be adjusted to obtain the adjusted training data, which is then used as the updated training data.

[0233] In this embodiment of the application, the methods for adjusting the updated training data may include the following two: (1) For symbol-level watermarking (i.e., the target update method is prompt word injection) If the value of the feature distribution difference data exceeds a preset threshold, content pruning can be automatically triggered: secondary supplementary explanatory text will be removed first, retaining only mandatory specification requirements; if it still exceeds the threshold, then... The visually reinforcing modifiers have degenerated into a mere patchwork of objective facts.

[0234] (2) For feature-level rehydration (i.e., the target update method is feature splicing): a scalable rehydration coefficient can be introduced into the cross-attention fusion formula. (Values ​​range from 0 to 1). Accordingly, the calculation formula for the updated training data is adjusted as follows:

[0235] If the value of the feature distribution difference data is greater than the preset threshold, then... Attenuation (e.g.) This reduces the proportion of weighted annotation rule feature injection.

[0236] This application embodiment extracts prior business context, such as review guidelines, warning information, and visual layout, from the front-end business interface (i.e., the page address of the annotation page corresponding to the training data) in real time using an intelligent agent, and then rehydrates it with the dehydrated data (i.e., training data) in the data warehouse. This method effectively eliminates the context loss problem of offline data and significantly improves the alignment accuracy of the model. Specifically, this mechanism allows the originally dry data containing only raw content and tags to regain complete business context information. The updated training data obtained after rehydration carries the same rule guidance as when human annotators make decisions, and the preset model can understand the specific business rules on which each sample data was generated during the learning process. Experiments show that in complex multimodal content review scenarios, the preset model fine-tuned using the data processing method provided in this application embodiment improves the judgment accuracy of samples with blurred boundaries by more than 30% compared to traditional methods, and the human-machine alignment consistency index improves by more than 25%.

[0237] Furthermore, the contextual redundancy mechanism in this embodiment does not rely on metadata records during the data acquisition phase. Instead, it enhances (updates) any historically labeled data (training data) collected by the intelligent agent after it obtains real-time warning information from the current business interface. This means that a large amount of dehydrated data (training data) already accumulated in the enterprise data warehouse can be context-completed through the data processing method provided in this embodiment, without modifying the original data acquisition system or requiring manual re-labeling. This feature gives the invention strong backward compatibility, fully activating the value of existing data assets and increasing data utilization efficiency by more than 40%.

[0238] Furthermore, embodiments of this application can extract the prior business context during the rehydration process ( The model is continuously stored in a structured format and associated with corresponding fine-tuning tasks. When it is necessary to explain the cause of a model's prediction result, the business rule context used during model training can be traced back to clarify the basis for specifying the model's learning objectives. Compared with traditional black-box fine-tuning methods, this traceability mechanism provides a clear evidence chain for model auditing, performance attribution analysis, and problem localization, meeting the compliance requirements for algorithm interpretability in highly regulated industries such as finance and healthcare.

[0239] Figure 12 A schematic diagram of the user operation flow corresponding to the data processing method provided in the embodiments of this application is shown. For example... Figure 12 As shown, the typical usage flow of the data processing method provided in this application embodiment can be divided into the following five steps: (1) Task creation: Users log in to the system console, enter the task creation page, and click to create a new task. In the task configuration form, users fill in the task name, select the target business line (i.e., the preset model to be trained), enter the corresponding front-end review page URL (i.e., the page address of the annotation page corresponding to the training data of the preset model), and specify the storage path of the training data in the data warehouse.

[0240] (2) Agent Preview: After the task is created, the system will schedule the agent to preview the specified front-end review page (annotation page). Users can view the business prior context summary (i.e., annotation rule information) extracted by the agent on the task configuration page, including the identified standard announcement content, warning text list, and interface layout type judgment results. Users can confirm or manually correct the extraction results.

[0241] (3) Rehydration configuration confirmation: The system displays a preview of the rehydration strategy on the task configuration page, including the context injection method (cue word injection method or feature concatenation method), hyperparameter adjustment suggestions (learning rate ratio of each encoder), etc. After the user confirms that the configuration is correct, the task is submitted to the execution queue.

[0242] (4) Training monitoring: During task execution, users can view training progress, loss curve changes, evaluation metric trends, and other information in real time on the task execution page. The system supports setting abnormal alarm rules, which will automatically notify users when abnormal situations such as gradient explosion or metric decline occur during training.

[0243] (5) Model Deployment: After training is completed, the system automatically evaluates the process and generates an evaluation report. If the evaluation metrics meet the preset deployment criteria, users can trigger the model deployment operation with one click. The system will push the new model weights to the online inference service and complete the gray-scale switching.

[0244] The data processing methods described above prioritize reducing the cognitive burden and operational complexity for users in their interaction design. The console employs a wizard-driven task creation process, guiding users step-by-step through the configuration process and avoiding the confusion caused by presenting too many parameter options at once. The business context (annotation rule information) extracted by the agent is displayed in the form of visual cards, allowing users to intuitively understand the system's perception results and perform necessary manual verification.

[0245] The data processing method described above also provides a task template function, allowing users to save frequently used business configurations as templates for direct reuse when creating similar tasks later, further improving operational efficiency. For fine-tuning tasks that need to be executed periodically, the system supports setting timed scheduling strategies to achieve fully unattended continuous model optimization.

[0246] Figure 13 A schematic diagram of the system architecture corresponding to the data processing method provided in the embodiments of this application is shown. For example... Figure 13 As shown, the data processing method provided in this application adopts a layered structure and end-to-end linkage design concept at the technical architecture level. The overall architecture can be divided into four core layers from top to bottom: front-end perception layer, data reconstruction layer, training control layer, and model service layer. Data flow and control information transmission are realized between each layer through standardized interfaces.

[0247] Among them, the front-end perception layer deploys a multimodal GUI intelligent agent as the core perception unit. This intelligent agent has the ability to autonomously access the web-based business system and can perform dual-channel analysis of DOM structure parsing and visual feature extraction on the target page.

[0248] The data reconstruction layer takes over the business prior context output by the perception layer and merges it with the original labeled data (training data) from the data warehouse to generate a reconstructed dataset (updated training data) carrying the complete business context.

[0249] The training control layer dynamically generates training hyperparameter configuration schemes based on the interface visual saliency analysis results provided by the perception layer, and schedules the underlying deep learning framework to perform fine-tuning tasks.

[0250] The model service layer is responsible for receiving the trained model weights and executing the hot update deployment of the online inference service.

[0251] The technical feature of the above architecture lies in establishing a cross-system closed loop from front-end pixel-level perception to back-end parameter-level control. Traditional automatic fine-tuning systems only achieve automation at the data level, while the embodiments of this application introduce a GUI intelligent agent as an intermediate perception layer, organically connecting the previously fragmented business front-end and training candidates. During system operation, the data flow between each level follows a unidirectional transmission principle, while control information supports reverse backtracking, ensuring that the training strategy can be adjusted in real time according to changes in front-end business.

[0252] This application employs a layered, modular architecture design, providing excellent system scalability and cross-platform compatibility. Each functional layer consists of decoupled, independent components that communicate via standardized interfaces. This allows the system to flexibly adapt to different front-end technology stacks (such as web-based B / S architecture systems and native control-based C / S architecture systems), and also to interface with different deep learning training frameworks (such as PyTorch, TensorFlow, and PaddlePaddle). When the front-end technology of the business system migrates or the underlying training framework needs upgrading, only the corresponding adapter components need to be replaced; the core refactoring logic and control flow remain unchanged, reducing system maintenance costs by more than 50%.

[0253] exist Figure 13 Based on the system architecture diagram shown, the core processing flow of the data processing method provided in this application embodiment follows a four-stage execution logic of "perception, rehydration, routing, and training," with strict dependencies and data links between each stage.

[0254] Figure 14 A schematic diagram illustrating the execution logic of the data processing method provided in this application embodiment is shown. For example... Figure 14 As shown, the data processing method provided in this application embodiment may include the following processing steps: (1) Task triggering and agent scheduling phase When the system detects an pending automatic fine-tuning task in a specific business queue, it first parses the task configuration information to obtain the front-end page URL (i.e., the page address of the annotation page corresponding to the training data of the preset model) and the data source path (i.e., the storage path of the training data) for the target business. Then, it sends a perception request to the GUI agent scheduling center, which allocates available agent instances based on the current load status of the agent resource pool. Upon receiving the target URL, the allocated agent instance initiates the page access and specific extraction process.

[0255] (2) Multimodal perception and context generation stage After the GUI agent completes page loading, it performs DOM structure parsing and visual feature extraction in parallel. The DOM parser traverses the page nodes, identifying and extracting key content such as review guidelines announcements, warning texts, and rule lists based on predefined semantic rules. The visual feature extractor performs encoding operations on the page screenshots, simultaneously calculating the area ratio of text regions to image regions. The context encoder fuses textual and visual information, outputting a structured business prior context (…). ) and interface type determination indicators ( ). (3) Data restoration and hyperparameter routing stage After receiving the business prior context (annotation rule information), the data reconstruction module pulls the original labeled data (training data) from the data warehouse and performs data fusion by selecting the appropriate rehydration strategy according to the task configuration. After fusion, the rehydrated dataset (updated training data) is serialized and stored in the training data directory. At the same time, the hyperparameter routing module calculates the learning rate configuration of each encoder based on the interface type determination index and generates a complete training hyperparameter configuration file.

[0256] (4) Automatic training and model deployment stage The training controller loads the rehydrated dataset and hyperparameter configuration file, initializes the deep learning training environment, and starts the fine-tuning loop. During training, the system continuously monitors the changing trends of the loss function value and evaluation metrics. After training is complete, the system automatically executes the model evaluation process, calculating the performance metrics of the new model on the validation set. The weights of the evaluated models are pushed to the model service layer for hot updates in the online inference service.

[0257] The data processing method provided in this application covers various data types, including front-end page data (annotation page), original annotation data (training data), business prior context (annotation rule information), rehydrated training data (updated training data), and model weights. The flow of each type of data within the system follows specific processing rules and storage specifications.

[0258] The front-end page data is input into the system in the form of a URL, and after being processed by the intelligent agent, it is converted into a business prior context.

[0259] The business prior context is stored in JSON (a lightweight data exchange format) format, which includes fields such as a list of text rules, visual feature vectors, and interface type determination results.

[0260] After the original labeled data is pulled from the data warehouse, it is fused with the business prior context in memory, and the resulting re-water training data is stored in the standard dataset format.

[0261] The model weights generated during the training process are saved in the framework's native format, along with metadata text recording the training configuration and evaluation results.

[0262] Key quality control steps in data processing engineering include context extraction result verification, reconstructed data integrity check, and training data distribution detection.

[0263] Among them, the context extraction result verification ensures the accuracy of the extracted content by combining rule matching and manual sampling.

[0264] The integrity check of the rehydrated data verifies whether each sample has been successfully injected into the business prior context.

[0265] The training data distribution detection uses statistical methods to detect changes in the data distribution before and after rehydration, ensuring that the rehydration operation does not introduce abnormal offsets.

[0266] In summary, the data processing method provided in this application offers significant cost-effectiveness in practical deployments, substantially reducing the overall operational costs of model fine-tuning and improving the efficiency of machine learning operations and maintenance. The cost of manually compiling corpora is reduced to zero, the operational costs of rule synchronization maintenance are significantly lowered, the trial-and-error costs of hyperparameter tuning are reduced, and the reuse value of historical data is increased. Taking a typical content moderation scenario as an example, after adopting the data processing method provided in this application, the model iteration cycle for a single business line is shortened from an average of two weeks to less than three days, the total annual model operation and maintenance cost is reduced by approximately 60%, and the accuracy and recall of the online model service are steadily improved.

[0267] As can be seen from the above, this embodiment of the application displays a task configuration page corresponding to a preset model. The task configuration page includes a task execution control and the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the agent. Then, the agent is invoked to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent in the annotation page on the task configuration page. The annotation rule information is used to update the training data. Finally, in response to the trigger operation of the task execution control, the task execution page is displayed. The task execution page includes the result of training the preset model with the updated training data.

[0268] This solution adds a front-end context collection step before initiating the training task for the preset model. The task configuration page displays the page address of the annotation page corresponding to the task execution control and the training data of the preset model. Specifically, the task execution control triggers the training task, the preset model is the object to be trained, the training data is the dehydrated data of the preset model, and the annotation page is the front-end graphical user interface displayed on the terminal when human annotators annotate the data. The annotation page contains rich business context information, and the page address is a link to the annotation page.

[0269] Building upon this, an intelligent agent can autonomously access the annotation page based on the page address, allowing the agent to identify annotation rule information within the annotation page. This annotation rule information is the structured business rule information extracted by the agent through parsing the annotation page; it can also be referred to as the business prior context. The annotation rule information reflects the explicit rules and implicit visual cues used by human annotators or operators when annotating the original content, encompassing the complete business context upon which human annotators or operators base their annotations on the training data. The task configuration page can display the annotation rule information identified by the agent.

[0270] Furthermore, this solution can update the training data using the annotation rule information obtained from agent recognition. This update process involves fusing and reconstructing the annotation rule information with the training data, essentially a context rehydration process. This update process ensures that the updated training data carries the same rule guidance as when human annotators or operators make decisions.

[0271] Subsequently, in response to the trigger operation of the task execution control, the training and fine-tuning task of the preset model is initiated. Because this solution rehydrates and reconstructs the dehydrated data before executing the training and fine-tuning task, the learning objective of the preset model can be kept consistent with the annotation rule information, thereby reducing alignment deviations caused by missing context.

[0272] Figure 15This illustration shows another application scenario of the data processing method provided in the embodiments of this application. For example... Figure 15 As shown, taking an example where the data processing device is integrated into an electronic device, and the electronic device is a server, the server can obtain the page address of the annotation page corresponding to the training data of the preset model; based on the page address, it accesses the annotation page to identify the annotation rule information in the annotation page; according to the annotation rule information, it updates the training data to obtain the updated training data; based on the updated training data, it trains the preset model to obtain the trained model.

[0273] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0274] This embodiment will be described from the perspective of a data processing device, which can be integrated into an electronic device, such as a server or a terminal. The terminal can include tablet computers, laptops, personal computers (PCs), wearable devices, virtual reality devices, or other smart devices that can generate image files.

[0275] A data processing method includes: obtaining the page address of the annotation page corresponding to the training data of a preset model; accessing the annotation page based on the page address to identify annotation rule information in the annotation page; updating the training data according to the annotation rule information to obtain updated training data; and training the preset model based on the updated training data to obtain a trained model.

[0276] Figure 16 Another flowchart illustrating the data processing method provided in an embodiment of this application is shown. Figure 16 As shown, the specific process of this data processing method is as follows: 201. Obtain the page address of the annotation page corresponding to the training data of the preset model.

[0277] The preset model can be any multimodal large language model specified by the user, and the preset model is the object for subsequent model training.

[0278] The training data for a pre-defined model can be understood as the labeled data used to train that model, i.e., the dehydrated data of that model. Training data may include raw content (such as at least one of text content and image content) and its labeled tags.

[0279] Since manual data annotators typically complete data annotation work in a complete graphical user interface environment, the annotation page can be understood as the front-end graphical user interface displayed on the terminal when manual data annotators are performing data annotation. The annotation page can contain rich business context information, such as review specification announcements, warning prompts, operation instructions, and interface visual layout.

[0280] A page address can be understood as the link address of a labeled page (Uniform Resource Locator, or URL for short). Using a page address, you can accurately locate and access the labeled page.

[0281] In this embodiment, the user can input the page address of the annotation page corresponding to the training data of the preset model through the terminal. The terminal establishes a communication connection with the server, and then the server can obtain the page address of the annotation page corresponding to the training data of the preset model from the terminal.

[0282] 202. Based on the page address, access the annotation page to identify the annotation rule information on the annotation page.

[0283] In this embodiment of the application, a multimodal GUI agent can be integrated into the server. The agent can access the annotation page and identify annotation rule information on the annotation page.

[0284] The annotation rule information can be understood as the structured business rule information extracted by the intelligent agent through parsing the annotation page, also known as the business prior context. The annotation rule information can include feature information such as review rule announcements, warning prompts, operation instructions, and interface visual layout displayed on the annotation page. The annotation rule information reflects the explicit rules and implicit visual cues used by human annotators or operators when annotating the original content.

[0285] Specifically, identifying annotation rule information on the annotation page can include: identifying at least one text content on the annotation page and extracting features from the content hierarchy relationship between the at least one text content to obtain text structure features; identifying at least one image content on the annotation page and extracting features from the at least one image content to obtain page visual features; and fusing the text structure features and page visual features to obtain annotation rule information corresponding to the training data.

[0286] The process of fusing text structural features and page visual features to obtain annotation rule information corresponding to the training data may include: identifying text regions corresponding to text content in the annotation page and determining the area of ​​the corresponding text region; identifying image regions corresponding to image content in the annotation page and determining the area of ​​the corresponding image region; determining the feature weights corresponding to feature fusion based on the area of ​​the text region and the area of ​​the image region; and fusing text structural features and page visual features based on the feature weights to obtain annotation rule information corresponding to the training data.

[0287] The process of determining the feature weights corresponding to feature fusion based on the text region area and the image region area can include: determining the total area of ​​the labeled page; calculating the ratio between the text region area and the total page area to obtain the text region area ratio; calculating the ratio between the image region area and the total page area to obtain the image region area ratio; and determining the feature weights corresponding to feature fusion based on the text region area ratio and the image region area ratio.

[0288] The process of determining the feature weights corresponding to feature fusion based on the ratio of text region area to image region area can include: adding the ratio of text region area to image region area to obtain the target area ratio; calculating the ratio between the text region area ratio and the target area ratio to obtain the relative ratio of text region area; determining the page type of the labeled page based on the relative ratio of text region area, and determining the feature weights corresponding to feature fusion based on the page type.

[0289] In this embodiment, the preset model may include a text encoding network and a visual encoding network. Based on this, after calculating the ratio between the text region area ratio and the target area ratio to obtain the relative ratio of the text region area, the model may further include: determining the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training based on the relative ratio of the text region area; and displaying the text learning rate and visual learning rate on the task configuration page. 203. Update the training data according to the annotation rules to obtain the updated training data.

[0290] The training data may include sample data of at least one training sample.

[0291] Specifically, updating the training data through the target update method and annotation rule information can include: when the target update method is the prompt word injection method, converting the annotation rule information into annotation rule prompt words; adding the annotation rule prompt words to the sample data to obtain the updated training data.

[0292] Specifically, updating the training data through the target update method and annotation rule information may also include: when the target update method is feature concatenation, extracting features from the annotation rule information to obtain annotation rule features, and extracting features from the sample data to obtain sample features; concatenating the annotation rule features and sample features to obtain the updated training data.

[0293] 204. Based on the updated training data, train the preset model to obtain the trained model.

[0294] During model training, the learning rate of each encoder is used to update the model parameters according to the learning rate calculated above.

[0295] After training is complete, the server can automatically evaluate the process and generate an evaluation report. If the evaluation metrics meet the preset deployment criteria, the new model weights are pushed to the online inference service and the phase-out is completed.

[0296] As can be seen from the above, the embodiments of this application obtain the page address of the annotation page corresponding to the training data of the preset model; then, based on the page address, the annotation page is accessed to identify the annotation rule information in the annotation page; then, the training data is updated according to the annotation rule information to obtain the updated training data; finally, the preset model is trained based on the updated training data to obtain the trained model.

[0297] This solution allows for the addition of a front-end context acquisition step before initiating the training task on the preset model. This step retrieves the page address of the annotation page corresponding to the training data of the preset model. Specifically, the task execution control triggers the training task, the preset model is the object subsequently trained, the training data is the dehydrated data of the preset model, the annotation page is the front-end graphical user interface displayed on the terminal when a human annotator annotates the data, the annotation page contains rich business context information, and the page address is a link to the annotation page.

[0298] Building upon this foundation, the agent can autonomously access the annotation page via its URL to identify annotation rule information. This annotation rule information, also known as business prior context, is the structured business rule information extracted by the agent through parsing the annotation page. It reflects the explicit rules and implicit visual cues used by human annotators or operators when annotating the original content, encompassing the complete business context upon which they annotate the training data. The task configuration page displays the annotation rule information identified by the agent.

[0299] Furthermore, this solution can update the training data using the identified annotation rule information. This update process involves fusing and reconstructing the annotation rule information with the training data, essentially a context rehydration process. This update process ensures that the updated training data carries the same rule guidance used by human annotators or operators in their decision-making.

[0300] Then, the pre-set model can be trained based on the updated training data to obtain the trained model. Because this approach rehydrates and reconstructs the dehydrated data before performing the training fine-tuning task, it can ensure that the learning objective of the pre-set model is consistent with the annotation rule information, thereby reducing alignment deviations caused by missing context.

[0301] To better implement the above methods, this application also provides a data processing device that can be integrated into a network device, such as a server or terminal. The terminal may include a tablet computer, a laptop computer, and / or a personal computer.

[0302] Figure 17 A schematic diagram of the structure of a data processing apparatus provided in an embodiment of this application is shown. Figure 17 As shown, the data processing device may include a task configuration unit 301, an agent invocation unit 302, and a task execution unit 303, as follows: (1) Task configuration unit 301; The task configuration unit 301 is used to display the task configuration page corresponding to the preset model. The task configuration page includes task execution controls and the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the agent.

[0303] For example, the task configuration unit 301 can be used to display a training task creation page, which includes a task creation control and a configuration information input control. In response to an input operation on the configuration information input control, the task configuration information entered through the configuration information input control is displayed on the training task creation page. The task configuration information includes the page address of the annotation page corresponding to the training data of the preset model. In response to a trigger operation on the task creation control, the task configuration page corresponding to the task configuration information is displayed.

[0304] For example, the configuration information input control includes a model selection sub-control and an address input sub-control. Based on this, the task configuration unit 301 can specifically be used to, in response to a trigger operation on the model selection sub-control, select the model corresponding to the trigger operation from the preset model set as the preset model; and in response to an input operation on the address input sub-control, obtain the web address corresponding to the input operation and use the web address as the page address of the annotation page for the training data of the preset model.

[0305] (2) Intelligent agent calling unit 302; The agent invocation unit 302 is used to invoke the agent to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent in the annotation page on the task configuration page. The annotation rule information is used to update the training data.

[0306] For example, the task configuration page also includes a control for invoking the agent. Based on this, the agent invoking unit 302 can specifically be used to respond to a trigger operation on the invoking control, invoking the agent to access the annotation page based on the page address, so that the agent can identify the annotation rule information corresponding to the training data on the annotation page; and display the annotation rule information on the task configuration page.

[0307] For example, the agent invocation unit 302 can be used to invoke an agent to access the annotation page based on the page address; the agent identifies at least one text content in the annotation page and extracts features from the content hierarchy relationship between the at least one text content to obtain text structure features; the agent identifies at least one image content in the annotation page and extracts features from the at least one image content to obtain page visual features; the text structure features and page visual features are fused to obtain annotation rule information corresponding to the training data.

[0308] For example, the agent invocation unit 302 can be used to identify the text region corresponding to the text content in the annotation page and determine the area of ​​the text region; identify the image region corresponding to the image content in the annotation page and determine the area of ​​the image region; determine the feature weights corresponding to feature fusion based on the area of ​​the text region and the area of ​​the image region; and fuse the text structure features and page visual features based on the feature weights to obtain the annotation rule information corresponding to the training data.

[0309] For example, the agent calling unit 302 can be used to determine the total area of ​​the labeled page; calculate the ratio between the text region area and the total page area to obtain the text region area ratio; calculate the ratio between the image region area and the total page area to obtain the image region area ratio; and determine the feature weights corresponding to feature fusion based on the text region area ratio and the image region area ratio.

[0310] For example, the agent calling unit 302 can be used to add the text region area ratio and the image region area ratio to obtain the target area ratio; calculate the ratio between the text region area ratio and the target area ratio to obtain the relative ratio of the text region area; determine the page type of the labeled page based on the relative ratio of the text region area, and determine the feature weights corresponding to feature fusion based on the page type.

[0311] (3) Task execution unit 303.

[0312] The task execution unit 303 is used to display a task execution page in response to a trigger operation on the task execution control. The task execution page includes the results of training a preset model with updated training data.

[0313] For example, the preset model includes a text encoding network and a visual encoding network. Based on this, the task execution unit 303 can be used to determine the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training based on the relative ratio of text region areas; and display the text learning rate and visual learning rate on the task configuration page.

[0314] For example, task execution unit 303 can specifically be used to calculate the text color contrast of a text region and the image color contrast of an image region on the annotation page, and to calculate the text level depth of a text node and the image level depth of an image node on the annotation page; to identify the page center coordinates on the annotation page, and to calculate the text position centrality of a text element in a text region from the page center coordinates, and the image position centrality of an image element in an image region from the page center coordinates; to determine the text modality saliency score corresponding to the text region based on the text region area, text color contrast, text level depth, and text position centrality, and to determine the visual modality saliency score corresponding to the image region based on the image region area, image color contrast, image level depth, and image position centrality; and to determine the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training based on the text modality saliency score and the visual modality saliency score.

[0315] For example, task execution unit 303 can be used to add the text modality saliency score and the visual modality saliency score to obtain the target saliency score; calculate the ratio between the text modality saliency score and the target saliency score to obtain the text modality relative ratio; and determine the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training based on the text modality relative ratio.

[0316] For example, task execution unit 303 can be used to freeze the network parameters of the text encoding network during training when the relative ratio of text modalities is less than the first relative ratio threshold of text modalities; and to freeze the network parameters of the visual encoding network during training when the relative ratio of text modalities is greater than the second relative ratio threshold of text modalities.

[0317] For example, task execution unit 303 can be used to obtain network capacity parameters of a preset model; determine the page type of the labeled page based on the relative ratio of text modalities; and allocate network capacity parameters based on the page type to obtain the text network capacity parameters corresponding to the text encoding network and the visual network capacity parameters corresponding to the visual encoding network.

[0318] For example, the task configuration page also includes an update method selection control. Based on this, the task execution unit 303 can specifically be used to, in response to a selection operation on the update method selection control, filter out the target update method corresponding to the selection operation from a preset set of update methods; in response to a trigger operation on the task execution control, execute the training task corresponding to the preset model to display the task execution page, which includes information about the execution process of the training task. The training task includes updating the training data using the target update method and annotation rule information, and training the preset model using the updated training data, text learning rate, and visual learning rate; when the completion of the training task is detected, the task execution result is displayed on the task execution page, indicating the result of training the preset model.

[0319] For example, the task execution unit 303 can be used to obtain at least one annotation rule information obtained by the agent accessing the annotation page at different times, and determine the rule information generation time of the annotation rule information and the training data generation time of the training data; calculate the time difference between the rule information generation time and the training data generation time, and based on the time difference, filter out the target annotation rule information corresponding to the training data from the annotation rule information; update the training data through the target update method and the target annotation rule information to obtain the updated training data.

[0320] For example, the training data includes sample data of at least one training sample. Based on this, the task execution unit 303 can be specifically used to convert the target annotation rule information into annotation rule prompts when the target update method is prompt word injection; and add the annotation rule prompts to the sample data to obtain the updated training data.

[0321] For example, task execution unit 303 can be used to identify target text structural features and target page visual features in target annotation rule information; generate text annotation rule prompts based on target text structural features; generate page visual rule prompts based on target page visual features; and fuse text annotation rule prompts and page visual rule prompts to obtain annotation rule prompts.

[0322] For example, the training data includes sample data of at least one training sample. Based on this, the task execution unit 303 can be specifically used to extract features from the target annotation rule information to obtain annotation rule features when the target update method is feature concatenation, and to extract features from the sample data to obtain sample features; and to concatenate the annotation rule features and sample features to obtain the updated training data.

[0323] For example, task execution unit 303 can be used to perform dot product calculation on sample features and annotation rule features to obtain the attention weight of sample features on annotation rule features; perform weighted calculation on annotation rule features according to attention weight to obtain weighted annotation rule features; and concatenate weighted annotation rule features with sample features to obtain updated training data.

[0324] For example, task execution unit 303 can be used to extract features from training data to obtain the original feature distribution data of training data; extract features from updated training data to obtain the updated feature distribution data of updated training data; calculate the feature distribution difference data between the original feature distribution data and the updated feature distribution data; adjust the updated training data according to the feature distribution difference data to obtain the adjusted training data, and use the adjusted training data as the updated training data.

[0325] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0326] As can be seen from the above, in this embodiment of the application, the task configuration unit 301 displays the task configuration page corresponding to the preset model. The task configuration page includes the page address of the annotation page corresponding to the training data of the preset model and is used to interact with the agent. Then, the agent calling unit 302 calls the agent to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent in the annotation page on the task configuration page. The annotation rule information is used to update the training data. Finally, the task execution unit 303 responds to the trigger operation of the task execution control and displays the task execution page, which includes the result of training the preset model with the updated training data.

[0327] This solution adds a front-end context collection step before initiating the training task for the preset model. The task configuration page displays the page address of the annotation page corresponding to the task execution control and the training data of the preset model. Specifically, the task execution control triggers the training task, the preset model is the object to be trained, the training data is the dehydrated data of the preset model, and the annotation page is the front-end graphical user interface displayed on the terminal when human annotators annotate the data. The annotation page contains rich business context information, and the page address is a link to the annotation page.

[0328] Building upon this, an intelligent agent can autonomously access the annotation page based on the page address, allowing the agent to identify annotation rule information within the annotation page. This annotation rule information is the structured business rule information extracted by the agent through parsing the annotation page; it can also be referred to as the business prior context. The annotation rule information reflects the explicit rules and implicit visual cues used by human annotators or operators when annotating the original content, encompassing the complete business context upon which human annotators or operators base their annotations on the training data. The task configuration page can display the annotation rule information identified by the agent.

[0329] Furthermore, this solution can update the training data using the annotation rule information obtained from agent recognition. This update process involves fusing and reconstructing the annotation rule information with the training data, essentially a context rehydration process. This update process ensures that the updated training data carries the same rule guidance as when human annotators or operators make decisions.

[0330] Subsequently, in response to the trigger operation of the task execution control, the training and fine-tuning task of the preset model is initiated. Because this solution rehydrates and reconstructs the dehydrated data before executing the training and fine-tuning task, the learning objective of the preset model can be kept consistent with the annotation rule information, thereby reducing alignment deviations caused by missing context.

[0331] To better implement the above methods, this application also provides a data processing device that can be integrated into a network device, such as a server or terminal. The terminal may include a tablet computer, a laptop computer, and / or a personal computer.

[0332] Figure 18 Another structural schematic diagram of the data processing apparatus provided in an embodiment of this application is shown. For example... Figure 18 As shown, the data processing device may include a page address acquisition unit 401, a page address access unit 402, a training data update unit 403, and a model training unit 404, as follows: (1) Page address acquisition unit 401; Page address acquisition unit 401 is used to acquire the page address of the labeled page corresponding to the training data of the preset model.

[0333] (2) Page address access unit 402; The page address access unit 402 is used to access the annotation page based on the page address in order to identify annotation rule information in the annotation page.

[0334] (3) Training data update unit 403; The training data update unit 403 is used to update the training data according to the annotation rule information to obtain the updated training data.

[0335] (4) Model training unit 404.

[0336] The model training unit 404 is used to train the preset model based on the updated training data to obtain the trained model.

[0337] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0338] As can be seen from the above, in this embodiment of the application, the page address acquisition unit 401 obtains the page address of the annotation page corresponding to the training data of the preset model; then, the page address access unit 402 accesses the annotation page based on the page address to identify the annotation rule information in the annotation page; the training data update unit 403 then updates the training data according to the annotation rule information to obtain the updated training data; finally, the model training unit 404 trains the preset model based on the updated training data to obtain the trained model.

[0339] This solution allows for the addition of a front-end context acquisition step before initiating the training task on the preset model. This step retrieves the page address of the annotation page corresponding to the training data of the preset model. Specifically, the task execution control triggers the training task, the preset model is the object subsequently trained, the training data is the dehydrated data of the preset model, the annotation page is the front-end graphical user interface displayed on the terminal when a human annotator annotates the data, the annotation page contains rich business context information, and the page address is a link to the annotation page.

[0340] Building upon this foundation, the agent can autonomously access the annotation page via its URL to identify annotation rule information. This annotation rule information, also known as business prior context, is the structured business rule information extracted by the agent through parsing the annotation page. It reflects the explicit rules and implicit visual cues used by human annotators or operators when annotating the original content, encompassing the complete business context upon which they annotate the training data. The task configuration page displays the annotation rule information identified by the agent.

[0341] Furthermore, this solution can update the training data using the identified annotation rule information. This update process involves fusing and reconstructing the annotation rule information with the training data, essentially a context rehydration process. This update process ensures that the updated training data carries the same rule guidance used by human annotators or operators in their decision-making.

[0342] Then, the pre-set model can be trained based on the updated training data to obtain the trained model. Because this approach rehydrates and reconstructs the dehydrated data before performing the training fine-tuning task, it can ensure that the learning objective of the pre-set model is consistent with the annotation rule information, thereby reducing alignment deviations caused by missing context.

[0343] This application also provides an electronic device, such as... Figure 19 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically: The electronic device may include components such as a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will understand that... Figure 19 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 501 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes various functions and processes data by running or executing software programs and / or modules stored in the memory 502, and by calling data stored in the memory 502. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 501.

[0344] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.

[0345] The electronic device also includes a power supply 503 that supplies power to various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 503 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0346] The electronic device may also include an input unit 504, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0347] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 501 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 502 according to the following instructions, and the processor 501 runs the applications stored in the memory 502 to realize various functions, as follows: Displays the task configuration page corresponding to the preset model. The task configuration page includes the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the agent. The agent is invoked to access the annotation page based on the page address to display the annotation rule information identified by the agent in the annotation page on the task configuration page. The annotation rule information is used to update the training data. In response to the trigger operation of the task execution control, the task execution page is displayed. The task execution page includes the results of training the preset model with the updated training data.

[0348] or, Obtain the page address of the annotation page corresponding to the training data of the preset model; based on the page address, access the annotation page to identify the annotation rule information in the annotation page; update the training data according to the annotation rule information to obtain the updated training data; train the preset model based on the updated training data to obtain the trained model.

[0349] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0350] As can be seen from the above, this embodiment of the application displays a task configuration page corresponding to a preset model. The task configuration page includes a task execution control and the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the agent. Then, the agent is invoked to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent in the annotation page on the task configuration page. The annotation rule information is used to update the training data. Finally, in response to the trigger operation of the task execution control, the task execution page is displayed. The task execution page includes the result of training the preset model with the updated training data.

[0351] This solution adds a front-end context collection step before initiating the training task for the preset model. The task configuration page displays the page address of the annotation page corresponding to the task execution control and the training data of the preset model. Specifically, the task execution control triggers the training task, the preset model is the object to be trained, the training data is the dehydrated data of the preset model, and the annotation page is the front-end graphical user interface displayed on the terminal when human annotators annotate the data. The annotation page contains rich business context information, and the page address is a link to the annotation page.

[0352] Building upon this, an intelligent agent can autonomously access the annotation page based on the page address, allowing the agent to identify annotation rule information within the annotation page. This annotation rule information is the structured business rule information extracted by the agent through parsing the annotation page; it can also be referred to as the business prior context. The annotation rule information reflects the explicit rules and implicit visual cues used by human annotators or operators when annotating the original content, encompassing the complete business context upon which human annotators or operators base their annotations on the training data. The task configuration page can display the annotation rule information identified by the agent.

[0353] Furthermore, this solution can update the training data using the annotation rule information obtained from agent recognition. This update process involves fusing and reconstructing the annotation rule information with the training data, essentially a context rehydration process. This update process ensures that the updated training data carries the same rule guidance as when human annotators or operators make decisions.

[0354] Subsequently, in response to the trigger operation of the task execution control, the training and fine-tuning task of the preset model is initiated. Because this solution rehydrates and reconstructs the dehydrated data before executing the training and fine-tuning task, the learning objective of the preset model can be kept consistent with the annotation rule information, thereby reducing alignment deviations caused by missing context.

[0355] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0356] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the data processing methods provided in embodiments of this application. For example, the instructions can execute the following steps: Displays the task configuration page corresponding to the preset model. The task configuration page includes the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the agent. The agent is invoked to access the annotation page based on the page address to display the annotation rule information identified by the agent in the annotation page on the task configuration page. The annotation rule information is used to update the training data. In response to the trigger operation of the task execution control, the task execution page is displayed. The task execution page includes the results of training the preset model with the updated training data.

[0357] or, Obtain the page address of the annotation page corresponding to the training data of the preset model; based on the page address, access the annotation page to identify the annotation rule information in the annotation page; update the training data according to the annotation rule information to obtain the updated training data; train the preset model based on the updated training data to obtain the trained model.

[0358] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0359] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0360] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the data processing methods provided in the embodiments of this application, the beneficial effects that any of the data processing methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0361] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in various alternative implementations of the data processing method described above.

[0362] The data processing method and related equipment provided in the embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A data processing method, characterized in that, include: Displays the task configuration page corresponding to the preset model. The task configuration page includes task execution controls and the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the agent. The agent is invoked to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent in the annotation page on the task configuration page. The annotation rule information is used to update the training data. In response to a trigger operation on the task execution control, a task execution page is displayed, the task execution page including the results of training the preset model with updated training data.

2. The data processing method according to claim 1, characterized in that, The task configuration page corresponding to the preset model includes: Display a training task creation page, which includes task creation controls and configuration information input controls; In response to an input operation to the configuration information input control, the task configuration information entered through the configuration information input control is displayed on the training task creation page. The task configuration information includes the page address of the annotation page corresponding to the training data of the preset model. In response to a trigger operation on the task creation control, the task configuration page corresponding to the task configuration information is displayed.

3. The data processing method according to claim 2, characterized in that, The configuration information input control includes a model selection sub-control and an address input sub-control. The step of displaying task configuration information entered through the configuration information input control on the training task creation page in response to an input operation on the configuration information input control includes: In response to a trigger operation on the model selection sub-control, the model corresponding to the trigger operation is selected from the preset model set as the preset model; In response to an input operation on the address input sub-control, the web address corresponding to the input operation is obtained, and the web address is used as the page address of the annotation page for the training data of the preset model.

4. The data processing method according to claim 1, characterized in that, The task configuration page also includes a control for invoking the intelligent agent. The step of invoking the agent to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent in the annotation page on the task configuration page, includes: In response to a trigger operation on the control, the agent is invoked to access the annotation page based on the page address, so that the agent can identify the annotation rule information corresponding to the training data in the annotation page. The annotation rule information is displayed on the task configuration page.

5. The data processing method according to claim 4, characterized in that, The step of invoking the agent to access the annotation page based on the page address, so as to identify the annotation rule information corresponding to the training data in the annotation page through the agent, includes: The intelligent agent is invoked to access the labeled page based on the page address; The agent identifies at least one text content in the labeled page and extracts features from the content hierarchy relationship between the at least one text content to obtain text structure features. The intelligent agent identifies at least one image content in the labeled page and extracts features from the at least one image content to obtain the page's visual features; The text structure features and the page visual features are fused to obtain the annotation rule information corresponding to the training data.

6. The data processing method according to claim 5, characterized in that, The step of fusing the text structural features and the page visual features to obtain the annotation rule information corresponding to the training data includes: Identify the text region corresponding to the text content in the annotation page, and determine the area of ​​the text region corresponding to the text region; Identify the image region corresponding to the image content on the annotation page, and determine the area of ​​the image region corresponding to the image region; The feature weights corresponding to the feature fusion are determined based on the area of ​​the text region and the area of ​​the image region. Based on the feature weights, the text structure features and the page visual features are fused to obtain the annotation rule information corresponding to the training data.

7. The data processing method according to claim 6, characterized in that, The step of determining the feature weights corresponding to the feature fusion based on the text region area and the image region area includes: Determine the total area of ​​the labeled page; Calculate the ratio between the area of ​​the text region and the total area of ​​the page to obtain the text region area ratio; Calculate the ratio between the area of ​​the image region and the total area of ​​the page to obtain the image region area ratio; The feature weights corresponding to the feature fusion are determined based on the ratio of the text region area and the ratio of the image region area.

8. The data processing method according to claim 7, characterized in that, The step of determining the feature weights corresponding to the feature fusion based on the ratio of the text region area and the ratio of the image region area includes: The target area ratio is obtained by adding the text region area ratio and the image region area ratio. Calculate the ratio between the text region area ratio and the target area ratio to obtain the relative ratio of the text region area; Based on the relative ratio of the text region areas, the page type of the labeled page is determined, and based on the page type, the feature weights corresponding to the feature fusion are determined.

9. The data processing method according to claim 8, characterized in that, The preset model includes a text encoding network and a visual encoding network. After calculating the ratio between the text region area ratio and the target area ratio to obtain the relative ratio of the text region areas, the method further includes: Based on the relative ratio of the text region areas, the text learning rate of the text encoding network and the visual learning rate of the visual encoding network are determined during the training process. The text learning rate and the visual learning rate are displayed on the task configuration page.

10. The data processing method according to claim 7, characterized in that, After calculating the ratio between the image region area and the total page area to obtain the image region area ratio, the method further includes: Calculate the text color contrast of the text region and the image color contrast of the image region on the annotation page, and calculate the text level depth of the text node and the image level depth of the image node on the annotation page. The center coordinates of the page are identified on the labeled page, and the text position centrality of the text elements in the text region from the center coordinates of the page and the image position centrality of the image elements in the image region from the center coordinates of the page are calculated. The text modality saliency score corresponding to the text region is determined based on the text region area, the text color contrast, the text hierarchy depth, and the text position centrality; and the visual modality saliency score corresponding to the image region is determined based on the image region area, the image color contrast, the image hierarchy depth, and the image position centrality. Based on the text modality saliency score and the visual modality saliency score, the text learning rate of the text encoding network and the visual learning rate of the visual encoding network are determined during the training process.

11. The data processing method according to claim 10, characterized in that, The step of determining the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training, based on the text modality saliency score and the visual modality saliency score, includes: The target saliency score is obtained by adding the text modality saliency score and the visual modality saliency score. Calculate the ratio between the text modality saliency score and the target saliency score to obtain the text modality relative ratio. Based on the relative ratio of the text modalities, the text learning rate of the text encoding network and the visual learning rate of the visual encoding network are determined during the training process.

12. The data processing method according to claim 10, characterized in that, After determining the text learning rate of the text encoding network and the visual learning rate of the visual encoding network during training based on the text modality ratio, the method further includes: When the relative ratio of the text modalities is less than the first relative ratio threshold of the text modalities, the network parameters of the text encoding network are frozen during the training process; When the relative ratio of the text modalities is greater than the second relative ratio threshold of the text modalities, the network parameters of the visual coding network are frozen during the training process.

13. The data processing method according to claim 11, characterized in that, After calculating the ratio between the text modality saliency score and the target saliency score to obtain the text modality relative ratio, the method further includes: Obtain the network capacity parameters of the preset model; The page type of the labeled page is determined based on the relative text modal values. Based on the page type, the network capacity parameters are allocated to obtain the text network capacity parameters corresponding to the text encoding network and the visual network capacity parameters corresponding to the visual encoding network.

14. The data processing method according to claim 9, characterized in that, The task configuration page also includes an update method selection control. The step of displaying the task execution page in response to a trigger operation on the task execution control includes: In response to a selection operation on the update method selection control, the target update method corresponding to the selection operation is filtered out from a preset set of update methods; In response to a trigger operation on the task execution control, the training task corresponding to the preset model is executed to display the task execution page. The task execution page includes the execution process information of the training task. The training task includes updating the training data through the target update method and the annotation rule information, and training the preset model through the updated training data, the text learning rate and the visual learning rate. When the completion of the training task is detected, the task execution result of the training task is displayed on the task execution page, and the task execution result indicates the result of training the preset model.

15. The data processing method according to claim 14, characterized in that, The step of updating the training data using the target update method and the annotation rule information includes: Obtain at least one annotation rule information obtained by the agent from accessing the annotation page at different times, and determine the rule information generation time of the annotation rule information and the training data generation time of the training data; Calculate the time difference between the rule information generation time and the training data generation time, and based on the time difference, filter out the target annotation rule information corresponding to the training data from the annotation rule information; The training data is updated using the target update method and the target annotation rule information to obtain the updated training data.

16. The data processing method according to claim 15, characterized in that, The training data includes sample data for at least one training sample. The step of updating the training data using the target update method and the target annotation rule information to obtain the updated training data includes: When the target update method is the prompt word injection method, the target annotation rule information is converted into annotation rule prompt words; The annotation rule prompts are added to the sample data to obtain the updated training data.

17. The data processing method according to claim 16, characterized in that, The step of converting the target annotation rule information into annotation rule prompt words includes: The target text structural features and target page visual features are identified from the target annotation rule information; Based on the structural features of the target text, generate text annotation rule prompts; Based on the visual features of the target page, generate visual rule prompts for the page; The text annotation rule prompts and the page visual rule prompts are merged to obtain the annotation rule prompts.

18. The data processing method according to claim 16, characterized in that, The step of updating the training data using the target update method and the target annotation rule information to obtain the updated training data includes: When the target update method is feature concatenation, feature extraction is performed on the target annotation rule information to obtain annotation rule features; Feature extraction is performed on the sample data to obtain sample features; The labeled rule features and the sample features are concatenated to obtain the updated training data.

19. The data processing method according to claim 18, characterized in that, The step of concatenating the labeled rule features and the sample features to obtain the updated training data includes: The dot product of the sample features and the annotation rule features is performed to obtain the attention weight of the sample features for the annotation rule features; The annotation rule features are weighted according to the attention weights to obtain the weighted annotation rule features; The weighted annotation rule features are concatenated with the sample features to obtain the updated training data.

20. The data processing method according to claim 14, characterized in that, After updating the training data using the target update method and the annotation rule information, the method further includes: Feature extraction is performed on the training data to obtain the original feature distribution data of the training data; Feature extraction is performed on the updated training data to obtain the updated feature distribution data of the updated training data; Calculate the feature distribution difference data between the original feature distribution data and the updated feature distribution data; The updated training data is adjusted based on the feature distribution difference data to obtain the adjusted training data, and the adjusted training data is used as the updated training data.

21. A data processing apparatus, characterized in that, include: The task configuration unit is used to display the task configuration page corresponding to the preset model. The task configuration page includes task execution controls and the page address of the annotation page corresponding to the training data of the preset model. The task configuration page is used to interact with the intelligent agent. The agent invocation unit is used to invoke the agent to access the annotation page based on the page address, so as to display the annotation rule information identified by the agent in the annotation page on the task configuration page, and the annotation rule information is used to update the training data; A task execution unit is used to display a task execution page in response to a trigger operation on the task execution control. The task execution page includes the results of training the preset model with updated training data.

22. An electronic device, characterized in that, It includes a processor and a memory, the memory storing an application program, and the processor running the application program within the memory to perform the steps of the data processing method according to any one of claims 1 to 20.

23. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the data processing method according to any one of claims 1 to 20.

24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the data processing method according to any one of claims 1 to 20.