Interactive big model domain preference data governance method and related device
By employing the interactive large-scale model domain preference data governance method of the Grado platform, the challenge of aligning the value preferences of professionals in specific domains with large-scale language models has been solved. This has enabled efficient and flexible preference data collection and governance, thereby enhancing the credibility and controllability of the model in specific domains.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, large-scale language models struggle to accurately align the value preferences and judgment criteria of professionals in specific domain tasks. Furthermore, the lack of efficient methods for collecting and managing preference data results in long feedback information transmission paths, severe signal loss, limited interaction methods, and insufficient flexibility.
It adopts an interactive large-model domain preference data governance method, and achieves deep integration of front-end display, back-end processing and user feedback through the Grado platform. It supports multimodal information input, displays inference results and allows for manual correction, supports multiple human-computer interaction methods, and realizes automated data collection and intelligent governance.
It improves the efficiency and consistency of preference data collection, realizes the automated collection and intelligent management of high-quality preference data, and enhances the credibility and controllability of the model in specific fields.
Smart Images

Figure CN121882248A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and data governance technology, and relates to an interactive large-scale model domain preference data governance method and related apparatus. Background Technology
[0002] In recent years, large-scale language models have demonstrated remarkable generalization capabilities in multi-task learning and complex semantic understanding, achieving cross-domain language, knowledge, and reasoning transfer through pre-training and fine-tuning on large-scale general data. However, when faced with specific domain tasks, general-purpose models often struggle to accurately align with the value preferences and judgment standards of professionals, resulting in a significant gap between "general intelligence" and "domain intelligence." To address this issue, preference learning has gradually become a key technological path driving the evolution of large models from "omnipotent brains" to "industry experts." Through mechanisms such as human feedback reinforcement learning, models can learn human value orientations, decision-making logic, and language styles in different contexts, thereby enhancing their credibility and controllability in specific domains.
[0003] However, domain preference data, as the core foundation for achieving refined alignment and specialization enhancement of large models, has long faced challenges such as data scarcity, complex governance, and high collection costs. Existing methods for collecting and governing preference data generally suffer from the following problems: 1. Lack of overall human-computer feedback design. Current systems often separate front-end display, back-end processing, and user feedback, resulting in long feedback information transmission paths, severe signal loss, and difficulty in forming a high-quality, traceable preference data loop; 2. Limited interaction methods and insufficient flexibility. Traditional preference collection interfaces support relatively simple feedback formats, which cannot adapt to the needs of different task types (such as sorting, comparison, error correction, semantic refinement, etc.) and lack effective support for the granularity of preference expression and semantic consistency; Therefore, there is an urgent need to build an efficient and interactive large-scale model domain preference data governance technology to achieve automated collection and intelligent governance of high-quality preference data, and to provide systematic support for the preference alignment and professional evolution of large-scale domain models. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide an interactive method and related apparatus for governing domain preference data in large models. This method and related apparatus have the characteristics of high human-computer interaction capability and high flexibility.
[0005] To achieve the above objectives, this invention discloses an interactive large-scale model domain preference data governance method, comprising: Input the information that needs to be fed into the model; The model performs reasoning based on the input information, obtains the reasoning result, and displays the reasoning result; The reasoning result is manually corrected, and the corrected result and the content of the correction are displayed.
[0006] Furthermore, the process of inputting information into the model is as follows: Enter the information that needs to be input into the model in the chat area. The information that needs to be input into the model includes user information, text commands, images and videos.
[0007] Furthermore, the reasoning results and the corrected results are displayed in the image display area.
[0008] Furthermore, it also includes: The reasoning results and the corrected results are saved.
[0009] Furthermore, the inference result is an image with detection boxes, used to identify objects in the image. For objects that cannot be detected in the image, the location information of the undetected objects is directly input, and detection boxes are marked on the undetected objects according to the input location information.
[0010] Furthermore, before the model performs inference based on the input information, it also includes: The input information is automatically cleaned and stored in a structured manner.
[0011] This invention discloses a Grado platform, comprising: The front-end interaction layer is used to input the information that needs to be input into the model, display the reasoning results, and display the corrected results and the corrected content. The backend processing layer is used by the model to perform reasoning based on the input information and obtain the reasoning result.
[0012] The user feedback layer is used to manually correct the inference results.
[0013] Furthermore, the backend processing layer is also used to save the inference results and the corrected results.
[0014] The present invention discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the interactive large model domain preference data governance method.
[0015] The present invention discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the interactive large model domain preference data governance method.
[0016] The present invention has the following beneficial effects: The interactive large-scale model domain preference data governance method and related device described in this invention, in specific operation, the model performs reasoning based on the input information, obtains reasoning results, displays the reasoning results, manually corrects the reasoning results, and displays the corrected results and the content of the correction. It has the characteristics of high human-computer interaction capability and flexibility, and is extremely practical. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a visualization of the running effect after user input in Example 2; Figure 3 This is a visualization of user feedback in Example 2; Figure 4 Visualization of the user interface; Figure 5 This is a schematic diagram illustrating the operating principle of the platform in this invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0021] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0022] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.
[0023] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0024] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0026] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.
[0027] Example 1 refer to Figure 1 The interactive large-model domain preference data governance method of the present invention includes the following steps: 1) Enter the information that needs to be input into the model in the chat area. The information that needs to be input into the model includes user information, text commands, images and videos, etc., which are required for multimodal reasoning. At the same time, the dialogue content of the model response is displayed in the chat area.
[0028] 2) The model performs inference based on the information input into the model, obtains the inference result, and displays the inference result in the image display area; at the same time, the inference result is manually corrected, and the manually corrected result is displayed in the image display area.
[0029] For example, the labeled image generated by the model is displayed in the upper part of the image display area, in the "Qwen Image" area. If the labeled image generated by the model has errors, the annotation is manually corrected, and then the manually corrected image is displayed in the lower part of the image display area, in the "Human Image" area.
[0030] 3) Users can export and upload annotations through the control area, and input commands such as upload, submit, retry, and clear history.
[0031] It should be noted that this invention is built on the Grado platform. Users can upload images via "Upload" or quickly load pre-generated results via the "Cache" button in offline mode. Users can enter task descriptions, such as target category, detection conditions, or semantic descriptions, via the "Input" box. After clicking "Submit", the system automatically calls the backend inference module for processing and presents the model output results in a visual form in the "Model" window.
[0032] In addition, in this embodiment, the process of manually correcting the annotations is as follows: Coordinates generated by user annotation operations (such as [x1, y1, x2, y2]) are captured in real time and stored in a formatted manner; The system automatically associates metadata such as task ID, timestamp, and user ID to achieve data traceability. It adopts a four-layer hierarchical structure of task-sample-feedback-annotation to achieve indexed and traceable management of preference data. It supports multiple rounds of modification and undoing to ensure the accuracy and repeatability of annotation results; The annotation information is synchronized and used to dynamically update the visualization interface, enabling real-time comparison between model output and human feedback.
[0033] In this embodiment, when inputting data, the data needs to be standardized, including field definitions, data types, annotation coordinate formats, and label description specifications. In addition, all samples will be automatically validated for format and semantic consistency before archiving to ensure that data between different tasks can be integrated, retrieved, and transferred.
[0034] This embodiment also includes: utilizing API-level data calls and reuse to facilitate preference alignment and incremental optimization across different domains, tasks, or models based on a shared data structure, thereby forming a sustainably evolving preference data ecosystem.
[0035] This invention has the following characteristics: This invention leverages the Grado platform to achieve deep integration of front-end display, back-end processing, and user feedback, creating an end-to-end preference information collection platform that enables visualization, real-time processing, and automation of the feedback process, significantly improving the efficiency and consistency of preference data collection.
[0036] This invention supports multiple data types such as images and text, as well as various human-computer interaction methods such as evaluation and error correction, allowing users to express preference information in a multi-dimensional and fine-grained manner, thereby improving the freedom and accuracy of data expression.
[0037] This invention establishes a unified data format standard and metadata description system, automatically cleans and structures the collected preference data, and ensures the traceability, compatibility and high-quality reusability of the data in subsequent modeling, alignment and optimization stages.
[0038] Example 2 In this embodiment, the background model is set to Qwen2.5-VL to perform fine-grained annotation for the visual localization task of a multimodal large model. The specific process is as follows: 1) User input and model inference; 11) Users input information through the Input field, which is the main entry point for user interaction with the model and supports the following operations: 111) Text input: The user inputs instructions, such as detecting all objects in an image and returning their locations and labels.
[0039] 112) After uploading the file, add text, for example, upload an image and enter the location of all the aircraft in the image.
[0040] 12) After clicking "Submit," the memory model will perform inference, refer to... Figure 2 The specific process is as follows: 121) Submit text instructions and images to the model; 122) Display text commands and the model's responses in the left-hand dialog area; 123) Display the image with the detection box in the QwenImage area on the right, which is the model output result; 124) Display the JSON format annotation results generated by the model below the input box.
[0041] 2) User annotation and feedback; For areas detected by the model but with insufficiently precise boundary annotations, users can manually make fine-grained adjustments based on the detected coordinates.
[0042] For targets that the model fails to detect, the user can directly enter the location of the corresponding target in the text box.
[0043] After optimizing the annotations, click the "Upload Refined" button for reference. Figure 3 ,but: 21) Automatically read the latest corrected annotation files and corresponding images.
[0044] 22) The corrected image will be displayed in the Human Image area on the right side of the updated interface.
[0045] 23) Save the corrected data to: / mnt / qwen / Qwen2.5-VL / results / positive / <timestamp> / ├── image.jpg └── response.txt 3) Save the results; The specific process of step 3) is as follows: 31) Automatic saving: The detection results generated by the model (Qwen labeled images + labeling information) will be automatically saved to the following path structure: results / └── negative / ├── 20250607_091809 / │ ├── image.jpg # Model detection image │ └── response.txt # JSON format annotation └── ... 32) One-click export of annotations: After clicking the "Export to LabelMe" button, then: Automatically save the current model detection results to: / mnt / qwen / Qwen2.5-VL / export_to_labelme / <timestamp> / ├── predicted.jpg └── predicted.json Example 3 refer to Figure 5 This embodiment discloses a Grado platform, including: The front-end interaction layer is used for task display, model output visualization, and user input collection, providing multiple types of task interfaces and feedback entry points; The backend processing layer is used for model invocation, data caching and storage, feedback recording and execution of reinforcement learning modules, realizing data governance, storage and optimization; The user feedback layer is used by users to evaluate, correct, or re-label the model's behavior through a graphical interface, generating positive and negative sample pairs to provide direct supervision signals for preference modeling.
[0046] Example 4 A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the deep learning-based generator PT fast-blowout early warning method. For example, the method includes: collecting historical runtime sequence data of the generator PT fast-blowout and its corresponding future time-series operating states; constructing a historical training dataset based on the historical runtime sequence data and its corresponding future time-series operating states; training a deep learning model using the historical training dataset to obtain an early warning model for generator PT fast-blowout; deploying the early warning model on a cloud server, collecting real-time runtime sequence data through the cloud server, and inputting the real-time runtime sequence data into the early warning model to obtain data recognition results; generating generator PT fast-blowout early warning information based on the data recognition results; and issuing an early warning based on the generator PT fast-blowout early warning information. The memory may include main memory, such as high-speed random access memory (RAM), or non-volatile memory, such as at least one disk storage device. The processor, network interface, and memory are interconnected via an internal bus, which may be an industry-standard architecture bus, a peripheral component interconnection standard bus, or an extended industry-standard architecture bus. The bus can be categorized as an address bus, data bus, or control bus. The memory stores programs; specifically, the program may include program code, which includes computer operation instructions. The memory may include main memory and non-volatile memory, and provides instructions and data to the processor.
[0047] Example 5 A computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the deep learning-based generator PT fast-blowout warning method. For example, the method includes: collecting historical runtime sequence data of the generator PT fast-blowout and its corresponding future time-series generator PT fast-blowout operating status; constructing a historical training dataset based on the historical runtime sequence data and its corresponding future time-series generator PT fast-blowout operating status; training a deep learning model using the historical training dataset to obtain a warning model for generator PT fast-blowout warning; deploying the warning model on a cloud server, collecting real-time runtime sequence data through the cloud server, and inputting the real-time runtime sequence data into the warning model to obtain data recognition results; generating generator PT fast-blowout warning information based on the data recognition results; and issuing a warning based on the generator PT fast-blowout warning information. Specifically, the computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory may include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.
[0048] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0049] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0050] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0051] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0052] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and disclosure of the invention. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.
[0053] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
[0054] The above description is merely a preferred embodiment of the present invention and does not constitute any limitation on the present invention. Any simple modifications, alterations, or equivalent structural changes made to the above embodiments based on the technical essence of the present invention shall still fall within the protection scope of the present invention.
Claims
1. An interactive large model domain preference data governance method, characterized in that, include: Input the information that needs to be fed into the model; The model performs reasoning based on the input information, obtains the reasoning result, and displays the reasoning result; The reasoning result is manually corrected, and the corrected result and the content of the correction are displayed.
2. The interactive, large model domain preference data governance method of claim 1, wherein, The process of inputting information into the model is as follows: Enter the information that needs to be input into the model in the chat area. The information that needs to be input into the model includes user information, text commands, images and videos.
3. The interactive, large model domain preference data governance method of claim 1, wherein, The reasoning results and the corrected results are displayed in the image display area.
4. The interactive, large model domain preference data governance method of claim 1, wherein, Also includes: The reasoning results and the corrected results are saved.
5. The interactive, large model domain preference data governance method of claim 1, wherein, The reasoning result is an image with detection boxes, used to identify objects in the image. For objects that cannot be detected in the image, the location information of the undetected objects is directly input, and detection boxes are marked on the undetected objects according to the input location information.
6. The interactive, large model domain preference data governance method of claim 1, wherein, Before the model performs inference based on the input information, it also includes: The input information is automatically cleaned and stored in a structured manner.
7. A Grado platform, characterized in that, include: The front-end interaction layer is used to input the information that needs to be input into the model, display the reasoning results, and display the corrected results and the corrected content. The backend processing layer is used by the model to perform inference based on the input information and obtain the inference result. The user feedback layer is used to manually correct the inference results.
8. The Grado platform according to claim 7, characterized in that, The backend processing layer is also used to save the inference results and the corrected results.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the interactive large model domain preference data governance method as described in any one of claims 1-6.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the interactive large-model domain preference data governance method as described in any one of claims 1-6.