Data processing method and device, equipment, medium and product
By optimizing the initial dataset and building a dedicated scenario operator library to process unstructured code data, the inefficiency problem in existing technologies is solved, achieving efficient code data processing and improved code generation performance.
CN121833682APending Publication Date: 2026-04-10CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by
Patent Information
- Application Number
- CN202511993379.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-10
AI Technical Summary
Technical Problem
Existing technologies struggle to efficiently process massive amounts of unstructured code data, resulting in low data processing efficiency and excessively high technical requirements for data processing personnel.
Method used
By optimizing the initial dataset, removing redundant, non-compliant, and messy information, a dedicated scene operator library is built. This library is then used to process the target dataset and output data that conforms to the format of the large code model.
Benefits of technology
It improves the processing efficiency of unstructured code data, reduces the technical requirements for data processing personnel, ensures the validity and standardization of data, and enhances code generation performance and adaptability.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN121833682A_ABST
Abstract
The embodiment of the invention discloses a data processing method and device, equipment, a medium and a product. The method comprises the following steps: determining an initial data set, and performing data optimization on the initial data set to obtain a target data set; wherein the target data set is a data set which does not contain redundancy, violation, mess and other information; determining a data format requirement of the code large model, and determining a scene operator library corresponding to the code large model based on the data format requirement; wherein the data format requirement is used for indicating a data format requirement in an application scene corresponding to the code large model; and processing the target data set through the scene operator library to obtain a processed data set, and adjusting the code large model through the processed data set to obtain a target code large model.
Need to check novelty before this filing date? Find Prior Art