Data processing method and device, equipment, medium and product

By optimizing the initial dataset and building a dedicated scenario operator library to process unstructured code data, the inefficiency problem in existing technologies is solved, achieving efficient code data processing and improved code generation performance.

CN121833682APending Publication Date: 2026-04-10CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD
Filing Date
2025-12-26
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently process massive amounts of unstructured code data, resulting in low data processing efficiency and excessively high technical requirements for data processing personnel.

Method used

By optimizing the initial dataset, removing redundant, non-compliant, and messy information, a dedicated scene operator library is built. This library is then used to process the target dataset and output data that conforms to the format of the large code model.

Benefits of technology

It improves the processing efficiency of unstructured code data, reduces the technical requirements for data processing personnel, ensures the validity and standardization of data, and enhances code generation performance and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833682A_ABST
    Figure CN121833682A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, equipment, a medium and a product. The method comprises the following steps: determining an initial data set, and performing data optimization on the initial data set to obtain a target data set; wherein the target data set is a data set which does not contain redundancy, violation, mess and other information; determining a data format requirement of the code large model, and determining a scene operator library corresponding to the code large model based on the data format requirement; wherein the data format requirement is used for indicating a data format requirement in an application scene corresponding to the code large model; and processing the target data set through the scene operator library to obtain a processed data set, and adjusting the code large model through the processed data set to obtain a target code large model.
Need to check novelty before this filing date? Find Prior Art