Data processing method and electronic equipment

By receiving the storage path and processing scope of the target document, using a large language model to parse the document to generate structured data, and combining predefined field mapping relationships and permission control, it solves the problems of low data accuracy and difficult version tracing in CBB data management, realizes efficient data management and intelligent analysis, and supports server R&D decisions.

CN120743886AActive Publication Date: 2025-10-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511225833.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-10-03
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

The existing CBB data management method relies on a hybrid mode of PDF and Excel, which leads to repeated data entry, confusing version information, and difficulty in classification and statistics. It cannot effectively support multi-version traceability and multi-dimensional statistical analysis, limiting the application value of data in R&D decision-making.

Method used

By receiving the storage path and processing scope of the target document, using a large language model to parse the document to generate structured data, combined with predefined field mapping relationships and permission control, it realizes intelligent classification and storage of data, and supports full version traceability and multi-dimensional statistical analysis.

Benefits of technology

It reduces manual operation errors, improves the accuracy and security of data management, supports full version traceability, and provides data support for server R&D decisions through intelligent classification and statistical analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743886A_ABST
    Figure CN120743886A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and electronic equipment, and relates to the technical field of data processing.The method comprises the steps that a storage path and a to-be-processed range of a target document are received, the document is analyzed through a large model, and structured data containing information such as codes, version numbers, items and multi-level classification of hardware modules to be developed are generated; after verification, various kinds of information are respectively stored and associated, and at the same time, systematic management and control are realized by combining authority control and version management. The problems of low data accuracy, repeated input, difficulty in version tracing, lack of multi-dimensional classification statistics and the like caused by dependence on a document and table mixed mode in data management of an existing hardware module to be developed can be solved, so that manual operation errors are reduced, the data management safety is improved, full-version tracing is supported, and the development efficiency is improved. And data support is provided for the research and development decision of the server through intelligent classification and statistical analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data processing method and electronic equipment. Background Art

[0002] In the field of server hardware R&D, Common Building Blocks (CBBs) are a core element for improving R&D efficiency, reducing development costs, and enhancing system flexibility and product consistency. They are widely used in key aspects such as motherboard design, storage architecture, I / O modules, and GPU platforms. With the diversification of server products and the continuous evolution of customer needs, the number and versions of CBB components have increased exponentially, and their management and analysis requirements have become increasingly complex. Related technologies have established a preliminary CBB data management process by recording CBB design information in PDF documents, assisting with data organization in Excel spreadsheets, and managing file paths using the SVN version control system. However, existing CBB data management methods directly utilize a hybrid model of PDF and Excel, without implementing automatic data parsing and structured storage. This can lead to problems such as duplicate data entry, confusing version information, and difficulty in classification and statistical analysis. Existing methods cannot effectively support multi-version traceability and multi-dimensional statistical analysis, limiting the application value of CBB data in R&D decision-making and failing to meet the deep-seated needs of modern server hardware development for data governance and intelligent management. Summary of the Invention

[0003] The present disclosure provides a data processing method and electronic device, the main purpose of which is to solve the problem of chaotic CBB data management in related technologies.

[0004] According to a first aspect of the present disclosure, there is provided a data processing method, comprising: Receive the storage path information of the target document in the file version library and the range of data to be processed input through the human-computer interaction interface, and obtain the target document based on the storage path information; Calling the data extraction interface to extract data from the document content within the data range to be processed in the target document, and inputting the extracted data into the pre-trained model so as to format the extracted data based on the preset prompt words to obtain the preset format data; Based on the predefined field mapping relationship, the preset format data is integrated into the preset target format data, where the target format data at least includes the code, version number, project to which it belongs, and multi-level classification information of the hardware module to be developed; Output and display the target format data for verification; In response to the target format data passing the verification, the code, version number, project information and multi-level classification information of the hardware module to be developed included in the verified target format data are stored separately and associated.

[0005] According to a second aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the data processing method described in the first aspect.

[0006] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the data processing method described in the first aspect.

[0007] According to a fourth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the data processing method as described in the first aspect above.

[0008] The present disclosure provides a data processing method and electronic device. The present disclosure receives the storage path and scope to be processed of a target PDF document, utilizes a large model to parse the document, generates structured data including CBB code, version number, project to which it belongs, and multi-level classification and other information, and stores and associates each type of information separately after verification. At the same time, a technical solution for realizing systematic management and control is implemented by combining authority control and version management. The method can solve the problems of low data accuracy, repeated entry, difficulty in version tracing, and lack of multi-dimensional classification statistics caused by reliance on a hybrid mode of PDF and Excel in existing CBB data management, thereby achieving the technical effect of reducing manual operation errors, improving data management security, supporting full version tracing, and providing data support for server R&D decisions through intelligent classification and statistical analysis.

[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure. Figure 1 A flowchart of a data processing method provided by an embodiment of the present disclosure; Figure 2A schematic diagram of the structure of a data processing device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0011] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0012] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0013] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the data processing method depends, the specific application environment architecture or specific hardware architecture is described here.

[0014] The data processing method and electronic device according to the embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0015] Figure 1 A flowchart of a data processing method provided by an embodiment of the present disclosure.

[0016] like Figure 1 As shown, the method comprises the following steps: Step 101: receiving storage path information of a target document in a file version library and a range of data to be processed inputted through a human-computer interaction interface, and acquiring the target document based on the storage path information.

[0017] In the embodiments of the present disclosure, during the CBB data management process of server development, when new CBB data needs to be processed, relevant information is first received through a human-computer interaction interface. This human-computer interaction interface is a visual interface for the CBB manager to operate. This interface includes a "Add" action entry. By triggering this action, the manager enters the information input phase. During this phase, the manager needs to enter information including the target document's storage path in the file version repository and the data range to be processed. In some possible embodiments, the target document specifically refers to a PDF document recording information related to CBB development. This document is generated after consensus is reached on the CBB discussion results and contains key information such as basic information required for CBB development, detailed descriptions of customer requirements, discussion details, the default storage location for the final CBB development file, a unique CBB code, and a version number. It is stored in the file version repository along with the final CBB development file. The file version repository utilizes the Subversion (SVN) version control system. Its core function is to specify the repository location, ensuring that the development team can collaboratively manage version changes for code and related files. Therefore, the storage path information refers to the PDF document's network address in SVN, which allows the target document to be accurately located. The range of data to be processed refers to the list of page numbers where the tables whose data needs to be extracted from the target PDF document are located. Specifying this range can effectively reduce the amount of text data that needs to be processed during the subsequent data parsing process, thereby improving processing efficiency. After the person in charge completes the input and confirmation of the above information, the system's background module will receive and identify the input information, and then automatically establish a connection with the SVN version library based on the acquired storage path information, locate the target PDF document according to the path, and perform a download operation to complete the acquisition of the target document. This step ensures the accuracy and pertinence of the target documents based on which subsequent data processing is based through clear information input and an automated document acquisition mechanism, laying a solid foundation for the intelligent filling and statistical classification process of the entire CBB data. Its beneficial effect is that by accurately receiving and acquiring target information, it reduces the tedious operations of manual search and acquisition of documents. At the same time, by specifying the range of data to be processed, it reduces the interference of invalid data, thereby improving the efficiency and accuracy of the overall process.

[0018] Step 102: call a data extraction interface to extract data from the document content within the data to be processed range in the target document, and input the extracted data into a pre-trained model so as to format the extracted data based on preset prompt words to obtain preset format data.

[0019] In an embodiment of the present disclosure, after obtaining the target document, the data extraction and format processing stage of the document content is entered. The target document here is a PDF file containing CBB-related information, and the range of data to be processed has been clearly defined as the page number list of the table where the data needs to be extracted. The system will call the data extraction interface, which can be but is not limited to the Tabula library based on Python to implement the function, and its core function is to accurately extract the table data within the specified page number range in the target PDF document. The Tabula library can recognize the table structure in PDF and convert the table content that originally existed in an unstructured form into structured data, such as text or table objects organized by rows and columns, to ensure that the extracted data retains the original row and column correspondence and content integrity of the table, and provide standardized basic data for subsequent processing.

[0020] After extracting the structured data, the system inputs it into a pre-trained model, which can be a preset large language model. The large language model adopts a local deployment mode, which can effectively ensure the security of the company's CBB data and avoid the risk of data leakage. While inputting data, the system will pass preset prompt words to the large language model. The prompt words are specially designed to explicitly require the model to format the input structured table data. For example, the prompt words can be set to "You are a data processing expert, please convert the input table data into a standard JSON format. Requirements: 1. Each table is processed separately 2. The original column name is retained 3. The numerical type is automatically identified 4. Missing values ​​are filled with null". After receiving the structured data and the preset prompt words, the large language model will process the data according to the requirements of the prompt words based on its understanding of natural language and data formats, and convert the structured data into the preset JSON format data. This ensures that the output data fields are clear and the format is uniform, which facilitates the integration and use of the data by subsequent systems.

[0021] This step achieves efficient conversion from PDF tables to structured data and then to preset format data through the synergy of automated data extraction interfaces and large language models. Its beneficial effect is that it reduces manual extraction and format conversion operations, lowers the risk of data errors caused by human operations, and at the same time, the unified data format facilitates subsequent data processing links, improving the efficiency and accuracy of the overall process.

[0022] Step 103 : Based on a predefined field mapping relationship, the preset format data is integrated into preset target format data, wherein the target format data includes at least the code, version number, project and multi-level classification information of the hardware module to be developed.

[0023] In the embodiment of the present disclosure, the process of integrating preset format data to obtain target format data first relies on predefined field mapping relationships. This mapping relationship is pre-set by the system during the design phase based on the CBB data management requirements. Specifically, it is the Chinese-English correspondence between the CBB data fields stored in the database and the system's internal field definitions. For example, for the original field "CBB code", the mapping relationship will correspond to the "code" field defined within the system; "first-level classification" corresponds to "firstCat", "second-level classification" corresponds to "secondCat", "third-level classification" corresponds to "thirdCat", "version number" corresponds to "version", "belonging project" corresponds to "project", etc., and also includes the mapping of custom fields under each category, such as "CPU platform" corresponds to "platform", "packaging method" corresponds to "package", etc. These mapping relationships ensure the accurate conversion from original data to system-recognizable data.

[0024] When the system backend receives data in a preset format (usually in JSON format, containing information such as "CBB code," "version number," "project," "first-level classification," "second-level classification," "third-level classification," and various custom fields) processed by a large language model, it invokes the preset integration logic to convert and reorganize the data based on the predefined field mapping relationships. During the integration process, not only are the original field names replaced with the system's internal field names, but structured data containing "label" and "value" are also constructed for each field. The "label" retains the original field's Chinese name, and the "value" is the specific data value corresponding to the field. For example, the "CBB code": "CBB_motherboard_CPU_AMD_001_V1.0" in the original JSON is integrated into "code": {"label": "CBB code", "value": "CBB_motherboard_CPU_XXX_001_V1.0"} (XXX is the manufacturer name). Similarly, the final target format data includes the code, version number, project (such as "Project 1", "Project 2", etc.) of the hardware module to be developed (i.e., CBB), and multi-level classification information (first-level classification such as "motherboard", second-level classification such as "CPU", third-level classification such as "XXX", etc.), as well as the custom field data corresponding to each classification.

[0025] Regarding ensuring the accuracy and efficiency of data processing by the large language model, for accuracy, on the one hand, clear pre-set prompts (such as requiring tabular data to be converted to a standard JSON format, retaining original column names, automatically identifying numeric types, and filling missing values ​​with null) guide the model to process data according to specifications. On the other hand, the model processing results are manually verified by the CBB manager, and data with parsing errors is manually modified to ensure the accuracy of the final data. Regarding efficiency, by limiting the scope of the data to be processed (i.e., specifying the page number list where the table is located), the amount of text data the model needs to process is reduced; a locally deployed large language model is used to reduce data transmission latency; and Python's Tabula library is used to efficiently extract tabular data, improving the processing speed before the data is input into the model, thereby improving overall processing efficiency.

[0026] This step achieves standardized data integration through predefined field mapping relationships, ensuring the system's unified identification and management of CBB data. Its beneficial effect is that it reduces processing obstacles caused by inconsistent data formats, lays a standardized foundation for subsequent data storage, statistics and classification, and improves the consistency and reliability of data processing.

[0027] Step 104: Output and display the target format data for verification.

[0028] In an embodiment of the present disclosure, the target format data generated in step 103 is reorganized and output to a front-end web interface developed using the Spring+Vue front-end and back-end separation model for display and verification by the CBB manager. During the integration process, the original JSON format data is converted into a structured display format containing "label" and "value" fields, combining the pre-stored Chinese-English correspondence between each field in the database. The "label" corresponds to the Chinese name of the field, and the "value" corresponds to the specific data value obtained by parsing. For example, the "code" field is displayed as {"label": "CBB code", "value": "CBB_motherboard_CPU_xxx (xxx is the manufacturer's name)_001_V1.0"}, and the "firstCat" field is displayed as {"label": "first-level category", "value": "motherboard"}. This format allows the CBB manager to more intuitively understand the meaning of each field and the corresponding data. The front-end webpage, the core interface for human-computer interaction, clearly displays all target format data, including the hardware module's code, version number, project affiliation, and multi-level classification information, ensuring that the person in charge can fully review the parsing results. This display and verification process is necessary because, while intelligent fill-in is possible when parsing PDF table data using a large model, the parsing results may be biased due to factors such as document format complexity and ambiguous table content. Therefore, the CBB person in charge must verify the accuracy of each field on the front-end webpage. If a field's "value" discrepancies with the actual content in "Table 1: CBB Property Table" in the PDF document (for example, "Number of CPUs" is parsed as "2S" when it should be "1S"), they can manually modify the data directly on the webpage until all data is confirmed to be correct. This verification process is critical for ensuring the reliability of subsequent data storage and management. This manual intervention ensures that the target format data is consistent with the original document information, laying the foundation for standardized CBB data management.

[0029] Step 105 : In response to the target format data passing verification, the code, version number, project information and multi-level classification information of the hardware module to be developed included in the verified target format data are stored separately and associated.

[0030] In an embodiment of the present disclosure, after the target format data passes verification by the CBB manager, the system will respond to this confirmation operation and classify and store the coding, version number, project information, and multi-level classification information of the hardware module to be developed (i.e., CBB) contained in the verified target format data, and realize the association between each piece of information through an identifier (ID). Specifically, the system backend will store different types of information in corresponding database tables: among them, the CBB code (such as "CBB_Motherboard_CPU_xxx (xxx is the manufacturer's name)_001_V1.0", used to uniquely identify the CBB) and version number (such as "V1.0", recording its version iteration status) will be stored in the CBB master data table as core information to support the basic information management of a single CBB; the project information (such as "Project 1", "Project 2", etc., identifying the customer's customized demand project corresponding to the CBB) will be stored in the project information table; multi-level classification information (including first-level classification such as "Motherboard", second-level classification such as "CPU", and third-level classification such as "xxx (xxx is the manufacturer's name)", reflecting the technical attributes and hierarchical affiliation of the CBB) will be stored in the classification information table.

[0031] During storage, the system automatically verifies whether this information already exists in the corresponding table. If the "General" item already exists in the project information table, the system directly retrieves its corresponding project ID. If not, the item is added to the project information table and a new project ID is generated. The storage logic for category information is consistent with this. For example, if the third-level category "xxx" (where xxx is the manufacturer's name) already exists in the category information table, its category ID is directly retrieved; otherwise, a new ID is generated. Ultimately, these project and category IDs, generated through verification or addition, are populated into the CBB master data table and linked to the CBB code and version number, forming a "code-version number-project ID-category ID" relationship. This ensures logical coherence between the various pieces of information and provides the data foundation for subsequent version traceability, category statistics, and other functions. This separate table storage and linkage approach ensures data standardization and independence while enabling information integration through ID linkage, avoiding data redundancy and improving the efficiency and accuracy of CBB data management.

[0032] The present disclosure provides a data processing method, which receives the storage path and scope to be processed of a target PDF document, uses a large model to parse the document to generate structured data containing information such as CBB code, version number, project and multi-level classification, and stores and associates each type of information separately after verification. At the same time, a technical solution for realizing systematic management and control is implemented by combining authority control and version management. The method can solve the problems of low data accuracy, repeated entry, difficulty in version tracing and lack of multi-dimensional classification statistics caused by reliance on a hybrid mode of PDF and Excel in existing CBB data management, thereby achieving the technical effect of reducing manual operation errors, improving data management security, supporting full version tracing, and providing data support for server R&D decisions through intelligent classification and statistical analysis.

[0033] In the embodiment of the present disclosure, for the operation of "storing the code, version number, project information and multi-level classification information of the hardware module to be developed included in the verified target format data separately and associating them", the specific implementation methods are diverse. For the sake of clarity, the following includes but is not limited to some implementation methods: storing the code and version number of the hardware module to be developed as a version record in the first data table; querying the second data table to determine whether the project information exists, if not, inserting a new record and obtaining the unique identifier of the project information, if yes, directly obtaining the unique identifier of the project information; querying the third data table to determine whether the multi-level classification information exists, if not, inserting a new record and obtaining the unique identifier of the multi-level classification information, if yes, directly obtaining the unique identifier of the multi-level classification information; associating the unique identifier of the version record, the unique identifier of the project information and the unique identifier of the multi-level classification information and storing them in a relational table to complete the association.

[0034] Specifically, the hardware module to be developed is a CBB (Common Building Block) used in server R&D. Its code (e.g., "CBB_Motherboard_CPU_xxx (xxx is the manufacturer's name)_001_V1.0," which uniquely identifies the CBB) and version number (e.g., "V1.0," which records its version iteration status) serve as core information and are stored in a first data table (the CBB master data table). This table primarily records the core identification and version information of a single CBB and serves as the foundation for linking other information. Project information is stored in a second data table (the project information table), which centrally manages all project names and related attributes to ensure the standardization and uniqueness of project information. Multi-level classification information (including first-level categories such as "Motherboard" and "Infrastructure," second-level categories such as "CPU" and "Storage," and third-level categories such as "xxx (xxx is the manufacturer's name)" and "SAS," reflecting the CBB's technical attributes and hierarchical affiliation) is stored in a third data table (the classification information table). This table records classification names and corresponding relationships by level, supporting hierarchical management of CBBs.

[0035] The three tables are linked through predetermined identification information (i.e., each table's unique identifier ID). When storing data, the system first verifies whether the project information to be stored already exists in the second data table. If so, the corresponding unique project ID is directly retrieved; if not, the project information is added to the second data table and a unique project ID is automatically generated. For the multi-level classification information in the third data table, the system similarly verifies whether the corresponding classification information already exists. If so, the classification ID is retrieved; otherwise, a new classification ID is created and generated. The system then populates the retrieved project and classification IDs into the first data table (the CBB master data table) and binds them to the CBB's code and version number, forming a "code-version number-project ID-classification ID" association. For example, for CBB data with the code "CBB_Motherboard_CPU_xxx (xxx is the manufacturer's name)_001_V1.0", version number "V1.0", project "General", and multi-level classification "Motherboard (Level 1)-CPU (Level 2)-xxx (xxx is the manufacturer's name) (Level 3)", after its code and version number are stored in the first data table, the project ID corresponding to the "General" project in the second data table and the classification ID corresponding to the "Motherboard-CPU-xxx (xxx is the manufacturer's name)" classification in the third data table will be recorded in the corresponding fields of the first data table, thereby achieving precise association between the three tables and providing structured data support for subsequent version tracing, classification statistics and other functions.

[0036] The first data table is the CBB master data table, which stores the code and version number of the hardware module to be developed (CBB). The second data table is the project information table, which stores project information. The third data table is the classification information table, which stores multi-level classification information. The predetermined identification information is the unique identifier (ID) recorded in each table. The target data table is the first data table (CBB master data table). When the verified target format data enters the storage phase, the system first verifies the second data table (project information table). If the project information to be stored (such as "General") already exists in the project information table, the system directly obtains the corresponding unique project ID. If not, the system adds the project information to the project information table and automatically generates a unique project ID. Similarly, the system verifies whether the multi-level classification information to be stored already exists in the third data table (classification information table). If so, the corresponding classification ID is obtained. If not, a new classification ID is added and generated. The system then stores the obtained project ID and classification ID (i.e., the predetermined identification information) in the corresponding fields of the first data table (CBB master data table). By storing project and category IDs in the CBB master data table (target data table), the CBB master data table can be directly linked to the project information table and category information table through these IDs. When querying the project details or category hierarchy of a specific CBB, the system can quickly locate the corresponding records in the project information table and category information table using the project ID and category ID stored in the CBB master data table. This enables linked query and management of data across multiple tables, ensuring the accuracy and efficiency of CBB data linkage.

[0037] In the embodiment of the present disclosure, in addition to the aforementioned contents, the data processing method also includes other specific implementation steps. In order to clearly present these components, the relevant specific implementation methods are described in detail below: before storing the code, version number, project information and multi-level classification information of the hardware module to be developed included in the verified target format data, the identity information of the currently logged-in user is obtained; according to the code of the hardware module to be developed, the information of the person in charge of the module is queried from the permission configuration table; it is verified whether the current user is the person in charge of the module or has administrator authority; if the verification fails, the storage process is terminated and a prompt message of insufficient authority is returned.

[0038] Specifically, before storing the verified target format data, a strict permission verification process must be implemented to ensure data security and standardized operations. First, the system automatically obtains the identity information of the currently logged-in user. This identity information typically includes a unique user identifier (such as user ID, username, etc.) to accurately identify the operating subject, which is the basis for permission verification. The hardware module to be developed is the CBB, and its code is unique, such as "CBB_Motherboard_CPU_AMD_001_V1.0." This code not only identifies the type and version of the CBB, but also serves as a key index for querying permission information. The system has a pre-set permission configuration table that stores the association information between each CBB code and the corresponding person in charge, including the person in charge's user ID and other content. The person in charge's information can be directly located through the CBB code. After obtaining the current user's identity information and retrieving the CBB's responsible person information, the system initiates a permissions verification mechanism: it compares the current user's identity information with the retrieved responsible person information to determine whether the current user is the designated responsible person for the CBB. The system also checks whether the current user has administrator privileges, a special privilege typically held by system administrators that allows the user to operate on all CBB data. If the current user is neither the responsible person for the CBB nor possesses administrator privileges, the verification fails. The system immediately terminates the subsequent data storage process and returns an insufficient permissions message to the user through the user-computer interface, clearly informing the user that they lack permission and cannot complete the data storage. This permission verification step has the beneficial effect of ensuring that only authorized personnel (responsible person or super user) can perform storage operations on CBB data through strict identity verification and permission control. This effectively prevents unauthorized users from arbitrarily modifying or storing data, ensuring the accuracy, security, and integrity of CBB data and minimizing the risk of data errors or leaks caused by conflicting permissions.

[0039] In the embodiment of the present disclosure, in addition to the aforementioned contents, the data processing method also includes other specific implementation steps. To clearly present these components, the relevant specific implementation methods are described in detail below: in response to a query instruction of the target hardware module to be developed, query the first data table according to the encoding of the target hardware module to be developed to obtain all version records of the target hardware module to be developed; sort the version records according to the size of the version number or the order of creation time; return a sorted version record list, in which each record in the list contains the version number, creation time and version record unique identifier.

[0040] Specifically, when querying the version record of the target hardware module to be developed (i.e., the target CBB), the system first responds to the query command. This query command is typically triggered by the user through the human-computer interaction interface. Specifically, this can be initiated by entering the target CBB's code (e.g., "CBB_Motherboard_CPU_AMD_001") in the query area of ​​the interface and clicking the query button. The uniqueness of the code ensures precise location of the query object. After receiving the query command, the system accesses the first data table based on the target CBB's code. The first data table is a database table specifically used to store basic information about each CBB version. Its structure includes fields such as the target CBB's code, version number (e.g., "V1.0," "V2.0," etc., using semantic versioning rules consisting of a major and minor version number. Changes in the major version number indicate significant functional changes), creation time (a timestamp accurate to the second, recording the time when the version data was first stored), and version record unique identifier (a system-generated string or numeric ID that uniquely distinguishes different version records of the same CBB). By matching the code in the query command with the "Code" field in the table, the system can filter all version records for the target CBB, covering the entire historical data from the initial version to the latest version. After obtaining all version records, the system needs to sort them. There are two sorting rules: one is to sort by version number, based on semantic versioning rules, comparing the major and minor version numbers. For example, "V2.1" is greater than "V2.0", and "V3.0" is greater than "V2.5". The other is to sort by creation time, usually in descending order, with the most recent version records listed first, so that users can view the latest versions first. In practice, the sorting method can be selected based on the user's settings during the query. If no specific setting is made, the default is to sort by creation time in descending order. After sorting is complete, the system returns the resulting list of version records to the user interface. Each record in the list clearly displays the version number, creation time, and unique identifier of the version record, allowing users to intuitively understand the version evolution of the target CBB.

[0041] In the embodiment of the present disclosure, in addition to the aforementioned contents, the data processing method also includes other specific implementation steps. In order to clearly present these components, the relevant specific implementation methods are described in detail below: receiving a statistical query request, the statistical query request includes a statistical dimension, and the statistical dimension is a project to which it belongs or a certain level in the multi-level classification information; according to the statistical dimension, aggregate query the second data table or the third data table, and calculate the number of modules under each dimension value and the percentage of the total.

[0042] Specifically, when processing a statistical query request, the system first receives a statistical query request initiated by the user. This request is usually triggered through the human-computer interaction interface. The user can select the required statistical dimension in the statistical function area of ​​the interface. The statistical dimension specifically includes the project to which it belongs or a certain level in the multi-level classification information. Among them, the project to which it belongs refers to the project name set by the company based on the customer's customized needs; the multi-level classification information is the system's hierarchical division of CBB, including first-level classification (such as "motherboard" and "infrastructure"), second-level classification (such as "CPU" and "security" under "motherboard", "storage" and "IO" under "infrastructure"), third-level classification, etc. Users can select any first-level classification as a statistical dimension according to their needs. After receiving the statistical query request and identifying the statistical dimension therein, the system will determine the corresponding data source table according to the dimension type for aggregate query. Specifically, if the statistical dimension is project, the system queries the second data table—this table specifically stores CBB project information, recording the association between project names and corresponding CBB data. If the statistical dimension is a level within a multi-level classification (such as a primary, secondary, or tertiary classification), the system queries the third data table—this table stores CBB multi-level classification information, including the names of each level and their associated IDs with the CBB data. During the aggregate query, the system extracts all associated CBB data records from the corresponding table based on the selected statistical dimension. A count operation is performed to calculate the number of CBB modules under each dimension value (i.e., the total number of CBBs belonging to the same project or classification level). For example, if the statistical dimension is "project," the system counts the number of CBBs under projects such as "Project 1" and "Project 2." If the statistical dimension is "primary classification," the system counts the number of CBBs under primary classifications such as "Mainboard" and "Infrastructure." The system then calculates the percentage of each dimension value relative to the total number of CBBs. This calculation is (number under the dimension value ÷ total number of CBBs) × 100%, rounding the result to the appropriate number of decimal places (e.g., integer percentage). Finally, the system organizes the statistical results into structured data containing dimension value labels (label), corresponding quantities (value) and proportions (percent), and returns them to the human-computer interaction interface for users to view.

[0043] In the disclosed embodiments, the operation of "integrating preset format data into target format data in a preset format based on predefined field mapping relationships" can be implemented in various ways. For clarity, the following examples include, but are not limited to, some implementations: Based on the predefined field mapping relationships, fields in the preset format data are converted to generate target format data containing field labels and field values. The preset format data and the target format data are data exchange formats.

[0044] Specifically, the database pre-stores the Chinese-English correspondences of CBB data fields (i.e., predefined field mappings). These mappings are set based on system design requirements. For example, "CBB code" corresponds to the English field "code," "version number" corresponds to "version," "project" corresponds to "project," "first-level category" corresponds to "firstCat," "second-level category" corresponds to "secondCat," and "third-level category" corresponds to "thirdCat." Custom fields within each third-level category (such as "platform" for Vendor 1's platform, "packaging method" corresponds to "package") are also mapped. The pre-formatted data is the standard JSON format (a data exchange format) returned by the large model after parsing a PDF spreadsheet. Its fields are presented with Chinese column names, including key names such as "CBB code," "version number," "project," and "first-level category," along with their corresponding values. During integration, the system converts each Chinese field in the pre-formatted data into its corresponding English field based on pre-defined field mappings in the database. For each converted field, the system constructs a structure consisting of a "label" and a "value." The "label" retains the original Chinese column name (used for field descriptions in front-end display), while the "value" is the specific data value corresponding to the field. For example, if the value corresponding to "CBB Code" in the pre-formatted data is "CBB_Motherboard_CPU_AMD_001_V1.0," the resulting "code" field in the target data will have the structure {"label": "CBB Code", "value": "CBB_Motherboard_CPU_xxx (xxx is the manufacturer's name)_001_V1.0"}. For the "First Category" field, the value in the pre-formatted data is "Motherboard." After conversion, the structure of the "firstCat" field in the target data will be {"label": "First Category", "value": "Motherboard"}. In this way, the integration of preset format data is completed, and the target format data finally generated is still in JSON format (data exchange format), which not only complies with the database's field naming specifications, but also meets the requirements of front-end page display through the "label" and "value" structure, providing a unified and standardized data foundation for subsequent manual verification and data storage.

[0045] In the embodiment of the present disclosure, in addition to the aforementioned content, the data processing method also includes other specific implementation steps. To clearly present these components, the relevant specific implementation methods are described in detail below: calling a text extraction interface to extract text from the target document to obtain text data; constructing summary prompt words, and the summary prompt words instruct the pre-trained model to generate a summary of the input text data; splicing the text data with the summary prompt words and inputting them into the preset large language model to obtain summary text; and generating retrieval information for the target document based on the summary text.

[0046] Specifically, the target document is a PDF document (such as "CBB Development Design Finalized.pdf") that records information related to the Common Building Block (CBB). This document contains basic information about the CBB development, a description of customer requirements, and discussion details. When performing text extraction on the target document, the system calls Python's Slate library (a PDF-based extension library that provides a simple API for extracting PDF text) to read the entire text content of the PDF document and obtain complete text data. This ensures that all non-tabular descriptive information in the document is included, providing a foundation for subsequent semantic analysis.

[0047] When performing semantic analysis on the extracted text data to generate summary information, the system passes the text data to the locally deployed large model and guides the large model to perform semantic understanding and refinement of the text data through preset prompt words (such as "Please output summary information of the above text content, the number of words in the summary is limited to 500"), and finally generates summary information that can summarize the core content of the document. This summary can concisely reflect the development background, key features and discussion points of CBB, making it easy to quickly grasp the core of the document.

[0048] When generating search information based on the target data table and summary information in the target file, the target data table refers to the table that stores CBB-related structured data (such as the CBB master data table that stores CBB code, version number, project, multi-level classification, and other information, as well as the associated project information table, classification information table, etc.). The system saves this table data (including CBB code, version, classification, project, and other structured fields) along with the generated summary text into the ElasticSearch search engine. As an engine that supports full-text search, ElasticSearch can combine the structured information of table data with the unstructured information of the summary to form search information. When users need to find existing CBB data through fuzzy query (such as searching for the development background or characteristics of related CBB based on keywords), the system can quickly match the search information in ElasticSearch and return the corresponding CBB data based on the degree of match, improving search efficiency and accuracy and providing support for rapid query of CBB data.

[0049] In the embodiment of the present disclosure, the target document is a portable document format file, and the hardware module to be developed is a common building block.

[0050] Specifically, the target document is a PDF document, generated after the results of the CBB (Common Building Block) discussion are confirmed. This document contains key information such as the basic information required for CBB development (such as unique CBB codes and version numbers), detailed descriptions of customer requirements, detailed discussion logs, and the default save location for the final CBB development files. It is stored in the same SVN directory as the final CBB development files and serves as the core data source for the system's data parsing and intelligent population. The hardware modules to be developed, known as Common Building Blocks (CBBs), are important shared components in server hardware R&D, playing a key role in improving R&D efficiency, reducing costs, enhancing product quality, and increasing system flexibility. CBBs are categorized into multiple categories based on application platforms, such as CPU type, storage, I / O modules, and GPUs. Their number and versions continue to increase with changing requirements and technological advancements, making them the core objects of data management, intelligent population, and statistical classification in this solution.

[0051] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these numbers do not limit the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not limit this.

[0052] Corresponding to the above-mentioned data processing method, the present disclosure also provides a data processing device. Since the device embodiment of the present disclosure corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment and will not be repeated in this disclosure.

[0053] Figure 2 A schematic diagram of a data processing device according to an embodiment of the present disclosure is shown in FIG. Figure 2 Shown, including: The receiving unit 21 is configured to receive the storage path information of the target document in the file version library and the range of data to be processed inputted through the human-computer interaction interface, and obtain the target document based on the storage path information; An extraction unit 22 is configured to call a data extraction interface to extract data from the document content within the data range to be processed in the target document, and input the extracted data into a pre-trained model so as to format the extracted data based on a preset prompt word to obtain data in a preset format; An integration unit 23 is configured to integrate the preset format data into a preset target format data based on a predefined field mapping relationship, wherein the target format data includes at least the code, version number, project to which the hardware module to be developed belongs, and multi-level classification information; A verification unit 24 is used to output and display the target format data for verification; The storage unit 25 is configured to store and associate the code, version number, project information and multi-level classification information of the hardware module to be developed included in the verified target format data in response to the target format data passing verification.

[0054] The present disclosure provides a data processing device, which receives the storage path and scope to be processed of a target PDF document, uses a large model to parse the document to generate structured data containing information such as CBB code, version number, project and multi-level classification, and stores and associates each type of information separately after verification. At the same time, it combines authority control and version management to realize a technical solution of systematic management and control. The device can solve the problems of low data accuracy, repeated entry, difficult version tracing and missing multi-dimensional classification statistics caused by relying on the hybrid mode of PDF and Excel in existing CBB data management, thereby achieving the technical effect of reducing manual operation errors, improving data management security, supporting full version tracing, and providing data support for server R&D decisions through intelligent classification and statistical analysis.

[0055] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and the principles are the same, which is not limited in this embodiment.

[0056] For the description of the features in the embodiments corresponding to the data processing device, reference can be made to the relevant description of the embodiments corresponding to the data processing method, which will not be repeated here.

[0057] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned data processing method embodiments.

[0058] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data processing method embodiments when run.

[0059] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0060] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned data processing method embodiments are implemented.

[0061] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned data processing method embodiments are implemented.

[0062] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0063] The above is a detailed introduction to a data processing method and electronic device provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A data processing method, characterized in that: include: Receive storage path information of a target document in a file version library and a range of data to be processed inputted through a human-computer interaction interface, and obtain the target document based on the storage path information; Calling a data extraction interface to extract data from the document content within the data to be processed range in the target document, and inputting the extracted data into a pre-trained model so as to format the extracted data based on a preset prompt word to obtain data in a preset format; Based on a predefined field mapping relationship, the preset format data is integrated into a preset target format data, wherein the target format data at least includes the code, version number, project to which the hardware module to be developed belongs, and multi-level classification information; Outputting and displaying the target format data for verification; In response to the target format data passing the verification, the code, version number, project information and multi-level classification information of the hardware module to be developed included in the verified target format data are stored separately and associated.

2. The data processing method according to claim 1, characterized in that: The code, version number, project information and multi-level classification information of the hardware module to be developed included in the verified target format data are stored separately and associated, including: storing the code of the hardware module to be developed and the version number as a version record in a first data table; Query the second data table to determine whether the project information exists. If not, insert a new record and obtain the unique identifier of the project information. If so, directly obtain the unique identifier of the project information. Querying the third data table to determine whether the multi-level classification information exists, if not, inserting a new record and obtaining the unique identifier of the multi-level classification information, if yes, directly obtaining the unique identifier of the multi-level classification information; The unique identifier of the version record, the unique identifier of the belonging project information and the unique identifier of the multi-level classification information are associated and stored in a relationship table to complete the association.

3. The data processing method according to claim 2, characterized in that: Before storing the code, version number, project information, and multi-level classification information of the hardware module to be developed included in the verified target format data, the method further includes: Get the identity information of the currently logged-in user; According to the code of the hardware module to be developed, query the information of the person in charge of the module from the authority configuration table; Verify whether the current user is the person in charge of the module or has administrator privileges; If the verification fails, the storage process is terminated and a prompt message indicating insufficient permissions is returned.

4. The data processing method according to claim 2, characterized in that: The method further comprises: In response to a query instruction of a target hardware module to be developed, querying the first data table according to the code of the target hardware module to be developed to obtain all version records of the target hardware module to be developed; Sorting the version records according to the size of the version numbers or the order of creation time; Returns a sorted list of version records. Each record in the list contains the version number, creation time, and unique identifier of the version record.

5. The data processing method according to claim 2, characterized in that: The method further comprises: receiving a statistical query request, wherein the statistical query request includes a statistical dimension, wherein the statistical dimension is a project or a level in multi-level classification information; According to the statistical dimension, the second data table or the third data table is aggregated and queried to calculate the number of modules under each dimension value and the percentage of the total.

6. The data processing method according to claim 1, characterized in that: The preset format data and the target format data are in a data exchange format; The step of integrating the preset format data into target format data in a preset format based on a predefined field mapping relationship includes: Based on the predefined field mapping relationship, the fields in the preset format data are converted to generate the target format data including field labels and field values.

7. The data processing method according to claim 1, characterized in that: The method further comprises: Calling a text extraction interface to extract text from the target document to obtain text data; Constructing a summary prompt word, wherein the summary prompt word instructs the pre-trained model to generate a summary of the input text data; splicing the text data with the summary prompt words and inputting the result into the preset large language model to obtain a summary text; Based on the summary text, retrieval information of the target document is generated.

8. The data processing method according to claim 1, characterized in that: The calling of the data extraction interface to extract data from the document content within the data range to be processed in the target document includes: Parsing the data range to be processed to determine the target page number where the document content is located; Calling the data extraction interface to read the table data in the target page number; The table data is converted into a structured data array containing column names and cell values ​​to complete data extraction.

9. The data processing method according to any one of claims 1 to 8, characterized in that: The hardware modules to be developed are common building blocks.

10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the data processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Natural language rule table information extraction system based on large model

    CN120031004A

  • HTML (Hypertext Markup Language) information extraction method, device and equipment based on multi-LoRA cascade strategy and medium

    CN120296275A

  • Project management software document uploading method based on OCR and large language model

    CN120526446A

  • Metadata generation system and method based on ensemble of multi open-source large language models

    KR102830267B1

  • Extraction of a nested hierarchical structure from text data in an unstructured version of a document

    US20210319039A1