Data processing method, electronic equipment and storage medium
Through the content direction prediction model and verification mechanism, data that meets the conditions is automatically identified and extracted, solving the problem of cumbersome data processing in the existing technology, and achieving efficient and accurate data processing.
Patent Information
- Application Number
- CN202510400854.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, when importing data to an AI agent, the user needs to clean and remove data with unclear content direction in advance, resulting in cumbersome data processing flow and the user needs to have the ability to distinguish the direction of data content.
By obtaining the pending file and entering the pre-trained content direction prediction model, determine the content direction of the file, verify the necessary file attributes and feature values in each direction for each piece of data, and extract valid data that meets the conditions.
A low-threshold and convenient data processing method is achieved, data processing efficiency and accuracy are improved, manual intervention is reduced, and data intervention is ensured that the data meets the expected content direction.
Smart Images

Figure CN120256423A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data processing method, an electronic device, and a storage medium. Background Art
[0002] An AI (artificial intelligence) intelligent agent is a combination of an AI large model and a toolset, and is an intelligent application or entity that can act autonomously, perceive the environment, make decisions, and interact with the environment. For an AI intelligent agent to achieve various functions, it relies on a knowledge base. When constructing a knowledge base, it is necessary to import available data into the knowledge base. To ensure the quality, compatibility, and effectiveness of the data, it is necessary to pre-process the data.
[0003] Currently, the data to be processed can be imported into an AI intelligent agent for identifying the content direction, and the AI intelligent agent can identify the content direction of the data and determine the available data. However, the AI intelligent agent requires the content direction of the data to be clear in order to accurately extract the available data. Therefore, when importing data into the AI intelligent agent, it is necessary for the user to pre-clean the data in advance to remove the data with unclear content direction, resulting in a more cumbersome data processing process and requiring the user to have the ability to distinguish whether the content direction of the data is clear.
[0004] Therefore, there is an urgent need for a low-threshold and convenient way to process data. Summary of the Invention
[0005] This application provides a data processing method, an electronic device, and a storage medium to solve the problem that when importing data into an AI intelligent agent, it is necessary for the operator to pre-clean the data in advance to remove the data with unclear content direction, resulting in a more cumbersome data processing process and requiring the user to have the ability to distinguish whether the content direction of the data is clear, and realizes a low-threshold and convenient way to process data.
[0006] In a first aspect, this application provides a data processing method, including:
[0007] Obtain a file to be processed, where the file to be processed includes one or more initial data, and the initial data includes one or more first file attributes and first feature values corresponding to the first file attributes;
[0008] Input the file to be processed into a pre-trained content direction prediction model to determine one or more content directions corresponding to the file to be processed; the content direction is used to indicate one or more necessary file attributes;
[0009] For each piece of the initial data, under each content direction, determine whether the one or more first file attributes include all the necessary file attributes in the content direction, and determine whether the second eigenvalue corresponding to each of all the necessary file attributes is a valid eigenvalue;
[0010] When the one or more first file attributes include all the necessary file attributes in the content direction and the second eigenvalue corresponding to each of all the necessary file attributes is a valid eigenvalue, extract the first file attributes with the first eigenvalue being a valid eigenvalue and the corresponding first eigenvalue as a piece of valid data in the content direction.
[0011] In an embodiment of the present application, a file to be processed is obtained. The file to be processed includes one or more pieces of initial data to facilitate the processing of the one or more pieces of initial data in the file to be processed. The file to be processed is input into a pre-trained content direction prediction model to determine one or more content directions corresponding to the file to be processed. This is to facilitate the verification of the one or more pieces of initial data according to the one or more content directions, avoiding verifying all content directions during the subsequent verification process of the initial data, which increases the ineffective processing process. Instead, the verification is targeted at the predicted content directions, thereby improving the efficiency of data processing for the file to be processed. For each piece of initial data, under each content direction, determine whether the one or more first file attributes include all the necessary file attributes in the content direction, and determine whether the second eigenvalue corresponding to each of all the necessary file attributes is a valid eigenvalue. Thus, by comparing the necessary file attributes and the corresponding second eigenvalues, it is possible to verify whether the initial data conforms to the corresponding content direction, improving the accuracy of verifying whether the initial data matches the corresponding content direction. When the one or more first file attributes include all the necessary file attributes in the content direction and the second eigenvalue corresponding to each of all the necessary file attributes is a valid eigenvalue, extract the first file attributes with the first eigenvalue being a valid eigenvalue and the corresponding first eigenvalue as a piece of valid data in the content direction. It is possible to extract all the file attributes and their eigenvalues that fully conform to the content direction and eliminate the file attributes that do not include valid eigenvalues. Thus, for each content direction, the verification of each piece of initial data is completed, and thus valid data can be extracted from the file to be processed under each content direction without human participation, enabling convenient, accurate, and fast data processing.
[0012] In a possible design, the method further includes:
[0013] If the one or more first file attributes do not include all the necessary file attributes in the content direction, or there is an invalid eigenvalue corresponding to the necessary file attribute, discard the initial data.
[0014] In a possible design, determining whether the one or more first file attributes include all necessary file attributes in the content direction includes:
[0015] For each piece of the initial data, in each content direction, score based on the one or more necessary file attributes to obtain first scoring results respectively corresponding to the one or more necessary file attributes;
[0016] According to one or more of the first scoring results, determine whether the one or more first file attributes include all necessary file attributes in the content direction.
[0017] In a possible design, the determining, according to one or more of the first scoring results, whether the one or more first file attributes include all necessary file attributes in the content direction includes:
[0018] When one or more of the first scoring results are greater than or equal to a first preset threshold, determine that the one or more first file attributes include all necessary file attributes in the content direction;
[0019] When there is a scoring result less than the first preset threshold among one or more of the first scoring results, determine that the one or more first file attributes do not include all necessary file attributes in the content direction.
[0020] In a possible design, the determining whether second eigenvalue respectively corresponding to all the necessary file attributes are all valid eigenvalues includes:
[0021] For each piece of the initial data, in each content direction, after determining that the one or more first file attributes include all necessary file attributes in the content direction, score based on the second eigenvalue respectively corresponding to all the necessary file attributes to obtain second scoring results respectively corresponding to the second eigenvalue respectively corresponding to all the necessary file attributes;
[0022] According to one or more of the second scoring results, determine whether the second eigenvalue respectively corresponding to all the necessary file attributes are all valid eigenvalues.
[0023] In a possible design, the determining, according to one or more of the second scoring results, whether the second eigenvalue respectively corresponding to all the necessary file attributes are all valid eigenvalues includes:
[0024] When one or more of the second scoring results are greater than or equal to a second preset threshold, determine that the second eigenvalue respectively corresponding to all the necessary file attributes are all valid eigenvalues;
[0025] When there is a scoring result less than the second preset threshold among one or more of the second scoring results, it is determined that there are invalid eigenvalue among the second eigenvalues corresponding to all the necessary file attributes respectively.
[0026] In a possible design, the extraction of the first file attributes with valid eigenvalues and the corresponding first eigenvalues is a piece of valid data in the content direction, including:
[0027] Obtain the scoring results corresponding to one or more of the first file attributes respectively;
[0028] Determine the necessary file attributes with scoring results greater than or equal to the third preset threshold among one or more of the first file attributes as the file attributes to be extracted, and determine the non-necessary file attributes with scoring results greater than or equal to the fourth preset threshold among one or more of the first file attributes as the file attributes to be extracted;
[0029] Obtain a piece of valid data according to all the file attributes to be extracted and the corresponding eigenvalues.
[0030] In a possible design, the method further includes:
[0031] Obtain one or more sample files and the content directions corresponding to the sample files respectively;
[0032] Train a preset model according to the file name of the sample file and the content at a specified position in the sample file to obtain the content direction prediction model.
[0033] In a possible design, the method further includes:
[0034] Obtain one or more pieces of the valid data and the content directions corresponding to the valid data respectively;
[0035] Update the content direction prediction model according to the file name of the valid data and the content at a specified position in the valid data to obtain the updated content direction prediction model.
[0036] In a possible design, the method further includes:
[0037] Store one or more pieces of the valid data into a knowledge base according to the content directions corresponding to the valid data, where the knowledge base is applied to an artificial intelligence (AI) intelligent agent, and the AI intelligent agent is used for environmental protection, social responsibility and corporate governance (ESG) management.
[0038] In a second aspect, the present application provides a data processing device, including: a module for executing the method in the first aspect and any possible design of the first aspect.
[0039] For the device provided in the second aspect above and each possible design of the second aspect, the beneficial effects can be referred to the beneficial effects brought by the first aspect above and each possible implementation manner of the first aspect, which will not be elaborated herein.
[0040] In a third aspect, the present application provides an electronic device, including: a processor;
[0041] The processor is configured to execute computer-executable programs or instructions in a memory, so that the electronic device executes the method in the first aspect and any possible design of the first aspect.
[0042] In a fourth aspect, the present application provides an electronic device, including: a memory and a processor; the memory is configured to store program instructions; the processor is configured to call the program instructions in the memory so that the electronic device executes the method in the first aspect and any possible design of the first aspect.
[0043] In a fifth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the electronic device is enabled to execute the method in the first aspect and any possible design of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of a data processing method provided by an embodiment of the present application.
[0045] Figure 2 It is a flowchart of a method for determining necessary file attributes provided by an embodiment of the present application.
[0046] Figure 3 It is a flowchart of a method for determining whether all necessary file attributes in a content direction are included provided by an embodiment of the present application.
[0047] Figure 4 It is a flowchart of a method for determining valid eigenvalue provided by an embodiment of the present application.
[0048] Figure 5 It is a flowchart of a method for determining whether all second eigenvalues corresponding to all necessary file attributes are valid eigenvalues provided by an embodiment of the present application.
[0049] Figure 6 It is a flowchart of a method for extracting valid data provided by an embodiment of the present application.
[0050] Figure 7 It is a flowchart of a method for training a content direction prediction model provided by an embodiment of the present application.
[0051] Figure 8Flowchart of a method for updating a content direction prediction model provided by an embodiment of the present application.
[0052] Figure 9 Flowchart of a data processing method provided by an embodiment of the present application.
[0053] Figure 10 Schematic structural diagram of a data processing apparatus provided by an embodiment of the present application.
[0054] Figure 11 Schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0055] Figure 12 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0056] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or similar expressions thereof refer to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a alone, b alone, or c alone may represent: a alone, b alone, c alone, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b, and c, where a, b, and c may be single or multiple. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0057] The orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "upper", "lower", "left", "right", "front", "rear", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation to the present application.
[0058] The terms "connected" and "coupled" should be understood in a broad sense. For example, the "connection" or "coupling" of a circuit structure can refer not only to a physical connection, but also to an electrical connection or a signal connection. For example, it can be a direct connection, i.e., a physical connection, or an indirect connection through at least one intermediate component, as long as the circuit is connected. It can also be the internal connection of two components; the signal connection can be made not only through a circuit for signal connection, but also through a media medium for signal connection. For example, radio waves. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0059] Exemplarily, this application provides a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Through a content direction prediction model, the content direction of a file to be processed is predicted, and the content direction of the file to be processed is verified through the file attributes and the characteristic values corresponding to the file attributes, so as to extract valid data from the file to be processed in each content direction without human participation, and can process data conveniently, accurately, and quickly.
[0060] Among them, the data processing method provided by this application is executed by a computer device, or by a data processing apparatus in the computer device. The data processing apparatus can be implemented by a combination of software and / or hardware. For example, the data processing apparatus can be an AI intelligent agent, an application (APP), a web page, or a public account, etc. For the sake of simplicity, the embodiments of this application are described by taking the execution of the data processing apparatus as an example.
[0061] Among them, the computer device can be a server, a desktop computer, a mobile phone, a tablet computer, a notebook computer, a wearable device, a vehicle-mounted device, or an augmented reality (AR) / virtual reality (VR) device, etc.
[0062] Next, the following embodiments of this application will be combined with Figures 1 to 9 , and the data processing method provided by this application will be elaborated in detail.
[0063] Please refer to Figure 1 , Figure 1 which is a flowchart of a data processing method provided by an embodiment of this application. As Figure 1 shown, the method includes:
[0064] S101. The data processing apparatus obtains a file to be processed.
[0065] Among them, the file to be processed includes one or more pieces of initial data.
[0066] Among them, the initial data includes one or more first file attributes and the first eigenvalue corresponding to the first file attribute.
[0067] The format of the file to be processed can be a document file, a picture file, a table file, etc.
[0068] Among them, the formats of document files include:.txt,.docx,.doc, etc. The formats of picture files include:.jpg,.jpeg,.png, etc. The formats of table files include:.xlsx,.xls, etc.
[0069] In some examples, the file to be processed can be a file uploaded by the user. In other examples, the file to be processed can be a file read by the data processing device from other devices or databases.
[0070] After obtaining the file to be processed, the data processing device can parse and segment the file to be processed through the embedding technology, map the data in the high-dimensional file to be processed into a low-dimensional space, and obtain one or more pieces of initial data. One piece of initial data is an N-dimensional real-valued vector, which is convenient for the data processing device to perform subsequent processing.
[0071] Among them, the content direction of one piece of initial data can include: one or more of the product carbon footprint (PCF) report of a certain product, the bill of materials (BOM) of a certain product, the environmental, social and governance (ESG) report of a certain enterprise, the enterprise risk and opportunity data of a certain enterprise, the supply chain relationship and the enterprise-product production relationship.
[0072] Among them, the first file attribute is the constituent attribute of the initial data, and the first eigenvalue corresponding to the first file attribute is the value of the constituent attribute of the initial data.
[0073] Taking one piece of initial data as the PCF report of a certain product as an example. This initial data includes:
[0074] "Product Chinese name / Product foreign name|X Product
[0075] Product carbon footprint result value|1234.5678
[0076] Product carbon footprint - Emission unit|kgCO2e
[0077] Product carbon footprint - Product measurement unit|unit
[0078] Life Cycle Analysis (LCA) Phase / Life Cycle Phase|Cradle to Gate
[0079] Data Representativeness|Industry Average
[0080] Geographical Representativeness|China
[0081] Data Statistical Time Period - Annual|2 years
[0082] Data Statistical Time Period - Start Date & End Date|March 27, 2022 to March 27, 2024
[0083] Manufacturer|Company X
[0084] Among them, "Product Chinese Name / Product Foreign Name", "Product Carbon Footprint Result Value", "Product Carbon Footprint - Emission Unit", "Product Carbon Footprint - Product Measurement Unit", "LCA Phase / Life Cycle Phase", "Data Representativeness", "Geographical Representativeness", "Data Statistical Time Period - Annual", "Data Statistical Time Period - Start Date & End Date", and "Manufacturer" are the first file attributes in this data entry. "Product X", "1234.5678", "kgCO2e", "unit", "Cradle to Gate", "Industry-level Data", "China", "2 years", "March 27, 2022 to March 27, 2024", and "Company X" are the first characteristic values corresponding to each of the first file attributes respectively.
[0085] Based on this, the data processing device obtains the file to be processed to facilitate the processing of one or more initial data in the file to be processed.
[0086] S102. The data processing device inputs the file to be processed into a pre-trained content direction prediction model to determine one or more content directions corresponding to the file to be processed.
[0087] Among them, the content direction is used to indicate one or more necessary file attributes.
[0088] The content directions include: PCF Report, Product BOM, ESG Report, Enterprise Risk and Opportunity Data, Supply Chain Relationship, and Enterprise-Product Production Relationship, etc. A file to be processed can correspond to one or more content directions.
[0089] Among them, the content direction prediction model can predict one or more content directions of the file by the file name and the content at the specified position in the file.
[0090] Among them, the content at the specified position is the starting content of the file, such as the first 50 characters at the start.
[0091] For a file to be processed, after the data processing device inputs the file to be processed into the content direction prediction model, the content direction prediction model can make a prediction based on the file name of the file to be processed and the content at the specified position of the file to be processed, and obtain a prediction result. The prediction result includes one or more content directions corresponding to the file to be processed, or there is no matching content direction.
[0092] If one or more content directions corresponding to the file to be processed are obtained through the content direction prediction model, then S103 can be continued. In addition, the data processing device can also inform the user of one or more content directions corresponding to the predicted file to be processed. For example, the data processing device can output a field of "The predicted content direction is XXX" to the user, so that the user can know the prediction result.
[0093] If it is known through the content direction prediction model that the file to be processed has no matching content direction, then the data processing device can inform the user that the file to be processed does not conform to any content direction and end the data processing process of the file to be processed. For example, the data processing device can output a field of "Does not conform to any content direction" to the user, so that the user can know the prediction result.
[0094] Among them, the necessary file attributes refer to the file attributes that must be included when conforming to a certain content direction.
[0095] For example, the necessary file attributes of the PCF report include: "Product Chinese Name / Product Foreign Name", "Product Carbon Footprint Result Value", "Product Carbon Footprint - Emission Unit", "Product Carbon Footprint - Product Measurement Unit", "LCA Stage / Life Cycle Stage", and "Data Representativeness". Then in a file, all these five necessary file attributes need to be included for the content direction of this file to be a PCF report.
[0096] The data processing device can pre-store all content directions and one or more necessary file attributes respectively indicated by these content directions in the cloud or the memory of the data processing device. After determining one or more content directions corresponding to the file to be processed according to the content direction prediction model, according to the one or more content directions, obtain one or more necessary file attributes corresponding to the one or more content directions.
[0097] Based on this, the data processing device can pre-determine one or more content directions for the file to be processed according to the content direction prediction model, so as to verify one or more initial data according to the one or more content directions, avoiding verifying all content directions during the subsequent verification process of the initial data, increasing the ineffective processing process, but verifying the predicted content directions in a targeted manner, thereby improving the efficiency of data processing for the file to be processed.
[0098] S103. For each piece of initial data, in each content direction, the data processing device determines whether one or more first file attributes include all the necessary file attributes in the content direction, and determines whether the second feature values corresponding to all the necessary file attributes are all valid feature values.
[0099] If one or more first file attributes include all the necessary file attributes in the content direction, and the second feature values corresponding to all the necessary file attributes are all valid feature values, the data processing device executes S104; if one or more first file attributes do not include all the necessary file attributes in the content direction, or there is an invalid feature value among the second feature values corresponding to the necessary file attributes, the data processing device executes S105. Herein, S105 is an optional step.
[0100] Specifically, the data processing device can first determine whether one or more first file attributes include all the necessary file attributes in the content direction, and then determine whether the second feature values corresponding to all the necessary file attributes are all valid feature values.
[0101] The data processing device can compare one or more first file attributes with each necessary file attribute respectively through methods such as semantic comparison, cosine similarity, and Jaccrad coefficient to determine whether one or more first file attributes include all the necessary file attributes in the content direction. For each second feature value, the data processing device can determine, for each necessary file attribute, whether the second feature value corresponding to the necessary file attribute conforms to a preset type. If it conforms to the preset type, it is determined as a valid feature value; if it does not conform to the preset type, or there is no feature value for the necessary file attribute, it is determined as an invalid feature value.
[0102] For example, the field in a piece of initial data is: "The product carbon footprint result value is 1234.5678". The product carbon footprint result value is a necessary file attribute in the PCF report. The feature value of the product carbon footprint result value should conform to a value of the numeric type. And the second feature value corresponding to the product carbon footprint result value in this piece of initial data is 1234.5678, and this feature value conforms to the numeric type, so it is a valid feature value. Another example, the field in a piece of initial data is: "The product carbon footprint result value is more than a thousand". The product carbon footprint result value is a necessary file attribute in the PCF report. The feature value of the product carbon footprint result value should conform to a value of the numeric type. And the corresponding second feature value is more than a thousand, and this feature value does not conform to the numeric type, so it is an invalid feature value.
[0103] Next, specific examples are used for illustration.
[0104] If the data processing device determines that the content directions of the file to be processed are content direction α and content direction β, the necessary file attributes corresponding to content direction α are file attribute A, file attribute B, and file attribute C, the necessary file attribute corresponding to content direction β is file attribute E, and the file to be processed includes 4 pieces of initial data, as shown in Table 1 below:
[0105] Table 1
[0106]
[0107] For initial data 1, in content direction α, initial data 1 includes file attribute A, file attribute B, and file attribute C, and the data processing device can determine that one or more first file attributes of initial data 1 include all the necessary file attributes in content direction α. Similarly, the data processing device can determine that one or more first file attributes of initial data 1 do not include all the necessary file attributes in content direction β.
[0108] For initial data 2, in content direction α, initial data 2 includes file attribute A and file attribute C, lacking file attribute B, and the data processing device can determine that one or more first file attributes of initial data 2 do not include all the necessary file attributes in content direction α. Similarly, the data processing device can determine that one or more first file attributes of initial data 2 do not include all the necessary file attributes in content direction β.
[0109] For initial data 3, in content direction α, initial data 3 does not include any necessary file attributes in content direction α, and the data processing device can determine that one or more first file attributes of initial data 3 do not include all the necessary file attributes in content direction α. Similarly, the data processing device can determine that one or more first file attributes of initial data 3 include all the necessary file attributes in content direction β.
[0110] For initial data 4, in content direction α, initial data 4 includes file attribute A, file attribute B, and file attribute C, and the data processing device can determine that one or more first file attributes of initial data 4 include all the necessary file attributes in content direction α. Similarly, the data processing device can determine that one or more first file attributes of initial data 4 do not include all the necessary file attributes in content direction β.
[0111] Thus, the data processing device can determine that in content direction α, one or more first file attributes corresponding to initial data 1 and initial data 4 respectively include all the necessary file attributes in the content direction. In content direction β, one or more first file attributes corresponding to initial data 3 include all the necessary file attributes in the content direction.
[0112] Based on this, further, the data processing device determines whether the second eigenvalue corresponding to each of all necessary file attributes is a valid eigenvalue.
[0113] For the initial data 1, in the content direction α, the second eigenvalues include eigenvalue ①, eigenvalue ②, and eigenvalue ③. Among them, if eigenvalue ② is an invalid eigenvalue, the data processing device can determine that there is an invalid eigenvalue among the second eigenvalues corresponding to the initial data 1 in the content direction α.
[0114] For the initial data 4, in the content direction α, the second eigenvalues include eigenvalue ⑧, eigenvalue ⑨, and eigenvalue ⑩. If eigenvalue ⑧, eigenvalue ⑨, and eigenvalue ⑩ are all valid eigenvalues, the data processing device can determine that the second eigenvalues corresponding to the initial data 4 in the content direction α are all valid eigenvalues.
[0115] For the initial data 3, in the content direction β, the second eigenvalue includes eigenvalue ⑦. If eigenvalue ⑦ is a valid eigenvalue, the data processing device can determine that the second eigenvalues corresponding to the initial data 3 in the content direction β are all valid eigenvalues.
[0116] Based on this, for each piece of initial data, in each content direction, the data processing device determines whether one or more first file attributes include all necessary file attributes in the content direction, and determines whether the second eigenvalues corresponding to all necessary file attributes are all valid eigenvalues. Thus, by comparing the necessary file attributes with the corresponding second eigenvalues, it verifies whether the initial data conforms to the corresponding content direction, improving the accuracy of verifying whether the initial data matches the corresponding content direction.
[0117] S104: If one or more first file attributes include all necessary file attributes in the content direction, and the second eigenvalues corresponding to all necessary file attributes are all valid eigenvalues, the data processing device extracts the first file attributes with the first eigenvalue being a valid eigenvalue and the corresponding first eigenvalue as a piece of valid data in the content direction.
[0118] Based on Table 1, as shown in Table 2, in the content direction α, one or more first file attributes corresponding to the initial data 4 include all necessary file attributes in the content direction, and the second eigenvalues corresponding to all necessary file attributes are all valid eigenvalues. In the content direction β, one or more first file attributes corresponding to the initial data 3 include all necessary file attributes in the content direction, and the second eigenvalues corresponding to all necessary file attributes are all valid eigenvalues.
[0119] Table 2
[0120]
[0121] The data processing device extracts file attribute A, file attribute B, file attribute C, and file attribute F, as well as eigenvalue ⑧, eigenvalue ⑨, eigenvalue ⑩, and eigenvalue from the initial data 4 in content direction α as a piece of valid data. In content direction β, file attribute E and eigenvalue ⑦ are extracted as a piece of valid data. In Table 2, there is no valid eigenvalue for file attribute G in the initial data 4, so the data processing device does not extract this file attribute.
[0122] In this way, based on the file to be processed shown in Table 1, the data processing device obtains two pieces of valid data shown in Table 3.
[0123] Table 3
[0124]
[0125]
[0126] Based on this, when all necessary file attributes in the content direction are included in one or more first file attributes, and the second eigenvalues corresponding to all necessary file attributes are valid eigenvalues, the data processing device extracts the first file attributes with valid first eigenvalues and the corresponding first eigenvalues as a piece of valid data in the content direction, which can extract all file attributes and their eigenvalues that fully conform to the content direction, and eliminate file attributes that do not include valid eigenvalues, thereby completing the verification of each piece of initial data for each content direction, and thus being able to extract valid data from the file to be processed in each content direction without human participation, and being able to process data conveniently, accurately, and quickly.
[0127] S105. If one or more first file attributes do not include all necessary file attributes in the content direction, or there is an invalid second eigenvalue corresponding to a necessary file attribute, the data processing device discards the initial data.
[0128] Based on Table 1, there is an invalid second eigenvalue corresponding to a necessary file attribute in the initial data 1, and one or more first file attributes corresponding to the initial data 2 do not include all necessary file attributes in the content direction, so the data processing device can discard the initial data 1 and the initial data 2.
[0129] Based on this, for the initial data that does not conform to the characteristics of a certain content direction, the data processing device performs a discard process, which can avoid determining this type of data as valid data, thereby improving the accuracy of determining valid data and the matching degree of valid data with the corresponding content direction.
[0130] In the embodiment of the present application, a data processing device obtains a file to be processed, where the file to be processed includes one or more pieces of initial data, so as to facilitate the processing of the one or more pieces of initial data in the file to be processed. The data processing device inputs the file to be processed into a pre-trained content direction prediction model to determine one or more content directions corresponding to the file to be processed. This is to facilitate the verification of the one or more pieces of initial data according to the one or more content directions, avoiding verifying all content directions during the subsequent verification process of the initial data, which increases the ineffective processing process. Instead, it specifically verifies the predicted content directions, thereby improving the efficiency of data processing for the file to be processed. For each piece of initial data, in each content direction, the data processing device determines whether one or more first file attributes include all necessary file attributes in the content direction, and determines whether the second feature values corresponding to all necessary file attributes are all valid feature values. Thus, by comparing the necessary file attributes with the corresponding second feature values, the data processing device can verify whether the initial data conforms to the corresponding content direction, improving the accuracy of verifying whether the initial data matches the corresponding content direction. If one or more first file attributes include all necessary file attributes in the content direction and the second feature values corresponding to all necessary file attributes are all valid feature values, the data processing device extracts the first file attributes with the first feature value being a valid feature value and the corresponding first feature value being a piece of valid data in the content direction. It can extract all file attributes and their feature values that completely conform to the content direction, and eliminate file attributes that do not include valid feature values. Thus, for each content direction, the verification of each piece of initial data is completed, and thus valid data can be extracted from the file to be processed in each content direction without human participation, enabling convenient, accurate, and fast data processing.
[0131] Based on the above exemplary description, a method for determining whether one or more first file attributes include all necessary file attributes in the content direction is introduced below.
[0132] Please refer to Figure 2 , Figure 2 which is a flowchart of a method for determining necessary file attributes provided in an embodiment of the present application. As Figure 2 shown, the method includes:
[0133] S201. For each piece of initial data, in each content direction, the data processing device scores based on one or more necessary file attributes to obtain one or more first scoring results corresponding to the one or more necessary file attributes.
[0134] Specifically, for each content direction, the data processing device can pre - know, through a preset column name knowledge base and a column name synonym library, what the necessary file attributes corresponding to this content direction are, and what the synonyms of these necessary file attributes are.
[0135] For each necessary file attribute, the data processing device can search in the initial data to check if it includes this necessary file attribute and score this necessary file attribute.
[0136] For example, taking a 10 - point full - score system as an example, for a necessary file attribute, if the initial data does not include this necessary file attribute, then the first scoring result for this necessary file attribute can be 0 points. If the initial data includes a first file attribute similar to this necessary file attribute, then the data processing device can score according to the similarity between the first file attribute and this necessary file attribute, and if they are exactly the same, it can be 10 points.
[0137] Taking an initial data of a PCF report of a certain product as an example. The content direction of this initial data is the PCF report. The necessary file attributes of the PCF report include: "Product Chinese Name / Product Foreign Name", "Product Carbon Footprint Result Value", "Product Carbon Footprint - Emission Unit", "Product Carbon Footprint - Product Measurement Unit", "LCA Stage / Life Cycle Stage", and "Data Representativeness". This initial data includes:
[0138] "Product Chinese Name / Product Foreign Name|X Product
[0139] Product Carbon Footprint Result Value|1234.5678
[0140] Product Carbon Footprint - Emission Unit|kgCO2e
[0141] Product Carbon Footprint - Product Measurement Unit|unit
[0142] LCA Stage / Life Cycle Stage|Cradle to Gate
[0143] Data Representativeness|Industry Average
[0144] Geographical Representativeness|China
[0145] Data Statistical Time Period - Annual|2 years
[0146] Data Statistical Time Period - Start Time & End Time|March 27, 2022 to March 27, 2024
[0147] Manufacturer|X Company"
[0148] The data processing device scores the initial data and obtains the first scoring results corresponding to one or more necessary file attributes of the initial data, as shown in Table 4 below:
[0149] Table 4
[0150] Product Chinese Name / Product Foreign Name 10 points Product Carbon Footprint Result Value 10 points Product Carbon Footprint - Emission Unit 10 points Product Carbon Footprint - Product Measurement Unit 10 points LCA Stage / Life Cycle Stage 10 points Data Representativeness 10 points
[0151] In some examples, the data processing device can enable the AI agent to master all the necessary file attributes in the content direction. Thus, the data processing device can also directly input each initial data into the AI, and the AI scores the necessary file attributes and obtains the first scoring results corresponding to one or more necessary file attributes respectively.
[0152] Based on this, for each initial data, the data processing device scores based on one or more necessary file attributes in each content direction and obtains the first scoring results corresponding to one or more necessary file attributes respectively. Thus, the first scoring results can be used to accurately determine whether the first file attribute in the initial data includes the necessary file attributes corresponding to the content direction in a content direction, which helps to ensure the accuracy of the final obtained valid data.
[0153] S202. The data processing device determines whether one or more first file attributes include all the necessary file attributes in the content direction according to one or more first scoring results.
[0154] For one or more first scoring results, the data processing device can determine whether one or more first file attributes include all the necessary file attributes in the content direction by comparing with a preset threshold or comparing the sum with a preset threshold and other methods.
[0155] Based on this, the data processing device can determine whether the initial data includes all the necessary file attributes in the content direction, thereby screening out available data from the necessary file attribute level, eliminating the initial data that does not include all the necessary file attributes, and avoiding spending more computing resources to reprocess this part of the initial data, improving the efficiency of data processing.
[0156] Based on the above exemplary description, a method for determining whether one or more first file attributes include all the necessary file attributes in the content direction according to one or more first scoring results is introduced below.
[0157] Please refer to Figure 3 , Figure 3 which is a flowchart of a method for determining whether all the necessary file attributes in the content direction are included provided by an embodiment of the present application. As Figure 3 shown, the method includes:
[0158] S301. The data processing device determines whether one or more first scoring results are all greater than or equal to a first preset threshold.
[0159] If one or more first scoring results are all greater than or equal to the first preset threshold, the data processing device executes S302; if there is a scoring result among one or more first scoring results that is less than the first preset threshold, the data processing device executes S303.
[0160] Taking the 10 - point system as an example, the first preset threshold can be 10 points. Thus, the data processing device can ensure that all necessary file attributes are included in the determined initial data, thereby improving the accuracy of data processing and ensuring that accurate valid data can be obtained. As shown in Table 4 above, if the first preset threshold is 10 points, the scores of all necessary file attributes in this initial data are all greater than or equal to 10 points. Therefore, for this initial data, the data processing device determines that one or more first file attributes corresponding to this initial data include all necessary file attributes in the content direction.
[0161] S302. The data processing device determines that one or more first file attributes include all necessary file attributes in the content direction.
[0162] S303. The data processing device determines that one or more first file attributes do not include all necessary file attributes in the content direction.
[0163] Based on the above - described example, a method for determining whether the second eigenvalue corresponding to each necessary file attribute is an effective eigenvalue is introduced below.
[0164] Please refer to Figure 4 , Figure 4 , which is a flowchart of a method for determining effective eigenvalues provided by an embodiment of the present application. As Figure 4 shown, the method includes:
[0165] S401. For each initial data, in each content direction, after determining that one or more first file attributes include all necessary file attributes in the content direction, the data processing device scores based on the second eigenvalues corresponding to all necessary file attributes, and obtains second scoring results corresponding to the second eigenvalues corresponding to all necessary file attributes respectively.
[0166] Specifically, the data processing device can pre - know the preset type of the eigenvalue corresponding to the necessary file attribute in each content direction through the value example knowledge base. For example, the eigenvalue of the necessary file attribute "product carbon footprint result value" should conform to the numeric type, such as 1234.5678.
[0167] After determining all necessary document attributes including the content orientation among one or more first document attributes, for each piece of initial data, the data processing device scores based on the second eigenvalue corresponding to each of the all necessary document attributes in each content orientation, and obtains the second scoring results corresponding to the second eigenvalues corresponding to each of the all necessary document attributes respectively.
[0168] For example, taking a full score of 10 as an example, for the second eigenvalue of a necessary document attribute, the data processing device can score based on the similarity between the type of this second eigenvalue and the preset type. If they are exactly the same, it can be 10 points.
[0169] Taking a PCF report of a certain product as an example for a piece of initial data. The content orientation of this initial data is the PCF report. The necessary document attributes of the PCF report include: The necessary document attributes of the PCF report include: "Product Chinese Name / Product Foreign Name", "Product Carbon Footprint Result Value", "Product Carbon Footprint - Emission Unit", "Product Carbon Footprint - Product Measurement Unit", "LCA Phase / Life Cycle Phase", and "Data Representativeness". This initial data includes:
[0170] "Product Chinese Name / Product Foreign Name|X Product
[0171] Product Carbon Footprint Result Value|More than 1000
[0172] Product Carbon Footprint - Emission Unit|kgCO2e
[0173] Product Carbon Footprint - Product Measurement Unit|Set
[0174] LCA Phase / Life Cycle Phase|Cradle to Gate
[0175] Data Representativeness|Industry Average
[0176] Geographical Representativeness|China
[0177] Data Statistical Time Period - Annual|2 years
[0178] Data Statistical Time Period - Start Time & End Time|From March 27, 2022 to March 27, 2024
[0179] Manufacturer|Company X”
[0180] The data processing device scores for this initial data, and the obtained second scoring results are shown in Table 5 below:
[0181] Table 5
[0182] Product Chinese Name / Product Foreign Name Product X 10 points Product Carbon Footprint Result Value More than 1000 5 points Product Carbon Footprint - Emission Unit kgCO2e 10 points Product Carbon Footprint - Product Measurement Unit unit 10 points LCA Stage / Life Cycle Stage Cradle to Gate 10 points Data Representativeness Industry Average 10 points
[0183] In some examples, the data processing device can enable the AI agent to master all the necessary file attributes in the content direction and the characteristic values corresponding to the necessary file attributes respectively. Thus, the data processing device can also directly input each piece of initial data into the AI agent, and the AI agent scores the second characteristic values corresponding to all the necessary file attributes respectively, and obtains the second scoring results corresponding to the second characteristic values corresponding to all the necessary file attributes respectively.
[0184] Based on this, for each piece of initial data, in each content direction, after determining that one or more first file attributes include all the necessary file attributes in the content direction, the data processing device scores based on the second characteristic values corresponding to all the necessary file attributes respectively, and obtains the second scoring results corresponding to the second characteristic values corresponding to all the necessary file attributes respectively. In order to use the second scoring results to accurately determine whether the second characteristic values corresponding to all the necessary file attributes in the initial data are valid characteristic values, which helps to ensure the accuracy of the finally obtained valid data.
[0185] S402. The data processing device determines whether the second characteristic values corresponding to all the necessary file attributes are all valid characteristic values according to one or more second scoring results.
[0186] For one or more second scoring results, the data processing device can compare with a preset threshold, or compare the sum with a preset threshold, etc., to determine whether the second characteristic values corresponding to all the necessary file attributes are all valid characteristic values.
[0187] Based on this, the data processing device can judge whether all the necessary file attributes in the initial data all have valid characteristic values, so as to screen out available data from the aspect of characteristic values, eliminate the initial data whose necessary file attributes do not have valid characteristic values, avoid spending more computing resources to reprocess this part of the initial data, improve the efficiency of data processing, and further ensure that the finally obtained valid data does not include data without valid characteristic values, improving the accuracy of data processing.
[0188] Based on the above exemplary description, the following introduces a method for determining whether the second characteristic values corresponding to all the necessary file attributes are all valid characteristic values according to one or more second scoring results.
[0189] Please refer to Figure 5 , Figure 5 which is a flowchart of a method for determining whether the second characteristic values corresponding to all the necessary file attributes are all valid characteristic values provided by an embodiment of the present application. As Figure 5 shown, the method includes:
[0190] S501. The data processing device determines whether one or more second scoring results are all greater than or equal to a second preset threshold.
[0191] If one or more second scoring results are all greater than or equal to the second preset threshold, the data processing device executes S502; if there is a scoring result among one or more second scoring results that is less than the second preset threshold, the data processing device executes S503.
[0192] Taking a 10 - point scale as an example, the second preset threshold can be 8 points. Thus, the data processing device can ensure that all the second eigenvalue corresponding to the necessary file attributes in the determined initial data are valid eigenvalues, thereby improving the accuracy of data processing and ensuring that accurate valid data can be obtained. As shown in Table 5 above, if the first preset threshold is 8 points, there is a second eigenvalue with a score less than 8 points among the scores of the second eigenvalues corresponding to all the necessary file attributes in this initial data. That is, the product carbon footprint result value |1000 is mostly 5 points. Therefore, for this initial data, the data processing device determines that there are invalid eigenvalues among the second eigenvalues corresponding to the necessary file attributes of this initial data. In addition, if there is no second eigenvalue for a necessary file attribute in an initial data, the data processing device can also determine that there are invalid eigenvalues among the second eigenvalues corresponding to the necessary file attributes of this initial data.
[0193] S502. The data processing device determines that all the second eigenvalues corresponding to the necessary file attributes are valid eigenvalues.
[0194] S503. The data processing device determines that there are invalid eigenvalues among all the second eigenvalues corresponding to the necessary file attributes.
[0195] Based on the above - described exemplary description, the method for extracting the first file attribute with the first eigenvalue as a valid eigenvalue and the corresponding first eigenvalue as the content of a valid data will be introduced below.
[0196] Please refer to Figure 6 , Figure 6 which is a flowchart of a method for extracting valid data provided by an embodiment of the present application. As Figure 6 shown, the method includes:
[0197] S601. The data processing device obtains the scoring results corresponding to one or more first file attributes respectively.
[0198] Among them, for an initial data, the scoring results corresponding to one or more first file attributes respectively include the second scoring results corresponding to the second eigenvalues corresponding to all the necessary file attributes respectively, and the third scoring results corresponding to the third eigenvalues of one or more first file attributes and all non - necessary file attributes other than the necessary file attributes.
[0199] The data processing device can score the third eigenvalue in the same way as obtaining the second scoring result to obtain the third scoring result.
[0200] For example, based on Table 2, for the initial data 4, as shown in Table 6 below, the data processing device obtains the necessary file attributes among them: the second scoring results corresponding to file attribute A, file attribute B, and file attribute C are: 10 points, 9 points, and 10 points respectively, and for the non-necessary file attributes: the third scoring results corresponding to file attribute F and file attribute G are: 8 points and 0 points respectively.
[0201] Table 6
[0202]
[0203] S602. The data processing device determines the necessary file attributes with scoring results greater than or equal to the third preset threshold among one or more first file attributes, and determines the non-necessary file attributes with scoring results greater than or equal to the fourth preset threshold among one or more first file attributes as the file attributes to be extracted.
[0204] Among them, the third preset threshold can be equal to the fourth preset threshold, for example, it is 8 points.
[0205] Based on Table 6, taking the 10-point system as an example, if both the third preset threshold and the fourth preset threshold are 8 points, for the initial data 4, the file attributes that the data processing device should extract include: file attribute A, file attribute B, file attribute C, and file attribute F.
[0206] S603. The data processing device obtains a valid piece of data based on all the file attributes to be extracted and the corresponding eigenvalues.
[0207] Based on Table 6, a valid piece of data obtained by the data processing device is shown in Table 7 below:
[0208] Table 7
[0209]
[0210]
[0211] Based on this, for each piece of initial data, in each content direction, when the data processing device can have all necessary file attributes in the content direction among one or more first file attributes, and the second eigenvalue corresponding to each of the all necessary file attributes is an effective eigenvalue, the data processing device extracts the first file attributes with the first eigenvalue being an effective eigenvalue and the corresponding first eigenvalue being a piece of effective data in the content direction, and eliminates the useless data, so as to accurately obtain the useful data, improve the accuracy of data processing, and make the obtained effective data conform to the content direction.
[0212] Based on the above exemplary description, the data processing device can perform the following Figure 7 shown method to train a preset model to obtain a content direction prediction model.
[0213] Please refer to Figure 7 , Figure 7 which is a flowchart of a method for training a content direction prediction model provided by an embodiment of the present application.
[0214] As Figure 7 shown, the method includes:
[0215] S701. The data processing device obtains one or more sample files and the content directions corresponding to the sample files respectively.
[0216] S702. The data processing device trains a preset model according to the file names of the sample files and the content at the specified positions in the sample files to obtain a content direction prediction model.
[0217] Specifically, by means of pre-input, the preset model is made to learn one or more sample files and the content directions corresponding to the sample files respectively.
[0218] During training, the preset model is made to analyze the content features of the sample files from two perspectives: the file name and the content at the specified position. Then, based on this, it is determined whether the input file conforms to the features of the file name and the content at the specified position in the content direction of the file. If not, the user is informed in text that "it does not conform to any content direction".
[0219] For example, when comparing the file name and the content at the specified position, the feature compliance can be used for comparison, scored according to 10 points, and a score higher than 4 points is regarded as a possible content direction, and the corresponding content direction is determined to participate in the subsequent data processing process.
[0220] Based on this, the content direction prediction model can preliminarily determine the content direction corresponding to the input file, simplify the data processing process, and prepare for the subsequent data processing.
[0221] Based on the above exemplary description, after obtaining the valid data, the data processing device can update the content direction prediction model in the following Figure 8 shown manner to obtain an updated content direction prediction model.
[0222] Please refer to Figure 8 , Figure 8 which is a flowchart of a method for updating a content direction prediction model provided by an embodiment of the present application.
[0223] As Figure 8 shown, the method includes:
[0224] S801. The data processing device obtains one or more pieces of valid data and the corresponding content directions of the valid data respectively.
[0225] S802. The data processing device updates the content direction prediction model according to the file name of the valid data and the content at the specified position in the valid data to obtain an updated content direction prediction model.
[0226] If the data processing device updates the content direction prediction model using the valid data, then the file name of the valid data, the content at the specified position, as well as the corresponding file attributes and the feature values corresponding to the file attributes will all be used to update the content direction prediction model, thereby improving the accuracy of the content direction prediction model.
[0227] Based on the above exemplary description, after obtaining the valid data, the data processing device can also store the valid data in the knowledge base for ESG management.
[0228] Please refer to Figure 9 , Figure 9 which is a flowchart of a data processing method provided by an embodiment of the present application.
[0229] As Figure 9 shown, the method further includes:
[0230] S106. Store one or more pieces of valid data in the knowledge base according to the corresponding content directions of the valid data.
[0231] Among them, the knowledge base is applied to the artificial intelligence AI agent.
[0232] Among them, the AI agent is used for ESG management.
[0233] Based on this, the data processing device can import one or more pieces of valid data into the knowledge base. The valid data are all data that conform to the corresponding content directions. Without the intervention of technical personnel, any structured data can be imported with a certain accuracy relying on this method. The data processing device determines the content direction based on the initial data and filters the valid data, thereby completing the structured import.
[0234] In addition, after determining one or more pieces of valid data, the data processing device may display one or more pieces of valid data to the user and ask the user to determine whether to add the valid data to the knowledge base. After receiving the user's confirmation instruction, the valid data is stored in the knowledge base. After storing the valid data in the knowledge base, the data processing device may also inform the user of the import result. For example, it outputs to the user "2 pieces of valid data reported by the PCF have been imported into the knowledge base", thereby improving the user experience.
[0235] Exemplarily, the present application provides a data processing device Figure 10 which is a schematic structural diagram of a data processing device provided in an embodiment of the present application. As Figure 10 shown, the device includes: an acquisition module 101, a prediction module 102, a determination module 103, and an extraction module 104.
[0236] The acquisition module 101 is configured to acquire a file to be processed, where the file to be processed includes one or more pieces of initial data, and the initial data includes one or more first file attributes and first feature values corresponding to the first file attributes;
[0237] The prediction module 102 is configured to input the file to be processed into a pre-trained content direction prediction model to determine one or more content directions corresponding to the file to be processed; the content direction is used to indicate one or more necessary file attributes;
[0238] The determination module 103 is configured to, for each piece of initial data, determine, under each content direction, whether one or more first file attributes include all necessary file attributes under the content direction, and determine whether second feature values corresponding to all necessary file attributes are all valid feature values;
[0239] The extraction module 104 is configured to, if one or more first file attributes include all necessary file attributes under the content direction, and second feature values corresponding to all necessary file attributes are all valid feature values, extract first file attributes with first feature values being valid feature values and the corresponding first feature values as a piece of valid data under the content direction.
[0240] It should be noted that the data processing device in the embodiment of the present application can be used to execute the technical solutions in the above method embodiments, and its implementation principles and technical effects are similar, which will not be elaborated here.
[0241] In some examples, the device further includes: a discard module;
[0242] A discard module, configured to discard initial data if one or more first file attributes do not include all necessary file attributes in the content direction, or if there are invalid eigenvalues corresponding to the necessary file attributes.
[0243] In some examples, the determination module 103 is specifically configured to, for each piece of initial data and in each content direction, score based on one or more necessary file attributes to obtain first scoring results respectively corresponding to the one or more necessary file attributes;
[0244] Based on the one or more first scoring results, determine whether the one or more first file attributes include all necessary file attributes in the content direction.
[0245] In some examples, the determination module 103 is specifically configured to determine that the one or more first file attributes include all necessary file attributes in the content direction when the one or more first scoring results are all greater than or equal to a first preset threshold;
[0246] When there is a scoring result less than the first preset threshold among the one or more first scoring results, determine that the one or more first file attributes do not include all necessary file attributes in the content direction.
[0247] In some examples, the determination module 103 is specifically configured to, for each piece of initial data and in each content direction, after determining that the one or more first file attributes include all necessary file attributes in the content direction, score based on the second eigenvalues respectively corresponding to the all necessary file attributes to obtain second scoring results respectively corresponding to the second eigenvalues respectively corresponding to the all necessary file attributes;
[0248] Based on the one or more second scoring results, determine whether the second eigenvalues respectively corresponding to the all necessary file attributes are all valid eigenvalues.
[0249] In some examples, the determination module 103 is specifically configured to determine that the second eigenvalues respectively corresponding to the all necessary file attributes are all valid eigenvalues when the one or more second scoring results are all greater than or equal to a second preset threshold;
[0250] When there is a scoring result less than the second preset threshold among the one or more second scoring results, determine that there are invalid eigenvalues corresponding to the second eigenvalues respectively corresponding to the all necessary file attributes.
[0251] In some examples, the extraction module 104 is specifically configured to obtain the scoring results respectively corresponding to the one or more first file attributes;
[0252] Determine the necessary file attributes with scoring results greater than or equal to the third preset threshold among one or more first file attributes, and determine the non-necessary file attributes with scoring results greater than or equal to the fourth preset threshold among one or more first file attributes as the file attributes to be extracted;
[0253] Obtain a valid piece of data based on all the file attributes to be extracted and their corresponding feature values.
[0254] In some examples, the device further includes: a training module;
[0255] The training module is used to obtain one or more sample files and the content directions corresponding to the sample files respectively;
[0256] Train a preset model based on the file names of the sample files and the content at specified positions in the sample files to obtain a content direction prediction model.
[0257] In some examples, the device further includes: an update module;
[0258] The update module is used to obtain one or more pieces of valid data and the content directions corresponding to the valid data respectively;
[0259] Update the content direction prediction model based on the file names of the valid data and the content at specified positions in the valid data to obtain an updated content direction prediction model.
[0260] In some examples, the device further includes: an import module;
[0261] The import module is used to store one or more pieces of valid data into the knowledge base according to the content directions corresponding to the valid data. The knowledge base is applied to an artificial intelligence AI agent, and the AI agent is used for environmental protection, social responsibility, and corporate governance ESG management.
[0262] Figure 11 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 11 shown, the electronic device may include: a processor 201. When the processor 201 executes the computer executable program or instruction in the memory, the data processing method shown in the embodiment of the present application is implemented. Figures 1 to 9 shown.
[0263] Figure 12 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 12 shown, the electronic device may include: a processor 301 and a memory 302. A computer program is stored in the memory 302. When the processor 301 executes the computer program, the data processing method shown in the embodiment of the present application is implemented. Figures 1 to 9 shown.
[0264] The electronic device of the present application can be used to implement the technical solutions of the foregoing method embodiments. The implementation principles and technical effects are similar. The operations implemented by each module can be further referred to the relevant descriptions of the method embodiments, which will not be elaborated here. The module here can also be replaced by a component or a circuit.
[0265] Exemplarily, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor to cause the electronic device to implement the method in the foregoing embodiments.
[0266] Exemplarily, the present application provides a computer program product, including: execution instructions, the execution instructions are stored in a readable storage medium, and at least one processor of the electronic device can read the execution instructions from the readable storage medium, and the at least one processor executes the execution instructions to cause the electronic device to implement the method in the foregoing embodiments.
[0267] Exemplarily, the present application further provides a chip. The chip is connected to a memory, or a memory is integrated on the chip. When the software program stored in the memory is executed, the method in the foregoing embodiments is implemented.
[0268] In the above embodiments, all or part of the functions can be implemented by software, hardware, or a combination of software and hardware. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0269] Those of ordinary skill in the art can understand all or part of the processes in implementing the methods of the above embodiments. The processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: various media that can store program codes such as a read-only memory (ROM) or a random access memory (RAM), a magnetic disk, or an optical disc.
[0270] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0271] Those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments, the combination of features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0272] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A data processing method, characterized in that, The method includes: Obtaining a file to be processed, where the file to be processed includes one or more pieces of initial data, and the initial data includes one or more first file attributes and first eigenvalue corresponding to the first file attribute; Inputting the file to be processed into a pre-trained content direction prediction model to determine one or more content directions corresponding to the file to be processed; the content direction is used to indicate one or more necessary file attributes; For each piece of the initial data, in each content direction, determining whether the one or more first file attributes include all necessary file attributes in the content direction, and determining whether second eigenvalues respectively corresponding to all the necessary file attributes are all valid eigenvalues; If the one or more first file attributes include all necessary file attributes in the content direction and second eigenvalues respectively corresponding to all the necessary file attributes are all valid eigenvalues, extracting first file attributes with first eigenvalue being a valid eigenvalue and the corresponding first eigenvalue as a piece of valid data in the content direction.
2. The method according to claim 1, wherein The method further includes: If the one or more first file attributes do not include all necessary file attributes in the content direction, or there is a second eigenvalue corresponding to the necessary file attribute that is an invalid eigenvalue, discarding the initial data.
3. The method according to claim 1 or 2, characterized in that, The determining whether the one or more first file attributes include all necessary file attributes in the content direction includes: For each piece of the initial data, in each content direction, scoring based on the one or more necessary file attributes to obtain first scoring results respectively corresponding to the one or more necessary file attributes; According to one or more of the first scoring results, determining whether the one or more first file attributes include all necessary file attributes in the content direction.
4. The method according to claim 3, characterized in that, The determining, according to one or more of the first scoring results, whether the one or more first file attributes include all necessary file attributes in the content direction includes: When one or more of the first scoring results are all greater than or equal to a first preset threshold, determining that the one or more first file attributes include all necessary file attributes in the content direction; When there is a scoring result less than the first preset threshold among one or more of the first scoring results, determining that the one or more first file attributes do not include all necessary file attributes in the content direction.
5. The method according to claim 1, characterized in that, The determining whether second eigenvalues respectively corresponding to all the necessary file attributes are all valid eigenvalues includes: For each piece of the initial data, in each content direction, after determining that the one or more first file attributes include all necessary file attributes in the content direction, scoring based on second eigenvalues respectively corresponding to all the necessary file attributes to obtain second scoring results respectively corresponding to second eigenvalues respectively corresponding to all the necessary file attributes; According to one or more of the second scoring results, determining whether second eigenvalues respectively corresponding to all the necessary file attributes are all valid eigenvalues.
6. The method according to claim 5, wherein Determining whether the second eigenvalue corresponding to each of all the necessary document attributes is a valid eigenvalue according to one or more of the second scoring results includes: When one or more of the second scoring results are greater than or equal to a second preset threshold, determining that the second eigenvalue corresponding to each of all the necessary document attributes is a valid eigenvalue; When there is a scoring result less than the second preset threshold among one or more of the second scoring results, determining that there is an invalid eigenvalue among the second eigenvalues corresponding to each of all the necessary document attributes.
7. The method according to claim 1, wherein Extracting the first document attribute whose first eigenvalue is a valid eigenvalue and the corresponding first eigenvalue as a piece of valid data in the content direction includes: Obtaining the scoring results corresponding to the one or more first document attributes respectively; Determining the necessary document attributes whose scoring results are greater than or equal to a third preset threshold among the one or more first document attributes as the document attributes to be extracted, and determining the non-necessary document attributes whose scoring results are greater than or equal to a fourth preset threshold among the one or more first document attributes as the document attributes to be extracted; Obtaining a piece of valid data according to all the document attributes to be extracted and the corresponding eigenvalues.
8. The method according to claim 1, characterized in that The method further includes: Obtaining one or more sample documents and the content directions corresponding to the sample documents respectively; Training a preset model according to the file names of the sample documents and the content at specified positions in the sample documents to obtain the content direction prediction model.
9. The method according to claim 1, characterized in that, The method further includes: Obtaining one or more pieces of the valid data and the content directions corresponding to the valid data respectively; Updating the content direction prediction model according to the file names of the valid data and the content at specified positions in the valid data to obtain an updated content direction prediction model.
10. The method according to claim 1, wherein The method further includes: Storing one or more pieces of the valid data into a knowledge base according to the content directions corresponding to the valid data, where the knowledge base is applied to an artificial intelligence AI agent, and the AI agent is used for environmental protection, social responsibility and corporate governance ESG management.
11. An electronic device, characterized in that, Including: A processor; The processor is configured to execute the computer executable program or instruction in the memory, so that the electronic device executes the method according to any one of claims 1-10.
12. An electronic device, characterized in that, Including: At least one memory and at least one processor; The memory is used for storing a computer executable program or instruction; The processor is configured to call the computer executable program or instruction in the memory, so that the electronic device executes the method according to any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer executable program or instruction, and the computer executable program or instruction is configured to execute the method according to any one of claims 1-10.