Component basic information data cleaning method

By employing a layered cleaning strategy and coding mapping relationships, the redundancy and inconsistency issues of basic component information data were resolved, achieving efficient and accurate data cleaning, improving the efficiency of component list organization, and providing high-quality data support for digital supply chain systems.

CN121597666APending Publication Date: 2026-03-03CHINA ACADEMY OF SPACE TECHNOLOGY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511561411.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

In existing technologies, the basic information data of components is redundant, inconsistent, and erroneous, resulting in low efficiency in data sharing and application, especially when dealing with complex and large amounts of data.

Method used

A hierarchical cleaning strategy is adopted to distinguish between data dictionary items and non-data dictionary items. A unified data standard and a comprehensive dictionary are used for automatic identification and normalization. Data cleaning is carried out in combination with coding mapping relationship. Deduplication is achieved through multi-condition comparison and keyword sorting and merging.

Benefits of technology

It improves the efficiency of component list sorting, ensures the accuracy and understandability of data relationships, reduces manual labor, enables data reuse and strategy unification, provides high-quality data support, and lays the foundation for digital supply chain management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121597666A_ABST
    Figure CN121597666A_ABST
Patent Text Reader

Abstract

The invention discloses a component basic information data cleaning method, and belongs to the technical field of data management. The method comprises the following steps: formulating basic information data standard specifications of components; constructing a data dictionary database; performing standardized preprocessing on the multi-source heterogeneous component data; performing basic attribute cleaning of uniqueness verification based on a dictionary rule mapping table; performing multi-condition comparison cleaning on the basic attributes of the non-dictionary class; merging, de-duplication and endowing unified codes; and data enrichment and system online are realized. According to the invention, the problems of multi-source isomerism, non-uniform standard and low cleaning efficiency of component data in the prior art are solved, efficient and accurate component basic information data cleaning is realized, and high-quality data support is provided for a digital supply chain system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for cleaning basic information data of electronic components, belonging to the field of data governance technology. Background Technology

[0002] In the development of aerospace products, components, as fundamental supporting materials, are crucial to product quality and reliability, and the accuracy and completeness of their basic information directly affect the product's quality and reliability. However, due to issues such as component data originating from multiple different research and development units, inconsistent standards, and non-standard attribute descriptions, there is a significant amount of redundancy, inconsistency, and error in the basic information data of components, severely impacting data sharing and application.

[0003] In existing technologies, data cleaning mainly relies on manual verification or simple rule matching, which is inefficient and prone to errors. This is especially true when dealing with materials like electronic components, which have complex technical attributes and large data volumes; traditional methods struggle to meet practical needs. Therefore, there is an urgent need for an efficient and accurate method for cleaning basic information data on electronic components to provide high-quality data support for digital supply chain systems.

[0004] Existing technologies suffer from problems such as heterogeneous and multi-source component data, inconsistent standards, and low cleaning efficiency. Summary of the Invention

[0005] The technical problem solved by this invention is to overcome the shortcomings of the prior art and provide a method for cleaning basic information data of components, which effectively improves the efficiency of component list sorting.

[0006] The technical solution of this invention is:

[0007] This invention discloses a method for cleaning basic information data of electronic components, including:

[0008] Determine the basic information attributes of the components that need to be cleaned;

[0009] Distinguish whether basic information attribute items belong to data dictionary items or non-data dictionary items;

[0010] Determine the dictionary base for the data dictionary items;

[0011] Determine the cleaning principles for non-data dictionary items;

[0012] Construct the header of the original data collection form;

[0013] Clearly define the requirements for raw data collection;

[0014] Collect component data according to the header of the original data collection form and the original data collection requirements;

[0015] Encode and preprocess the collected components to obtain the unique identification code of the collected components and the standardized component information;

[0016] According to the dictionary library of the data dictionary items, match the corresponding standard values for the data dictionary items in the standardized component information;

[0017] Based on the non-data dictionary item cleaning principle, clean the non-data dictionary items in the standardized component information to obtain the cleaned standardized data;

[0018] Merge and deduplicate the cleaned standardized data to obtain the cleaned data;

[0019] Encode the cleaned data to obtain the unique identification code of the cleaned data;

[0020] Establish the coding mapping relationship between the unique identification code of the collected components and the unique identification code of the cleaned data.

[0021] Further, in the above method, the basic information attribute items of the components to be cleaned include: large category, medium category, small category, fine category, product name, model or series, model specification, quality grade, packaging form, manufacturer, detailed specification number, special instructions, external dimensions, general specification, measurement unit, domestic or imported, country of origin and supply status.

[0022] Further, in the above method, the method for distinguishing whether the basic information attribute items belong to data dictionary items is specifically: the large category, medium category, small category, fine category, quality grade, packaging form, manufacturer, measurement unit, domestic or imported, country of origin and supply status belong to data dictionary items, and the product name, model or series, model specification, detailed specification number, general specification, special instructions and external dimensions belong to non-data dictionary items.

[0023] Further, in the above method, the method for determining the dictionary library of the data dictionary items is specifically:

[0024] The dictionary libraries of the large category, medium category, small category and fine category are respectively: Level I of Q / QJA40, Level II of Q / QJA40, Level III of Q / QJA40, Level IV of Q / QJA40;

[0025] The dictionary libraries of the quality grade, packaging form and manufacturer are respectively: 317 items in the quality grade dictionary, 1828 items in the packaging form dictionary, and 755 items in the manufacturer dictionary;

[0026] The dictionary library of the measurement unit is: only or meter or set or piece;

[0027] The dictionary library of domestic or imported is: domestic or imported;

[0028] The dictionary library of the country of origin is: 22 items in the country of origin dictionary;

[0029] The dictionary for supply status is: normal, new, discontinued, or about to be discontinued.

[0030] Furthermore, in the above method, the cleaning principle for determining non-data dictionary items specifically includes:

[0031] The cleaning principles for each model and specification are as follows:

[0032] The model and specifications should not include quality grade information;

[0033] Select English and half-width format for all special characters;

[0034] Remove any extra spaces before or after the content of the fields you enter;

[0035] The general cleaning principles are:

[0036] The specific parameter for transistor products is the device amplification factor.

[0037] The specific parameter for cable assembly products is the length of the crimped cable.

[0038] If there is no such line, write a short horizontal bar and a half-width character.

[0039] The cleaning principle for external dimensions is:

[0040] Components with standard package designations already filled in;

[0041] Fill in the dimensions according to the specifications in the product manual;

[0042] When no unit is specified, the default value is millimeters (mm).

[0043] Use metric units;

[0044] If there is no such thing, write a short horizontal bar.

[0045] Furthermore, the above method specifies the requirements for raw data collection, specifically:

[0046] Components of Class I, II, III, and IV shall be handled in accordance with "Q / QJA 40.1A-2023 Classification and Code of Supporting Materials for Aerospace Models Part 1: Electronic, Electrical and Electromechanical Components";

[0047] The quality grade, packaging form, manufacturer, and unit of measurement should be filled in according to the dictionary values.

[0048] When filling in the content of each attribute item, use the half-width input method;

[0049] Use a minus sign to connect range values;

[0050] Required fields: If none, fill in a short hyphen (half-width). Optional fields: If none, leave blank.

[0051] Furthermore, in the above method, the collected components are encoded, identified, and preprocessed to obtain standardized component information, specifically as follows:

[0052] Each collected component is assigned a unique identifier, resulting in the identified component.

[0053] The labeled components undergo standardized preprocessing, including removing leading and trailing spaces, unifying capitalization, standardizing bracket format, and removing redundant spaces, to obtain standardized component information.

[0054] The advantages of this invention over the prior art are as follows:

[0055] (1) This invention provides a reliable basis for data cleaning by formulating unified data standards and building a complete data dictionary library. It can automatically identify and standardize non-standard information of components based on data dictionary rules, ensure the correctness and understandability of the relationship between data, avoid data redundancy, reduce repetitive manual labor, realize the reuse of the results of component data cleaning, the accumulation of experience, and the unification of strategies, and effectively improve the efficiency of component list sorting.

[0056] (2) The present invention adopts a layered cleaning strategy, first processing dictionary-type attributes and then processing non-dictionary-type attributes, which solves the low-level data quality problem while improving cleaning efficiency and accuracy. It uses the coding mapping relationship to link the originally scattered data, strengthens the practical application of data dictionary and data standard, and ensures the information processing of component list "from acquisition to governance, from deconstruction to matching".

[0057] (3) This invention uses a multi-condition comparison method and keyword sorting to merge and remove duplicates. It compares the corresponding fields of the requirement list and the standard data source to perform data association matching, which effectively solves the problem of standardization of component information caused by inconsistent reporting sources and differences in the technical level of reporting personnel.

[0058] (4) This invention realizes full-process management from data cleaning to system launch, provides a convenient component basic information cleaning tool for various engineer roles related to the filling, maintenance and statistical analysis of component list information, meets the needs of rapid batch processing of BOM list-level component basic information, unifies component material codes and basic data items, lays the foundation for digital collaborative development, and provides high-quality component data support and efficient governance technology for digital management of the supply chain. Attached Figure Description

[0059] Figure 1This is a flowchart of the component basic information data cleaning method according to an embodiment of the present invention. Detailed Implementation

[0060] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] like Figure 1 As shown, the present invention provides a method for cleaning basic information data of electronic components, comprising the following steps:

[0062] This invention takes the data cleaning of basic component information in the construction of an information system of a certain unit as an example. According to the requirements, it completed the cleaning of more than 1 million pieces of component data, providing high-quality data support for the digital supply chain system.

[0063] Step 1: Determine the basic information attributes of the components that need to be cleaned.

[0064] In this embodiment, the basic information attributes for determining the components that need to be cleaned include: category, medium category, minor category, subcategory, product name, model or series, model specifications, quality grade, packaging form, manufacturer, detailed specification number, special instructions, external dimensions, general specifications, unit of measurement, domestic / imported, country of origin, and supply status.

[0065] Step 2: Distinguish whether the basic attribute items belong to the data dictionary items.

[0066] In this embodiment, the categories, intermediate categories, minor categories, subcategories, quality grades, packaging forms, manufacturers, units of measurement, domestic / imported, country of origin, and supply status are defined as dictionary items, while product names, models or series, model specifications, detailed specification numbers, general specifications, special instructions, and external dimensions are not defined as data dictionary items.

[0067] Step 3: Determine the dictionary base for data dictionary items and the cleaning principles for non-data dictionary items.

[0068] Sub-step S31: Determine the dictionary database for major categories, intermediate categories, minor categories, subcategories, quality grades, packaging forms, manufacturers, units of measurement, domestic / imported, country of origin, and supply status of data dictionary items, as shown in Table 1 below;

[0069] Table 1. Dictionary database information for dictionary entries.

[0070]

[0071]

[0072] Sub-step S32: Determine the cleaning principles for non-data dictionary items such as product name, model or series, model specifications, detailed specification number, general specifications, special instructions, and external dimensions. See Table 2 below.

[0073] Table 2 Cleaning Principles for Non-Dictionary Items

[0074]

[0075]

[0076]

[0077] Step 4: Construct the header of the raw data collection form and clarify the raw data collection requirements.

[0078] Sub-step S41: Construct the header of the original data collection form as shown in the table below:

[0079] Table 3. Header of the raw data collection form

[0080]

[0081] Sub-step S42: Clarify the requirements for raw data collection as shown in Table 4 below.

[0082] Table 4 Requirements for Raw Data Collection

[0083]

[0084] Step 5: Summarize, code, and perform standardized preprocessing on the original collected component data.

[0085] Sub-step S51: Summarize the basic data of original components collected from multiple sources;

[0086] Table 5. Basic data of original components collected from multiple sources.

[0087]

[0088] Sub-step S52: Establish a unique identification code for the summarized raw data;

[0089] Table 6. Correspondence between unique identifiers and data

[0090] Unique Identifier data WZ0100000000001 Data 1 WZ0100000000002 Data 2 WZ0100000000003 Data 3 WZ0100000000004 Data 4 WZ0100000000005 Data 5 WZ0100000000006 Data 6 WZ0100000000007 Data 7

[0091] Sub-step S53: Perform normalization preprocessing on the summarized raw data, including removing leading and trailing spaces, unifying capitalization, standardizing bracket format, and removing redundant spaces;

[0092] Table 7 Examples of raw and preprocessed data

[0093] Raw data Preprocessed data ab CD / (20-50)-12-30F / SP ABCD / (20-50)-12-30F / SP

[0094] Step 6: Based on the dictionary rule mapping table, match the data dictionary items with standardized values, and based on the non-data dictionary item cleaning principle, perform multi-condition comparison and standardization on the non-data dictionary items.

[0095] Sub-step S61: Based on the dictionary rule mapping table, match the data dictionary items with the relevant standard specifications for major category, medium category, minor category, sub-category, quality grade, packaging form, manufacturer, unit of measurement, domestic / imported, country of origin, and supply status according to various basic information of components, such as detailed component specifications, legal entity registration information of component manufacturers, component packaging standards, component quality grade standards, component classification standards, etc.

[0096] Table 8. Dictionary Item Data Cleaning Diagram

[0097]

[0098] Sub-step S62: Based on the non-data dictionary item cleaning principle, perform multi-condition comparison and standardization of non-data dictionary item product name, model or series, model specifications, detailed specification number, general specifications, special instructions, and external dimensions.

[0099] Table 9. Cleaning Results of Non-Data Dictionary Items

[0100]

[0101]

[0102]

[0103] Step 7: Merge and deduplicate the cleaned and normalized data.

[0104] Unique Identifier Raw data After standardization deal with WZ0100000000001 Data 1 Data 1 reserve WZ0100000000002 Data 2 Data 1 delete WZ0100000000003 Data 3 Data 1 delete WZ0100000000004 Data 4 Data 1 delete

[0105] Step 8: Establish the encoding mapping relationship between the original data and the cleaned data.

[0106] Table 10 Unique Identifiers for Data After Cleaning

[0107]

[0108]

[0109] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solutions of the present invention by utilizing the methods and techniques disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the protection scope of the technical solutions of the present invention.

[0110] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.

[0111] The contents not described in detail in this specification are common knowledge to those skilled in the art.

Claims

1. A method for cleaning basic information data of electronic components, characterized in that, including: Determine the basic information attribute items of components to be cleaned; Distinguish whether the basic information attribute items belong to data dictionary items or non-data dictionary items; Determine the dictionary library of data dictionary items; Determine the cleaning principles of non-data dictionary items; Construct the header of the original data collection form; Clarify the requirements for collecting original data; Collect component data according to the header of the original data collection form and the requirements for collecting original data; Encode and preprocess the collected components to obtain the unique identification code of the collected components and the standardized component information; Match the data dictionary items in the standardized component information with the corresponding standard values according to the dictionary library of the data dictionary items; Clean the non-data dictionary items in the standardized component information based on the cleaning principles of non-data dictionary items to obtain the cleaned standardized data; Merge and deduplicate the cleaned standardized data to obtain the cleaned data; Encode the cleaned data to obtain the unique identification code of the cleaned data; Establish the coding mapping relationship between the unique identification code of the collected components and the unique identification code of the cleaned data.

2. The component basic information data cleaning method according to claim 1, characterized in that, The basic information attribute items of the components to be cleaned include: major category, medium category, minor category, detailed category, product name, model or series, model specification, quality grade, packaging form, manufacturer, detailed specification number, special instructions, external dimensions, general specification, measurement unit, domestic or imported, country of origin and supply status.

3. The component basic information data cleaning method according to claim 1, characterized in that, The specific method for distinguishing whether the basic information attribute items belong to data dictionary items is: the major category, medium category, minor category, detailed category, quality grade, packaging form, manufacturer, measurement unit, domestic or imported, country of origin and supply status belong to data dictionary items, and the product name, model or series, model specification, detailed specification number, general specification, special instructions and external dimensions belong to non-data dictionary items.

4. The component basic information data cleaning method according to claim 1, characterized in that, The specific method for determining the dictionary library of data dictionary items is: The dictionary libraries of the major category, medium category, minor category and detailed category are respectively: Grade I of Q / QJA40, Grade II of Q / QJA40, Grade III of Q / QJA40, Grade IV of Q / QJA40; Quality grade dictionary, packaging form dictionary, manufacturer dictionary; The dictionary library of the measurement unit is: only or meter or set or piece; Domestic or imported dictionary library; The dictionary library of the supply status is: normal or new product or out of production or about to be out of production.

5. The component basic information data cleaning method according to claim 1, characterized in that, The specific method for determining the cleaning principles of non-data dictionary items is: The cleaning principle of the model specification is: The model specification should not contain quality grade information; All special symbols are selected in English and half-width format; Delete the extra spaces before and after the content of the filled fields; The cleaning principle of the general specification is: The selected parameter of the transistor product is the device amplification factor parameter; The selected parameter of the cable assembly product is the specific length of the crimped cable; Write a short dash in half-width if there is none; The cleaning principle of the external dimensions is: Components with standard package codes filled in; Fill in according to the external dimensions specified in the product manual; When not writing the dimension, it is defaulted to millimeters (mm); Use the metric dimension; Write a half-width short dash if there is none.

6. The component basic information data cleaning method according to claim 1, characterized in that: Clarify the requirements for collecting original data, specifically: Components of Class I, II, III, and IV shall be handled in accordance with "Q / QJA 40.1A-2023 Classification and Code of Supporting Materials for Aerospace Models Part 1: Electronic, Electrical and Electromechanical Components"; The quality grade, packaging form, manufacturer, and unit of measurement should be filled in according to the dictionary values. When filling in the content of each attribute item, use the half-width input method; Use a minus sign to connect range values; Required fields: If none, fill in a short hyphen (half-width). Optional fields: If none, leave blank.

7. The component basic information data cleaning method according to claim 1, characterized in that, The collected components are coded, identified, and preprocessed to obtain standardized component information, specifically: Each collected component is assigned a unique identifier, resulting in the identified component. The labeled components undergo standardized preprocessing, including removing leading and trailing spaces, unifying capitalization, standardizing bracket format, and removing redundant spaces, to obtain standardized component information.

Citation Information

Patent Citations

  • Material directory-based material code generation and management method

    CN107122947A

  • Electronic component data cleaning method and equipment

    CN110647521A

  • Data processing method and system based on electronic component product data dictionary

    CN120087893A

  • Artificially intelligent master data management

    US20220207007A1