Electronic material automatic classification method and system based on multi-level feature recognition
By employing a multi-level feature recognition method, the problem of automatic classification and coding of unstructured electronic material information was solved, achieving efficient and accurate material information processing, generating standard production codes, and improving the robustness and accuracy of the automation solution.
Patent Information
- Application Number
- CN202511437868.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Existing technologies struggle to effectively process unstructured electronic material information, resulting in inefficient and error-prone manual coding. Automated solutions lack deep understanding and context awareness, making it difficult to handle complex synonyms and diverse descriptions.
A multi-level feature recognition method is adopted, including unstructured material information acquisition, material information normalization and preprocessing, main category identification and feature word tagging, structured feature extraction and parsing, normalized coding and code segment generation, and finally generating standard production codes.
It enables efficient and accurate automatic classification and coding of electronic materials, replacing inefficient and error-prone manual operations. It has strong analytical capabilities and robustness, and solves the fundamental problem of describing diversity.
Smart Images

Figure CN120929989B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of material classification, and more specifically, to an automatic classification method and system for electronic materials based on multi-level feature recognition. Background Technology
[0002] With the rapid development of electronic information technology, the types and quantities of electronic components have exploded, posing significant challenges to the research, development, production, procurement, and warehousing management of electronic products. To achieve refined and automated management of massive amounts of electronic materials, constructing a scientific and unified material coding system is crucial. However, in actual business processes, material information often exists in unstructured text form, such as engineers' design documents, suppliers' bills of materials, or material descriptions in procurement systems. These descriptions exhibit significant inconsistencies in format, terminology, and units; for example, 10uF25V and Cap 10mF 25Volt may refer to the same material. This heterogeneity of information sources and the arbitrariness of descriptions make it difficult for computers to directly process and identify material information, heavily relying on manual classification and coding.
[0003] Currently, the industry commonly uses manual coding or simple semi-automatic coding methods based on keyword matching to solve this problem. Manual coding is not only inefficient and costly, but also highly susceptible to errors due to differences in operator experience and subjective judgment. Even a minor coding mistake can lead to serious consequences such as procurement errors and production line shutdowns. Existing simple automation solutions mostly use fixed rules or templates for keyword matching. While this method improves efficiency to some extent, its core flaw lies in its lack of deep understanding of material characteristics and contextual awareness. It struggles to handle complex synonyms, abbreviations, and inverted parameter orders. When encountering new material types or non-standard descriptions, the system frequently malfunctions or fails to recognize them, resulting in extremely poor robustness and scalability. Therefore, the fundamental reason why existing technologies have failed to popularize truly fully automated classification solutions is the inability to effectively solve the challenge of intelligent recognition and parsing of unstructured text into structured features.
[0004] To completely solve the aforementioned technical pain points and achieve efficient, accurate, and automatic classification and encoding of unstructured electronic material information, it is urgent to propose a new technical solution that can simulate expert experience and deeply analyze the inherent logic of material descriptions. Summary of the Invention
[0005] To address the aforementioned fundamental problems, according to one aspect of this application, an automatic classification method for electronic materials based on multi-level feature recognition is provided, which includes: acquiring unstructured material information.
[0006] Unstructured material information is normalized and preprocessed to obtain preprocessed material information data.
[0007] The preprocessed material information data is subjected to main category identification and feature word tagging to obtain the tagged material information data.
[0008] Structured features are extracted and parsed from the labeled material information data to obtain the structured features of the material information.
[0009] The structured features of material information are subjected to normalized encoding and code segment generation to obtain a set of encoded material information segments.
[0010] The material information is encoded, integrated, and verified to output the final production code.
[0011] According to another aspect of this application, an automatic electronic material classification system based on multi-level feature recognition is provided, which includes: an unstructured material information acquisition module for acquiring unstructured material information.
[0012] The material information data preprocessing module is used to normalize and preprocess unstructured material information to obtain preprocessed material information data.
[0013] The material information data identification and labeling module is used to identify the main category and label the feature words of the preprocessed material information data to obtain labeled material information data.
[0014] The material information parsing module is used to extract and parse the structured features of the tagged material information data to obtain the structured features of the material information.
[0015] The material information coding module is used to perform standardized coding and code segment generation on the structured features of material information to obtain a set of coded material information segments.
[0016] The production coding output module is used to encode, integrate, and verify the set of coded segments of material information to obtain the final production code.
[0017] Compared with existing technologies, this application provides an automatic classification method and system for electronic materials based on multi-level feature recognition. It constructs a progressively deeper automated processing flow to address the unstructured information processing challenges mentioned in the background technology. Specifically, it acquires unstructured material information and, through information normalization and preprocessing, first solves the data heterogeneity problem caused by inconsistencies in format, units, and terminology. Subsequently, through main category identification and feature term tagging, it achieves accurate positioning of the core attributes of the material, surpassing the shallow recognition of traditional keyword matching and effectively overcoming its inability to handle complex semantics. Based on this, it further extracts and parses the labeled features in a structured manner, transforming them into standardized features that can be deeply understood by machines. Finally, through encoding integration and verification, it automatically generates a standard, unique production code, such as the SPYY.XBBBVVVV structure. This replaces inefficient and error-prone manual operations and, with its powerful parsing capabilities and robustness, solves the fundamental problem that existing automated solutions struggle to handle diverse descriptions. Attached Figure Description
[0018] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings.
[0019] Figure 1 This is a flowchart of an automatic classification method for electronic materials based on multi-level feature recognition according to an embodiment of this application.
[0020] Figure 2 This is a schematic diagram of the data flow of an automatic electronic material classification method based on multi-level feature recognition according to an embodiment of this application.
[0021] Figure 3 This is a flowchart of step 2 in the automatic classification method for electronic materials based on multi-level feature recognition according to an embodiment of this application.
[0022] Figure 4 This is a block diagram of an automatic electronic material classification system based on multi-level feature recognition according to an embodiment of this application. Detailed Implementation
[0023] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. It should be understood that the drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0024] It is worth mentioning that this application is based on the following coding rules. Specifically, it uses a structured and information-rich electronic material coding system, the core of which is a 12-digit code called the production code, in the format SPYY.XBBBVVVV. This coding system aims to intuitively reflect the key characteristics of components through the code itself, facilitating collaborative work among engineers in selection, procurement, and production. The code consists of two main parts: the main coding segment XBBBVVVV and the sub-coding segment SPYY. The main coding segment defines the core attributes of the material, where X is a single letter representing the major category of the component, such as R for resistors and C for capacitors; BBBVVVV represents the two most important values of the material, with BBB representing power, withstand voltage, and other capability values, and VVVV representing resistance, capacitance, and other main parameter values. These two fields mostly use scientific notation, with the first few significant digits and the last digit representing the exponent. The sub-coding segment SPYY describes the auxiliary characteristics of the material. S and P together define the physical form and specific packaging of the material. The first 'S', when a number, indicates a common industry material and distinguishes between surface mount and through-hole mounting methods; when a letter, it indicates a proprietary material within the company. The final 'YY' field represents the key characteristics of the device, and its meaning dynamically changes according to the device category (X). For example, for resistors and capacitors, it typically indicates the accuracy class; while for integrated circuits, it indicates the operating temperature range, such as commercial or industrial grade. Through this layered and segmented fine-grained definition, this coding rule achieves a highly structured and easily parsed material identification system.
[0025] Based on this, this application proposes an automatic classification method for electronic materials based on multi-level feature recognition. Figure 1 This is a flowchart of an automatic classification method for electronic materials based on multi-level feature recognition according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow in an automatic electronic material classification method based on multi-level feature recognition according to an embodiment of this application. Figure 1 and Figure 2 As shown, the automatic classification method for electronic materials based on multi-level feature recognition according to an embodiment of this application includes: Step 1, acquiring unstructured material information; Step 2, performing material information normalization and preprocessing on the unstructured material information to obtain preprocessed material information data; Step 3, performing main category identification and feature word tagging on the preprocessed material information data to obtain tagged material information data; Step 4, performing structured feature extraction and parsing on the tagged material information data to obtain structured features of material information; Step 5, performing normalized encoding and code segment generation on the structured features of material information to obtain a set of encoded segments of material information; Step 6, performing encoding integration and verification output on the set of encoded segments of material information to obtain the final production code.
[0026] In step 1, unstructured material information is acquired. It should be understood that in the supply chain management of the electronics industry, material information is the core data flow that runs through all stages, including R&D, procurement, production, and warehousing. However, this information often exists in a highly heterogeneous and unstructured form at the source, scattered in engineers' design documents, supplier-provided spreadsheets, or text fields in enterprise resource planning systems. These descriptions not only lack a unified format, but their terminology, units, and parameter order are also highly arbitrary, posing a significant obstacle to automated processing. Therefore, in order to effectively and systematically analyze and classify this chaotic and disordered raw data, the unstructured material information is acquired first. These diverse and formatted raw material descriptions are captured and aggregated from their respective isolated data sources to form an initial data flow that can be analyzed by subsequent intelligent processing programs.
[0027] In an exemplary embodiment of this application, step 1 operates as follows: Obtaining unstructured material information involves establishing an interface capable of accessing and extracting original material descriptions from various external data sources. Unstructured material information specifically refers to text strings or data objects that have not undergone any processing and retain their original format and content; it reflects the appearance of the material at the time of initial recording. Obtaining this information is not a single action but a configurable data access process. This process can interface with an enterprise's internal material management system, database, or a shared folder storing supplier material lists through a pre-defined interface adapter.
[0028] Specifically, a data source configuration file is predefined first. This file specifies the types of information sources to be accessed, such as databases, file systems, or application programming interfaces (APIs). For each source, specific access parameters must be set. Taking retrieving information from a file as an example, the configuration file will specify the file's path, filename, file encoding format (e.g., UTF-8), and the column name or index containing the key fields for material descriptions. For instance, a configuration item can be set to access a file named "Supplier A Bill of Materials.csv" and specify that data should be extracted from the column named "Material Description".
[0029] When the automatic classification process starts, the data acquisition module will perform the corresponding reading operations according to the above configuration. It will access the specified path, open the target file, and scan line by line. When it scans a certain line, it will accurately locate the cell containing the specific information based on the configured column name and material description. For example, when processing the description "Resistor,SMD,0603,10k Ohm,±1%,1 / 8W,Thick Film" from the supplier list, the acquisition module will read this string of characters completely, and the output will be an independent unstructured material information object loaded into the processing flow. This object internally encapsulates the raw text string that was just read, namely "Resistor,SMD,0603,10k Ohm,±1%,1 / 8W,ThickFilm".
[0030] In step 2, the unstructured material information undergoes material information normalization and preprocessing to obtain preprocessed material information data. Correspondingly, the original data originates from unstructured text from various sources, inevitably containing a large amount of noise, such as mixed uppercase and lowercase letters, arbitrary switching between Chinese and English punctuation, and the coexistence of multiple synonyms, abbreviations, and colloquialisms for the same physical concept. More complexly, the writing formats for numerical values and units representing physical quantities vary widely, ranging from engineering abbreviations like 10 k Ohm to purely numerical expressions like 0.001 uF, and fractional forms like 1 / 8W. This high heterogeneity and non-standardization of the data constitute a fundamental obstacle to subsequent accurate feature recognition and intelligent classification. Without processing, any subsequent analysis algorithm will frequently err due to the inability to compare and judge on a unified benchmark, resulting in low recognition accuracy and poor model robustness. Therefore, in order to build a robust and reliable automated classification process, it is necessary to deeply clean and standardize these raw and messy text data to lay a clean, consistent and machine-friendly data foundation for subsequent advanced processing steps such as main category identification, feature extraction and encoding generation.
[0031] In one exemplary embodiment of this application, Figure 3 This is a flowchart of step 2 in the automatic classification method for electronic materials based on multi-level feature recognition according to an embodiment of this application. Figure 3 As shown, step 2 involves material information normalization and preprocessing of unstructured material information to obtain preprocessed material information data, including: step 21, obtaining and integrating material descriptions from unstructured material information to obtain original text strings; step 22, performing text normalization processing on the original text strings to obtain normalized text strings; and step 23, performing structured encapsulation and output on the normalized text strings to obtain the preprocessed material information data.
[0032] In the exemplary embodiment described above, step 2 operates as follows: First, step 21 is executed to accurately extract the core material description text from the input unstructured material information encapsulation object. During implementation, the input object is accessed, and a predefined target field is read, which stores the original material description. For a single descriptive string, such as "Resistor,SMD,0603,10k Ohm,±1%,1 / 8W,Thick Film" mentioned above, it is directly extracted. The input processed is the aforementioned unstructured material information object; after the extraction operation, a pure text string, i.e., the original text string, is output. In this example, the output original text string is: "Resistor,SMD,0603,10k Ohm,±1%,1 / 8W,Thick Film".
[0033] Next, step 22 is performed to normalize the original text string. In an exemplary embodiment of this application, step 22, normalizing the original text string to obtain a normalized text string, includes: step 221, performing a format conversion operation on the original text string to obtain a lowercase text string; step 222, performing symbol filtering and synonym replacement on the lowercase text string to obtain filtered and replaced text; and step 223, normalizing the units and values of the filtered and replaced text to obtain the normalized text string.
[0034] Specifically, the normalization process first executes step 221, the format conversion operation. This operation eliminates matching failures that may be caused by inconsistent capitalization in the text. It calls a standard text processing function to iterate through the entire original text string, such as "Resistor,SMD, 0603, 10k Ohm, ±1%,1 / 8W,Thick Film", converting all English characters to lowercase. After the conversion to lowercase, a lowercase text string is output. For example, the output becomes: "resistor,smd, 0603, 10k ohm, ±1%,1 / 8w, thick film". Next, based on the lowercase text string output in the previous step, step 222, symbol filtering and synonym replacement, is executed. The first step is symbol filtering, which relies on a pre-defined set of symbol rules. This set of rules is defined by domain experts based on industry conventions for describing electronic components, clearly specifying which symbols should be retained and which should be removed or replaced. For example, the rule set can be configured to: retain symbols representing key parameters such as "%", ".", and "-", as well as specific unit symbols such as "ω" and "μ"; while uniformly replacing symbols used as separators or modifiers such as commas, semicolons, and ± signs with spaces, and ensuring that multiple consecutive spaces are compressed into one. When processing the string "resistor,smd, 0603, 10k ohm,±1%,1 / 8w, thick film", all commas, "", and plus / minus signs "±" are converted to spaces, resulting in the intermediate result "resistor smd 0603 10k ohm 1% 1 / 8w thick film". Next, synonym replacement is performed, which loads a pre-defined thesaurus of synonyms and abbreviations. This dictionary, also maintained by experts, maps common non-standard terms, full names, or abbreviations to standard expressions. For example, the dictionary contains the mapping rule {"ohm": "ω","resistor": "res", "capacitor": "cap"}. The processor performs token matching on the filtered text, replacing "ohm" with its standard unit symbol ω and "resistor" with its standard abbreviation "res". The final output is the filtered and replaced text: "res smd 0603 10k ω 1% 1 / 8w thick film". Finally, based on the filtered and replaced text, step 223, unit and numerical standardization, is performed. All numerical values and units representing physical quantities in the text are uniformly converted to standard scientific notation or decimal format, completely eliminating the diversity of expression forms. This operation relies on a unit and numerical standardization rule base, which consists of a series of regular expressions for matching specific patterns and corresponding conversion logic.For example, one rule is specifically designed to identify and convert power values in fractional form. It matches patterns like 1 / 8w using regular expressions and performs internal calculations to convert it to the decimal form 0.125w. Another rule might handle other non-standard units, such as normalizing uf to μf. When processing the filtered and replaced text "ressmd 0603 10kω 1% 1 / 8w thick film", the fractional power 1 / 8w is successfully converted to 0.125w. Relatively normalized expressions like 10kω and 1% remain unchanged at this stage. The final output is the normalized text string, such as "res smd 0603 10kω 1% 0.125w thick film".
[0035] Finally, step 23 is performed for structured encapsulation. This process first performs text segmentation. This operation takes a normalized text string as direct input. It parses the input string based on a preset delimiter, which in this embodiment is set to a single space character. A segmentation module scans the string "ressmd 0603 10kω 1% 0.125w thick film" from beginning to end, extracting the preceding continuous character sequence as an independent word whenever a space is encountered. This process divides the continuous text stream into an ordered sequence of multiple string elements. The output of this operation is a word array, specifically: ["res","smd","0603","10kω","1%","0.125w","thick","film"]. After generating the word array, the data object construction stage begins. This stage gathers all the key data from this preprocessing stage, including the original text string, the normalized text string, and the newly generated word array. A data construction module creates a new, structured data object and assigns values according to a predefined key-value pair format. Specifically, it creates three keys: the first key is the raw text, with the corresponding value being the raw text string "Resistor,SMD, 0603, 10k Ohm, ±1%,1 / 8W,Thick Film"; the second key is the normalized text, with the corresponding value being the normalized text string "res smd 0603 10kω 1% 0.125w thick film"; and the third key is the lexicon array, with the corresponding value being the lexicon array generated in the previous step. Thus, the final output preprocessed material information data is a well-structured and complete data object. The object is a structured data set containing three key-value pairs, with the following contents: "Original Text": "Resistor,SMD, 0603, 10k Ohm, ±1%, 1 / 8W, Thick Film", "Normalized Text": "res smd 0603 10kω1% 0.125w thick film", and "Term Array": ["res","smd","0603","10kω","1%","0.125w","thick","film"]. This object completely encapsulates the entire transformation result from the original information to an analyzable term sequence.
[0036] In step 3, the preprocessed material information data undergoes main category identification and feature term labeling to obtain labeled material information data. It's understandable that after the previous stage of processing, although the original, chaotic material description has been transformed into a clean, standardized term array, these terms themselves are still independent strings lacking inherent semantic connections. For example, the term 0603 could be a specific package size or part of a parameter value; and 10kω, as a string, has not yet revealed its identity as the core material parameter "resistance value." This lack of semantics makes subsequent structured data extraction and encoding impossible. Therefore, before proceeding to specific numerical analysis and encoding, the main category identification and feature term labeling steps are necessary. That is, firstly, the fundamental attributes of the material are determined through key identification words, i.e., which main category it belongs to (e.g., resistance, capacitance). Completing this step provides a crucial contextual environment for all subsequent analysis work. Subsequently, guided by this context, each lexical unit in the lexical array with specific engineering significance is precisely assigned its corresponding feature identity label, such as specific packaging, resistance value, and power, thereby transforming an undifferentiated lexical sequence into a labeled data structure with rich semantic annotations that can be deeply understood by machines.
[0037] In one exemplary embodiment of this application, Figure 4 This is a flowchart of step 3 in the automatic classification method for electronic materials based on multi-level feature recognition according to an embodiment of this application. Figure 4 As shown, step 3, which involves performing main category identification and feature word tagging on the preprocessed material information data to obtain tagged material information data, includes: step 31, extracting a material information word array from the preprocessed material information data; step 32, performing main category identification and context rule loading on the material information word array to obtain the device category X code, context rule set, and initial tagged word list; step 33, performing word traversal and rule application on the material information word array based on the context rule set to obtain an intermediate tagged object sequence; step 34, merging the intermediate tagged object sequence and the initial tagged word list to obtain a final tagged word list; and step 35, encapsulating the preprocessed material information data, the final tagged word list, and the device category X code into a data structure to obtain the tagged material information data.
[0038] In the above exemplary embodiment, step 3 is performed as follows: First, step 31 is executed to extract the material information token array. This process is a direct data extraction operation that separates the token array from the preprocessed material information data. Specifically, the input preprocessed material information data object is parsed, and the value corresponding to the field with the key name "tokens" is read. The output is a pure, ordered string array, i.e., the material information token array. For example, the output material information token array is: ["res","smd","0603","10kω","1%","0.125w","thick","film"].
[0039] Next, step 32 is executed to perform main category identification and context rule loading. This step begins with the construction and loading of the category dictionary. A static, pre-built category dictionary is loaded for querying. This dictionary is a key-value pair data structure, constructed based on Table 1 and completed by domain experts according to industry-standard and internal coding specifications. The dictionary keys are normalized keywords or standard abbreviations representing various electronic components, while the values are the unique component category X code used for final coding. For example, the category dictionary might contain the following mapping entries: {"res":"R","resistor":"R","cap":"C","capacitor":"C","inductor":"L","ic":"U"}. This dictionary is the knowledge foundation of the entire identification system, ensuring the accuracy and consistency of category judgment.
[0040] Table 1. Values for Device Category (X):
[0041]
[0042] After the category dictionary is prepared, a sequential scan and matching operation is performed. A matching program strictly follows the index order of the input material information term array, starting from the first term (index 0), and traverses the array one by one. In each iteration, the program uses the term at the current index as the query key to search the loaded category dictionary. In this example, when the first term `res` at index 0 is reached, the program searches the category dictionary with `res` as the key. Since an entry with the key `res` exists in the dictionary, the query is successful, and the corresponding value `R` is returned. Immediately afterwards, the first match locking mechanism is triggered. This mechanism is designed to ensure the uniqueness and determinism of the category determination. Once the first term matching the key in the category dictionary is successfully found during the sequential scan, the entire scan and matching process immediately terminates. This is to prevent ambiguity or secondary information in the material description from interfering with the determination of the primary category, ensuring that the material category is uniquely determined by the first and most explicit keyword in the description. In the current example, since the first term *res* is successfully matched, the locking mechanism is triggered. The program records the matched term *res* as the matched category term, confirms its corresponding value *R* in the dictionary as the device category X code, and stops scanning subsequent terms such as *smd*, *0603*, etc. After the device category X code is uniquely determined to be *R*, the context rule loading operation is performed. The core of this operation is to access a preset, hierarchical main rule base. This main rule base is a much larger knowledge set, which uses all possible device category X codes, such as "R", "C", "L", etc., as top-level indices, and stores a dedicated set of context rules for each category. The program uses the newly determined device category X code *R* as the query key to retrieve and load the context rule set specifically for resistive materials from the main rule base. This loaded rule set contains a series of rules for identifying various specific characteristics of resistive materials, such as the main parameter resistance value, the capability parameter power, the specific package, accuracy, secondary classification, etc. Each rule is itself a structured object, containing a regular expression for matching terms, a list of keywords, and a unique feature label assigned to the term upon successful matching. Specifically, taking the resistor category as an example, when its context rule set is loaded, the rules and their matching patterns can be more clearly defined. For instance, a rule for identifying the main parameter resistance value, with the feature label "main parameter," can have its matching pattern P_main defined as: Match(T,P_main)=true if and only if the structure of term T satisfies the form T=N[M_unit]U_unit. In this definition, N represents the numerical part consisting of one or more digits, M_unit represents an optional order of magnitude belonging to the set {k,M}, and U_unit is the fixed physical unit ω.Similarly, the rule for identifying precision, P_precision, can be defined as T=N%, where a numerical part is followed by a percentage sign. The rule for identifying capability parameters (power), P_capability, can be defined as T=Nw, where a numerical part is followed by the power unit 'w'. For identifying specific packages, P_package can be defined as T=D1D2D3D4, where D is any number from 0 to 9, representing a strict four-digit combination. For identifying secondary categories (materials), P_subcategory can be defined as T∈{"thick","film","carbon"}, which checks if the term T exists in a preset keyword set. This preset keyword set is a set of terms extracted and solidified by domain experts through systematic analysis of industry standards and mainstream device specifications, clearly referring to the core material of the material. Finally, the initial tagging operation is performed. An empty list data structure, the initial tagging term list, is created. Subsequently, an initial tagging object is constructed and added to this list. This tag object records the results of this category identification in detail. It contains at least three key pieces of information: the matched category term itself, the term: res, the index of the term in the original term array: index: 0, and a tag that clearly identifies its identity: "device category". This ultimately completes and generates three independent, well-defined outputs. The first output is the device category X code, which in this example is the string R. The second output is the context rule set, a complete set of rules loaded from the main rule base specifically for identifying various characteristics of the resistor. The third output is the initial tag term list, which is a list containing only one element: [{term: "res", index: 0, tag: "device category"}].
[0043] Next, step 33 is executed. The input to this process consists of two parts: the first is a material information lexicon array containing ["res","smd","0603","10kω","1%","0.125w","thick","film"]; the second is a loaded set of context rules specifically for resistive materials. This process is driven by a rule application engine, which first initializes an empty list to store the output of this step, i.e., the sequence of intermediate labeled objects. The engine then iterates through the input material information lexicon array, strictly following its index order, from the first lexicon index 0 to the last lexicon index 7. For each lexicon in the array, the engine executes a standardized set of processing logic. The first sub-step of processing is skipping already labeled lexicons. Before processing any lexicon, the engine checks whether the lexicon has already been assigned a device category label in the initial labeled lexicon list generated in step 32. In this example, when the engine processes the lexicon "res" at index 0, it finds that the lexicon already exists in the initial labeled list. Therefore, the engine completely skips all rule matching operations for that term and directly proceeds to the processing flow of the next term. This ensures that keywords already used as the basis for category judgment are not repeatedly marked or incorrectly assigned other feature identities. For terms that are not skipped, processing proceeds to the second sub-step, namely priority rule matching. The engine applies rules one by one to the current term according to the rule priority order pre-defined in the context rule set. This priority setting is pre-defined by domain experts based on experience, and its purpose is to resolve potential matching ambiguities. For example, a term representing resistance value, 1000, also conforms to the pattern of a four-digit specific package. To prevent it from being incorrectly identified as a specific package, the matching rule for the primary parameter is given a higher priority than the specific package rule. In this example, when the engine processes the term smd with index 1, it applies rules in priority order. First, the primary parameter rule is applied, and the pattern does not match. Then, the precision rule is applied, and it does not match. Then, the capability parameter (power) rule is applied, and it does not match. Next, the specific package rule is applied, and it does not match. Finally, the secondary classification (material) rule is applied, and it still does not match. After traversing all rules, if a word element does not find any matching rules, a no-match processing mechanism is triggered. The engine creates a token object for that word element, but its label is assigned a special value: "no match". Therefore, the token object generated for "smd" is {word element: "smd", index: 1, label: "no match"}. When the engine processes word element 0603 with index 2, it also starts matching according to priority. It first attempts to match the primary parameter rule, which fails. Then it attempts to match the precision and capability parameter (power) rules, both of which fail. Next, when applying the specific encapsulation rule, its pattern of a four-digit number successfully matches 0603. At this point, the first match and labeling mechanism is triggered.Once the first successfully matched rule is found, the engine immediately stops applying any subsequent, lower-priority rules to the current term. It immediately creates a token object for that term, recording its content, index, and the matched rule label. Therefore, the token object generated for 0603 is {term: "0603", index: 2, label: "concrete encapsulation"}. Next, the engine processes the term 10kω with index 3. When applying the first and highest-priority primary parameter rule, its pattern (numerical value + optional order of magnitude + resistance unit ω) perfectly matches the term 10kω. The first match triggers the tagging mechanism, and the engine immediately generates a token object {term: "10kω", index: 3, label: "primary parameter"} and terminates subsequent rule matching for that term. For the term 1% with index 4, after trying higher-priority primary parameter rules that did not match, the engine applies the precision rule, and its pattern (numerical value + percent sign) successfully matches. Therefore, the tag object {lexicon: "1%", index: 4, tag: "precision"} is generated. For the lexicon 0.125w with index 5, after trying and excluding the primary parameter and precision rules, it successfully matches the pattern (numerical value + power unit w) of the capability parameter (power) rule. Therefore, the tag object {lexicon: "0.125w", index: 5, tag: "capability parameter"} is generated. For the lexicon thick with index 6, after excluding all higher priority numerical and formatting rules, it successfully matches the pattern of the secondary category (material) rule, which exists in the keyword list {"thick", "film", "carbon"}. Therefore, the tag object {lexicon: "thick", index: 6, tag: "secondary category"} is generated. Finally, for the lexicon film with index 7, the processing is similar to that of thick, and it also successfully matches the secondary category (material) rule, generating the tag object {lexicon: "film", index: 7, tag: "secondary category"}. After traversing all the terms in the material information term array, all the tag objects generated for terms that were not skipped are sequentially stored in the list initialized at the beginning of this step. This list containing all the newly generated tag objects is the intermediate tag object sequence. In this example, the specific content of the sequence is [{term: "smd", "..." not matched"}, {term: "0603", "..." specific encapsulation"}, {term: "10kω", "..." main parameter"}, {term: "1%", "..." precision"}, {term: "0.125w", "..." capability parameter"}, {term: "thick", "..." secondary classification"}, {term: "film", "..." secondary classification"}].
[0044] Next, proceed to step 34. This process takes the two lists mentioned above as input and appends all the marked objects from the intermediate marked object sequence completely and sequentially to the end of the initial marked term list. This is a list concatenation or merging operation designed to combine two datasets into one. After processing, a complete list containing all terms and their corresponding labels from the original material description is formed. To ensure the orderliness of the final list, it can be selectively sorted once based on the index value contained in each marked object after merging, thereby ensuring that the marked objects in the final list are arranged strictly in the order they appear in the original material description. By performing the merging operation, the output is the final marked term list. In this example, the list will contain the tagging information for all eight terms, and its complete content is: [{term: "res", index: 0, tag: "device category"}, {term: "smd", index: 1, tag: "not matched"}, {term: "0603", index: 2, tag: "specific package"}, {term: "10kω", index: 3, tag: "main parameter"}, {term: "1%", index: 4, tag: "precision"}, {term: "0.125w", index: 5, tag: "capability parameter"}, {term: "thick", index: 6, tag: "secondary category"}, {term: "film", index: 7, tag: "secondary category"}]. This list comprehensively reflects all the results after a complete semantic analysis of the original material information term array.
[0045] Finally, step 35 is executed. This process involves object construction and data injection. First, a new, empty data object is created; this new object represents the tagged material information data. Next, a data inheritance operation is performed, accessing the input preprocessed material information data object and completely copying all its key-value pairs—the original text, normalized text, and term arrays and their corresponding values—to the newly created tagged material information data object. This operation ensures that all information from the upstream processing stage is preserved, guaranteeing data integrity and traceability. After data inheritance, a new data injection operation is performed. It adds two entirely new key-value pairs to the tagged material information data object to inject the core analytical results of this stage. The first injected key is the device category X code, and its corresponding value is set to the input device category X code, i.e., the string R. The second injected key is a tagged term, and its corresponding value is set to the input final tagged term list, i.e., the complete list containing all terms and their corresponding labels. Through a series of operations including object construction, data inheritance, and new data injection, a complete and information-rich labeled material information data object is constructed. Its specific form is a structured data set, as shown below: {"Original Text": "Resistor, SMD, 0603, 10k Ohm, ±1%, 1 / 8W, Thick Film", "Normalized Text": "ressmd 0603 10kω 1% 0.125w thick film", "Term Array": ["res","smd","0603","10kω","1%","0.125w","thick","film"], "Device Category X Code": "R", "Tagged Term": [{Term: "res", Index: 0, Tag: "Device Category"}, {Term: "smd", Index: 1, Tag: "Not Matched"}, {Term: "0603 {{lexical: "10kω", index: 3, tag: "primary parameter"}, {lexical: "1%", index: 4, tag: "precision"}, {lexical: "0.125w", index: 5, tag: "capability parameter"}, {lexical: "thick", index: 6, tag: "secondary classification"}, {lexical: "film", index: 7, tag: "secondary classification"}]}.
[0046] In step 4, structured feature extraction and parsing are performed on the tagged material information data to obtain structured features of the material information. It should be understood that in the tagged material information data, each key term, such as 10kω or 0603, has been successfully assigned its unique engineering identity, such as a key parameter or specific encapsulation. However, these tagged feature values themselves, at the data level, still exist in the form of raw, unparsed strings. While a string "10kω" is correctly identified, the quantitative value "10000" it implies, which can be directly calculated or compared by the machine, has not yet been extracted. This gap between tagged text and standardized numerical values is the final obstacle to achieving fully automated processing and encoding of material information. Therefore, after completing all semantic tagging, this structured feature extraction and parsing step must be performed. Further, a more refined set of parsing rules oriented towards numerical and qualitative content needs to be applied to each labeled feature word to transform its string representation into standard, directly usable quantitative values or normalized qualitative descriptions. Finally, all these independent features after deep parsing are aggregated into a single, flat, structured feature set.
[0047] In one exemplary embodiment of this application, such as Figure 4 As shown, step 4 involves extracting and parsing structured features from the labeled material information data to obtain structured features of the material information, including: step 41, loading a numerical parsing rule set based on the device category X encoding; step 42, performing quantitative word iteration and parsing on each word in the labeled material information data based on the numerical parsing rule set to obtain quantitative parsing features of material information words; step 43, performing qualitative and physical feature parsing on each word in the labeled material information data to obtain qualitative parsing features of material information words; and step 44, aggregating the quantitative parsing features and qualitative parsing features of material information words to obtain the structured features of the material information.
[0048] In the above exemplary embodiment, step 4 operates as follows: First, step 41 is executed. The device category X code in this application is the string "R". The core of this process relies on a pre-built master parsing rule base maintained by domain experts. This rule base is a structured knowledge collection, organized as a mapping table with all possible device category X codes, such as R, C, L, etc., as top-level keys. Each top-level key corresponds to a specific set of numerical parsing rules as its value. This rule set itself is also a structured data object, with feature labels defined in step 3, such as primary parameters, capability parameters, and precision, as secondary keys. The value corresponding to each secondary key is a specific parsing algorithm or parameter. This set of algorithms or parameters defines in detail how to transform a string-form word that conforms to the label into structured data containing standardized values and units. When this step is executed, a parsing rule loading module receives the input device category X code "R". It uses "R" as the query key to search in the master parsing rule base and retrieves the entire set of numerical parsing rules associated with the resistor category R. For example, for the resistor category, the retrieved rule set would contain the following targeted parsing logic: For the main parameter label, the rules define the order-of-magnitude multiplier k (representing 10). 3 ) and M (representing 10) 6 The rules define the identification and conversion of the unit w for capability parameter labels and specify the final physical unit as w. For precision labels, the rules define the identification of the percentage sign % and specify the final unit as %.
[0049] Next, step 42 is executed. The core operation of this engine is to traverse each term in the tagged material information data object. For each tagged object in the list, the engine first checks the value of its label field. Subsequent parsing operations are only triggered if the label value belongs to a predefined set of quantitative features, i.e., primary parameters, capability parameters, or precision; for other tagged terms, they are skipped in this step. When a quantitative feature term is processed, rule matching is performed first. The engine uses the label of the current term as the key to select the corresponding sub-rule from the input set of numerical parsing rules. For example, when traversing to the term 10kω, labeled as a primary parameter, the engine loads a rule specifically for parsing resistance values. Subsequently, numerical and unit extraction is performed. The selected sub-rule contains one or more regular expressions applied to the string value of the term. For 10kω, the rule successfully separates it into three parts: the numerical part 10, the multiplier part k, and the unit part ω. Next, numerical calculation is performed. The numerical part 10 is converted into a standard floating-point number. For the multiplier part k, the engine queries a pre-defined multiplier mapping table contained within a rule set. This mapping table is pre-defined based on recognized standards in the field of electronic components and the specific conventions of the coding specifications of this invention (e.g., {"k":1000,"M":1000000}), and obtains its corresponding multiplier 1000. Then, the value is multiplied by the multiplier to obtain the final calculation result: 10 × 1000 = 10000. After the calculation is completed, unit standardization is performed. The engine confirms that the extracted unit part ω is a standard physical unit. Finally, structured encapsulation is performed. The calculated final value 10000 and the standard unit ω are encapsulated into a structured object, such as {value:10000, unit:ω}. Similarly, for the term 1% with a precision of 1%, the parsing process extracts the value 1 and the unit %, encapsulating them as {value:1, unit:%}. For the term 0.125w, which is a capability parameter, the parsing result is {value: 0.125, unit: w} since there is no multiplier. All successfully parsed quantitative features are added to the initially initialized dataset, with their tags as keys and encapsulated structured objects as values. After iteration, this dataset containing all quantitative parsing results is the material information term quantitative parsing feature, such as {main parameter: {value: 10000, unit: ω}, precision: {value: 1, unit: %}, capability parameter: {value: 0.125, unit: w}}.
[0050] Then, step 43 is executed. This process first performs feature extraction and preliminary assignment. A qualitative parsing engine traverses the input tagged material information data. During the traversal, the engine performs preliminary processing based on the tag field of each tagged object. When the engine processes a tagged object with index 2 and the tag "specific encapsulation", it extracts its term value 0603 and assigns it to a temporary encapsulation value variable. When processing a tagged object with index 6 and the tag "secondary category", it extracts its term "thick" and adds it to a temporary secondary category list. Subsequently, when processing a tagged object with index 7 and the tag "secondary category", it extracts its term "film" and appends it to the same secondary category list, at which point the list contains ["thick", "film"]. When processing a tagged object with index 1 and the tag "unmatched", the engine checks whether its term "smd" conforms to a preset dictionary containing common morphological keywords such as "smd", "dip", and "axial". Since "smd" is in this dictionary, the engine records its value in a temporary morphological factor hint variable. Other tag terms, such as device category and main parameters, are ignored in this step. After traversing the entire list, the process enters the attribute inference and integration stage. This stage performs a series of post-processing and logical inferences on the temporary variables generated in the previous stage. First, the form factor is determined. This determination process is based on a preset priority rule: information provided by the form factor hint variable has the highest priority. In this example, since the value of the form factor hint variable is smd, the final form factor is directly determined to be SMD. If the hint variable is empty, secondary inference logic is enabled. For example, based on the package value 0603, a lookup is performed in a preset package-form mapping table. This mapping table is pre-constructed by domain experts based on industry-standard electronic component packaging, systematically classifying various specific packages, such as 0603 and SOP-8, into preset form factor categories (S-segment) according to their physical mounting methods and manufacturing process characteristics, and assigning them unique package codes (P-segment). Thus, its form factor is inferred to be SMD, i.e., surface mount device.
[0051] Next, the secondary categories are merged. The engine processes the list of secondary categories, concatenating all the string elements "thick" and "film" using a connector such as a space to form a single, complete string "thick film," which is then used as the final secondary category value. "Thick Film" indicates a thick film resistor. The final output is a material information lexical qualitative parsing feature object. This object is a structured data collection that stores all parsed qualitative features in key-value pairs. In this example, the specific content of the output object is: {package value:"0603", morphology factor:"SMD", secondary category value:"thick film"}.
[0052] Finally, step 44 is executed. First, a new, empty data object is created to hold the final result: the material information structured feature. After object creation, a category assignment operation is performed. This module assigns the input device category X code value, "R", to a pre-defined key in the material information structured feature object that identifies the fundamental attributes of the material, such as device category. This operation first establishes the category of the final output object. Next, the core feature merging operation is performed. The two input feature sets are processed sequentially. First, it iterates through all key-value pairs in the material information lexical quantitative parsing feature set, completely copying and merging them into the material information structured feature object being constructed. For example, the key main parameter and its corresponding structured value {value: 10000, unit: ω} are added, followed by the key-value pairs of precision and capability parameters. After merging the quantitative features, the module processes the material information lexical qualitative parsing feature set in the same way, copying key-value pairs such as encapsulation, morphology factors, and secondary classification into the material information structured feature object. Thus, a structured feature object for aggregated material information is constructed, which comprehensively describes all the key features of the original material information. Its specific form is a structured data set, as shown below: {"Device Category": "R","Main Parameter": { "Value": 10000, "Unit": "ω"},"Precision": { "Value": 1, "Unit": "%"},"Capability Parameter": { "Value": 0.125, "Unit": "w"},"Specific Package":"0603","Secondary Category": "thick film","Form Factor": "SMD"}.
[0053] In a preferred exemplary embodiment of this application, step 4, extracting and parsing structured features from the labeled material information data to obtain structured features of the material information, includes: step 4-1, loading a numerical parsing rule set based on the device category X encoding; step 4-2, generating and co-optimizing parsing hypotheses on the labeled material information data based on the numerical parsing rule set to obtain quantitative and qualitative parsing features of material information terms; and step 4-3, aggregating the quantitative and qualitative parsing features of material information terms into structured features to obtain the structured features of the material information. Specifically, steps 4-1 and 4-3 are implemented in the same way as steps 41 and 44 in the above exemplary embodiment, and therefore will not be described in detail. Here, the specific implementation of step 4-2 will be described in detail.
[0054] Specifically, after the preceding feature token labeling step is completed, some tokens in the labeled material information data may correspond to multiple reasonable parsing paths due to their inherent ambiguity or the irregularities in the original descriptive text. For example, in the resistor section of this application, an independent token "100" may refer to a resistance value of 100 ohms or a power of 100 milliwatts. A traditional method of parsing tokens in isolation, one by one, lacks a global perspective and struggles to make the optimal decision among multiple seemingly reasonable local parsings. It is highly susceptible to the failure of parsing the entire material information due to misjudgment of a single token. More importantly, all the features of an effective electronic material are not randomly combined, but rather collectively constitute a high-probability, physically achievable whole, i.e., possessing feature coherence. Therefore, this invention introduces this parsing hypothesis generation and collaborative optimization step to construct a decision-making framework that can simulate the common sense of domain experts. To systematically generate all possible complete parsing schemes, i.e. hypotheses, and through a quantitative scoring mechanism that combines local grammatical matching degree and global physical coherence, the unique and most practical quantitative and qualitative parsing features of material information terms are selected from numerous possibilities, thereby fundamentally overcoming the challenges brought about by data noise and parsing ambiguity.
[0055] Based on this, step 4-2, based on the numerical analysis rule set, performs analytical hypothesis generation and collaborative optimization on the labeled terms in the labeled material information data to obtain quantitative and qualitative analytical features of the material information terms. This includes: performing Cartesian product combination of ambiguous terms in the labeled material information data to obtain a hypothesis set. It should be understood that when one or more terms have multiple analytical possibilities, these possibilities may be interdependent, and making decisions on each ambiguous point in isolation is insufficient. For example, for the description "Resr 1000603", the term 100 might be assumed to be resistance Ha1 or power Ha2, while 0603 might be assumed to be package Hb1 or an irrelevant number Hb2. For global evaluation, all possible complete interpretations need to be constructed. Specifically, the first step is to identify all ambiguous labeled terms and generate a set of basic hypothesis branches for each term. Subsequently, by performing a Cartesian product operation on all these basic hypothesis branches, a complete and mutually exclusive set of hypotheses {H1, H2, ..., Hn} containing all possible combinations is generated. Each hypothesis Hi is an end-to-end, complete interpretation of the original material description. For example, for a somewhat unstandardized description "Res 100 1 / 8W 0603", the parser will be ambiguous about the term 100 after recognizing Res as a resistor category. At this point, a basic hypothesis branch will be generated for this ambiguous term 100, containing two possibilities: hypothesis Ha1, that is, 100 is resolved as a primary parameter, corresponding to the feature {resistance: 100Ω}; and hypothesis Ha2, that is, 100 is resolved as a capability parameter, corresponding to the feature {power: 100mW}. Meanwhile, other terms such as 1 / 8W and 0603 are explicitly resolved to {Power: 0.125W} and {Specific Package: 0603}, respectively, each forming a branch with only one option. Subsequently, by performing a Cartesian product operation, these branches are combined to generate two complete, mutually exclusive global hypotheses. The first hypothesis, H1, is the result of combining Ha1 with other terms, and its content is: {Category: Resistor, Resistance: 100Ω, Power: 0.125W, Specific Package: 0603}. The second hypothesis, H2, is the result of combining Ha2 with other terms, and its content is: {Category: Resistor, Power: 100mW, Power: 0.125W, Specific Package: 0603}. Thus, a fuzzy input is transformed into two clear, complete, and evaluable candidate solutions. H1 is physically coherent, while H2 is highly unreasonable due to containing two conflicting power values.
[0056] The feature coherence score and local parsing score for each hypothesis in the hypothesis set are calculated to obtain the feature coherence score set and the local parsing score set. Accordingly, an optimal parsing result should not only be as faithful as possible to the literal information of the original text, but also conform to the basic physical laws and industry conventions of the electronic components field. Therefore, it is necessary to calculate the scores for these two dimensions separately. The calculation of the local parsing score (local(Hi)) aims to measure the degree of matching between the features in the hypothesis and its original lexical units, i.e., the syntactic matching degree. For example, the confidence of parsing the lexical unit 10kω as {numerical value: 10000, unit: ω} can be assigned a value based on the matching strength of a predefined regular expression, usually resulting in a high score to ensure that the parsing result does not deviate from the original text. The calculation of the feature coherence score (Hi) is used to quantify the reasonableness of the combination of all features in the hypothesis. Its core is to use the point mutual information (PMI) formula log{P(fa,fb) / [P(fa)P(fb)]} to measure it. The joint probability P(fa,fb) and marginal probabilities P(fa) and P(fb) are statistically derived from a large-scale, pre-built knowledge base containing massive amounts of valid material data, for example, by crawling millions of publicly available component datasheets. For example, the PMI value of the feature pair (specific package: 0603, power: 0.125W) will be significantly positive, while the PMI value of (specific package: 0603, power: 100W) will be significantly negative. Assume that the coherence (Hi) score of Hi is the sum of the PMI values of all its internal feature pairs. Specifically, first, calculate the local parsing score. For hypothesis H1, its features {resistance: 100Ω} and {power: 0.125W} are directly parsed from the terms 100 and 1 / 8W respectively, with extremely high grammatical matching. Therefore, its local parsing score local(H1) may be assigned a near-perfect value, such as 0.9. For hypothesis H2, although its individual features also originate from the original lexical units, it contains two conflicting power attributes, which creates redundancy and contradiction in its grammatical structure. Therefore, its local parsing score, local(H2), is penalized, resulting in a lower score, such as 0.5. Subsequently, the feature coherence score is calculated. For hypothesis H1, the sum of the point mutual information of all its feature pairs needs to be calculated, for example, PMI(resistance: 100Ω, power: 0.125W) + PMI(resistance: 100Ω, specific package: 0603) + PMI(power: 0.125W, specific package: 0603). Since these feature combinations are extremely common in the pre-built knowledge base, each PMI value is positive, and the sum will result in a high positive score, such as 1.5. Hypothesis H2 contains a physically impossible feature pair (power: 100mW, power: 0.125W).The joint occurrence probability P(fa,fb) of this feature pair in the knowledge base is almost zero, resulting in a very large negative PMI value. Even if the PMIs of other feature pairs are positive, this huge negative value will still make the final sum of coherence(H2) a significantly negative score, such as -3.8. Through calculation, hypothesis H1 achieved a high local resolution score (0.9) and a high feature coherence score (1.5), while hypothesis H2 performed poorly in both dimensions, with scores of 0.5 and -3.8, respectively.
[0057] A weighted total score is calculated based on the feature coherence score and local analytical score of each group in the feature coherence score set and local analytical score set to obtain the hypothesis total score set. It should be understood that in practice, these two scores may conflict. For example, a hypothesis with a low local analytical score due to a typo in the original text may have a very high coherence of its feature combinations. To make the correct decision in such cases, a mechanism that can balance the importance of the two is needed. This step is implemented by applying a pre-defined scoring function Score(Hi) = w1*local(Hi) + w2*coherence(Hi) to each hypothesis Hi in the hypothesis set. Here, w1 and w2 are pre-defined weight hyperparameters, whose specific values are determined by common machine learning methods such as grid search or gradient optimization on a labeled validation dataset. These two weights reflect the balance between fidelity to the original text and conformity to common sense in the final decision. For example, w1 can be set to 0.4 and w2 to 0.6 to slightly emphasize the actual coherence of the feature combinations. Based on this, the total score for the two hypotheses can be calculated. For the physically consistent hypothesis H1, the total score Score(H1) is calculated as 0.4*0.9+0.6*1.5=1.26. For the hypothesis H2, which has an inherent contradiction, the total score Score(H2) is 0.4*0.5+0.6*(-3.8)=-2.08.
[0058] Based on the hypothesis with the highest total score in the hypothesis score set, quantitative and qualitative analytical features of material information terms are obtained. That is, the hypothesis score set generated in the previous step is traversed to find the highest score value, and the hypothesis Hi corresponding to that score value is determined. This highest-scoring hypothesis is considered the best explanation for the original ambiguous description. For example, comparing the total score of H1 (1.26) with the total score of H2 (-2.08), since 1.26 is the only highest score, the final quantitative and qualitative analytical features of material information terms are directly taken from the inherent structure of this winning hypothesis H1. Specifically, this step outputs the complete feature set contained in H1: where the quantitative analytical feature is {resistance: 100Ω, power: 0.125W}, and the qualitative analytical feature is {category: resistor, specific package: 0603}. Furthermore, even if some terms have errors, as long as most features are coherent, a correct overall explanation can still be found, thus improving robustness. For example, even if the power of 1 / 8W is incorrectly interpreted as 18W, the PMI value of 18W, along with characteristics such as the 0603 package and 10kΩ resistance, is extremely low. Therefore, the total score of the hypothesis including this incorrect interpretation will be very low, leading to the selection of the correct hypothesis. In this way, an initially ambiguous input converges into a single, globally verified, and structured final interpretation result through competition and collaborative optimization of multiple hypotheses, providing accurate input for subsequent coding steps.
[0059] In step 5, the structured features of the material information are subjected to normalized encoding and code segment generation to obtain a set of encoded material information segments. Accordingly, after the preceding structured feature extraction and parsing stages, all the key attributes of a material have been transformed into a fully structured, machine-readable, but still verbose feature set. This set describes the various characteristics of the material in a human-understandable key-value pair format, such as "Main Parameter": {"Value": 10000, "Unit": "ω"}. However, this descriptive format is not the highly compact, standardized normalized encoding required for enterprise resource planning, material management, and automated production systems. These downstream systems require a coded string with extremely high information density, fixed length, and precise, unambiguous meaning for each bit or segment. Therefore, after completing the comprehensive analysis of the features, the normalized encoding and code segment generation steps need to be performed to convert this rich but loosely structured feature object into a series of standardized and independent encoding fields (i.e. code segments) according to a set of preset and rigorous encoding specifications, laying the foundation for finally splicing and generating a unique material code string that meets production requirements.
[0060] In an exemplary embodiment of this application, step 5, which involves performing standardized encoding and code segment generation on the structured features of the material information to obtain a set of encoded segments of the material information, includes: step 51, extracting the X code segment from the structured features of the material information; step 52, performing quantitative parameter encoding on the structured features of the material information to obtain the BBB code segment and the VVVV code segment; step 53, performing qualitative feature encoding on the structured features of the material information to obtain the S code segment, the P code segment, the Y1 code segment, and the Y2 code segment; and step 54, performing code segment aggregation on the BBB code segment, the VVVV code segment, the S code segment, the P code segment, the Y1 code segment, and the Y2 code segment to obtain the set of encoded segments of the material information.
[0061] In the exemplary embodiment described above, step 5 operates as follows: First, step 51 is performed. This is the first step in encoding generation, aimed at determining the fundamental category code of the material. The input material information structured feature object is directly accessed, and the value corresponding to the device category key is extracted. In this example, the value is the string R. According to the encoding specification, the device category code segment (X code segment) directly adopts this value. Therefore, the output is an X code segment with the value R.
[0062] Next, step 52 is executed. The core of this process is a repeatedly invoked numerical encoding subroutine. This subroutine is designed to receive a parameter object to be encoded, an identifier pointing to a specific rule table, a specified number of significant digits for precision, and a device category code as input. Its internal processing logic is as follows: First, based on the input device category code and rule table identifier, the base unit used for encoding the parameter is determined by querying the preset encoding specifications, as shown in Tables 2 and 3. Next, the numerical values in the input parameter object are uniformly converted to base values expressed in this base unit. Subsequently, this base value is transformed into scientific notation form m*10^e, where the mantissa m is adjusted to an integer with a specified number of significant digits to meet the input requirements, and e is the corresponding integer power. Then, in the exponent-encoding mapping section of the specified rule table, i.e., Table 2, the corresponding single encoded character is found based on the calculated exponent e. Finally, the mantissa m is formatted into a string of a specified number of digits, padded with leading zeros if necessary, and concatenated with the found exponent encoded character to generate the final code segment string.
[0063] Table 2. Ability Value BBB Multiplier and Correspondence with Common Units:
[0064]
[0065] Table 3. Correspondence between parameter values VVVV ratio and commonly used units:
[0066]
[0067] The generation process of the VVVV code segment is executed first. An encoding generation module calls the aforementioned numerical encoding subroutine, passing it the main parameter object {"value":10000,"unit":"ω"}, the rule table identifier for the VVVV code segment (Table 3), the preset number of significant digits (3), and the device category R. When the subroutine executes, it first determines the base unit of the main resistance parameter as Ω based on the rule table. Since the unit of the input value 10000 is already Ω, and ω is represented in lowercase, no unit conversion is needed. Next, the value 10000 is converted into scientific notation with 3 significant digits, i.e., 100 × 10⁻⁶. 2 The mantissa m is 100 and the exponent e is 2. Then, the code corresponding to exponent 2 is looked up in Table 3, which is the character 2. Finally, the mantissa 100 is concatenated with the exponent code 2 to generate the VVVV code segment string 1002. Next, the BBB code segment generation process is executed. The encoding generation module calls the numerical encoding subroutine again, this time passing the capability parameter object {"value":0.125,"unit":"w"}, the rule table identifier for the BBB code segment (Table 2), the preset number of significant digits (2), and the device category R. When the subroutine executes, it first determines the base unit of resistance power as mW according to Table 2. Therefore, the input value 0.125w is normalized to 125mW. Since the encoding requires two significant digits, according to the one-to-one elimination principle in the encoding specification, 125 is processed to obtain 120. Then, 120 is converted to scientific notation, i.e., 12 × 10⁻¹⁰. 1 We obtain the mantissa m as 12 and the exponent e as 1. Next, we look up the code corresponding to exponent 1 in Table 2, which is the character 1. Finally, we concatenate the mantissa 12 with the exponent code 1 to generate the BBB code segment string 121. The final output is the BBB code segment 121 and the VVVV code segment 1002.
[0068] Next, step 53 is executed. This process is performed step-by-step by a qualitative coding engine, generating a corresponding coded string for each target code segment. First, the generation of the S-code segment is performed. The input to this process is the morphological factor attribute in the material information structured feature object, with a value of SMD and a specific package name of 0603. The coding engine uses the meaning of the surface mount device represented by the value SMD as the query key to perform a search in a preset table, constructed by domain experts based on industry production process standards, as shown in Table 4. In this table, the engine matches the entry for "common surface mount" and obtains its corresponding coded value, resulting in S-code segment string 2.
[0069] Table 4. Meaning of General Material S (Number Segment):
[0070]
[0071] Next, the generation of the P-segment is performed. The input to this process is the encapsulation value of 0603 from the material information structured feature object and the S-segment string value of 2. The S-segment value serves as crucial contextual information, limiting the search range. The encoding engine uses 0603 as the primary lookup key and, based on the S-segment value of 2, locates the column specifically for S=2 (i.e., regular patch) in a more complex, pre-defined table, as shown in Table 5. In this column, the engine finds the row that perfectly matches 0603 and extracts the encoded character defined in that cell, resulting in the P-segment string 6.
[0072] Table 5. Meaning of P in common packages for discrete components:
[0073]
[0074]
[0075]
[0076]
[0077] Subsequently, the generation of the Y1 code segment is performed. The input to this process is the precision attribute from the material information structured feature object, with a value of {"value":1,"unit":"%"} and the device category attribute value of R. The encoding engine first queries a preset encoding meaning rule table based on the device category R, as shown in Table 6, to determine that for resistive materials, the Y1 field means precision. Then, the engine extracts the precision value of 1% and uses it as the query key to match it in a preset table, as shown in Table 7. In this table, the engine finds the entry corresponding to ±1% and obtains its corresponding encoded character, resulting in the Y1 code segment string F.
[0078] Table 6. Specific Meanings of Code Segments for Various Devices:
[0079]
[0080] Table 7. Values for the Y1 field when representing precision:
[0081]
[0082] Finally, the Y2 code segment is generated. The input to this process is the secondary classification attribute (value: thick film) and the device category attribute (value: R) from the material information structured feature object. The encoding engine first loads a secondary classification encoding rule table specifically for resistors based on the device category R, as shown in Table 8. Then, the engine uses the input value "thick film" to perform keyword or semantic matching in this table. In this example, "thick film" matches the entry for ordinary thick film resistors in the table, and the engine then obtains the encoded character corresponding to that entry. The output of this step is the Y2 code segment string "0".
[0083] Table 8. Definition of Secondary Classification (Y2) of Resistor Devices:
[0084]
[0085] The final output consists of four independent, normalized encoded strings: S segment 2, P segment 6, Y1 segment F, and Y2 segment 0.
[0086] Next, step 54 is executed. Its primary operation is object initialization, which involves creating a new, empty data object to hold the final result; this object is the set of encoded material information segments. Following this, the module performs a data injection operation. It sequentially accesses all the independent code segments generated in the previous steps and adds them completely to the newly created set object, using a preset string identifying each code segment as the key and the segment's own value as the value. Specifically, it stores code segment R (X) with the key X; it stores code segment 121 (BBB) and code segment 1002 (VVVV) with the keys BBB and VVVV respectively; finally, it stores code segment 2 (S), code segment 6 (P), code segment F (Y1), and code segment 0 (Y2) with their corresponding keys S, P, Y1, and Y2 respectively. The final material information is then constructed using the encoded segment set. In this example, the specific content of the output object is: {X:"R",BBB:"121",VVVV:"1002",S:"2",P:"6",Y1:"F",Y2:"0"}.
[0087] In step 6, the encoded segment set of material information is integrated and verified to obtain the final production code. That is, in the previous steps, standardized independent code segments representing various characteristics of the material have been successfully generated, but they still exist as a structured data set and have not yet formed a single, continuous code string that can be directly applied to the production and procurement processes. This discrete data structure cannot be directly recognized and utilized by downstream systems. Therefore, after generating and aggregating all code segments, an integrated and verified output step is required to precisely assemble all the independent code segments in the encoded segment set into a complete final production code that meets the format requirements, according to a preset final encoding paradigm. Finally, a completeness and compliance verification is performed to ensure the uniqueness and usability of the output.
[0088] In an exemplary embodiment of this application, step 6, which encodes, integrates, and verifies the set of encoded segments of material information to obtain the final production code, includes: step 61, concatenating the S code segment and the P code segment to obtain the SP code segment; step 62, concatenating the Y1 code segment and the Y2 code segment to obtain the YY code segment; step 63, concatenating the BBB code segment and the VVVV code segment to obtain the BBBVVVV code segment; and step 64, aggregating the SP code segment, the YY code segment, the X code segment, and the BBBVVVV code segment to obtain the final production code.
[0089] In the exemplary embodiment described above, step 6 operates as follows: the process strictly follows a preset paradigm that defines the final encoding structure, namely "SPYY.XBBBVVVV". This paradigm is predetermined and specifies the exact position, order, and delimiters of each code segment in the final string.
[0090] First, step 61 is executed to concatenate the S and P code segments. The value 2 corresponding to key S and the value 6 corresponding to key P are extracted from the input set. Then, a string concatenation operation is performed, linking the two segments in the order S first, then P, to generate a temporary SP code segment with a value of 26.
[0091] Next, step 62 is executed to concatenate the Y1 and Y2 code segments. In the same way, the Y1 code segment F and the Y2 code segment 0 are extracted from the input set and concatenated in sequence to generate a temporary YY code segment with the value F0.
[0092] Subsequently, step 63 is executed, where the BBB code segment and the VVVV code segment are concatenated. The BBB code segment 121 and the VVVV code segment 1002 are extracted and concatenated in sequence to generate a temporary BBBVVVV code segment with a value of 1211002.
[0093] After all sub-segments are concatenated, step 64 is executed for final aggregation. Following the order specified by the final encoding paradigm SPYY.XBBBVVVV, the generated temporary code segments and the remaining independent code segments in the input set are concatenated. It sequentially extracts SP code segment 26, YY code segment F0, a preset character "." as a separator, X code segment R, and BBBVVVV code segment 1211002, and concatenates them strictly in this order to generate a candidate final production encoded string: 26F0.R1211002.
[0094] Before output, the generated string undergoes format and integrity checks. It checks if the total string length is 13 characters (including separators), confirms if the 5th character is a ".", the 6th character is a letter, and the remaining characters are either letters or numbers, ensuring no empty segments exist. In this example, the generated string 26F0.R1211002 fully conforms to all preset validation rules. After passing the validation, the string is confirmed as the final production code. This production code serves as a unique, standardized, and machine-readable identifier for electronic materials throughout their entire lifecycle, including R&D, procurement, warehousing, and production. It replaces the previously ambiguous and varied text descriptions, thereby enabling efficient, accurate, and automated flow of material information between various systems.
[0095] It is worth mentioning that although this application uses resistive materials as a specific example for detailed description, its core technical solution has universal applicability. The framework constructed by this method based on category recognition-rule loading-segmented encoding only needs to configure and load dedicated rule sets (such as main parameters, capability parameters, secondary classifications, etc.) for other device categories such as capacitors and inductors to achieve automated and standardized encoding of their unstructured information. Therefore, the scope of protection of this application is not limited to resistive components.
[0096] In summary, the automatic classification method for electronic materials based on multi-level feature recognition, as described in this application, is elucidated. It constructs a progressively deeper automated processing flow to address the unstructured information processing challenges mentioned in the background art. Specifically, it acquires unstructured material information and, through information normalization and preprocessing, first solves the data heterogeneity problem caused by inconsistencies in format, units, and terminology. Subsequently, through main category identification and feature term tagging, it achieves precise positioning of the core attributes of the material, surpassing the shallow identification of traditional keyword matching and effectively overcoming its inability to handle complex semantics. Based on this, it further extracts and parses the tagged features in a structured manner, transforming them into machine-understandable paradigmatic features. Finally, through encoding integration and verification, it automatically generates a standard, unique production code, such as the SPYY.XBBBVVVV structure. This replaces inefficient and error-prone manual operations and, with its powerful parsing capabilities and robustness, solves the fundamental problem that existing automated solutions struggle to handle diverse descriptions.
[0097] Figure 4 This is a block diagram of an automatic electronic material classification system based on multi-level feature recognition, according to an embodiment of this application. Figure 4 As shown, the electronic material automatic classification system 100 based on multi-level feature recognition according to an embodiment of this application includes: an unstructured material information acquisition module 110, used to acquire unstructured material information; a material information data preprocessing module 120, used to perform material information normalization and preprocessing on the unstructured material information to obtain preprocessed material information data; a material information data identification and marking module 130, used to perform main category identification and feature word marking on the preprocessed material information data to obtain marked material information data; a material information parsing module 140, used to extract and parse structured features from the marked material information data to obtain structured features of the material information; a material information encoding module 150, used to perform normalized encoding and code segment generation on the structured features of the material information to obtain a set of encoded segments of the material information; and a production code output module 160, used to perform encoding integration and verification output on the set of encoded segments of the material information to obtain the final production code.
[0098] Here, those skilled in the art will understand that the specific operations of each step in the above-described automatic electronic material classification system based on multi-level feature recognition have been referenced above. Figures 1 to 3 The automatic classification method for electronic materials based on multi-level feature recognition has been described in detail, and therefore, its repeated description will be omitted.
Claims
1. An automatic classification method for electronic materials based on multi-level feature recognition, characterized in that, include: Obtain information about unstructured materials; Unstructured material information is normalized and preprocessed to obtain preprocessed material information data; The process involves: identifying the main category and tagging feature terms in preprocessed material information data to obtain tagged material information data; extracting a material information term array from the preprocessed material information data; identifying the main category and loading context rules onto the material information term array to obtain a device category X code, a context rule set, and an initial tagged term list; performing term traversal and rule application on the material information term array based on the context rule set to obtain an intermediate tagged object sequence; merging the intermediate tagged object sequence and the initial tagged term list to obtain a final tagged term list; encapsulating the preprocessed material information data, the final tagged term list, and the device category X code into a data structure to obtain the tagged material information data; extracting and parsing structured features from the tagged material information data to obtain structured features of the material information; performing normalized encoding and code segment generation on the structured features of the material information to obtain a set of encoded segments of the material information; and integrating and verifying the set of encoded segments of the material information to output the final production code.
2. The automatic classification method for electronic materials based on multi-level feature recognition according to claim 1, characterized in that, The unstructured material information is normalized and preprocessed to obtain preprocessed material information data, including: obtaining and integrating material descriptions from the unstructured material information to obtain raw text strings; performing text normalization on the raw text strings to obtain normalized text strings; and performing structured encapsulation and output on the normalized text strings to obtain the preprocessed material information data.
3. The automatic classification method for electronic materials based on multi-level feature recognition according to claim 2, characterized in that, The text normalization process for the original text string to obtain a normalized text string includes: performing a format conversion operation on the original text string to obtain a lowercase text string; performing symbol filtering and synonym replacement on the lowercase text string to obtain filtered and replaced text; and performing unit and numerical standardization on the filtered and replaced text to obtain the normalized text string.
4. The automatic classification method for electronic materials based on multi-level feature recognition according to claim 1, characterized in that, The process involves extracting and parsing structured features from labeled material information data to obtain structured features, including: loading a numerical parsing rule set based on device category X encoding; performing quantitative word iteration and parsing on each word in the labeled material information data based on the numerical parsing rule set to obtain quantitative parsing features of material information words; performing qualitative and physical feature parsing on each word in the labeled material information data to obtain qualitative parsing features of material information words; and aggregating the quantitative and qualitative parsing features of material information words to obtain the structured features of the material information.
5. The automatic classification method for electronic materials based on multi-level feature recognition according to claim 1, characterized in that, The material information structured features are subjected to paradigmatic encoding and code segment generation to obtain a set of encoded material information segments, including: extracting X code segments from the material information structured features; performing quantitative parameter encoding on the material information structured features to obtain BBB code segments and VVVV code segments; performing qualitative feature encoding on the material information structured features to obtain S code segments, P code segments, Y1 code segments, and Y2 code segments; and aggregating the BBB code segments, VVVV code segments, S code segments, P code segments, Y1 code segments, and Y2 code segments to obtain the set of encoded material information segments.
6. The automatic classification method for electronic materials based on multi-level feature recognition according to claim 1, characterized in that, The material information is encoded, integrated, and verified to output the final production code. This includes: concatenating the S and P code segments to obtain the SP code segment; concatenating the Y1 and Y2 code segments to obtain the YY code segment; concatenating the BBB and VVVV code segments to obtain the BBBVVVV code segment; and aggregating the SP, YY, X, and BBBVVVV code segments to obtain the final production code.
7. An automatic classification system for electronic materials based on multi-level feature recognition, characterized in that, include: The unstructured material information acquisition module is used to acquire unstructured material information. The material information data preprocessing module is used to normalize and preprocess unstructured material information to obtain preprocessed material information data. The material information data identification and labeling module is used to identify the main category and label the feature words of preprocessed material information data to obtain labeled material information data. This includes: extracting a material information word array from the preprocessed material information data; identifying the main category and loading context rules onto the material information word array to obtain the device category X code, a context rule set, and an initial labeled word list; traversing words and applying rules to the material information word array based on the context rule set to obtain an intermediate labeled object sequence; merging the intermediate labeled object sequence and the initial labeled word list to obtain a final labeled word list; and encapsulating the preprocessed material information data, the final labeled word list, and the device category X code into a data structure to obtain the labeled material information data. The material information parsing module is used to extract and parse the structured features of the labeled material information data to obtain structured features of the material information. The material information encoding module is used to perform normalized encoding and code segment generation on the structured features of the material information to obtain a set of encoded segments of the material information. The production encoding output module is used to integrate and verify the encoded segments of the material information and output the final production code.
Citation Information
Patent Citations
Standardized unified code generation method and system for building material names and specifications
CN120611173A