Electronic component product data intelligent collection and automatic input method and system

By using an intelligent scheduling engine and semantic matching technology, electronic component data is automatically collected and entered, solving the problem of low efficiency in manual operation and achieving efficient and accurate data management.

CN122364347APending Publication Date: 2026-07-10SHANGTONG ELECTRONICS (NANJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGTONG ELECTRONICS (NANJING) CO LTD
Filing Date
2026-04-02
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In the existing technology, the collection and entry of electronic component data relies on manual operation, which is inefficient, has a high error rate, and the inconsistent naming of different data sources makes data matching difficult.

Method used

An intelligent scheduling engine is used to automatically obtain component information from multiple external data sources, and through semantic matching and cross-validation, automated data collection and entry are achieved.

Benefits of technology

It improves the efficiency and accuracy of component data collection and entry, reduces manual operations, solves the problem of inconsistent naming of different data sources, and ensures the integrity and consistency of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122364347A_ABST
    Figure CN122364347A_ABST
Patent Text Reader

Abstract

This application relates to a method and system for intelligent collection and automatic entry of electronic component product data, belonging to the field of data collection and entry. The method includes: receiving unstructured text input by a user; parsing the unstructured text to determine the user's query intent; determining and executing the corresponding target processing flow; the corresponding target processing flow includes: identifying key identification information of the target electronic component from the unstructured text; using a preset intelligent scheduling engine to obtain raw data corresponding to the target electronic component from multiple external data sources, and cross-validating the raw data; semantically matching the raw data with pre-stored standardized data to determine the matching standardized data items; and entering the raw data into the data entry corresponding to the target electronic component in a preset information database based on the matching standardized data items. This application improves the efficiency and accuracy of the data collection and entry process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data collection and entry, and particularly to a method and system for intelligent collection and automatic entry of electronic component product data. Background Art

[0002] In the electronic component trading industry, business personnel need to handle a large amount of component data entry work daily. Currently, the collection and entry of component data mainly rely on manual operations, and its typical process is as follows: First, business personnel obtain the customer's demand information for components, which is usually presented in the form of natural language text, such as "Need 10 ADI's AD8232BRZ, industrial grade" or "Is there any stock of STM32F103C8T6?". Then, business personnel need to manually identify and split out key information, such as component brand ("ADI"), part number ("AD8232BRZ"), etc., based on their own experience from these unstructured natural language descriptions. However, due to differences in the expression habits of different people, the same component may be described in multiple forms (for example, "Texas Instruments" may be abbreviated as "TI" or "Deyi"). This uncertainty in the expression method makes it difficult for manual extraction of key information, not only inefficient, but also prone to human errors such as missing or misjudging; Next, after extracting the key information, business personnel need to obtain the detailed technical data of the component. This process usually requires manually opening multiple browser tabs, accessing different supplier websites (such as Mouser, DigiKey), industry databases or search engines one by one, and manually entering the part number for retrieval. This decentralized query method of multiple data sources is extremely time-consuming and laborious; In addition, the naming method of the component data obtained externally is often inconsistent with the internal system standards of the enterprise where the business personnel are located. For example, the external data source may use the full brand name "STMicroelectronics", while the enterprise internal system is used to using the abbreviation "ST"; the product catalog may be shown as "microcontroller", while the internal system classifies it as "MCU". This difference in the naming method requires business personnel to manually compare and judge between the obtained external data and the standardized data pre-stored in the enterprise based on subjective experience. Not only is the accuracy difficult to guarantee, but when the enterprise data volume increases or new brands are encountered, it is easy to generate matching errors, affecting the efficiency of subsequent order processing and inventory management.

[0003] Finally, after a series of tedious manual operations, business personnel can manually copy and paste the compiled component parameters into the corresponding fields in the enterprise ERP or inventory management system to complete the final data entry. The entire process involves multiple steps, including manual identification, multi-site retrieval, manual comparison, and copying and pasting. The operation is cumbersome, time-consuming, and the possibility of errors accumulates with each additional step; therefore, it needs improvement. Summary of the Invention

[0004] To improve the efficiency and accuracy of data collection and entry, this application provides a method and system for intelligent collection and automatic entry of electronic component product data.

[0005] Firstly, this application provides a method for intelligent collection and automatic entry of electronic component product data, including: Receive unstructured text input by the user, parse the unstructured text, and determine the user's query intent; Based on the query intent, determine and execute the corresponding target processing flow; The query intent includes at least a query for the entity information of the target electronic component, and the target processing flow corresponding to the query for the entity information of the target electronic component includes: Basic steps: Identify key identification information of the target electronic component from the unstructured text; based on the key identification information, use a preset intelligent scheduling engine to automatically obtain original data corresponding to the target electronic component from multiple predetermined external data sources, and cross-validate the original data obtained from different data sources; Ontology information query steps: Semantically match the cross-validated raw data with the pre-stored standardized data to determine the matching standardized data items; based on the matching standardized data items, enter the raw data into the data entry corresponding to the target electronic component in the preset information database to complete the information filing or update of the target electronic component and end the target processing flow.

[0006] By adopting the above technical solutions, users do not need to input according to a fixed format and can directly use everyday language to describe their needs; the system can automatically obtain component information from multiple sources and perform cross-validation to ensure the accuracy and integrity of the data; by semantic matching, external data is aligned with the enterprise's internal standardized data, solving the problem of inconsistent naming of different data sources; ultimately, the automatic filing or updating of component information is realized, greatly reducing the workload of manual data entry and improving data entry efficiency and accuracy.

[0007] Optionally, the step of automatically acquiring raw data corresponding to the target electronic component from multiple predetermined external data sources using a preset intelligent scheduling engine, and cross-validating the raw data acquired from different data sources, includes: The system utilizes a pre-defined intelligent scheduling engine to call at least one pre-defined third-party API interface in parallel to obtain the first raw data corresponding to the target electronic component. When the first original data is empty or does not meet the preset integrity requirements, the intelligent scheduling engine automatically triggers the preset targeted web page crawling module, so that the targeted web page crawling module can crawl the second original data corresponding to the target electronic component from the web pages of at least one target website based on the preset crawling rules. The acquired first and second raw data are fused together, and the fusion process includes at least the following: When the same data item has multiple sources, cross-validate the consistency of the data values ​​from all sources corresponding to the data item, and retain the data value with the highest confidence. When data items from different sources complement each other, the data from each source are merged to form complete original data.

[0008] By adopting the above technical solutions, the efficiency of data acquisition is improved by calling multiple API interfaces in parallel; when API data is insufficient, crawlers are automatically triggered as a supplement to ensure the comprehensiveness of the data; multi-source data is fused and processed, and when the same data item has multiple sources, the most reliable data value is selected through cross-validation and confidence assessment; when data from different sources are complementary, they are merged to form a complete dataset, thereby obtaining more accurate and complete component information than a single data source.

[0009] Optionally, the step of semantically matching the cross-validated raw data with pre-stored standardized data to determine the matching standardized data items includes: The target field and its text information are extracted from the cross-validated raw data. The text information is then input into a pre-trained semantic vector model and transformed into a first semantic vector. Each standard information corresponding to the target field in the pre-stored standardized data is input into the semantic vector model and transformed into the corresponding second semantic vector; Calculate the similarity between the first semantic vector and each of the second semantic vectors; When the similarity exceeds a preset threshold, the corresponding standard information is determined as a standardized data item that matches the target field.

[0010] By adopting the above technical solution, text information is transformed into semantic vectors, which can capture the semantic association between the full name and abbreviation of a brand, as well as expressions in different languages. Intelligent matching is achieved by calculating vector similarity, which solves the problem that traditional keyword matching cannot handle synonyms and near-synonyms. When the similarity exceeds a preset threshold, the standardized data item to be matched is automatically determined, so that non-standard data obtained from the outside can be accurately mapped into the enterprise's internal standardized system, laying the foundation for subsequent data entry and correlation analysis.

[0011] Optionally, the query intent also includes querying alternative component information for the target electronic component. The corresponding target processing flow includes: the basic steps and the alternative information query steps. The alternative information query steps include: extracting key parameter information of the target electronic component from the acquired original data; obtaining parameter data of multiple candidate electronic components belonging to the same product category as the target electronic component from the external data source; calculating the similarity between the key parameter information of the target electronic component and the parameter data of each candidate electronic component to obtain parameter similarity; when the parameter similarity exceeds a preset alternative similarity threshold, determining the corresponding candidate electronic component as a substitute for the target electronic component, and storing the information of the substitute in association under the data entry corresponding to the target electronic component in the preset information database. The query intent also includes querying information on supporting components for the target electronic component. The corresponding target processing flow includes: the basic steps and the supporting information query steps. The supporting information query steps include: the intelligent scheduling engine triggers the targeted web crawling module to crawl application documents related to the target electronic component; the crawled application documents are parsed to extract the identification information of other electronic components that co-occur with the target electronic component; the co-occurrence frequency of each other electronic component is counted, and other electronic components whose co-occurrence frequency exceeds a preset frequency threshold are identified as supporting components of the target electronic component; the information on the supporting components and the supporting relationship are associated and stored with the information on the target electronic component in the preset information database.

[0012] By adopting the above technical solutions, in the alternative query, by extracting the key parameter information of the target component and calculating the parameter similarity with the candidate components of the same category, the alternative models are automatically discovered, and the alternative relationships are associated and stored, providing users with alternative solutions when there is a shortage of stock or when cost optimization is needed; in the matching query, by crawling application documents and statistically analyzing the co-occurrence frequency of components, the matching usage relationships between components are automatically discovered, providing a reference for system design and selection, and helping users quickly find complete solutions.

[0013] Optionally, the query intent also includes a selection intent based on the requirement description, and the target processing flow corresponding to the selection intent based on the requirement description includes: The unstructured text is parsed to extract at least one constraint description; the constraint description includes at least a functional requirement description. Each constraint is described and mapped to a corresponding filtering condition. Based on the filtering condition, candidate electronic components that meet the condition are retrieved from the preset information database, and a selection recommendation list containing all the candidate electronic components is output.

[0014] By adopting the above technical solution, users do not need to know the specific model; they only need to describe their functional requirements to receive recommendations. The system extracts constraints through requirement parsing, maps natural language descriptions into executable filtering conditions, and retrieves candidate components that meet the conditions from a preset information database. This method lowers the selection threshold, helps users unfamiliar with specific models to quickly find components that meet their needs, and improves selection efficiency and accuracy.

[0015] Optionally, before automatically acquiring the original data corresponding to the target electronic component from multiple predetermined external data sources using a preset intelligent scheduling engine, a data source strategy determination step is also included: The product category of the target electronic component is determined based on the key identification information; Based on the query intent and the product category, a preset data source strategy library is queried to determine a data collection strategy that matches the current query; the data collection strategy includes a list of APIs to be called, a list of target websites to be crawled, and the calling priority of each data source; Specifically, the system utilizes a pre-defined intelligent scheduling engine to automatically acquire raw data corresponding to the target electronic component from multiple predetermined external data sources, and the intelligent scheduling engine executes the acquisition according to the determined data acquisition strategy.

[0016] By adopting the above technical solutions, the system matches the optimal data collection strategy from the preset data source strategy library according to different product categories and query intentions; the coverage and accuracy of different product categories of components vary on different data sources, and selecting data sources by category can improve the query success rate; different query intentions require different types of data, and selecting data sources by intention can obtain the required information more effectively; calling data sources according to priority ensures that the most reliable data source is called first, improving the overall query efficiency.

[0017] Optionally, the method further includes: The system receives user monitoring settings for target electronic components in a preset information database. The monitoring settings include monitoring frequency and monitoring indicators. The monitoring indicators include at least one or more of the following: price changes, inventory status, and life cycle status. According to the monitoring frequency, the basic steps are periodically performed on the monitored target electronic components to obtain the latest raw data. The newly acquired raw data is compared with the stored data to detect changes in monitoring indicators. When changes in monitored metrics exceed preset warning thresholds, warning information is generated and pushed to the user.

[0018] By adopting the above technical solution, users can set monitoring tasks for components of interest, and the system will periodically obtain the latest data according to the set frequency. By comparing the latest data with the stored data, the system can automatically detect important information such as price changes, inventory status changes, and life cycle status changes. When the changes exceed the warning threshold, the system will actively push the warning to the user, enabling the user to keep abreast of changes in the status of components and take precautions against the impact of price fluctuations, stockout risks, or product discontinuation, thereby reducing procurement risks and cost losses caused by information lag.

[0019] Secondly, this application provides an intelligent collection and automatic input system for electronic component product data, including: The semantic parsing module is used to receive unstructured text input by the user, parse the unstructured text, and determine the user's query intent; The process processing module is used to determine and execute the corresponding target processing flow based on the query intent; The query intent includes at least a query for the entity information of the target electronic component, and the target processing flow corresponding to the query for the entity information of the target electronic component includes: Basic steps: Identify key identification information of the target electronic component from the unstructured text; based on the key identification information, use a preset intelligent scheduling engine to automatically obtain original data corresponding to the target electronic component from multiple predetermined external data sources, and cross-validate the original data obtained from different data sources; Ontology information query steps: Semantically match the cross-validated raw data with the pre-stored standardized data to determine the matching standardized data items; based on the matching standardized data items, enter the raw data into the data entry corresponding to the target electronic component in the preset information database to complete the information filing or update of the target electronic component and end the target processing flow.

[0020] Thirdly, this application provides an intelligent collection and automatic input device for electronic component product data, including a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed as described in any of the methods in the first aspect.

[0021] Fourthly, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described in any of the first aspects.

[0022] In summary, this application includes at least one of the following beneficial technical effects: Traditional solutions rely on general NER models (accuracy of 60-70%) or regular expressions, while this application uses a domain-optimized BERT model combined with a rule engine, which can automatically parse the input text and accurately extract the part number and brand, replacing the existing technology that relies on manual splitting. This avoids the problems of manually missing characters and confusing brand abbreviations, thus improving extraction efficiency.

[0023] Furthermore, existing systems use a single data source for querying (such as API only or crawler only). The innovative multi-level scheduling engine in this application dynamically selects the optimal data source path by monitoring API response time, cost and success rate in real time, which can quickly and comprehensively supplement the attribute information of components.

[0024] Furthermore, traditional keyword matching (such as hard-coded tables like "TI=Texas Instruments") cannot handle new brands and complex directory relationships. The vector space model constructed in this application uses 256-dimensional semantic embedding to convert the queried brand and directory text into vectors, and then calculates the cosine similarity (≥0.8 for automatic matching) with the system's pre-trained vector library, solving the pain point of existing technologies where full names and abbreviations, as well as different expressions, cannot be matched when manually compared subjectively. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart illustrating the intelligent collection and automatic entry method for electronic component product data disclosed in the embodiments of this application. Detailed Implementation

[0027] The following is in conjunction with the appendix Figure 1 This application will be described in further detail.

[0028] This application discloses a method for intelligent collection and automatic entry of electronic component product data (hereinafter referred to as the information entry method), the execution subject of which is an intelligent collection and automatic entry system for electronic component product data (hereinafter referred to as the information entry system). The following will be combined with... Figure 1 This section elaborates on the specific execution process of the information entry method in the information entry system.

[0029] S101: Receive unstructured text input by the user, parse the unstructured text, and determine the user's query intent.

[0030] S102, Based on the query intent, determine and execute the corresponding target processing flow.

[0031] The query intent includes at least a query for the entity information of the target electronic component, and the target processing flow corresponding to the query for the entity information of the target electronic component includes: Basic steps: Identify key identification information of target electronic components from unstructured text; based on key identification information, use a preset intelligent scheduling engine to automatically obtain raw data corresponding to the target electronic components from multiple predetermined external data sources, and cross-validate the raw data obtained from different data sources. Ontology information query steps: Semantically match the cross-validated raw data with the pre-stored standardized data to determine the matching standardized data items; based on the matching standardized data items, enter the raw data into the data entry corresponding to the target electronic component in the preset information database to complete the information filing or update of the target electronic component and end the target processing flow.

[0032] The aforementioned basic step of "automatically acquiring raw data corresponding to the target electronic component from multiple predetermined external data sources using a preset intelligent scheduling engine, and cross-validating the raw data acquired from different data sources" specifically includes the following sub-steps: The system utilizes a pre-defined intelligent scheduling engine to call at least one pre-defined third-party API interface in parallel to obtain the first raw data corresponding to the target electronic component. When the first raw data is empty or does not meet the preset integrity requirements, the intelligent scheduling engine automatically triggers the preset targeted web crawling module, so that the targeted web crawling module can crawl the second raw data corresponding to the target electronic component from the web pages of at least one target website based on the preset crawling rules. The acquired first and second raw data are fused together. The fusion process includes at least the following: When the same data item has multiple sources, cross-validate the consistency of the data values ​​from all sources corresponding to the data item, and retain the data value with the highest confidence. When data items from different sources complement each other, the data from each source are merged to form complete original data.

[0033] The aforementioned ontology information query step of "semantically matching the cross-validated raw data with pre-stored standardized data to determine the matching standardized data items" specifically includes the following sub-steps: Extract the target field and its text information from the cross-validated raw data, input the text information into a pre-trained semantic vector model, and transform it into a first semantic vector; Each standard information corresponding to the target field in the pre-stored standardized data is input into the semantic vector model and transformed into the corresponding second semantic vector; Calculate the similarity between the first semantic vector and each of the second semantic vectors; When the similarity exceeds a preset threshold, the corresponding standard information is determined as a standardized data item that matches the target field.

[0034] In implementation, the information entry system first receives unstructured text input by the user. This unstructured text refers to text information expressed in natural language without fixed format constraints. Its forms are diverse, including: instant messaging messages such as "Please check the STM32F103C8T6 part" sent via WeChat or DingTalk; email content such as "We need 10K ADI AD8232s, industrial grade"; speech-to-text results such as "Looking for a transistor that can replace the IRF3205"; chat log snippets such as "Do you have that ST MCU, model STM32F103, in stock?"; and text lines in Excel / PDF files such as "STmicroelectronics STM32F103C8T6" copied from a quotation. These texts share the common characteristic that key information (brand, part number, requirements, etc.) is mixed within the natural language, without fixed positions or formats, requiring semantic understanding for extraction.

[0035] The information entry system parses the received unstructured text to determine the user's query intent. Query intent refers to the core purpose of the user's query. In this embodiment, it includes at least: ontological information query (i.e., querying detailed information about the target electronic component itself (such as parameters, specifications, packaging, etc.)), alternative component information query (i.e., finding substitutes for the target electronic component), matching component information query (i.e., finding other components that are used in conjunction with the target electronic component), and selection intent based on requirement description (i.e., recommending suitable components based on functional requirement description).

[0036] The information entry system can employ an intent recognition method based on a large language model to parse unstructured text. Specifically, an intent classification model is pre-built; in this embodiment, the intent classification model uses a BERT-based text classification model. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model capable of capturing the contextual semantic information of text, making it very suitable for intent classification tasks. Specifically, a classification head is added to the pre-trained BERT model, taking the vector representation of the [CLS] token output by BERT as input and mapping it to the intent category space through a fully connected layer and a softmax function. To train the intent classification model, a large-scale labeled dataset needs to be constructed. The construction of the training data includes the following steps: First, we collected genuine user query texts from electronic component business scenarios. Sources included: historical chat logs (after anonymization); email correspondence; query records in system operation logs; and typical query statements manually simulated.

[0037] Then, domain experts were invited to annotate the collected query texts with intent. The annotation categories included: ontology information query (the corresponding annotation example is "Please help me find information about STM32F103C8T6"), alternative component information query (the corresponding annotation example is "Are there any alternative models for AD8232?"), matching component information query (the corresponding annotation example is "What power supply chip is usually used with STM32F103?"), and selection intent based on requirement description (the corresponding annotation example is "Find a 5V power supply, 100kHz bandwidth op amp, priced under 3 yuan").

[0038] After annotation, a training dataset containing tens of thousands of samples is formed (each sample includes a specific annotated example and query intent). To improve the model's generalization ability, data augmentation is performed on the original data, including: Synonym replacement: replacing keywords in the query with synonyms, such as "check it out" → "search" → "help me find it". Sentence transformation: changing declarative sentences into interrogative sentences, or interrogative sentences into imperative sentences. Noise injection: randomly inserting or deleting irrelevant words to simulate colloquial expressions in real-world scenarios.

[0039] The constructed training dataset is divided into training, validation, and test sets. The training process is as follows: Each query text is encoded according to the BERT input format: a [CLS] marker is added at the beginning of the text; a [SEP] marker is added at the end of the text; the text is segmented into tokens and converted into corresponding token IDs; an attention mask is generated. The encoded input is fed into the BERT model, and the output vector at the [CLS] position is extracted. This vector aggregates the semantic information of the entire input text. Then, this vector is fed into the classification head, and the probability distribution of each intent category is obtained through a fully connected layer and a softmax function. The cross-entropy loss function is used to calculate the difference between the predicted probability and the true label: Loss = -∑y_i*log(p_i); where y_i is the one-hot encoding of the true label, and p_i is the probability of the i-th category predicted by the model. The gradient is calculated using the backpropagation algorithm, and the BERT model parameters and classification head parameters are updated using the Adam optimizer. Training is performed for multiple epochs until the model's accuracy on the validation set no longer improves.

[0040] When the information entry system is actually running, the trained intent classification model is used to recognize the intent of the unstructured text input by the user. The specific process is as follows: The intent classification model is trained using unstructured text input from the user. The model performs forward computation and outputs the probability distribution of each intent category.

[0041] For example: Input: "Please help me look up the information for STM32F103C8T6" → Output: Main component information query: 0.97; Alternative component information query: 0.01; Matching component information query: 0.01; Selection intent based on requirement description: 0.01 The information entry system ultimately selects the category with the highest probability as the recognition result, which is the ontology information query.

[0042] After identifying the query intent, the information entry system determines the corresponding target processing flow based on a pre-stored mapping table of query intents and target processing flows. An example is shown below: When the query intent is for ontology information, the corresponding target processing flow is the basic steps plus an ontology information query step. When the query intent is for alternative component information, the corresponding target processing flow is the basic steps plus an alternative information query step (see below). When the query intent is for matching component information, the corresponding target processing flow is the basic steps plus a matching information query step (see below). When the query intent is for selection based on a requirement description, the corresponding target processing flow is the selection processing flow (see below).

[0043] In this embodiment, when the query intent is identified as "querying the entity information of a target electronic component," the information entry system determines to execute the target processing flow of "basic steps + entity information query steps." The specific execution process of this "basic steps + entity information query steps" will be described in detail below: First, the information entry system identifies key identification information (the key identification information being the representation of the target electronic component) from the raw unstructured text. This key identification information includes at least the brand (e.g., "ADI", "ST", "Texas Instruments") and part number (e.g., "AD8232BRZ", "STM32F103C8T6"). Specifically, the information entry system can employ an entity extraction method based on a large language model. This method leverages the powerful language understanding capabilities and rich domain knowledge of large language models (LLMs) to guide the model through carefully designed prompts to complete the information extraction task. Large language models (such as the GPT series, GLM series, and LLaMA series) are pre-trained on massive amounts of internet text data, having learned rich world knowledge and language patterns. Specifically in the field of electronic components, large language models have already encountered a large amount of text containing component information during pre-training, thus possessing the following capabilities: Large language models can recognize electronic component brands because their pre-training corpora contain a wealth of brand-related context, such as: The training corpus is: "Analog Devices launches the new generation of high-performance op-amp AD8232"; the model learns the knowledge that: ADI = Analog Devices = brand.

[0044] The training corpus is: "STMicroelectronics' STM32 series MCUs are widely popular"; the model learns the following knowledge: ST = STMicroelectronics = brand.

[0045] The training corpus was: "Texas Instruments power chip TPS5430"; the model learned the following knowledge: TI = Texas Instruments = brand.

[0046] The training corpus is: "PIC series microcontrollers of Microchip Technology"; the knowledge learned by the model is: Microchip = Microchip Technology = brand.

[0047] Through training with a massive amount of such corpus, the model established the concept of "brand name" and memorized a large number of specific brands and their common expressions (full name, abbreviation, alternative name).

[0048] Electronic component part numbers typically follow specific naming rules, and large language models learn these patterns through pre-training, such as: The part number format is a combination of letters and numbers, such as "AD8232BRZ" and "STM32F103C8T6". The model has learned that the part number usually starts with a letter and is followed by numbers.

[0049] The part number format is: series + suffix, such as "TPS5430" and "LM358"; the model learned the rule that series of parts have common prefixes.

[0050] The part number format is: package suffix, such as "-BRZ" or "-C8T6"; the model learned that the suffix often indicates the package and temperature rating.

[0051] The part number pattern is: common length, such as 6-12 characters; the model learned that the part number is usually not long.

[0052] In addition, the model also learns the concept of "part number" through context. For example, in the sentence "Help me check the part number AD8232BRZ", "AD8232BRZ" follows "check it" and is referred to by "this part number", so the model can infer that this is a part number.

[0053] In the training corpus, the co-occurrence pattern of "brand + product number" is very common, such as: The co-occurrence pattern is: brand + space + part number, such as: "ADI AD8232BRZ".

[0054] The co-occurrence pattern is: brand + part number, such as: "ADI's AD8232BRZ".

[0055] The co-occurrence pattern is: part number + from + brand, such as: "AD8232BRZ from Analog Devices".

[0056] The co-occurrence pattern is: brand + part number + other description, such as: "ST's STM32F103C8T6 microcontroller".

[0057] By learning these patterns, the model is able to accurately identify the boundaries between brand and part number in new text.

[0058] While large language models already possess the aforementioned capabilities, appropriate prompts are needed to guide the model to focus on specific tasks and output standardized results. For example: "You are an electronic component information extraction assistant. Please extract the brand and part number from the following text. The brand refers to the manufacturer's name, and the part number refers to the specific model number of the component. If the brand is not explicitly mentioned in the text, leave the brand field blank. Please output in JSON format." After the user inputs text, the information entry system concatenates the user's input text with the above instructions and inputs it into the large language model. The large language model outputs, for example, {"Brand": "ADI", "Part Number": "AD8232BRZ"}. For example, for the input text "Help me find ADI's AD8232BRZ, industrial application", the information entry system extracts the brand "ADI" and the part number "AD8232BRZ".

[0059] After extracting key identification information, the information entry system automatically retrieves raw data corresponding to the target electronic component from multiple pre-defined external data sources using a pre-set intelligent scheduling engine. The intelligent scheduling engine performs the raw data retrieval operation according to the following steps: Step 1: The information entry system is pre-configured with multiple third-party API interfaces from the component industry, for example: The third-party API interface is the DigiKey API, and its corresponding data source is the DigiKey official website. The query method is to query by part number. The example returned data is: brand, manufacturer part number, description, inventory, price, and specification link.

[0060] The third-party API interface is the Mouser API, and its corresponding data source is the Mouser official website. The query method is to query by part number. The example returned data is: brand, manufacturer part number, technical parameters, package, and inventory.

[0061] The third-party API interface is the Octopart API, and its corresponding data source is the Octopart search engine website. The query method is to search by part number or brand. The example returned data is: prices, inventory, alternatives, and specifications from multiple suppliers.

[0062] The third-party API interface is the Electronic Components Trading Network API, and its corresponding data source is the domestic electronic components trading platform. The query method is to query by part number. The example returned data is: domestic spot inventory, price, and brand.

[0063] The intelligent scheduling engine calls the aforementioned third-party API interfaces in parallel to obtain the first raw data corresponding to the target part number. Parallel calls can improve query efficiency, and multiple API parallel requests can usually be completed within 2-3 seconds.

[0064] Step 2: The information entry system performs a completeness assessment on the initial raw data returned by the third-party API interface. The assessment indicators include: Is it empty?: Whether the dataset returned by the third-party API interface is null, whether the returned JSON object is an empty object {}, whether the returned array is an empty array [], and whether the returned string is an empty string "".

[0065] Key field missing status: The information entry system pre-defines a list of key fields for each component category. Key fields refer to the most representative parameters necessary to describe the component category. For example, when the component category is an integrated circuit (IC), the corresponding key field list includes brand, part number, description, package, operating temperature range, and power supply voltage. For a query intent regarding the component's intrinsic information, the information entry system requires at least four core fields: brand, part number, description, and package. If any one of these is missing, it is considered a missing key field. Optionally, the information entry system can calculate the key field missing rate = (number of missing key fields) / (total number of preset key fields) × 100%. When the missing rate exceeds a preset threshold (e.g., 30%), it is considered "severely missing key fields".

[0066] Is the data volume sufficient? The information entry system presets a minimum parameter count threshold for different types of components. This threshold is based on the statistical analysis of the typical parameter count in the datasheets of that component type. For example, when the component type is a simple passive component (resistor), the typical parameter count ranges from 5 to 10 (such as brand, part number, resistance value, accuracy, package, and power). The minimum parameter count threshold is 5. If the total number of returned data items is lower than the minimum parameter count threshold for that type, it is considered that the data volume is insufficient. For example, for operational amplifiers (typically 15-25 items), if the API only returns 4 items (brand, part number, description, and package), which is significantly lower than the minimum threshold of 10 items, it is considered that the data volume is insufficient, and the targeted web crawling module needs to be triggered to supplement it.

[0067] The preset integrity requirements can be set according to business needs. For example, it may be required that at least three core fields, namely brand, description, and packaging, be returned; otherwise, the integrity requirements are considered not met.

[0068] Step 3: The intelligent scheduling engine will automatically trigger the preset targeted web crawling module when any of the following conditions are met: Condition 1: All third-party API interfaces return empty data. Condition 2: The returned data does not meet the preset integrity requirements.

[0069] The targeted web crawling module is an automated web scraping tool built on automated testing frameworks such as Playwright. Playwright can simulate real browser behavior, execute JavaScript, handle dynamically loaded content, simulate clicks and scrolling, and is suitable for crawling modern web pages. Crawling rules are pre-defined crawling strategies for each target website, including the following rule items: 1. Target URL template, such as: https: / / www.digikey.com / products / en?keywords={part_number}

[0070] 2. Waiting conditions. For example: wait for the ".product-details" element to finish loading.

[0071] 3. Extract XPath / CSS selectors from data. For example: Brand: span[data-testid='manufacturer']; Package: td:contains('Package / Case')+td.

[0072] 4. Page turning rules. If the search results have multiple pages, click the "Next Page" button to continue crawling.

[0073] 5. Anti-scraping measures. Such as random User-Agent, delayed requests, and proxy IP rotation.

[0074] For example, for the DigiKey website, the crawling rules might be defined as follows: First, visit https: / / www.digikey.com / products / en?keywords=AD8232BRZ. Then wait for the product details page to load. Next, use a CSS selector to extract the brand: [data-testid='manufacturer-value']. ​​Then use XPath to extract the encapsulation: / / td[contains(text(),'Package / Case')] / following-sibling::td. Finally, use a CSS selector to extract the description: [itemprop='description'].

[0075] The targeted web crawling module extracts data from the web pages of the target website according to the crawling rules, and obtains the second set of raw data.

[0076] The information entry system merges the first set of raw data (from the API) and the second set of raw data (from the targeted web crawling module) to generate complete raw data. The fusion process includes two scenarios: Scenario 1: When the same data item (e.g., "encapsulation") appears in multiple data sources simultaneously, the information entry system performs consistency cross-validation on the data values ​​from these sources and selects the most reliable data value based on confidence level assessment. The basic process of specific cross-validation is as follows: The comparison checks whether data values ​​from different sources are consistent. "Consistency" here includes: complete matching (strings match exactly, such as "LFCSP-8" and "LFCSP-8"), semantic equivalence (different expressions but the same meaning, such as "LFCSP-8" and "8-LeadLFCSP"), and convertible equivalence (identical after unit conversion or format standardization, such as "85°C" and "85℃"). The information entry system performs string normalization processing (removing spaces, standardizing capitalization, and standardizing units) before comparison.

[0077] If the data is consistent, the credibility of the data value is enhanced; multiple independent sources corroborating each other indicate that the data value is accurate and reliable. If the data is inconsistent, it is necessary to determine which source is more reliable, and proceed to the confidence assessment process. The confidence assessment process is as follows: The information entry system pre-sets a basic confidence score for each data source, reflecting the inherent credibility of that data source. The basic score is set according to the data source type: for example, if the data source is the original manufacturer's official website (such as the manufacturer's official website), its basic confidence score is 95 points; if the data source is an authorized distributor (such as DigiKey, Mouser, Element14), its basic confidence score is 85 points.

[0078] In each query, the information entry system calculates scores for four dynamic factors based on the quality of the data obtained.

[0079] Factor 1: Data source authority (weight 40%). The authority score is determined based on the type and level of the data source, similar to the basic score but more detailed: such as the original manufacturer's official website (authority score of 100) > authorized distributors (authority score of 90) > third-party aggregation platforms (authority score of 70).

[0080] Factor Two: Historical Success Rate (Weight 30%). The historical success rate reflects the reliability of the data source in past queries. A query is considered successful only if it meets the following conditions simultaneously: the API returns an HTTP 200 status code, with no timeouts or exceptions; the returned dataset is not null and contains at least one valid data item; the returned data includes brand and part number fields (even if other fields are missing); and the returned data is in the correct format and can be parsed by the system. If all the above conditions are met, it is counted as a success; otherwise, it is counted as a failure. Accordingly, the historical success rate = (number of successes) / (number of successes + number of failures) × 100%. The information entry system maintains a query statistics table for each data source: for example, the historical success rate for the DigiKeyAPI data source is 95%; the historical success rate for a crawler of a factory's official website is 96%. Historical success rate score = historical success rate × 100 (e.g., 95% → 95 points). Factor 3: Data Completeness (Weight 20%). Data completeness reflects the coverage ratio of key fields in the returned data. The list of key fields is preset by product category (as described above). Data Completeness = (Number of key fields returned this time) / (Total number of key fields preset for this product category) × 100%. For example, for the MCU product category, there are 15 preset key fields, and only 12 were returned this time, so the completeness score = 12 / 15 × 100 = 80 points.

[0081] Factor 4: Format Compliance (Weight 10%). Format compliance reflects whether the returned data conforms to industry standards or expected formats. Assessment methods include: When the field type is encapsulated, its standardization check is whether it conforms to the JEDEC standard format; for example, "LFCSP-8" conforms, while "LFCSP8" may not.

[0082] When the field type is brand, its standardization check standard is whether it conforms to the standard name or common abbreviation; for example, "LFCSP-8" conforms, while "LFCSP8" may not.

[0083] When the field type is part number, its standardization check is whether it contains illegal characters; for example, it is standard to contain only letters, numbers, and hyphens.

[0084] When the field type is a parameter value, its standardization check is whether it contains a unit; for example, "85°C" is standard, while "85" is not.

[0085] If fully compliant with the standard: 100 points; slightly non-compliant (e.g., missing units, non-standard format but still parseable): 70 points; seriously non-compliant (e.g., garbled text, completely incorrect format): 30 points.

[0086] The information entry system uses a weighted summation method to calculate the overall confidence level of each data source in this query: Overall Confidence Level = (Authority Score × 40%) + (Historical Success Rate Score × 30%) + (Data Completeness Score × 20%) + (Format Compliance Score × 10%). Wherein, Historical Success Rate Score = Historical Success Rate × 100, and Data Completeness Score = Data Completeness × 100.

[0087] Scenario 2: When data items from different data sources can complement each other, the information entry system first verifies whether these data point to the same target electronic component. After confirming this, the data is merged to form a more complete dataset. An exemplary verification method is to compare the part number fields returned by each data source: if all data sources return completely identical part numbers, it is directly determined to be the same component. For example: DigiKey returns "AD8232BRZ", the original manufacturer's website returns "AD8232BRZ", and Mouser returns "AD8232BRZ" → completely identical, verification passed.

[0088] Sometimes, different data sources may have slight differences in how part numbers are represented, requiring normalization before comparison: Remove special characters: Normalize "AD8232-BRZ" to "AD8232BRZ". Unify case: Convert all to uppercase. Handle prefixes and suffixes: For example, "STM32F103C8T6" and "STM32F103C8T6TR" may be considered different models, requiring judgment based on business rules. Part number matching degree = string similarity (Source 1 part number, Source 2 part number). Edit distance (Levenshtein Distance) or Jaccard similarity can be used for calculation. When the matching degree exceeds a threshold (e.g., 95%), it is determined to be the same component.

[0089] When part number information is incomplete or ambiguous, the system performs joint verification using multiple fields. Verification field combinations can include: brand + part number + package, brand + part number + descriptive keywords, and brand + key parameters (e.g., for op-amps: number of channels + supply voltage + bandwidth). The information entry system assigns weights to each field combination (e.g., the brand field has a weight of 30%, with matching rules based on full / abbreviated semantic matching; the part number field has a weight of 40%, with matching rules based on exact matching or high similarity; the package field has a weight of 20%, with matching rules based on semantic equivalence matching; and the key parameter field has a weight of 10%, with matching rules based on similar data or consistent unit conversion). The overall matching degree is calculated as ∑(field matching degree × field weight). When the overall matching degree exceeds a preset threshold (e.g., 90%), the components are determined to be the same.

[0090] For example, the DigiKey API provides brand, description, package, and operating temperature range. The original manufacturer's website crawler provides electrical characteristics (supply voltage, quiescent current) and typical application circuits. The Mouser API provides inventory quantity and price tiers. These data items are mutually exclusive and can be merged into a complete dataset containing brand, description, package, operating temperature, supply voltage, quiescent current, typical applications, inventory, and price. The criterion for determining complementarity is: when the intersection of the data item sets returned by the two data sources is less than a preset threshold (e.g., 30%), and the union covers the key fields, they are considered complementary.

[0091] Finally, after obtaining complete raw data, the information entry system needs to match this raw data with the standardized data pre-stored within the information entry system, so as to map the non-standardized data obtained from the outside into the enterprise's internal standardized system.

[0092] Standardized data refers to a unified and standardized data dictionary accumulated by an enterprise over long-term business operations. This includes: a standardized brand library, such as the full standard names of "STMicroelectronics," "Texas Instruments," and "Analog Devices," along with their corresponding commonly used abbreviations; and a standardized product catalog library, such as standard classifications like "Microcontroller (MCU)," "Operational Amplifier (Op Amp)," and "Power Management IC (PMIC)."

[0093] The information entry system extracts target fields that need to be standardized and matched from the cross-validated raw data. In this embodiment, the target fields include at least: Brand name: such as "ST", "ADI", "TI" and other possible expressions. Product catalog name: such as "microcontroller", "MCU", "single-chip microcomputer" and other synonyms or near-synonyms. For example, the brand text extracted from the raw data may be "STMicroelectronics", while the product catalog text may be "STM32Series Microcontrollers".

[0094] The information entry system uses a pre-trained semantic vector model to convert text information into vector representations. In this embodiment, the BERT (Bidirectional Encoder Representations from Transformers) model is selected as the base model. BERT is a pre-trained language model based on the Transformer architecture, capable of mapping text to fixed-length semantic vectors, where texts close in distance within the vector space are semantically similar. The training process for this model is as follows: The BERT model is pre-trained using a large-scale general corpus (such as Wikipedia and book corpora) to enable it to understand general language. Then, the model is fine-tuned using specialized corpora in the field of electronic components to make it better suited for processing component-related text. The specialized corpus includes: component specification sheets, lists of brand and product catalog names, technical discussions in industry forums, and manually annotated data from historical query logs. Positive and negative sample pairs are constructed to further optimize the model's discriminative ability: positive sample pairs include synonyms such as "STMicroelectronics" and "ST", and "microcontroller" and "MCU". Negative sample pairs include ambiguous expressions such as "STMicroelectronics" and "Texas Instruments," and "microcontroller" and "power chip." Through comparative learning, the vector distance between positive sample pairs is reduced, while the vector distance between negative sample pairs is increased, thereby improving the model's accuracy in the standardized matching task.

[0095] The text information of the target field (e.g., "STMicroelectronics") is input into the semantic vector model to obtain the first semantic vector, denoted as V_query. Each standard information corresponding to the target field from the pre-stored standardized data is input into the same semantic vector model to obtain the corresponding second semantic vector set, denoted as {V_std1, V_std2, ..., V_stdn}. For example, for the brand field, the standardized brand library contains: “STMicroelectronics” → Vector V_std1.

[0096] "Texas Instruments" → Vector V_std2.

[0097] “Analog Devices” → Vector V_std3.

[0098] “Microchip Technology” → Vector V_std4.

[0099] ... Then, the cosine similarity between the first semantic vector and the second semantic vector is calculated using the formula: similarity=cos(θ)=(V_query·V_std) / (||V_query||×||V_std||). Here, the symbol · represents the dot product (also called the inner product or dot product) of the vectors, and the symbol || represents the Euclidean length of the vectors.

[0100] The cosine similarity value ranges from -1 to 1. A value closer to 1 indicates that the two vectors are more aligned in direction, meaning they are more semantically similar. For example: Similarity between V_query and V_std1 (STMicroelectronics): 0.95.

[0101] Similarity between V_query and V_std2 (Texas Instruments): 0.45.

[0102] Similarity between V_query and V_std3 (Analog Devices): 0.38.

[0103] Similarity between V_query and V_std4 (Microchip Technology): 0.42.

[0104] The information entry system presets a similarity threshold, which is set to 0.8 in this embodiment. When the highest similarity exceeds this threshold, a match is considered successful. In the example above, the highest similarity is 0.95 > 0.8, therefore "STMicroelectronics" is identified as the standardized data item that matches the brand field in the original data.

[0105] The threshold is set based on the following: 0.8 is an empirical value derived from a large amount of experimental data, which can achieve a good balance between matching accuracy (reducing false matches) and recall (reducing missed matches). If the business has extremely high requirements for accuracy, the threshold can be appropriately increased (e.g., 0.9); if more emphasis is placed on coverage, the threshold can be appropriately decreased (e.g., 0.7).

[0106] Finally, the information entry system enters the original data into the data entry corresponding to the target electronic component in the preset information database according to the matched standardized data items.

[0107] A pre-built database is an internal database used by an enterprise to store component information. It can be an ERP system, PLM system, inventory management system, or a dedicated component database. Its structure is typically organized by component, with each component corresponding to an independent data entry containing the following fields: Identification fields: such as internal ID, part number, brand.

[0108] Description fields: such as product name, description, product catalog.

[0109] Technical parameters: such as package, operating temperature, supply voltage, bandwidth, etc.

[0110] Supply chain fields: such as supplier, inventory, price, MOQ, etc.

[0111] Document fields: such as specification sheet links, application notes, and reference designs.

[0112] Example of data entry: The information entry system will enter the following values ​​from the original data: Package field value "8-Lead LFCSP (4mm x 4mm)" → into the Package field. Operating temperature range "-40°C to +125°C" → into the Operating Temperature field. Supply voltage "2.7V to 6V" → into the Supply Voltage field. Brand (already matched as "STMicroelectronics") → into the Brand field (standardized). Product catalog (already matched as "Microcontroller (MCU)") → into the Product Catalog field (standardized).

[0113] If the target electronic component is being entered for the first time, the information entry system creates a new data entry in the preset information database, completing the information filing. If the target electronic component already exists in the preset information database, the system compares the newly acquired data with the existing data, supplements missing fields, updates changed fields, and retains an update log, completing the information update. This concludes the process of querying and processing the target electronic component's intrinsic information.

[0114] Optionally, the query intent also includes querying alternative component information for the target electronic component. The corresponding target processing flow includes: the basic steps and the alternative information query steps, wherein the alternative information query steps include: Extract key parameter information of the target electronic component from the acquired raw data; obtain parameter data of multiple candidate electronic components belonging to the same product category as the target electronic component from external data sources; calculate the similarity between the key parameter information of the target electronic component and the parameter data of each candidate electronic component to obtain parameter similarity; when the parameter similarity exceeds a preset substitution similarity threshold, the corresponding candidate electronic component is identified as a substitute for the target electronic component, and the information of the substitute is associated and stored in the data entry corresponding to the target electronic component in the preset information database; The query intent also includes querying information on supporting components for the target electronic component. The corresponding target processing flow includes: basic steps and supporting information query steps. The supporting information query steps include: The intelligent scheduling engine triggers the targeted web crawling module to capture application documents related to the target electronic component; it performs text parsing on the captured application documents to extract the identification information of other electronic components that co-occur with the target electronic component; it counts the co-occurrence frequency of each other electronic component and identifies other electronic components whose co-occurrence frequency exceeds a preset frequency threshold as matching components of the target electronic component; and it associates and stores the information of the matching components and their matching relationships with the information of the target electronic component in a preset information database.

[0115] The target processing flow corresponding to the selection intent based on the requirements description includes: The unstructured text is parsed to extract at least one constraint description; the constraint description includes at least a functional requirement description; each constraint description is mapped to a corresponding filtering condition; based on the filtering condition, candidate electronic components that meet the condition are retrieved from a preset information database, and a selection recommendation list containing all candidate electronic components is output.

[0116] In implementation, when the query intent identified by the information entry system is not a query for the entity itself, but rather another type of query intent, the information entry system will execute the corresponding target processing flow. In this embodiment, other query intents include at least: queries for alternative component information for a target electronic component, queries for supporting component information for a target electronic component, and selection intents based on requirement descriptions. The target processing flow corresponding to each query intent is described in detail below.

[0117] I. When the information entry system identifies that the user's query intent is for alternative component information of a target electronic component, the target processing flow executed by the information entry system includes: a basic step and an alternative information query step. The basic step is exactly the same as the basic step in the previous text information query, including: identifying key identification information (brand and part number) of the target electronic component from unstructured text, using an intelligent scheduling engine to obtain the original data of the target electronic component from multiple external data sources, and performing cross-validation and fusion processing on the obtained data. Specific implementation details have been described in detail above and will not be repeated here. After the basic step is completed, the information entry system enters the alternative information query step, which includes the following sub-steps: The information entry system extracts key parameter information for substitute matching from the acquired raw data of the target electronic component. Key parameter information refers to the set of core parameters that uniquely characterize the electrical and physical properties of the electronic component. The key parameter information differs for electronic components of different product categories, as substitute matching requires comparison based on the most essential technical characteristics of that category. For example, the key parameter information for the operational amplifier product category includes supply voltage range, bandwidth, slew rate, quiescent current, number of channels, input bias current, noise density, and package; the key parameter information for the microcontroller (MCU) product category includes core architecture, clock speed, Flash capacity, RAM capacity, number of I / Os, communication interface type, operating voltage range, and package. Accordingly, the information entry system, based on the product category of the target electronic component, obtains the list of key parameter information corresponding to that category from a pre-defined "category-key parameter mapping table," and then extracts the values ​​of these parameters from the raw data. If some key parameters are missing from the raw data, the information entry system can mark them as "unknown" or attempt to supplement them from other data sources. Next, the information entry system retrieves parameter data of multiple candidate electronic components of the same category from an external data source based on the category information of the target electronic component. Here, product category refers to the classification hierarchy of electronic components according to their function and application field. In this embodiment, the product category adopts a multi-level classification system. For example, integrated circuits (ICs) are a first-level category, which includes second-level categories such as analog ICs, data ICs, and power ICs. Analog ICs include operational amplifiers (such as AD8232, LM358, OPA2376, and MCP6002) as a third-level category; data ICs include microcontrollers (MCUs) (such as STM32F103, PIC16F877A, and ATMEGA328P) as a third-level category; and power ICs include DC-DC converters (such as TPS5430, LM2596, and MP2307) as a third-level category.

[0118] In the basic steps, the information entry system has acquired the original data of the target electronic component through the intelligent scheduling engine and performed cross-validation and fusion processing on the data. At this point, the information entry system has complete information on the target component, including its product catalog information. The sources of the product catalog information include: fields such as "product category" and "classification" directly extracted from the original data, category keywords parsed from the component description text, and category information mapped from industry standard classification codes. Although the current query intent does not include a complete ontology information query step, in order to discover substitutes later, the information entry system needs to perform lightweight standardization on the product catalog information of the target component to ensure the accuracy of subsequent queries within the same category. The lightweight standardization is implemented using rule-based keyword matching, with a preset rule table: containing "op amp", "operational amplifier", "op amp" → category = "operational amplifier (Op Amp)"; containing "MCU", "microcontroller", "single-chip microcomputer" → category = "microcontroller (MCU)"; containing "DC-DC", "power conversion" → category = "DC-DC converter".

[0119] Alternatively, the product catalog text can be input into a semantic vector model, and cosine similarity can be calculated between it and pre-stored standardized category name vectors. The standardized category with the highest similarity can be taken as the result.

[0120] After the information entry system determines the standardized category of the target electronic component, it needs to decide at which level to search for candidate electronic components. This embodiment adopts a layer-by-layer expansion strategy: priority is given to exact matching (first searching for candidate components in the same third-level category), if no results are found, it expands to the next higher level (if there are no results or too few results in the same third-level category, it expands to the second-level category), and if there are still no results in the second-level category, it expands to the first-level category.

[0121] The information entry system, based on a defined product category and selected level, initiates queries to external data sources to obtain parameter data for other electronic components (i.e., candidate electronic components) within the same category. Query methods include: API query: calling a third-party API interface to filter and return a list of candidate electronic components by category; Database query: querying candidate electronic components of the same category from the system's internal pre-set information database; Targeted web crawling module: crawling candidate electronic components of the same category from industry websites.

[0122] The acquired candidate electronic component parameter data should include the same set of key parameters as the target electronic component to facilitate subsequent similarity comparison. Example: The target electronic component is AD8232 (operational amplifier). The information entry system should retrieve candidate operational amplifiers of the same category, including: LM358: Supply voltage 3-32V, bandwidth 1MHz, slew rate 0.3V / µs, quiescent current 500µA, number of channels 2, package DIP-8.

[0123] OPA2376: Supply voltage 2.2-5.5V, bandwidth 5.5MHz, slew rate 2V / µs, quiescent current 760µA, number of channels 2, package MSOP-8.

[0124] MCP6002: Supply voltage 1.8-6V, bandwidth 1MHz, slew rate 0.6V / µs, quiescent current 100µA, number of channels 2, package SOIC-8.

[0125] TL082: Supply voltage ±3-18V, bandwidth 3MHz, slew rate 13V / µs, quiescent current 1.4mA, number of channels 2, package DIP-8.

[0126] The information entry system calculates the similarity between the key parameter information of the target electronic component and the parameter data of each candidate electronic component to obtain the parameter similarity. The parameter similarity calculation method is as follows: Since the key parameters contain multiple data types (numerical, Boolean, enumerated, and text), different similarity calculation strategies are required for different types: For numerical parameters (such as supply voltage, bandwidth, quiescent current, etc.), an appropriate calculation method needs to be selected based on the specific form of the parameter. In this embodiment, numerical parameters can be divided into two categories: single-valued parameters and range parameters.

[0127] For parameters with a single numerical value (e.g., bandwidth 2MHz, quiescent current 620µA), a range-based relative difference calculation is used: Numerical similarity = max(0, 1 - |v_target - v_candidate| / (max_range × tolerance_factor)); where v_target refers to the parameter value of the target electronic component, v_candidate is the parameter value of the candidate electronic component, and max_range is the preset typical value range of the parameter in the category (i.e., the difference between the possible maximum and minimum values ​​of the parameter in this category). max_range is calculated based on the parameters of all common components in this category. For example, for the bandwidth parameter of an operational amplifier: common operational amplifier bandwidth range: 10kHz~100MHz; max_range = 100MHz - 10kHz ≈ 100MHz. tolerance_factor is a preset tolerance factor, reflecting the allowable deviation of the parameter, and its value range is usually 0.1~0.3. For parameters requiring high precision (such as reference voltage), take a smaller value (such as 0.05); for parameters with less stringent requirements (such as bandwidth), take a larger value (such as 0.2).

[0128] For parameters with a range (e.g., a supply voltage range of 2.7-6V), a single value cannot be directly substituted into the formula; the range overlap must be calculated first. The range parameter consists of a minimum value v_min and a maximum value v_max, indicating that the component can operate within this voltage range. Range overlap = overlap interval length / target interval length; where: overlap interval length = max(0, min(v_target_max, v_candidate_max) - max(v_target_min, v_candidate_min)); target interval length = v_target_max - v_target_min.

[0129] For example, the power supply voltage range parameters are as follows: target AD8232 power supply range: v_target_min=2.7V, v_target_max=6V; candidate LM358 power supply range: v_candidate_min=3V, v_candidate_max=32V. The lower limit of the overlapping interval = max(2.7,3)=3V; the upper limit of the overlapping interval = min(6,32)=6V; the length of the overlapping interval = 6-3=3V; the length of the target interval = 6-2.7=3.3V; the range overlap = 3 / 3.3 = 0.91; for range parameters, the numerical similarity = range overlap, i.e., 0.91.

[0130] For Boolean parameters (such as "whether it is rail-to-rail output") or enumeration parameters (such as encapsulation type), exact matching or semantic matching is used: if there is an exact match, the similarity is 1.0; if they are semantically equivalent (such as "LFCSP-8" and "8-LFCSP"), the similarity is 0.8; if there is no match, the similarity is 0. The information entry system performs a weighted summation of the similarities of each parameter, resulting in parameter similarity = ∑(parameter similarity_i × weight_i); where weight_i reflects the importance of the parameter in the substitution judgment. Weight_i can be set through expert experience or learned from historical substitution data through machine learning methods, and the sum of the weights of all parameters is 100%.

[0131] Finally, the information entry system compares the calculated parameter similarity with a preset substitution similarity threshold. The substitution similarity threshold is set according to business requirements; in this embodiment, it is set to 0.8. This means that only candidate electronic components with a parameter similarity of 80% or higher are considered qualified substitutes.

[0132] The information entry system associates and stores the information of identified substitutes with the target electronic components. Specific operations include: adding a "substitute" field to the data entries for the target electronic components in a pre-defined database, recording the substitute's identification information (such as internal ID, part number, brand); simultaneously, adding a "substituted" field to each substitute's data entry, recording the information of the target electronic component being substituted; and storing the confidence level (i.e., parameter similarity) of the substitution relationship.

[0133] II. When the system identifies that the user's query intent is for information on matching components for a target electronic component, the target processing flow executed by the information entry system includes: basic steps (as before) and matching information query steps. The matching information query steps include the following sub-steps: The information entry system utilizes an intelligent scheduling engine to trigger a targeted web crawling module, retrieving application documents related to the target electronic components from multiple sources. Application documents refer to technical documents containing information such as actual application scenarios, reference designs, and typical circuits for electronic components. Common types include: Application Notes (sourced from the manufacturer's official website; an example document is "AN-1234: Application of AD8232 in ECG"), Reference Designs (sourced from the manufacturer's official website and design communities; an example document is "Motor Control Reference Design Based on STM32F103"), Evaluation Board User Guides (sourced from the manufacturer's official website; an example document is "EVAL-AD8232 Evaluation Board User Manual"), Open Source Hardware Schematics (sourced from GitHub and Hackaday; an example document is "Arduino Uno Schematic (including compatible components)"), and Technical Forum Posts (sourced from EEBVBlog, 21ic, and StackExchange; an example document is "Seeking Recommendations for Power Chips Compatible with STM32F103").

[0134] The intelligent scheduling engine constructs search keywords based on the brand and part number of the target electronic component and initiates crawling requests to preset target websites. For example, for AD8232, it crawls a list of application notes from the target website: Analog Devices official website; and crawls technical articles from the target website: DigiKey technology library.

[0135] The information entry system performs text parsing on the captured application documents to extract the identification information of other electronic components that co-occur with the target electronic component. The specific text parsing process is as follows: text preprocessing (removing HTML tags, script code, and sample code; removing irrelevant content such as advertisements, navigation bars, and footers; converting the document to plain text format), and component identification information recognition (using a recognition method similar to the large language model entity extraction method described earlier to identify component identification information from the text; designing specific extraction instructions tailored to the characteristics of the application documents, such as "You are an electronic component identification assistant. Please identify all the mentioned electronic components from the following technical document text and extract the brand and part number of each component. If the brand is not explicitly mentioned, it can be left blank. Please output in JSON array format."). In addition to identifying individual components, the information entry system also determines the relationship between these components and the target electronic component. If the following pattern exists in the text, it is determined to be "co-occurring": If the same circuit diagram pattern is used, such as "the circuit consists of AD8232 and STM32F103", it is determined that they appear together.

[0136] If a matching recommendation pattern is used, such as "AD8232 is usually used with LM358", then it is determined to be a common occurrence.

[0137] Reference design patterns, such as "This reference design uses AD8232 as the front end and STM32F103 as the controller", are considered to co-occur.

[0138] In Bill of Materials (BOM) mode, if AD8232 and TPS5430 are both included in the BOM, they are considered to appear together.

[0139] The information entry system performs statistical analysis on commonly occurring components extracted from multiple application documents, calculating the co-occurrence frequency of the target electronic component with each other electronic component. Co-occurrence frequency refers to the proportion of times a certain other electronic component and the target electronic component co-occur in the same application document out of the total number of documents. Co-occurrence frequency = (Number of documents containing both the target and other electronic components) / (Total number of documents) × 100%.

[0140] Next, the information entry system compares the calculated co-occurrence frequency with a preset frequency threshold. The frequency threshold is preset according to business needs. In this embodiment, it is set to 20%, which means that only other electronic components with a co-occurrence frequency of 20% or higher are considered to be supporting devices for the target electronic component.

[0141] Finally, the information entry system associates and stores the data and matching relationships of the identified supporting components with the target electronic components. Specific operations include: adding a "Supporting Components" field to the data entries of the target electronic components in the preset information database, recording the identification information of the supporting components and their frequency of occurrence; simultaneously, adding an "Application Scenarios" field to each supporting component's data entry, recording which components it is commonly used with.

[0142] Third, when the information entry system identifies the user's query intent as a selection intent based on the requirement description, the target processing flow executed by the information entry system is an independent selection processing flow, which does not include basic steps (because there are no specific target components in the input).

[0143] The information entry system parses the unstructured text input by the user to extract at least one constraint description. The constraint description describes the user's requirements regarding the function, performance, price, etc., of the required electronic components. In this embodiment, the constraint description includes at least functional requirements, and may also include price constraints, packaging constraints, etc. The specific requirement parsing method is as follows: The information entry system adopts an intent parsing method based on the Large Language Model (LLM). It leverages the powerful language understanding capabilities and domain knowledge of the Large Language Model (LLM) to automatically identify and extract various constraints from the user's natural language input through carefully designed parsing instructions.

[0144] Large language models (such as the GPT series, GLM series, and LLaMA series) are pre-trained on massive amounts of internet text data, and have learned rich world knowledge and language patterns. Specifically, in the scenario of electronic component selection, large language models have already encountered a large amount of text containing descriptions of component requirements during pre-training, thus possessing the following capabilities: (1) Ability to recognize functional requirements: The large language model can recognize functional requirements such as "op-amp", "MCU", and "power chip" because the pre-training corpus contains a large number of such terms and their contexts. For example, if the training corpus is "I need a low-power op-amp for signal conditioning", the model learns that "op-amp" = operational amplifier, which is a functional requirement. If the training corpus is "Choose an MCU as the main controller", the model learns that "MCU" = microcontroller, which is a functional requirement. Through training with a large amount of such corpus, the model establishes the concept of functional requirements and memorizes a large number of specific functional terms and their synonyms (such as "op-amp" = "operational amplifier" = "op amp").

[0145] (2) Ability to recognize performance parameters: The large language model can recognize performance parameters such as "bandwidth 100kHz" and "5V power supply" because the model has learned the combination pattern of parameters and values: the parameter type is bandwidth, and the corresponding common expression pattern is "bandwidth 100kHz", "100kHz bandwidth", "bandwidth ≥ 100k", and the corresponding rule learned by the model is: "bandwidth" + value + unit. The parameter type is power supply voltage, and the corresponding common expression pattern is "5V power supply", "power supply voltage 5V", "operating voltage 5V", and the corresponding rule learned by the model is "power supply" + value + "V".

[0146] (3) Ability to recognize price constraints: The large language model can recognize price constraints such as "within 3 yuan" and "price below 5 yuan" because the model has learned price-related expression patterns. For example, when the price pattern is an upper limit constraint (such as "within 3 yuan", "not exceeding 3 yuan", "below 3 yuan"), the corresponding pattern learned by the model is: "within / not exceeding / below" + value + "yuan / yuan". When the price pattern is a range constraint ("2-3 yuan", "between 2 and 3 yuan"), the corresponding pattern learned by the model is: value 1 + "-" + value 2 + "yuan".

[0147] (4) Ability to recognize package constraints: The large language model can recognize package requirements such as "surface mount package", "LQFP", and "SOIC-8" because the model has learned package-related terms. For example, if the package type is the mounting method (such as "surface mount package", "through-hole", "SMD"), the corresponding rule learned by the model is: mounting method terms such as "surface mount / through-hole". If the package type is a specific package (such as "LQFP-48", "SOIC-8", "QFN"), the corresponding rule learned by the model is: package code + number of pins.

[0148] Design specialized parsing instructions. These instructions tell the model which constraints need to be extracted from the text, specifying the type of constraints to extract (e.g., functional constraints, performance constraints, price constraints, packaging constraints, etc.), and requiring output in JSON format. A few examples will help the model understand the task. For example: "You are an electronic component selection assistant, specifically helping engineers and purchasing personnel extract selection constraints from natural language requirements. Constraints include: functional requirements (e.g., 'op-amp', 'MCU', 'power chip'), performance parameters (e.g., 'bandwidth 100kHz', '5V power supply'), price constraints (e.g., 'within 3 yuan'), packaging requirements (e.g., 'surface mount package'), etc. Please output in JSON format." The information entry system maps the extracted constraint descriptions to filter conditions that can be used for database queries. Specifically, the information entry system maintains a "constraint-filter condition mapping table" to convert natural language descriptions into structured query conditions. For example: When the constraint type is functional requirement, the corresponding example natural language description is "op-amp" or "operational amplifier", and the corresponding mapped filter condition is: product catalog = "operational amplifier (Op Amp)".

[0149] When the constraint type is power supply voltage, the corresponding example natural language description is "5V power supply" or "power supply voltage 5V", and the corresponding mapped filtering condition is: the power supply voltage range includes 5V.

[0150] The constraint type is bandwidth, and the corresponding example natural language description is "bandwidth 100kHz" or "100kHz bandwidth". The corresponding mapped filtering condition is: bandwidth ≥ 100kHz.

[0151] The constraint type is price, and the corresponding example natural language description is "within 3 yuan" and "price less than 3 yuan". The corresponding mapped filtering condition is: price ≤ 3.0.

[0152] ... For descriptions with comparative relationships (such as "above", "within", "around"), the information entry system maps them to numerical range conditions. For example, if the description is "above 100kHz", the corresponding numerical range condition is ≥100kHz; if the description is "within 3 yuan", the corresponding numerical range condition is ≤3.0.

[0153] The information entry system retrieves candidate electronic components that meet the filtering criteria obtained through mapping from a preset information database. The retrieval strategy includes: Exact match: For enumerated conditions (such as functional requirements or package types), perform an exact match query. For example: Product catalog = "Operational Amplifier (Op Amp)".

[0154] Range matching: For numerical conditions (such as bandwidth and price), perform a range query. For example: bandwidth ≥ 100kHz AND price ≤ 3.0.

[0155] Combined query: Multiple conditions are linked using "AND" logic, meaning all conditions must be met simultaneously. Product catalog = "Operational amplifier" AND power supply voltage range including 5V AND bandwidth ≥ 100kHz AND price ≤ 3.0.

[0156] The information entry system retrieves the set of alternative electronic components and outputs a selection recommendation list.

[0157] Optionally, the basic step of "automatically acquiring raw data corresponding to the target electronic component from multiple predetermined external data sources using a preset intelligent scheduling engine" also includes a data source strategy determination step: Determine the product category of the target electronic components based on key identification information; Based on the query intent and product category, the system queries the preset data source strategy library to determine the data collection strategy that matches the current query. The data collection strategy includes a list of APIs to be called, a list of target websites to be crawled, and the call priority of each data source. Among them, the preset intelligent scheduling engine automatically obtains the original data corresponding to the target electronic components from multiple predetermined external data sources, and the intelligent scheduling engine executes the data acquisition according to the determined data acquisition strategy.

[0158] In implementation, when the information entry system prepares to retrieve raw data of target electronic components from external data sources using the intelligent scheduling engine, it does not simply call all data sources in a fixed order. Instead, it first executes a data source strategy determination step, dynamically selecting the optimal combination of data sources and the calling order based on the specific circumstances of the current query. This data source strategy determination step occurs within the basic steps, specifically after "identifying key identifier information from unstructured text" and before "retrieving raw data using the intelligent scheduling engine." Its core process is as follows: First, the information entry system determines the product category of the target electronic component based on the identified key identification information. Then, it queries the preset data source strategy library in conjunction with the query intent and product category. Next, it obtains the data acquisition strategy that matches the current query, which includes a list of APIs to be called, a list of target websites to be crawled, and the call priority of each data source. Finally, the intelligent scheduling engine executes the data acquisition operation according to the determined data acquisition strategy.

[0159] Product categories are hierarchical classifications of electronic components based on their function and application areas. This embodiment adopts a three-level classification system, which is the same as the product category system described in step S4 above, and will not be repeated here. In short, the first-level category is such as "integrated circuits", the second-level category is such as "analog ICs", and the third-level category is such as "operational amplifiers". Specific components such as AD8232 belong to the third-level category "operational amplifiers".

[0160] The information entry system determines product categories using several methods: First, based on part number lookup. The system pre-builds and maintains a "part number-category mapping table," recording the category information corresponding to common part numbers. For example, part number AD8232 corresponds to the third-level category "operational amplifier," and part number STM32F103C8T6 corresponds to the third-level category "microcontroller (MCU)." When the queried part number exists in the mapping table, its category information can be directly obtained. Second, based on brand and model pattern inference. For part numbers not found in the mapping table, the system infers categories based on part number naming rules and brand characteristics. For example, components with the brand "ADI" and model numbers starting with "AD" are mostly analog ICs; components with the brand "STM" and model numbers starting with "STM32" are mostly microcontrollers (MCUs). The system uses a pre-defined rule base for this inference. Third, by calling an external API. When the above methods fail to determine the product category, the information entry system can query the detailed information of the part number through a third-party API (such as Octopart) and extract the product category information from the returned data.

[0161] The information entry system selects the appropriate category granularity based on the query intent and the amount of information available. Specifically, for queries on the subject matter, the most precise third-level category is usually used; for queries on substitutions, the third-level category is the primary focus, but it can be expanded to the second-level category if necessary; for queries on complementary items, the second-level or first-level category may be required; and for queries on selection, the category granularity is determined based on the functional categories specified in the user's requirements description.

[0162] The data source strategy library is a pre-built knowledge base for storing data source invocation strategies for different query scenarios. It records under what circumstances which data sources should be invoked, and in what order. At the core of the data source strategy library is a strategy mapping table, where each strategy record contains the following fields: Query Intent Field: Records the type of query intent to which this strategy applies, such as "ontology information query", "alternative query", "complementary query" or "selection query".

[0163] Primary Category Field: Records the primary category of components to which this strategy applies, such as "Integrated Circuits (IC)".

[0164] Secondary Category Field: Records the secondary categories of components to which this strategy applies, such as "Analog IC".

[0165] The third-level category field records the third-level classification of components to which this strategy applies, such as "operational amplifier".

[0166] Data source list field: Records the data sources applicable to this scenario, their priority, calling parameters, and other information.

[0167] Validity period field: Records the valid period of the policy, such as from January 1, 2024 to December 31, 2024.

[0168] Confidence field: Records the reliability of the strategy, represented by a value between 0 and 1.

[0169] For each data source, record the following details in the data source list field: Data Source Identifier: A unique identifier for the data source, such as "digikey_api" or "ti_website_crawler". Data Source Type: Identifies whether the data source is an API interface or a web crawling module. Priority: Identifies the calling order of this data source in the current strategy; the smaller the value, the higher the priority. Call Parameters: Records the specific parameters required to call this data source, such as the API endpoint address, request parameter template, and crawler URL template.

[0170] The data source strategy library is built and continuously optimized in the following ways: First, domain experts, based on experience, preset initial strategies for common product categories and query intents. For example, for queries about operational amplifiers, experts can prioritize the original manufacturer's official website (due to its high authority), followed by APIs from authorized distributors such as DigiKey and Mouser (due to their comprehensive data). For queries about domestically produced components, domestic trading platforms can be prioritized, followed by the original manufacturer's official website. Second, the information entry system records the execution results of each query, including indicators such as success rate, response time, and data completeness, and dynamically optimizes the strategy through statistical analysis. For example, if the information entry system finds that the success rate of a certain API interface for querying a certain type of component is consistently low, it will automatically lower its priority. Finally, the data source strategy library is updated regularly, such as monthly, to include new data sources, eliminate invalid data sources, and adjust priority rankings.

[0171] Once the information entry system determines the "query intent" and "product category" of the current query, it uses these two dimensions as query conditions to retrieve matching data collection strategies from the data source strategy library.

[0172] The data entry system uses a step-by-step matching approach to find the most suitable strategy, with the following specific rules: First, the system attempts to find a strategy that simultaneously satisfies both the "query intent" and "third-level category" conditions. If such a strategy exists, it is used directly. If an exact match fails, the system searches for a strategy that satisfies both the "query intent" and "second-level category" conditions. If no match is found, the system searches for a strategy that satisfies both the "query intent" and "first-level category" conditions. If no match is found in either of these cases, the system uses a general strategy corresponding to the intent type, which is not limited to a specific category.

[0173] For example, consider an ontology information query for operational amplifiers. Assume the query intent is "ontology information query," and the product category is the third-level category "operational amplifier." The data entry system first attempts an exact match, searching for a strategy that combines "ontology information query + operational amplifier." Since this combination has a pre-defined strategy, the data entry system directly uses that strategy. This strategy might include the following data source list: the highest priority ADI website crawler, the second highest priority DigiKey API, the third highest priority Mouser API, and other backup data sources.

[0174] The priority of records in the data source policy library is determined based on the following three static factors: The first factor: Authority, with a score range of 60-100 points, including the original manufacturer's official website (90-100 points), authorized distributors (80-89 points), aggregation platforms (70-79 points), and forums (60-69 points).

[0175] The second factor is data richness, which is scored from 60 to 100 points. It refers to the number and quality of fields in the returned data, that is, the coverage ratio of key fields in the returned data.

[0176] The third factor is cost, with a score range of 60-100 points. This refers to the API response time or the time it takes for the web crawler to capture data; the faster the response, the higher the score.

[0177] For example, the data source score is calculated as follows: Authority score × 50% + Data richness score × 30% + Cost score × 20%. The higher the data source score, the higher its priority.

[0178] Optionally, the information entry method may also include the following steps: Receive user monitoring settings for target electronic components in a preset information database. The monitoring settings include monitoring frequency and monitoring indicators. The monitoring indicators include at least one or more of the following: price changes, inventory status, and life cycle status. According to the monitoring frequency, perform basic steps regularly on the monitored electronic components to obtain the latest raw data. The newly acquired raw data is compared with the stored data to detect changes in monitoring indicators. When changes in monitored metrics exceed preset warning thresholds, warning information is generated and pushed to the user.

[0179] In practice, after the initial entry of electronic component data, the component information is not static. Prices may fluctuate, inventory may change, and products may be discontinued. Failure to promptly grasp these changes can lead to procurement decision errors, production delays, or cost overruns. Therefore, this application provides a component lifecycle monitoring step to achieve continuous tracking and proactive early warning of components of interest. The component lifecycle monitoring step is an independent, optional enhancement. After completing basic component information queries and entry, users can set monitoring tasks for components already existing in a preset database. The information entry system will periodically and automatically execute the basic steps according to the user-defined monitoring frequency, obtain the latest component data, compare it with the stored data, and proactively push early warning information to the user when significant changes are detected.

[0180] Users can initiate monitoring task settings in two ways: Method 1: Setting up on the component details page. When a user views detailed information about a component in the preset information library, the page provides a "Set Monitoring" button, which the user can click to enter the monitoring settings interface. Method 2: Using the batch setting function. On the component list page, users can select multiple components that need to be monitored and then batch set the same monitoring parameters.

[0181] When setting up monitoring tasks, users need to configure the following parameters: Monitoring frequency defines the cycle in which the information entry system checks changes in component information. The information entry system provides several preset frequencies for users to choose from. Monitoring frequencies include real-time monitoring (daily checks, suitable for general components with frequent price fluctuations); weekly monitoring (e.g., checks every Monday, suitable for regular components); monthly monitoring (e.g., checks on the 1st of each month, suitable for components with stable usage); or user-defined monitoring frequencies.

[0182] Monitoring metrics define the types of information that users are interested in, and users can select one or more metrics to monitor. For example, monitoring metrics could be price changes (i.e., monitoring changes in component prices), inventory status (monitoring changes in inventory quantity), or package / specification changes (monitoring changes in technical parameters, such as changes in package descriptions or parameter corrections).

[0183] For each monitoring indicator, users can set an alert threshold, which is triggered when the change exceeds the threshold. For example, when the monitoring indicator is price change, the corresponding alert threshold can be a price fluctuation exceeding 10%. When the monitoring indicator is inventory status, the corresponding alert threshold can be "out of stock" or "production halted".

[0184] After the user sets up the monitoring task, the system stores the monitoring task information in the monitoring task table. The monitoring task table includes the following fields: Task ID (a unique identifier for the task, such as task_10001), Component ID (the internal ID of the monitored component, such as comp_8232), Monitoring Frequency, Monitoring Indicators, Warning Threshold, Last Execution Time (i.e., the time when the last monitoring was executed), and Next Execution Time (the planned time for the next execution based on the monitoring frequency).

[0185] The information entry system maintains a scheduled task operator that periodically scans the monitoring task table, finding all tasks with a "next execution time" less than or equal to the current time, and adds them to the execution queue. For example, at 2:00 AM every day, the system scans all monitoring tasks with a status of "enabled" and a "next execution time" that is on or before the current day, triggering their execution in batches. For each triggered monitoring task, the information entry system performs basic steps on the monitored target electronic components to obtain the latest raw data. It should be noted that the basic steps performed by the monitoring task are exactly the same as those during the initial query, but do not include subsequent semantic matching and data entry steps (unless otherwise configured by the user). This design aims to obtain the latest raw data for comparison with the stored data.

[0186] The information entry system compares the latest acquired raw data with the component data stored in the preset information database, focusing on user-defined monitoring indicators. For example, price change comparison: price data typically includes multiple dimensions: tiered prices for different purchase quantities, quotes from different suppliers, etc. The information entry system uses the following rules for comparison: when the price type is a standard unit price, a direct numerical comparison rule is used; when the price type is a price tier, a comparison rule comparing the prices of each tier is used (e.g., the price for 100 pieces increases from 2.35 yuan to 2.50 yuan); when the price type is a supplier quote, a comparison rule comparing the lowest quote is used (e.g., the lowest quote increases from 2.3 yuan to 2.45 yuan). The price change margin is calculated as: |New Price - Old Price| / Old Price × 100%.

[0187] The information entry system compares the detected degree of change with the user-defined warning threshold. When the degree of change exceeds the threshold, an warning is triggered, generating a warning message. The warning message includes at least the component identifier (brand and part number), the change indicator (the monitored indicator that has changed), the change details (old value → new value), and the magnitude of the change. Finally, the warning message is pushed to the user (e.g., via email / SMS notification).

[0188] This application also discloses an intelligent collection and automatic input system for electronic component product data. It includes: The semantic parsing module is used to receive unstructured text input by the user, parse the unstructured text, and determine the user's query intent; The process processing module is used to determine and execute the corresponding target processing flow based on the query intent; The query intent includes at least a query for the entity information of the target electronic component, and the target processing flow corresponding to the query for the entity information of the target electronic component includes: Basic steps: Identify key identification information of target electronic components from unstructured text; based on key identification information, use a preset intelligent scheduling engine to automatically obtain raw data corresponding to the target electronic components from multiple predetermined external data sources, and cross-validate the raw data obtained from different data sources. Ontology information query steps: Semantically match the cross-validated raw data with the pre-stored standardized data to determine the matching standardized data items; based on the matching standardized data items, enter the raw data into the data entry corresponding to the target electronic component in the preset information database to complete the information filing or update of the target electronic component and end the target processing flow.

[0189] This application also discloses an intelligent collection and automatic input device for electronic component product data. The device includes a memory and a processor. The memory stores a computer program that can be loaded by the processor and executed as described above for the intelligent collection and automatic input method for electronic component product data.

[0190] This application also discloses a computer-readable storage medium that stores a computer program that can be loaded by a processor and executed as described above in the method for intelligent collection and automatic entry of electronic component product data. The computer-readable storage medium includes, for example, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0191] It should be noted that in this paper, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0192] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit the scope of protection of the application. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on these embodiments, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

Claims

1. A method for intelligent collection and automatic input of electronic component product data, characterized in that, include: Receive unstructured text input by the user, parse the unstructured text, and determine the user's query intent; Based on the query intent, determine and execute the corresponding target processing flow; The query intent includes at least a query for the entity information of the target electronic component, and the target processing flow corresponding to the query for the entity information of the target electronic component includes: Basic steps: Identify key identification information of the target electronic component from the unstructured text; based on the key identification information, use a preset intelligent scheduling engine to automatically obtain original data corresponding to the target electronic component from multiple predetermined external data sources, and cross-validate the original data obtained from different data sources; Ontology information query steps: Semantically match the cross-validated raw data with the pre-stored standardized data to determine the matching standardized data items; based on the matching standardized data items, enter the raw data into the data entry corresponding to the target electronic component in the preset information database to complete the information filing or update of the target electronic component and end the target processing flow.

2. The method for intelligent collection and automatic entry of electronic component product data according to claim 1, characterized in that, The process involves automatically acquiring raw data corresponding to the target electronic component from multiple predetermined external data sources using a pre-set intelligent scheduling engine, and cross-validating the raw data acquired from different data sources, including: The system utilizes a pre-defined intelligent scheduling engine to call at least one pre-defined third-party API interface in parallel to obtain the first raw data corresponding to the target electronic component. When the first original data is empty or does not meet the preset integrity requirements, the intelligent scheduling engine automatically triggers the preset targeted web page crawling module, so that the targeted web page crawling module can crawl the second original data corresponding to the target electronic component from the web pages of at least one target website based on the preset crawling rules. The acquired first and second raw data are fused together, and the fusion process includes at least the following: When the same data item has multiple sources, cross-validate the consistency of the data values ​​from all sources corresponding to the data item, and retain the data value with the highest confidence. When data items from different sources complement each other, the data from each source are merged to form complete original data.

3. The method for intelligent collection and automatic entry of electronic component product data according to claim 1, characterized in that, The step of semantically matching the cross-validated raw data with pre-stored standardized data to determine the matching standardized data items includes: The target field and its text information are extracted from the cross-validated raw data. The text information is then input into a pre-trained semantic vector model and transformed into a first semantic vector. Each standard information corresponding to the target field in the pre-stored standardized data is input into the semantic vector model and transformed into the corresponding second semantic vector; Calculate the similarity between the first semantic vector and each of the second semantic vectors; When the similarity exceeds a preset threshold, the corresponding standard information is determined as a standardized data item that matches the target field.

4. The method for intelligent collection and automatic entry of electronic component product data according to claim 1, characterized in that, The query intent also includes querying alternative component information for the target electronic component. The corresponding target processing flow includes: the basic steps and the alternative information query steps. The alternative information query steps include: extracting key parameter information of the target electronic component from the acquired original data; obtaining parameter data of multiple candidate electronic components belonging to the same product category as the target electronic component from the external data source; calculating the similarity between the key parameter information of the target electronic component and the parameter data of each candidate electronic component to obtain parameter similarity; when the parameter similarity exceeds a preset alternative similarity threshold, determining the corresponding candidate electronic component as a substitute for the target electronic component, and storing the information of the substitute in association under the data entry corresponding to the target electronic component in the preset information database. The query intent also includes querying information on supporting components for the target electronic component. The corresponding target processing flow includes: the basic steps and the supporting information query steps. The supporting information query steps include: the intelligent scheduling engine triggers the targeted web crawling module to crawl application documents related to the target electronic component; the crawled application documents are parsed to extract the identification information of other electronic components that co-occur with the target electronic component; the co-occurrence frequency of each other electronic component is counted, and other electronic components whose co-occurrence frequency exceeds a preset frequency threshold are identified as supporting components of the target electronic component; the information on the supporting components and the supporting relationship are associated and stored with the information on the target electronic component in the preset information database.

5. The method for intelligent collection and automatic entry of electronic component product data according to claim 1, characterized in that, The query intent also includes a selection intent based on the requirement description, and the target processing flow corresponding to the selection intent based on the requirement description includes: The unstructured text is parsed to extract at least one constraint description; the constraint description includes at least a functional requirement description. Each constraint is described and mapped to a corresponding filtering condition. Based on the filtering condition, candidate electronic components that meet the condition are retrieved from the preset information database, and a selection recommendation list containing all the candidate electronic components is output.

6. The method for intelligent collection and automatic entry of electronic component product data according to claim 2, characterized in that, Before automatically acquiring the original data corresponding to the target electronic component from multiple predetermined external data sources using a preset intelligent scheduling engine, the process also includes a data source strategy determination step: The product category of the target electronic component is determined based on the key identification information; Based on the query intent and the product category, query the preset data source strategy library to determine the data collection strategy that matches the current query; The data collection strategy includes a list of APIs to be called, a list of target websites to be crawled, and the calling priority of each data source; Specifically, the system utilizes a pre-defined intelligent scheduling engine to automatically acquire raw data corresponding to the target electronic component from multiple predetermined external data sources, and the intelligent scheduling engine executes the acquisition according to the determined data acquisition strategy.

7. The method for intelligent collection and automatic entry of electronic component product data according to claim 1, characterized in that, The method further includes: The system receives user monitoring settings for target electronic components in a preset information database. The monitoring settings include monitoring frequency and monitoring indicators. The monitoring indicators include at least one or more of the following: price changes, inventory status, and life cycle status. According to the monitoring frequency, the basic steps are periodically performed on the monitored target electronic components to obtain the latest raw data. The newly acquired raw data is compared with the stored data to detect changes in monitoring indicators. When changes in monitored metrics exceed preset warning thresholds, warning information is generated and pushed to the user.

8. A smart system for collecting and automatically entering product data of electronic components, characterized in that, include: The semantic parsing module is used to receive unstructured text input by the user, parse the unstructured text, and determine the user's query intent; The process processing module is used to determine and execute the corresponding target processing flow based on the query intent; The query intent includes at least a query for the entity information of the target electronic component, and the target processing flow corresponding to the query for the entity information of the target electronic component includes: Basic steps: Identify key identification information of the target electronic component from the unstructured text; based on the key identification information, use a preset intelligent scheduling engine to automatically obtain original data corresponding to the target electronic component from multiple predetermined external data sources, and cross-validate the original data obtained from different data sources; Ontology information query steps: Semantically match the cross-validated raw data with the pre-stored standardized data to determine the matching standardized data items; based on the matching standardized data items, enter the raw data into the data entry corresponding to the target electronic component in the preset information database to complete the information filing or update of the target electronic component and end the target processing flow.

9. A device for intelligent collection and automatic input of electronic component product data, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 7.