Vehicle-mounted chip industry chain monitoring system and method based on NLP technology
The vehicle chip industry chain monitoring system based on NLP technology solves the problem of low efficiency of traditional analysis methods, and realizes efficient and accurate analysis and real-time monitoring of the vehicle chip industry chain, thereby improving analysis efficiency and the objectivity of results.
Patent Information
- Application Number
- CN202511063969.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional chip supply chain analysis methods rely on manual analysis, which is inefficient and prone to errors, and cannot meet the needs of modern enterprises for real-time and accurate analysis. Furthermore, chip supply chain information is scattered across different platforms, making data acquisition and integration complex.
The vehicle chip industry chain monitoring system based on NLP technology includes data acquisition, classification and grading, and visualization output modules. It uses deep learning and KNN algorithms for data analysis and grading, combines RPA technology to automatically collect and update the feature library, cleans the data through the SimHash algorithm, constructs a multi-dimensional feature model, and performs sentiment analysis and visualization.
It enables efficient and accurate analysis of the automotive chip industry chain, avoiding the time-consuming and error-prone nature of manual analysis, improving analysis efficiency and the objectivity of results, and meeting enterprises' needs for real-time information.
Smart Images

Figure CN120996343A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of chip industry chain information management in the automobile industry, and particularly relates to a vehicle-mounted chip industry chain monitoring system and method based on NLP technology. BACKGROUND
[0002] The automobile industry is facing the transformation demand of new energy and networking, and in this process, the vehicle needs to be greatly adjusted in the vehicle three-electricity controller and the vehicle machine system. Compared with traditional fuel vehicles, new energy vehicles need to arrange a large number of body situation perception controllers and their energy controllers, and the performance and stability of the controller determine the safety and reliability of the vehicle model. The core key component of the high-performance controller is the high-performance vehicle-grade chip.
[0003] Chips involve design, manufacturing, packaging testing, materials and equipment, and fluctuations in any link may trigger the "butterfly effect". And the chip and its related controller are high-priced and high-precision parts, and their cost can be as high as one-third of the cost of the vehicle in some high-end vehicle models. In the current design definition, stability and reliability of chip logistics procurement are the key links in the development process of new energy vehicles. If the chip supply shortage problem occurs, it will directly affect the continuity of vehicle production, because it is necessary to monitor the supply chain information in real time. Through real-time monitoring of supplier capacity and logistics status, enterprises can adjust procurement strategies in advance and reduce the risk of single dependence.
[0004] The traditional industry chain analysis method mainly relies on manual analysis, which is low in efficiency and easy to make mistakes, and cannot meet the real-time and accurate analysis needs of modern enterprises. Since the chip industry chain information data belongs to different enterprises and different data platforms, each data platform has a corresponding database and a corresponding data form, which makes the information acquisition and integration more complex and increases the difficulty of analysis. SUMMARY
[0005] In order to solve the problems proposed in the background art, the application provides a vehicle-mounted chip industry chain monitoring system and method based on NLP technology.
[0006] One of the purposes of the application is a vehicle-mounted chip industry chain monitoring system based on NLP technology, comprising: A data acquisition module is used to use NLP technology to deeply mine the vehicle-mounted chip industry knowledge base, extract high-frequency features related to vehicle-mounted chips to form a local professional feature library, and collect vehicle-mounted chip industry chain information according to the local professional feature library to form an external database. The classification and grading module is used for semantic extraction of data in the external database based on NLP technology; a multi-dimensional feature model is constructed based on the extracted semantic information by using a deep learning algorithm; the semantic information in the multi-dimensional feature model is classified and graded; a recommendation system is created based on the multi-dimensional feature model that has been classified and graded by using a KNN algorithm; and the recommendation system is used for classifying and grading the collected data about the vehicle chip industry chain. The visual output module is used for visualizing the classified and graded data output by the recommendation system.
[0007] Further, the local professional feature library is dynamically updated, and the updating method comprises: When the occurrence frequency of the newly added high-frequency feature is greater than or equal to the set frequency, the cosine similarity between the newly added high-frequency feature and each high-frequency feature in the local professional feature library is calculated. If the cosine similarity between the newly added high-frequency feature and a high-frequency feature in the local professional feature library is greater than or equal to the first similarity threshold, the newly added high-frequency feature is directly included in the local professional feature library. If the cosine similarity between the newly added high-frequency feature and all high-frequency features in the local professional feature library is less than the second similarity threshold, the newly added high-frequency feature is stored in the “observation pool”. If the cumulative occurrence frequency of the newly added high-frequency feature is greater than or equal to the set number within a set time period, the similarity between the newly added high-frequency feature and each high-frequency feature in the local professional feature library is recalculated, and the above process is repeated, otherwise the newly added high-frequency feature is removed from the “observation pool”.
[0008] Further, the local professional feature library is dynamically updated, and the updating method comprises: If the cosine similarity between the newly added high-frequency feature and all high-frequency features in the local professional feature library is less than the first similarity threshold, and the cosine similarity between the newly added high-frequency feature and a high-frequency feature in the local professional feature library is between the first similarity threshold and the second similarity threshold, an artificial review process is triggered, and the newly added high-frequency feature is included in the local professional feature library after the review is passed.
[0009] Further, the external database is subjected to data cleaning, and one of the cleaning methods comprises: The long text is divided into several text blocks by using a double-block attention technology; The query vector and the key vector of each text block are determined, the attention weight is obtained according to the similarity between the query vector and the key vector, and the attention weight is normalized by using a Softmax function; The text block is filtered according to the attention weight, the words with a weight lower than a set value are deleted, and the cleaned data is obtained.
[0010] The second cleaning method comprises: The SimHash algorithm is used to calculate the text fingerprints of the text data in the external database, and when the fingerprint similarity of two text data is greater than or equal to a set value, the two text data are merged.
[0011] Further, the data cleaned external database is further classified, and the primary classification method comprises: Based on the key features related to the vehicle chip industry chain, the K nearest neighbor algorithm is used to calculate the close matching distance of each data to be classified in the data cleaned external database and the key features; and the data in the data cleaned external database is divided into multiple categories according to the close matching distance.
[0012] Further, the method for constructing a multi-dimensional feature model comprises: The data is input into a model based on a deep learning algorithm according to technical parameter dimensions, market dynamic dimensions, application dimensions, and supply chain relationship dimensions; The model extracts combined features by training and learning the potential relationship of the above-mentioned dimensions, and forms a multi-dimensional feature model.
[0013] Further, the method for classifying and grading the semantic information in the multi-dimensional feature model comprises: The NLP technology is used to perform sentiment analysis on each semantic information in the multi-dimensional feature model to obtain a sentiment analysis result; and the semantic information is labeled as positive and negative according to the sentiment analysis result; The cosine similarity of the semantic information and a set core theme in the local professional feature library is calculated, and when the similarity is greater than a set value, the semantic information is labeled as directly related, otherwise it is labeled as indirectly related; If the event described by the semantic information will cause a single vehicle type to stop production risk ≥ a set number of days or affect multiple mainstream vehicle enterprises, the semantic information is labeled as high-risk impact, otherwise it is labeled as general impact; From the application point of view, the semantic information is labeled as self-defined, followed and externally sampled. Self-defined means that the classified and labeled data corresponding to the chip are customized by the enterprise for its specific needs; followed means that the enterprise using the classified and labeled data corresponding to the chip continues to use the existing mature chip solution without major changes; externally sampled means that the classified and labeled data corresponding to the chip are purchased from external suppliers; A vehicle chip industry chain monitoring method based on NLP technology is used to achieve the second purpose of the application, comprising: The data acquisition module is used to use NLP technology to deeply mine the vehicle chip industry knowledge base, extract high-frequency features related to vehicle chips, and construct a local professional feature library; and vehicle chip industry chain information is collected according to the local professional feature library to form an external database; The classification grading module is used for semantic extraction of data in an external database based on NLP technology; a multi-dimensional feature model is constructed based on the extracted semantic information by using a deep learning algorithm; information classification and hierarchical labeling are performed on the semantic information in the multi-dimensional feature model; a recommendation system is created based on the multi-dimensional feature model that has been subjected to information classification and hierarchical labeling by using a KNN algorithm; and the recommendation system is used for classifying and grading the collected data about the vehicle chip industry chain. The visual output module is used for visualizing the data classified and graded by the recommendation system.
[0014] A non-transitory computer-readable storage medium for achieving the third purpose of the present application has a computer program stored thereon, and the computer program is executed by a processor to implement the steps of the vehicle chip industry chain monitoring method based on NLP technology.
[0015] A computer program product for achieving the fourth purpose of the present application includes a computer program / instruction, and the computer program / instruction is executed by a processor to implement the steps of the vehicle chip industry chain monitoring method based on NLP technology.
[0016] The beneficial effects of the present application include: 1. By establishing a dynamic feature library of multi-channel professional users, combining an internal professional knowledge base and RPA technology, automatic collection and analysis of high-frequency features related to vehicle chips are achieved, avoiding the time-consuming and error-prone problems of traditional manual analysis, and improving the analysis efficiency; 2. By using KNN algorithm and K nearest neighbor algorithm, the external information is analyzed by key feature matching calculation, and the precise classification and analysis of complex industry chain data are achieved, overcoming the problem of difficult data integration in the prior art; 3. By using the sentiment analysis and entity recognition capabilities in NLP technology, long text is simplified through attention mechanism, data is cleaned and classified, and the problem of random and personal tendency of analysis results is effectively avoided, improving the objectivity of the analysis results; 4. By using the visual information summary display technology, combined with BI tools, real-time pushing and intuitive display of information are achieved, meeting the needs of enterprise management in various fields for precise difference analysis and positioning of similar information from different angles in a short period of time. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is a flowchart of an embodiment of the method of the present application. DETAILED DESCRIPTION
[0018] The following specific embodiments are used to explain the technical solutions of the present application, so that those skilled in the art can understand the present application. The protection scope of the present application is not limited to the following specific embodiments. The technical solutions of the present application are also included in the protection scope of the present application which are different from the following specific embodiments made by those skilled in the art.
[0019] A vehicle-mounted chip industry chain monitoring method based on NLP technology 1. Data acquisition and preprocessing 1.1. Constructing a local professional feature library Using the word segmentation technology of NLP (Natural Language Processing), the network professional feature library and the internal professional knowledge base are deeply mined to extract high-frequency features related to vehicle-mounted chips. For example, in the network professional feature library, through the word segmentation processing of various chip technology forums, industry reports and other text data, high-frequency features related to vehicle-mounted chips such as "7nm process" and "vehicle-grade MCU performance parameters" are extracted; and a local professional feature library is constructed according to the high-frequency features. The feature library will serve as the basis for subsequent data collection, classification and labeling reference, and the core training data of the recommendation system.
[0020] 1.2. Updating the local professional feature library Set a daily update mechanism: when the frequency of a new high-frequency feature (such as "XX chip is subject to export control") is greater than or equal to a set frequency (such as 5 times / day), the new high-frequency feature is automatically supplemented to the local professional feature library, and the cosine similarity is used to calculate the correlation distance (range 0-1, the closer to 1, the higher the correlation) between the new keyword and each high-frequency feature in the local professional feature library. The method for updating the local professional feature library using the cosine similarity includes: If the similarity between the new high-frequency feature and a certain high-frequency feature in the local professional feature library is greater than or equal to a first similarity threshold, the new high-frequency feature is directly included in the local professional feature library. If the cosine similarity between the new high-frequency feature and all high-frequency features in the local professional feature library is less than the first similarity threshold, and the similarity between the new high-frequency feature and a certain high-frequency feature in the local professional feature library is between the first similarity threshold and a second similarity threshold, a review process (such as a model review process, the model is trained based on the chip field to confirm the relevance of each high-frequency feature to the chip industry chain) is triggered, and the new high-frequency feature is included in the local professional feature library after the review is passed. When the similarity between the newly added high-frequency feature and all high-frequency features in the local professional feature library is less than the second similarity threshold, the newly added high-frequency feature is stored in the "to-be-observed pool". If the frequency accumulation is greater than or equal to the set number (such as 30 times) within a set time period (such as 30 days), the similarity is automatically recalculated, and the above process of calculating the cosine similarity and updating the local professional feature library is triggered. Otherwise, the newly added high-frequency feature is removed from the "to-be-observed pool".
[0021] Through the above rules, the dynamic feature library can automatically absorb high-value new features, and avoid irrelevant information redundancy through manual intervention, ensuring the strong relevance and timeliness of the features in the library to the vehicle chip industry chain.
[0022] The first similarity threshold is greater than the second similarity threshold; for example, the first similarity threshold is 0.7, and the second similarity threshold is 0.4.
[0023] 1.3, Multi-source data collection According to the local professional feature library, RPA digital robots (such as UiPath tools) are used to automatically collect chip industry chain information texts related to high-frequency features stored in the local professional feature library from multiple data sources; the chip industry chain information texts constitute an external database.
[0024] The data sources include but are not limited to chip supplier websites, industry information websites, professional databases, etc. For example, for chip supplier websites, RPA technology is used to simulate manual operations, automatically log in to supplier websites, and grab chip product parameters, price information, inventory status, etc. For industry information websites, regularly collect the latest chip technology trends, market trend analysis, etc. For example: Obtain the chip configuration of the vehicle model on the vertical platform; query the corresponding supplier business information of the chip through the relevant website; obtain the chip production capacity announcement through the chip manufacturer's website; obtain the chip production capacity report and price index through the industry database; obtain the "unreliable entity list" or "entity list" through the policy channel, etc.
[0025] 1.4, Data cleaning De-duplication and merging SimHash algorithm is used to calculate the text fingerprint of the external database composed of chip industry chain information texts. When the similarity between two text fingerprints is greater than or equal to 90%, it is determined as duplicate data and merged.
[0026] Simplify long text through NLP attention mechanism, delete duplicate or meaningless words. Including: 1. Long text chunking: Using DCA (Dual Chunk Attention) technology, long text is divided into several smaller "chunks". For example, a long report on the chip industry is divided into paragraphs or fixed number of words, each paragraph or fixed number of words is a small block, which is convenient for subsequent processing of each small block, reducing the complexity of calculation.
[0027] 2. Calculate attention weight: For each small block, determine the query (Query), key (Key) and value (Value) vector. For example, when processing text blocks related to chip technology parameters, the current focus on technology parameter description is used as the query vector, the representation of other words in the text block is used as the key vector, and the specific semantic content of the corresponding word is used as the value vector. By calculating the similarity between the query vector and the key vector (such as using dot product operation), the attention weight distribution is obtained. For example, calculate the dot product of the "chip process" query vector and the key vector of other words in the text block to measure their relevance. Use the Softmax function to normalize the attention weight, so that the weight sum is 1, to determine the relative importance of each word in the current calculation.
[0028] 3. Simplify long text and remove duplicates: According to the calculated attention weight, filter the words in each small block. Words with a weight less than a set threshold are likely to be repeated or meaningless words that contribute less to understanding key information, which will be deleted; the set threshold can be dynamically adjusted according to the characteristics of the chip industry chain text, for example, initially set to 5%, determine repeated or meaningless words that contribute less to understanding key information. For example, adverbs, redundant adjectives unrelated to the core content of the chip, etc. Organize the remaining words in the text block and merge expressions with the same meaning to form a simplified text block. For example, "semiconductor chip" and "chip (semiconductor material)" are unified as "semiconductor chip".
[0029] 2 Data classification Based on the key features related to the vehicle-mounted chip industry chain, through the KNN algorithm, the similarity matching distance of each data to be classified in the external database after data cleaning is calculated, and the distance is classified into 5 levels, 5 is the highest direct correlation, and 1 is weak correlation. The technical effects include: because the number of samples obtained is large and miscellaneous, data samples need to be cleaned and classified for subsequent steps of using NLP for secondary deep interpretation, i.e. secondary classification and grading. Through classification to determine the importance level, through importance to push to the engineer for processing.
[0030] 3 Secondary classification and grading of information based on NLP technology This step aims to conduct further semantic analysis and secondary classification of data classified by KNN algorithm through natural language processing (NLP) technology, providing more accurate data support for subsequent decision-making. The specific operation and example are as follows: a) Deep semantic information extraction based on NLP Use sentiment analysis and semantic recognition, part-of-speech tagging and other technologies in NLP to clean up information data and extract its deep semantic information.
[0031] Sentiment analysis: used to judge the emotional tendency expressed in the text, such as positive, negative or neutral. Algorithms such as Bag of Words, TF-IDF (Term Frequency-Inverse Document Frequency) and Support Vector Machine (SVM) are used. For example, for the text "a well-known chip manufacturer announced to expand the production capacity of automotive-grade chips to meet market demand", through sentiment analysis algorithm, it identifies positive words such as "expand production capacity" and "meet demand", and determines that the text sentiment is positive.
[0032] Semantic recognition: used to understand the meaning of words, phrases and sentences in the text, and identify the topics, fields and other aspects involved in the text. For example, for the text "the newly released automotive-grade AI chip performs well in the autonomous driving scenario, with significant algorithmic power improvement", the semantic recognition algorithm can identify that the text describes around "automotive-grade AI chip", "autonomous driving" and "algorithmic power improvement", thus clarifying its relevance to the automotive chip industry chain.
[0033] Part-of-speech tagging: used to tag the part of speech of each word in the text, such as noun, verb, adjective, etc. For example, in the sentence "chip supply is tight, affecting the progress of automobile production", through part-of-speech tagging algorithm, "chip" is tagged as a noun, "supply" as a verb, and "tight" as an adjective. This helps to further understand the grammatical structure and semantic relationship of the text, providing a basis for subsequent analysis.
[0034] b) Building multi-dimensional feature model based on deep learning algorithm Use deep learning models such as Convolutional Neural Network (CNN) or Recurrent Neural Network (RNN) to analyze data from multiple angles and build multi-dimensional feature models of data. Deep learning algorithms can automatically learn complex patterns and features from large amounts of data. For example, in the technology dimension, the chip's process technology (such as 7nm, 5nm, etc.), core frequency, transistor number, etc. are used as input features of the model, and deep learning algorithms are used to mine the potential relationships and rules between these features, thus quantitatively evaluating the technology level of the chip. By encoding the data according to different feature dimensions and inputting it into the deep learning model, the model can automatically extract and combine these features to form a multi-dimensional feature representation of the data.
[0035] The feature dimensions include: Technical parameter dimension: Analyze the performance parameters, process technology, and other characteristics of the chip, and extract relevant technical parameter features such as process technology (7nm, 14nm, etc.), computing power (number of operations per second), power consumption (watts), etc. For example, for the data describing "a new car-grade chip uses 7nm process technology, computing power reaches 100 billion operations per second, and power consumption is only 5 watts", the algorithm can extract these key technical parameters as feature dimensions.
[0036] Market dynamics dimension: Includes information such as chip price fluctuations, market demand changes, supplier capacity adjustments, etc. For example, "the price of car-grade MCU chips has risen by 10% in recent times, and market demand remains strong", the model will take the price increase rate and demand strength as features of the market dynamics dimension.
[0037] Application dimension: Study the application scenarios and demand characteristics of chips in different fields (such as automobiles, communications, artificial intelligence, etc.). Through learning and modeling of these dimensions, a multi-dimensional feature model that can comprehensively reflect the data features is constructed Supply chain relationship dimension: Involves the cooperation between chip suppliers and automobile manufacturers, supply stability, etc. For example, "a chip supplier has signed long-term cooperation agreements with multiple mainstream automobile manufacturers, and the supply stability is high", at this time the algorithm identifies the cooperation between the supplier and the manufacturer and the supply stability as features of this dimension.
[0038] c) Information classification and grading labeling of external databases according to local professional feature library The data extracted through multi-dimensional feature extraction is labeled by the local professional feature library. The local professional feature library takes the local chip industry features as the reference standard, and this feature library is built based on a large number of labeled and practically verified data, covering typical features and classification standards in all aspects of the chip industry. The multi-dimensional feature model built through deep learning is compared and analyzed with the local feature model, and information is classified and labeled according to multiple dimensions such as positive / negative, direct / indirect correlation, high / low risk, self-defined, followed, and externally sourced.
[0039] Information classification includes: Positive / Negative: Based on sentiment analysis results, information is classified as positive or negative. For example, "a well-known chip manufacturer announced the expansion of car-grade chip production capacity to meet market demand" is positive information, while "a chip manufacturing plant in a certain region was shut down due to a fire, affecting chip supply" is negative information.
[0040] Directly related / indirectly related: Determine the degree of direct or indirect association between information and the vehicle chip industry chain. By calculating the cosine similarity between the text and the core topics in the feature library (such as "chip supply for vehicle" and "vehicle production"), a threshold of ≥0.6 is considered directly related, avoiding manual judgment bias. For example, "a certain automobile manufacturer is forced to reduce production due to chip shortage" directly involves the impact of chip supply on automobile production, which is directly related information; while "the global price of semiconductor raw materials has risen, which may affect chip manufacturing costs" is indirectly related to the chip industry chain, as the impact is more indirect.
[0041] Classification includes: High risk / general impact: Assess the impact of information on the vehicle chip industry chain. For example, "a key chip supplier may be forced to stop supplying for more than a month due to force majeure factors", which has a significant impact on automobile production, and is considered a high-risk impact; while "a certain niche chip price fluctuates slightly, with little impact on the mainstream market", which is considered a general impact. In some embodiments, the high-risk judgment standard is: causing a single vehicle model to stop production for ≥7 days or affecting ≥3 major car manufacturers.
[0042] Application angle classification includes: Customized: Indicates that the chip corresponding to the classification annotation data is customized by the enterprise for its specific needs, such as "a car manufacturer customizes a special chip to achieve unique autonomous driving capabilities".
[0043] Follow-up: Indicates that the enterprise using the chip corresponding to the classification annotation data continues to use the existing mature chip solution without major changes, such as "a car manufacturer continues to use the chip architecture of the previous generation in the new vehicle model".
[0044] External procurement: Indicates that the chip corresponding to the classification annotation data is purchased by the enterprise from external suppliers, such as "a car manufacturer signs a procurement contract with a chip supplier to purchase a large number of automotive-grade chips".
[0045] For example, for an information about a serious quality problem of a certain chip, after comparing with the local professional feature library, it is judged that the sentiment is negative, which belongs to direct correlation (because it directly affects the use of the chip), and the classification is high risk (quality problems may lead to product recall, damage to the reputation of the enterprise, etc. serious consequences), the application angle needs to be judged according to the specific situation whether it involves the quality problem of the customized chip, etc., so as to complete the classification and classification annotation of the information in each dimension.
[0046] d) Create a recommendation system based on K-nearest neighbor algorithm and complete 5-level hierarchical classification annotation On the basis of completing the above classification grading annotation, the K nearest neighbor (KNN) algorithm is used to calculate the similarity distance of the annotated information and the grading standard in the local professional feature library, and a recommendation system for accurate classification and annotation of chip industry chain related information is created according to the similarity distance between the data. The recommendation system classifies the data in 5 levels, completes the final data annotation, and provides support for industry chain analysis and decision-making to assist enterprises in better managing the chip industry chain.
[0047] The K nearest neighbor algorithm calculates the distance (such as Euclidean distance) between the data to be classified and the known category data based on the feature vector of the data. The K nearest neighbor data is selected, and the category of the data to be classified is determined by voting and other methods according to the category distribution of the K nearest neighbor data. For example, in the "directly related" category, the most suitable information is classified as level 5 according to the matching degree of the information and the core features, and the information with weak correlation is classified as level 4 to level 1, and finally the data is finely annotated, making the classification result more hierarchical and distinguishable, and providing more accurate basis for subsequent visualization and decision analysis.
[0048] The recommendation system is used to assist accurate information classification and annotation. During the 5-level classification and annotation of information, the recommendation system can provide quantitative and objective basis for the classification of information according to the similarity calculation results of new information and standard samples in the local professional feature library. When a new chip industry chain related information is obtained, the recommendation system quickly calculates the similarity of the information with the standard samples in the library. If the similarity with the 5-level standard sample (which usually represents information that is highly directly related to the vehicle chip industry chain and has a significant impact) exceeds a certain threshold, such as 90%, the new information is recommended as level 5. In this way, subjective bias in the manual classification and annotation process is reduced, and the accuracy and consistency of information classification are improved. For example, for the information "a global leading chip manufacturer announces the termination of a core vehicle chip R&D project", the recommendation system compares it with the standard samples and judges that it has a high similarity with the 5-level standard samples about chip supply interruption and major R&D changes, so it is recommended as a 5-level information, indicating that the information has a significant impact on the vehicle chip industry chain.
[0049] In some embodiments, the 5-level classification is as follows: Level 5: Highly matched with the local professional feature library, and belongs to directly related, positive, and high-risk impact information. For example, "a leading chip manufacturer successfully develops a new generation of high-performance vehicle AI chip, and has reached cooperation intentions with multiple leading vehicle enterprises, which will greatly improve the performance of autonomous driving", such information has a significant positive impact on the vehicle chip industry chain and has a high direct correlation, and is labeled as level 5.
[0050] Level 4: High matching degree with local professional feature library, directly related, positive, general impact or indirectly related, positive, high risk impact information. For example, "a mainstream chip supplier announced that it will increase the yield rate of automotive-grade chips, which will help reduce costs." Although it does not reach the impact of level 5 information, it still has a positive effect on the industry chain and can be marked as level 4.
[0051] Level 3: Medium matching degree, information with relatively moderate impact in various category combinations. For example, "a certain region introduces policies to encourage the development of the chip industry, which may attract relevant enterprises to settle down," which is indirectly related, positive, and general impact information, marked as level 3.
[0052] Level 2: Low matching degree, indirectly related, negative, and general impact information. For example, "a small chip design company is experiencing financial difficulties, which may affect the progress of some of its R&D projects," which has a small and indirect impact on the overall industry chain and is marked as level 2.
[0053] Level 1: Low matching degree with local professional feature library, indirectly related, negative, and weak impact information. For example, "a foreign consumer electronics chip manufacturer releases a new mobile phone chip, which has little relevance to the automotive chip market," which has a very small impact on the automotive chip industry chain and is marked as level 1.
[0054] Through the application of the above K-nearest neighbor algorithm, fine 5-level classification labeling of data is achieved, providing more detailed and accurate data support for subsequent data analysis and decision-making.
[0055] In some embodiments, the generation method of the recommendation system is: Collect various types of information from the local professional feature library with information classification and grading labeling and the external database processed by NLP technology, including but not limited to chip technical parameters (such as process technology, computing power, power consumption, etc.), market dynamics (price fluctuations, supply and demand relationships, etc.), supplier information (production capacity, supply stability, etc.), and cooperation with automobile manufacturers, etc. Perform de-duplication, cleaning, and other preprocessing operations on these data to ensure data accuracy and consistency. For example, check and integrate parameter information about a specific chip model obtained from different sources, and remove duplicate or incorrect data records.
[0056] Extract features that are important for recommendations. For example, by analyzing historical data, extract features such as chip usage frequency and failure rate in different application scenarios. For supplier information, extract features such as supply timeliness and product qualification rate. For example, if a certain chip is frequently used in a large number of autonomous driving-related projects, this high usage frequency feature can be used as an important basis for its strong applicability in the autonomous driving field in subsequent recommendations.
[0057] K-Nearest Neighbors algorithm as the basis for building a recommendation model. First, determine the value of K in the model, the selection of K value is usually determined by cross-validation and other methods to balance the accuracy and computational efficiency of the model. For example, by comparing the classification accuracy and recall rate of known data under different K values through multiple experiments, select the K value that makes the comprehensive performance of the two optimal. The data processed by feature engineering is used as the input of the model, and the model calculates the similarity between the data. In calculating the similarity, appropriate distance measurement methods such as Euclidean distance, cosine similarity, etc. are used. For a new information about a chip, the model calculates the distance between it and the existing standard sample information in the local professional feature library, and finds the K nearest samples. For example, for a piece of information "a new 28nm automotive-grade chip with high computing power and low power consumption characteristics", the model will calculate the distance between this information and other chip information in the library in the process, computing power, power consumption and other feature dimensions, and find the K most similar chip information records.
[0058] A large amount of historical data is used to train the recommended model. In the training process, the parameters of the model are constantly adjusted to improve the accuracy of the model in classifying and recommending information. At the same time, new data is used regularly to verify and optimize the model, ensuring that the model can adapt to changing information data. For example, collect new chip market dynamic data every month, retrain and evaluate the model, and if the model's accuracy in recommending some types of information decreases, analyze the reasons and adjust the model parameters or feature selection accordingly.
[0059] 4 Visual output Combine visualization information aggregation display technology and BI (Business Intelligence) tools to visually display the data about automotive chips processed by the recommendation system. Set three warning thresholds: price fluctuation ≥10% triggers yellow warning, supplier capacity reduction ≥30% triggers orange warning, and core chip supply risk triggers red warning. For example, use the Echarts chart library to generate chip price trend line chart, different supplier chip supply share pie chart, chip technology development trend radar chart, etc. Through these visual charts, enterprise managers can clearly and intuitively understand the key information of the chip industry chain. At the same time, set up warning function, when the chip price fluctuation exceeds the preset threshold, the supplier inventory is lower than the safety level, etc. The system automatically sends a warning signal to remind relevant personnel to pay attention and handle it in time.
[0060] Establish an information real-time pushing mechanism, according to the user demand and setting, important chip industry chain information is pushed to the relevant personnel in time. For example, for the personnel responsible for purchasing, when the price of a certain chip concerned decreases to a certain amplitude, the supplier has a new preferential activity, the system pushes the relevant information to the purchasing personnel through the short message, the email or the enterprise internal instant communication tool, so that it adjusts the purchasing strategy in time; for the technical research and development personnel, when there is new chip technology breakthrough, the information such as industry standard update, timely push, help it understand the latest industry trend, provide reference for the research and development work.
[0061] Push three kinds of charts in real time through Power BI: Supplier risk heat map (horizontal axis: region, vertical axis: capacity proportion, color: risk level); Chip parameter matching matrix (row: vehicle demand, column: chip model, cell: matching degree %); Price fluctuation line chart (time axis: recent 30 days, curve: different chip price index).
[0062] In some embodiments, it also includes: difference list generation; according to the "existing data" and "newly added data", automatically output the difference items; such as "the price of a certain chip model increases by 5% compared with yesterday, the correlation degree is 5 levels, it is suggested to start the alternative supplier, and trigger the early warning, and the high-risk items are pushed to the purchasing director in real time.
[0063] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the application.
[0064] The embodiment of the application also provides a non-transitory computer readable storage medium, the computer readable storage medium stores a computer program, the computer program includes program instructions, the program instructions are executed by the processor to realize each step of the method of the application, which will not be repeated here.
[0065] The computer readable storage medium can be the data transmission device or the internal storage unit of the computer device provided in any of the preceding embodiments, for example, the hard disk or the memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card and the like.
[0066] Further, the computer readable storage medium can also include both the internal storage unit of the computer device and the external storage device. The computer readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer readable storage medium can also be used to temporarily store data to be output or data that has been output.
[0067] Those skilled in the art will appreciate that embodiments of the present application can be supplied as a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code thereon.
[0068] The present application is described in reference to the flowchart and / or block diagrams of the method, apparatus (system) and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagrams, as well as a combination of flows and / or blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce the functions specified in the flowchart and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for functionally implementing the flow or flows and / or blocks in one or more blocks.
[0069] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including instruction means, which functionally implement the flowchart and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for functionally implementing the flow or flows and / or blocks in one or more blocks.
[0070] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operational steps are performed on the computer or other programmable data processing apparatus to produce a computer implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide the functions specified in the flowchart and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for functionally implementing the flow or flows and / or blocks in one or more blocks.
[0071] The contents not described in detail in the specification belong to the prior art known to those skilled in the art.
Claims
1. A vehicle-mounted chip supply chain monitoring system based on NLP technology, characterized in that, include: Data acquisition module: Used to perform in-depth mining of the automotive chip industry knowledge base using NLP technology, and extract high-frequency features related to automotive chips to form a local professional feature library; Information on the automotive chip industry chain is collected based on the local professional feature library to form an external database; The classification and grading module is used to perform semantic extraction based on NLP technology from data in external databases; construct a multi-dimensional feature model based on the extracted semantic information using deep learning algorithms; classify and grade the semantic information in the multi-dimensional feature model; and create a recommendation system based on the multi-dimensional feature model with information classification and grading annotation using the KNN algorithm. The recommendation system is used to classify and grade the collected data on the automotive chip industry chain. Visualization output module: Used to visualize the classified and graded data output by the recommendation system.
2. The vehicle chip supply chain monitoring system based on NLP technology as described in claim 1, characterized in that, It also includes dynamically updating the local professional feature library, and the update method includes: When the frequency of a newly added high-frequency feature is greater than or equal to a set frequency, the cosine similarity between the newly added high-frequency feature and each high-frequency feature in the local professional feature library is calculated. If the cosine similarity between the newly added high-frequency feature and a certain high-frequency feature in the local professional feature library is greater than or equal to the first similarity threshold, then the newly added high-frequency feature will be directly included in the local professional feature library. If the cosine similarity between the newly added high-frequency feature and all high-frequency features in the local professional feature library is less than the second similarity threshold, the newly added high-frequency feature is stored in the observation pool. If the cumulative frequency of the newly added high-frequency feature within the subsequent set time period is greater than or equal to the set number of times, the similarity between the newly added high-frequency feature and each high-frequency feature in the local professional feature library is recalculated and the above process is repeated. Otherwise, it is removed from the observation pool.
3. The vehicle chip supply chain monitoring system based on NLP technology as described in claim 1, characterized in that, It also includes dynamically updating the local professional feature library, and the update method includes: If the cosine similarity between the newly added high-frequency feature and all high-frequency features in the local professional feature library is less than the first similarity threshold; and the cosine similarity between the newly added high-frequency feature and a certain high-frequency feature in the local professional feature library is between the first similarity threshold and the second similarity threshold, then the review process is triggered, and after the review is passed, the newly added high-frequency feature is included in the local professional feature library.
4. The vehicle chip supply chain monitoring system based on NLP technology as described in claim 1, characterized in that, It also includes data cleaning of the external database, and the cleaning methods include: The long text is divided into several text blocks using a two-block attention technique; Determine the query vector and key vector for each text block, and obtain the attention weight based on the similarity between the query vector and the key vector; use the Softmax function to normalize the attention weight; The text blocks are filtered based on attention weights, and words with weights lower than the set value are deleted to obtain cleaned data.
5. The vehicle chip supply chain monitoring system based on NLP technology as described in claim 4, characterized in that, It also includes an initial classification of the cleaned external database, the initial classification method of which includes: Based on key features related to the automotive chip industry chain, the K-nearest neighbor algorithm is used to calculate the proximity matching distance between each piece of data to be classified in the cleaned external database and the key features; the data in the cleaned external database is divided into multiple categories according to the proximity matching distance.
6. The vehicle chip supply chain monitoring system based on NLP technology as described in claim 5, characterized in that, Methods for constructing multi-dimensional feature models include: The data is input into a deep learning algorithm-based model according to the dimensions of technical parameters, market dynamics, application, and supply chain relationships. The model learns the potential relationships between the above dimensions through training and extracts combined features to form a multi-dimensional feature model.
7. The vehicle chip supply chain monitoring system based on NLP technology as described in claim 6, characterized in that, The methods for classifying and hierarchically labeling the semantic information in the multi-dimensional feature model include: NLP technology is used to perform sentiment analysis on each semantic information in the multi-dimensional feature model to obtain sentiment analysis results; based on the sentiment analysis results, the semantic information is labeled as positive or negative. Calculate the cosine similarity between the semantic information and the set core topics in the local professional feature library. If the similarity is greater than the set value, the semantic information is marked as directly related; otherwise, it is marked as indirectly related. If the event described by the semantic information will lead to a production halt risk of a single model ≥ the set number of days or affect multiple mainstream automakers, the semantic information will be marked as a high-risk impact; otherwise, it will be marked as a general impact. From an application perspective, semantic information is labeled as custom, reused, and externally sourced.
8. A method for monitoring the automotive chip supply chain based on NLP technology according to the system described in claim 1, characterized in that, include: We use NLP technology to deeply mine the knowledge base of the automotive chip industry and extract high-frequency features related to automotive chips to form a local professional feature library. Information on the automotive chip industry chain is collected based on the local professional feature library to form an external database; Semantic extraction based on NLP technology is performed on data from an external database; a multi-dimensional feature model is constructed based on the extracted semantic information using a deep learning algorithm; the multi-dimensional feature model is then classified and graded against a local professional feature library; a recommendation system is created based on the classified and graded multi-dimensional feature model using the KNN algorithm; the recommendation system is used to classify and grade the collected data on the automotive chip industry chain. Visualize the classified and graded data output by the recommendation system.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the vehicle chip industry chain monitoring method based on NLP technology as described in claim 8.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the vehicle chip industry chain monitoring method based on NLP technology as described in claim 8.
Citation Information
Cited By
Automobile chip selection method and related device
CN122692621A