Intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph

By using a distributed crawler system and knowledge graph technology, product data is obtained from multiple platforms, and a multi-dimensional knowledge graph is constructed to compare product technical features. This solves the problems of semantic gaps and delayed risk assessment in product selection, and realizes the systematization and automation of intelligent product selection and infringement warning.

CN122347437APending Publication Date: 2026-07-07IMC DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
IMC DIGITAL TECH CO LTD
Filing Date
2026-04-07
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Existing technologies suffer from a semantic gap between consumer and technical language in product selection and patent infringement risk assessment, leading to biased data analysis. The lack of cross-platform correlation and systematic design variants results in delayed risk assessment and reliance on human experience.

Method used

Product data is acquired from multiple platforms through a distributed crawler system. Natural language processing and time series analysis are used to extract product technical features. A multi-dimensional knowledge graph is constructed for comparison, generating a patent infringement risk assessment value. Based on the graph path analysis, a design avoidance solution is generated.

Benefits of technology

It enables cross-platform data synchronization and mining, improves data reliability and the comprehensiveness of product selection analysis, quantifies patent infringement risks, generates scientific product design avoidance solutions, and improves product selection efficiency and compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122347437A_ABST
    Figure CN122347437A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent product selection, and particularly relates to an intelligent product selection and infringement early warning method based on cross-platform mining and a knowledge graph, comprising: obtaining product description texts on multiple platforms, user interaction behavior sequences and market dynamic indicators; using natural language processing to analyze the product texts and extract a technical feature vector, combining time series analysis to generate a product heat index, and calculating market saturation based on market dynamics to form a product market potential evaluation indicator; constructing a multi-dimensional product knowledge graph based on the technical feature vector, and selecting candidate products according to the potential indicator; mapping the candidate product technical vector to the knowledge graph to quantitatively generate a patent infringement risk evaluation value; when the risk exceeds a threshold value, generating a design avoidance scheme based on the technical substitution path and design variant scheme in the graph; and outputting intelligent product selection decision support information. The present application improves platform product selection efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent product selection technology, specifically to an intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph. Background Technology

[0002] In the fields of e-commerce and product development, product selection and patent infringement risk assessment have always been significant challenges for enterprises. Existing technologies often focus on analyzing data from a single platform or simple patent comparisons, neglecting a crucial but frequently overlooked issue in the extraction of product technical features: the semantic gap between consumer language and technical language.

[0003] Traditional product selection methods primarily rely on sales data and market research, analyzing historical sales trends to predict future best-selling products. However, this approach has significant limitations: consumers often use colloquial, non-technical descriptive language in e-commerce platform reviews and social media discussions, which differs greatly from the precise technical terminology used in patent documents and technical documentation. This semantic gap leads to serious biases in existing technologies when extracting accurate technical features from user-generated content, thereby affecting the accuracy of subsequent patent infringement risk assessments.

[0004] Existing technologies fail to establish a closed-loop connection from market signals to technological implementation and then to legal protection. When a certain technical feature shows a growth trend on multiple platforms simultaneously, it should trigger an active scan of existing patent layouts along that technical path, rather than simply passively comparing candidate products. Moreover, the generation of design variant solutions often relies on human experience and lacks a systematic derivation mechanism based on technical alternative paths in knowledge graphs.

[0005] Therefore, it is necessary to design an intelligent product selection and infringement warning method based on cross-platform mining and knowledge graphs to solve the problems existing in the current technology. Summary of the Invention

[0006] In view of this, the present invention proposes an intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph, aiming to solve the problems of semantic fragmentation of multi-source data and delayed risk warning.

[0007] This invention proposes an intelligent product selection and infringement warning method based on cross-platform mining and knowledge graphs, including:

[0008] Based on a distributed crawler system and platform API interfaces, product-related data is obtained from several heterogeneous platforms. The product-related data includes product description text, user interaction behavior sequences, and market dynamic indicators.

[0009] Semantic analysis of product description text is performed using natural language processing technology to extract product technical feature vectors; user interaction behavior sequences are processed using time series analysis methods to generate product popularity indicators; market saturation indicators are calculated based on the market dynamic indicators; and the product popularity indicators and market saturation indicators are merged to generate product market potential assessment indicators.

[0010] Based on the product's technical feature vector, a multidimensional product knowledge graph is constructed using a semantic relationship extraction algorithm; candidate products with market growth potential are screened according to the product's market potential assessment indicators.

[0011] The technical feature vectors of the candidate products are mapped to a multi-dimensional product knowledge graph, and a multi-dimensional comparison is performed with the technical nodes in the knowledge graph. Based on the multi-dimensional comparison results, a patent infringement risk assessment value is generated by risk quantification.

[0012] When the patent infringement risk assessment value exceeds the preset risk threshold, a product design avoidance scheme is generated based on the technology alternative paths and design variant schemes in the multidimensional product knowledge graph using the graph path analysis algorithm.

[0013] The candidate products, patent infringement risk assessment values, and product design avoidance schemes are integrated and analyzed to generate intelligent product selection decision support information.

[0014] Furthermore, when performing semantic parsing on the product description text using natural language processing techniques to extract product technical feature vectors, the following is included:

[0015] The product description text is segmented and part-of-speech tagged to identify core product nouns and technical adjectives;

[0016] Based on a technical dictionary and domain terminology database, extract key technologies and design elements of the product;

[0017] The extracted key technologies and design elements of the product are vectorized to generate product technical feature vectors.

[0018] Furthermore, by combining time series analysis methods to process user interaction behavior sequences, a product popularity index is generated, and a market saturation index is calculated based on the aforementioned market dynamics index, including:

[0019] The user interaction behavior sequence is cleaned to remove abnormal clicks and invalid browsing records; the user interaction frequency per unit time is calculated to generate a user interaction frequency sequence; the user interaction frequency sequence is smoothed using a moving average algorithm; the growth rate of the smoothed user interaction frequency sequence is calculated to generate a product popularity index.

[0020] Obtain market supply and demand data for similar products; calculate the ratio of market supply to market demand to generate a market supply-demand ratio index; calculate the current market growth rate based on historical market data; and non-linearly combine the market supply-demand ratio index with the market growth rate to generate a market saturation index.

[0021] Furthermore, when integrating the product popularity index and the market saturation index to generate a product market potential assessment index, the following steps are included:

[0022] The product popularity index and the market saturation index are normalized; weight coefficients for the product popularity index and the market saturation index are set based on industry characteristics; and a weighted average algorithm is used to merge the normalized product popularity index and the market saturation index to generate a product market potential assessment index.

[0023] Furthermore, when constructing a multidimensional product knowledge graph based on the product technology feature vector using a semantic relationship extraction algorithm, the process includes:

[0024] The product technical feature vector is mapped to the node space of the knowledge graph to generate product nodes; the technical similarity between the product nodes and existing technical nodes is calculated; when the technical similarity is greater than a preset similarity threshold, a technical association is established between the product nodes and the technical nodes; the weight value of the technical association is set according to the magnitude of the technical similarity; product nodes with technical association are clustered to form technical theme clusters; a multidimensional product knowledge graph is constructed based on the technical theme clusters; the multidimensional product knowledge graph includes product nodes, technical nodes, design element nodes, and patent nodes.

[0025] Furthermore, when screening candidate products with market growth potential based on the aforementioned product market potential assessment indicators, the following steps are included:

[0026] Set a screening threshold for the product market potential assessment index; compare the product market potential assessment index of all products with the screening threshold; and select products whose product market potential assessment index is greater than the screening threshold as candidate products with market growth potential.

[0027] Furthermore, mapping the technical feature vectors of the candidate products to a multi-dimensional product knowledge graph and performing multi-dimensional comparisons with the technical nodes in the knowledge graph includes:

[0028] Locate the technical node in the multidimensional product knowledge graph that is most similar to the technical feature vector of the candidate product; calculate the technical semantic similarity between the technical feature vector of the candidate product and the located technical node.

[0029] Analyze the degree to which the technical feature vectors of the candidate products cover the technical node features, and generate technical feature coverage.

[0030] Retrieve patent nodes associated with technical nodes and evaluate the matching degree between the technical features of the candidate products and the patent claims;

[0031] Multi-dimensional comparison results are generated based on the technical semantic similarity, the technical feature coverage, and the matching degree.

[0032] Furthermore, based on the aforementioned multi-dimensional comparison results, when generating a patent infringement risk assessment value through risk quantification, the following are included:

[0033] The indicators of each dimension in the multi-dimensional comparison results are normalized; risk weight coefficients of each indicator are set based on legal practice; a weighted summation algorithm is used to calculate a preliminary risk score; the preliminary risk score is corrected according to the legal status and geographical scope to generate a patent infringement risk assessment value.

[0034] Furthermore, when generating product design avoidance solutions using graph path analysis algorithms based on the technology alternative paths and design variants in a multi-dimensional product knowledge graph, the following are included:

[0035] Analyze the correlation strength between the technical feature vector of the candidate product and each technical node in the multidimensional product knowledge graph; identify technical nodes whose correlation strength with the technical feature vector of the candidate product is greater than a preset correlation strength threshold; calculate the contribution to the patent infringement risk assessment value; and filter technical nodes whose contribution is greater than a preset contribution threshold based on the contribution to form a set of technical nodes to be avoided.

[0036] Search for adjacent technology nodes in the multidimensional product knowledge graph that are adjacent to each technology node in the set of technology nodes to be avoided; calculate the technology compatibility index between the adjacent technology nodes and the candidate products; filter technology nodes whose technology compatibility index is greater than a preset compatibility threshold according to the technology compatibility index to form a set of alternative technology nodes;

[0037] The design variant schemes corresponding to each technology node in the set of alternative technology nodes are analyzed; the technical effect retention rate of each design variant scheme is calculated; the implementation difficulty coefficient of each design variant scheme is evaluated; a comprehensive score is given to each design variant scheme based on the technical effect retention rate and the implementation difficulty coefficient; design variant schemes with a comprehensive score higher than a preset score threshold are identified as product design avoidance schemes; the product design avoidance schemes include feasible design avoidance schemes, technical substitution suggestions, and design modification guidance.

[0038] Furthermore, when integrating and analyzing the candidate products, patent infringement risk assessment values, and product design avoidance schemes to generate intelligent product selection decision support information, the following are included:

[0039] The candidate products are categorized and organized according to product type and market potential; the candidate products are labeled with risk levels based on the patent infringement risk assessment value; the product design avoidance schemes are associated with the corresponding candidate products; a comprehensive score of market potential and infringement risk for each candidate product is calculated; the candidate products are sorted according to the comprehensive score to generate a candidate product recommendation list; the candidate product recommendation list, risk labeling information, and product design avoidance schemes are integrated into intelligent product selection decision support information.

[0040] Compared with existing technologies, the advantages of this invention are as follows: By using distributed crawlers and platform API interfaces to acquire product description text, user interaction behavior sequences, and market dynamic indicators in parallel, it can cover product information across multiple platforms and dimensions, achieving cross-platform data synchronous mining, avoiding biases caused by single-platform data, and improving data reliability and the comprehensiveness of product selection analysis. Natural language processing technology is used to perform semantic parsing of product descriptions, combined with time series analysis to generate product popularity indicators, and market saturation indicators are constructed based on market supply and demand and growth rates. Then, through normalization and weighted fusion, product market potential assessment indicators are formed, realizing digital quantitative analysis from text to market, making potential product identification more scientific and forward-looking. Based on technical feature vectors and semantic relationship extraction technology, a multi-dimensional association network between product nodes, technology nodes, design element nodes, and patent nodes is automatically constructed, which can intuitively present technological evolution relationships, design dependencies, and patent protection scope, realizing structured knowledge storage and visual association judgment, facilitating subsequent rapid retrieval and automated analysis. By comparing the semantic similarity, feature coverage, and matching degree with patent claims of the technical vectors of candidate products and technical nodes in the knowledge graph, a multi-dimensional technical comparison result is formed. Risk correction is then applied based on legal factors, enabling the quantification and generation of a more objective, transparent, and accurate patent infringement risk assessment value, reducing the uncertainty of traditional manual analysis. Key technical nodes requiring avoidance are identified, and alternative technical nodes are screened through technical compatibility calculations. Corresponding design variants are analyzed, and a comprehensive score is generated based on the retention rate of technical effects and the difficulty of implementation, thus generating the optimal product design avoidance solution. This solution not only avoids patent risks but also maintains product functionality, improving product compliance and innovation capabilities. The market potential, infringement risk level, and avoidance solutions of candidate products are integrated and sorted according to comprehensive scores, automatically generating a product selection recommendation list and decision information. This automates and intelligently integrates the entire product selection process from data acquisition, risk identification, solution generation to final recommendation, improving the platform's product selection efficiency and accuracy while reducing the cost of manual judgment. Attached Figure Description

[0041] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0042] Figure 1 The flowchart illustrates the intelligent product selection and infringement warning method based on cross-platform mining and knowledge graph provided in this embodiment of the invention. Detailed Implementation

[0043] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0044] Existing technologies generally suffer from drawbacks such as limited data sources, insufficient semantic parsing capabilities, and incomplete analysis chains. On the one hand, product selection prediction mainly relies on sales data from a single platform or simple trend analysis, making it difficult to reflect real market signals across platforms. On the other hand, everyday expressions in user comments and social discussions cannot be accurately mapped to the technical terminology in patent documents; existing methods lack cross-semantic mapping capabilities, leading to biases in technical feature extraction. Existing technologies cannot construct a unified closed loop between market trends, technological implementation, and patent protection, nor can they automatically derive feasible design alternatives based on technological paths. This results in delayed infringement risk assessment and avoidance strategies relying on human experience, lacking systematicity and reliability.

[0045] For example, when reviewing popular products like "folding phone stands," consumers often use everyday descriptions such as "smoother hinges," "more stable angles," and "less prone to wobbling." However, these expressions might actually correspond to technical terms like "damped hinge structure" or "multi-stage locking mechanism." Existing technology struggles to accurately map user language to actual technical structures, making it impossible to determine whether these features fall within the scope of existing patents. Furthermore, if discussions about similar stands surge across multiple platforms, existing technology cannot proactively trigger patent scans for related technical paths. Merchants often only discover they've entered a high-density patent zone after product selection or production, lacking systematic alternative design solutions based on knowledge graphs, thus facing a high risk of infringement.

[0046] If the above problems are not addressed, enterprises will continue to face high risks and costs in product selection and R&D. On the one hand, due to the inability to accurately identify the true technical characteristics behind user language, product technology positioning will be biased, potentially leading to misjudgments of market demand or omissions of key technological trends. On the other hand, the lack of a proactive linkage mechanism between market signals and patent risks means that enterprises often only discover the involvement of core patents after products have entered production or even been launched, thus facing significant losses such as infringement disputes, forced withdrawal from the market, or redesign. Designing avoidance solutions relies on human experience, making it difficult to respond promptly to rapidly changing markets and leaving enterprises in a passive position in competition. If these issues are not resolved, they will seriously affect enterprises' product selection efficiency, compliance capabilities, and innovative competitiveness.

[0047] For this, please refer to Figure 1 As shown, this application proposes an intelligent product selection and infringement warning method based on cross-platform mining and knowledge graphs, including:

[0048] S100: Based on a distributed crawler system and platform API interface, it obtains product-related data from several heterogeneous platforms. The product-related data includes product description text, user interaction behavior sequences, and market dynamic indicators.

[0049] S200: Semantically analyzes product description text using natural language processing technology to extract product technical feature vectors; processes user interaction behavior sequences using time series analysis methods to generate product popularity indicators; calculates market saturation indicators based on market dynamic indicators; and integrates product popularity indicators and market saturation indicators to generate product market potential assessment indicators.

[0050] S300: Based on product technical feature vectors, a multi-dimensional product knowledge graph is constructed using semantic relationship extraction algorithms; candidate products with market growth potential are selected based on product market potential assessment indicators.

[0051] S400: Map the technical feature vectors of candidate products to a multi-dimensional product knowledge graph and compare them with the technical nodes in the knowledge graph in multiple dimensions; based on the multi-dimensional comparison results, generate a patent infringement risk assessment value through risk quantification.

[0052] S500: When the patent infringement risk assessment value exceeds the preset risk threshold, a product design avoidance scheme is generated based on the technology alternative paths and design variant schemes in the multi-dimensional product knowledge graph using the graph path analysis algorithm.

[0053] S600: It integrates and analyzes candidate products, patent infringement risk assessment values, and product design avoidance solutions to generate intelligent product selection decision support information.

[0054] Specifically, product-related data is collected from multiple heterogeneous platforms through a distributed crawler system and platform API interfaces. The distributed crawler system, built on the Scrapy framework and deployed across at least three servers forming a distributed crawling cluster, enables high-concurrency crawling of massive amounts of platform data. Platform APIs include Amazon Product Advertising API, Taobao Open Platform API, Douyin Open Platform API, and Google Trends API, which can respectively obtain structured or semi-structured data such as product titles, product descriptions, user reviews, and search volume trends. This method selects at least two e-commerce platforms, one social media platform, and one search engine as data sources, ensuring that the collected data covers different stages of the product lifecycle. The collected data includes product description text (such as titles, product detail page text, and user reviews), user interaction behavior sequences (such as clickstreams, dwell time, and sharing / forwarding records), and market dynamic indicators (such as search volume changes, supply, and demand). Natural language processing techniques are used to semantically parse the product description text, and the BERT model is used to extract deep semantic features. A BiLSTM-CRF model is then used for term recognition and entity annotation, ultimately generating technical feature vectors that characterize the product's key technical attributes. In the user behavior analysis section, time-series modeling of user interaction behavior sequences is performed based on moving average and exponential smoothing algorithms. Product popularity indicators are formed by calculating the frequency, trend changes, and growth rate of behavior per unit time. In market analysis, dynamic market indicators are constructed based on product search volume, supply and demand of similar products, and the market supply-demand ratio and growth rate are calculated to form a market saturation indicator. After normalizing these two types of indicators, they are weighted and fused according to set weights to obtain a product market potential assessment indicator. Based on this assessment indicator, a screening threshold is set and products are sorted by value to select candidate products with high market growth potential. For the technical risk assessment of candidate products, their technical feature vectors are mapped to a pre-constructed multi-dimensional product knowledge graph. This knowledge graph is constructed using rule-based semantic relationship extraction methods and deep learning relationship extraction models, and includes different types of knowledge entities such as product nodes, technology nodes, design element nodes, and patent nodes. Within the knowledge graph, candidate products are compared with technical nodes across multiple dimensions, including semantic similarity calculation (based on vector cosine similarity or deep semantic matching models), technical feature coverage analysis (determining the coverage of technical features involved in the candidate product within existing technical nodes), and patent claim matching assessment. Based on the multi-dimensional comparison results, a risk quantification model is used for weight integration, combined with patent legal status, geographical scope, and authorization status, to ultimately generate a patent infringement risk assessment value for the candidate product. When the infringement risk of a candidate product exceeds a preset threshold, circumvention design analysis is conducted based on technical alternative paths and design variants within the knowledge graph.The process involves identifying key technological nodes that contribute significantly to high-risk outcomes, then selecting highly compatible alternative paths from adjacent technological nodes. Feasible alternatives are assessed for their retention of technical effectiveness and implementation difficulty, resulting in actionable design mitigation solutions, such as technical structure adjustment suggestions, alternative component solutions, or design modification guidelines. Finally, candidate product information, infringement risk assessment values, and corresponding mitigation solutions are integrated and analyzed to generate intelligent product selection decision support information. This includes a candidate product recommendation list, risk level markings, and available technical mitigation solutions, providing enterprises with data-driven, technically guided, and highly reliable product selection support.

[0055] This invention addresses the shortcomings of existing technologies in technology identification, market potential assessment, patent risk evaluation, and design avoidance by combining cross-platform multi-source data acquisition, natural language processing, and time series analysis; by using a multi-dimensional product knowledge graph for structured technology comparison and patent mapping to achieve quantitative risk assessment; and by automatically generating feasible avoidance strategies based on technology substitution paths and design variants in the knowledge graph.

[0056] The working process and principle of this application are as follows: A distributed crawler cluster based on Scrapy and multi-platform API interfaces are used to continuously collect product-related data from multiple heterogeneous channels such as e-commerce platforms, social media platforms, and search engines. The collected data includes textual information such as product titles, details, and user reviews, as well as user clickstreams, dwell time, and sharing / forwarding behavior sequences. It also includes market dynamic indicators such as product search volume, supply and demand of similar products. All data is cleaned and formatted before entering a unified data center, providing a reliable data foundation for subsequent analysis. Natural language processing models are used to semantically analyze product description text, extracting technical feature vectors that characterize the core functions and structural features of the product. Time series analysis is performed on user interaction data to extract product popularity indicators that reflect changes in market attention trends. Simultaneously, market saturation-related indicators are generated based on market dynamic data. These indicators are normalized and integrated to obtain market potential assessment indicators that characterize the product's market growth capacity and competitive intensity at the current stage, which serve as the basis for selecting candidate products with growth potential. For the selected candidate products, in-depth technical comparisons are conducted based on the previously constructed multi-dimensional product knowledge graph. This knowledge graph comprises numerous product technology entities, functional components, relationships, and patent information. It includes both rule-based semantic relationships and complex technical associations automatically extracted by deep learning models. The technical feature vectors of candidate products are mapped to corresponding nodes in the graph, and comparisons are performed with relevant technical nodes across multiple dimensions, including semantic similarity, technical coverage, and claim matching. Based on the comparison results, and considering factors such as patent legal status, geographical scope, and patent validity, a patent infringement risk assessment value is generated for each candidate product to measure the likelihood and severity of potential patent conflicts. When the risk assessment value of a candidate product exceeds a set risk threshold, the avoidance analysis module is automatically activated. This module searches for alternative paths along technical relationships within the knowledge graph, identifies technical nodes that can achieve similar functions but do not fall within the scope of patent protection, and further analyzes the feasibility, compatibility, and design variation space of these alternatives. Based on these analysis results, product design avoidance solutions are generated, providing R&D personnel with feasible technical replacement suggestions, structural modification directions, or functional refactoring strategies to reduce infringement risks from the source. The system cross-integrates market potential assessments of candidate products, patent infringement risk assessments, and corresponding avoidance strategies to generate intelligent product selection decision support information. Output includes a recommended list of candidate products, potential risk markers, explanations of feasible avoidance strategies, and comprehensive product selection recommendations.

[0057] As a preferred embodiment, the specific implementation of this application is as follows: Taking a "portable waterproof Bluetooth speaker" as a specific application example, the operation process and effect of this method in actual product selection and infringement warning are fully demonstrated. A distributed crawler cluster based on Scrapy (at least three servers) crawls product titles, details pages, and user reviews from the two major e-commerce platforms, Amazon and Taobao, in parallel. At the same time, it calls the Douyin Open Platform to obtain relevant short video comments and forwarding data, and pulls the search popularity time series of this category through Google Trends. The raw data obtained from crawling and API is deduplicated, cleaned, and time-aligned before entering the text and time series analysis modules. The text part uses BERT for deep semantic encoding and combines it with BiLSTM-CRF for entity and technical term extraction. Technical elements such as "waterproof rating", "passive noise reduction cavity", "charging interface sealing design", and "detachable hook" are identified from the comments and quantified. The behavioral sequence is denoised using moving average and exponential smoothing, and the trend changes of click volume, dwell time, and sharing are extracted to generate a popularity curve. The search volume, supply and demand of similar products obtained from the market side are used to assess the current market saturation. After integrating popularity and saturation indicators, the Bluetooth speaker was determined to have medium-to-high market potential and was selected as a candidate product. Next, the technical feature vector of this candidate product was mapped onto a constructed multi-dimensional product knowledge graph. The graph was used to locate the most similar technical nodes (e.g., "IPX7 waterproof structure" and "plug-in Type-C sealing design") and retrieve relevant patent nodes for semantic similarity comparison, technology coverage analysis, and claim matching evaluation. The comparison results showed that the candidate product highly overlapped with several valid patents in the "charging interface sealing structure" and "hook fixing mechanism" categories. Based on this, the risk assessment module generated a high patent infringement risk value and triggered the avoidance process. The avoidance module searched the graph for alternative technical paths adjacent to the high-risk nodes, identified several alternative solutions, such as changing the existing plug-in seal to a magnetic seal to avoid coverage by existing patent claims, or replacing the detachable hook with an integrated folding snap-on structure, and provided specific implementation suggestions (material selection, connector dimensions, reinforcement schemes for key stress points, and performance impact estimation). Finally, the dashboard outputs a list of candidate products for the speaker, including its market potential score, patent risk level (high / medium / low), detailed matching descriptions of risky technologies, and several priority-based design avoidance solutions. It also generates reports and actionable design guidelines for collaborative decision-making by product managers, legal professionals, and engineers. User selections and subsequent implementation feedback are fed back as training samples for the model and graph, used to continuously optimize weights, thresholds, and alternative strategies, thereby continuously improving the hit rate and reducing the probability of infringement in subsequent product selections.

[0058] Through the above-mentioned approach, this application establishes a precise mapping from consumer language to technical semantics by combining cross-platform multi-source data collection with natural language processing, bridging the semantic gap between everyday descriptions and patent terminology. By integrating time series analysis and market dynamics, it constructs market potential assessment indicators, enabling early and scientific screening of potential best-selling products. Based on a multi-dimensional product knowledge graph, it conducts structured technical comparisons and patent matching, quantifying and promptly warning of infringement risks before product initiation. When risks exceed thresholds, it automatically mines alternative paths along the graph and generates feasible design avoidance solutions, reducing legal and redesign costs caused by infringement. Finally, it integrates market potential, risk assessment, and avoidance solutions to form an operational intelligent product selection decision support system.

[0059] This application further proposes methods for semantic parsing of product description text using natural language processing techniques to extract product technical feature vectors, including:

[0060] Perform word segmentation and part-of-speech tagging on product description text to identify core product nouns and technical adjectives;

[0061] Based on a technical dictionary and domain terminology database, extract key technologies and design elements of the product;

[0062] The extracted key technologies and design elements of the product are vectorized to generate product technical feature vectors.

[0063] Specifically, semantic parsing of product description text to extract product technical feature vectors requires word segmentation and part-of-speech tagging of the original text. The Jieba word segmentation tool is used to perform fine-grained word segmentation on product titles, details page content, and user reviews, with a custom dictionary configured to ensure accurate identification of industry-specific terms, brand names, and technical terms. The Hidden Markov Model (HMM) algorithm is then used to perform part-of-speech tagging on the segmentation results, thereby identifying key semantic units in the text. This step accurately locates core product nouns and technical adjectives. Core nouns include product names, product component names, and function names, while technical adjectives include words describing product performance, structure, or functional characteristics, providing foundational information for subsequent feature extraction. Key technologies and design elements are extracted based on a pre-built technical dictionary and domain terminology database. The technical dictionary includes IEEE standard terminology and industry-wide general technical vocabulary, covering key technical information such as functional technologies, structural technologies, and material technologies. The domain terminology database covers industry-specific terms for electronic products, home furnishings, and clothing and footwear, used to extract product appearance design elements, structural design elements, and interaction design elements. By matching and semantically expanding technical dictionaries and domain terminology databases, core technical points and design features of products can be extracted from text, achieving comprehensive coverage of technical and design information. The extracted key technologies and design elements are vectorized to generate a complete product technical feature vector. Vectorization employs the Word2Vec model, with a vector dimension set to 300 dimensions to fully capture semantic relationships and contextual information. The generated product technical feature vector consists of two sub-vectors: a technical dimension feature sub-vector, used to express the semantic information of product functional technology, structural technology, and material technology; and a design dimension feature sub-vector, used to express the semantic information of product appearance, structure, and interaction design features.

[0064] Through the aforementioned technical solution, this application achieves efficient transformation from unstructured text to quantifiable technical features by systematically semantically parsing product description text. By using word segmentation and part-of-speech tagging, it accurately identifies core product nouns and technical adjectives, resolving the semantic gap between colloquial language and professional technical terminology. Based on a technical dictionary and domain terminology database, it extracts key product technologies and design elements, achieving comprehensive coverage of functional technologies, structural technologies, material technologies, and appearance, structure, and interaction design features. The extracted technical and design information is vectorized to generate feature sub-vectors for both technical and design dimensions, ensuring that the product's technical feature vectors express both technical and design semantic information.

[0065] This application further proposes combining time series analysis methods to process user interaction behavior sequences, generating product popularity indicators, and calculating market saturation indicators based on market dynamic indicators, including:

[0066] The user interaction behavior sequence is cleaned to remove abnormal clicks and invalid browsing records; the user interaction frequency per unit time is calculated to generate a user interaction frequency sequence; the user interaction frequency sequence is smoothed using a moving average algorithm; the growth rate of the smoothed user interaction frequency sequence is calculated to generate a product popularity index.

[0067] Obtain market supply and demand data for similar products; calculate the ratio of market supply to market demand to generate a market supply-demand ratio index; calculate the current market growth rate based on historical market data; and non-linearly combine the market supply-demand ratio index with the market growth rate to generate a market saturation index.

[0068] Specifically, time series analysis is used to process user interaction behavior sequences to generate product popularity metrics, and market saturation metrics are calculated in conjunction with market dynamics indicators. The entire process includes multiple stages such as data cleaning, behavior frequency calculation, smoothing, growth rate calculation, and market supply and demand analysis. The collected user interaction behavior sequences, including clicks, dwell times, shares, and comments, undergo rigorous data cleaning to remove abnormal clicks and invalid browsing records. Abnormal clicks are defined as events with a click time of less than 100 milliseconds, and invalid browsing records are operations with a browsing time of less than 2 seconds, thus ensuring the reliability and authenticity of the base data for calculation. User interaction frequency is statistically analyzed hourly, including the number of clicks, shares, and comments per hour, generating a complete user interaction frequency sequence. To eliminate noise from short-term fluctuations, a moving average algorithm is used to smooth the user interaction frequency sequence, with a sliding window set to 7 days (i.e., 168 hours), allowing the sequence to more accurately reflect the overall trend of user behavior. Based on the smoothed sequence, the growth rate of user interaction frequency is calculated, and the percentage change in popularity is generated by comparing the current period value with the previous period value, thus forming the product popularity metric. Popularity indicators are divided into short-term and long-term indicators. Short-term indicators reflect product popularity fluctuations over the past 7 days, while long-term indicators reflect popularity changes over the past 30 days, comprehensively depicting market attention at different time scales. For market saturation analysis, supply and demand data for similar products are first obtained. Supply includes inventory and new listings, while demand includes product search volume and actual sales. Based on this data, the market supply-demand ratio is calculated—the ratio of market supply to market demand—reflecting the degree of market supply and demand tension. Simultaneously, the current market growth rate is calculated using historical market data from the past 6 months. By comparing the change in the current month's value with the previous month's value, the overall market expansion or contraction trend is assessed. The market supply-demand ratio and market growth rate are non-linearly combined to generate a market saturation indicator. The combination uses a logarithmic function, weighting the natural logarithm of the supply-demand ratio with the market growth rate. The combination coefficients α and β control the influence weights of the supply-demand ratio and growth rate, respectively, with values ​​ranging from 0.3 to 0.7. This method generates a market saturation indicator that reflects market competition intensity and growth trends.

[0069] Through the aforementioned technical solution, this application quantifies and evaluates product market performance by systematically analyzing user interaction behavior sequences and market dynamic indicators. Abnormal clicks and invalid browsing data are cleaned to ensure the reliability of behavioral data. Then, user interaction frequency is calculated hourly and smoothed using a moving average algorithm to generate short-term and long-term product popularity indicators, reflecting changes in market attention at different time scales. Supply and demand data for similar products are obtained, the market supply-demand ratio is calculated, and combined with historical market growth rates, a market saturation indicator is generated through non-linear combination, comprehensively reflecting the intensity of market competition and growth potential.

[0070] This application further proposes that when integrating product popularity indicators and market saturation indicators to generate product market potential assessment indicators, the following should be included:

[0071] The product popularity index and market saturation index are normalized; weight coefficients for the product popularity index and market saturation index are set based on industry characteristics, and the normalized product popularity index and market saturation index are merged using a weighted average algorithm; thus, a product market potential assessment index is generated.

[0072] Specifically, the raw popularity and saturation data from various platforms and sources are preprocessed, including outlier identification and removal, missing value imputation, and time alignment. Seasonal and abrupt change corrections are applied to the time series data to ensure the stability of subsequent calculations. Then, product popularity and market saturation indicators are normalized to map indicators of different dimensions and magnitudes to comparable measurement intervals, eliminating the impact of dimensional differences and extreme values ​​on the fusion results (normalization methods can employ min-maximum mapping, Z-score standardization, or scaling transformation based on robust statistics to adapt to data distribution characteristics). After normalization, two types of data are pre-set or dynamically adjusted based on typical characteristics of the target industry (e.g., industry life cycle, product iteration speed, competition intensity, user attention stability, etc.) and product category attributes. The weighting coefficients of the indicators—weights can be set by domain experts' experience, determined by historical backtesting, or refined through automated calibration methods (such as optimizing the fusion effect with historical validation sets) to ensure greater sensitivity to popularity in emerging market scenarios and greater sensitivity to saturation in mature market scenarios. Then, a weighted average fusion strategy is used to linearly merge the normalized popularity and saturation according to their weights to obtain a comprehensive market potential score. At the same time, confidence assessment and robustness checks can be introduced during the fusion process (such as sensitivity analysis of input weights or normalization methods) to indicate the reliability of the fusion results. Finally, the generated market potential assessment indicators are post-processed and archived, including grading by hierarchical thresholds, comparing with historical benchmarks to detect abnormal increases or decreases, and recording the indicators, their sources, and weight settings as auditable metadata.

[0073] The process involves normalization, mapping the original product popularity index to the [0,1] interval via linear transformation. For the market saturation index, it is first inverted (because lower market saturation indicates greater opportunity) before normalization. During normalization, outliers are automatically identified and handled. For example, if a product popularity index is significantly higher than the 99th percentile of historical data, it is limited to a reasonable range to prevent extreme values ​​from excessively influencing the overall assessment. Regarding weighting coefficients, an industry-specific database is maintained, storing the sensitivity parameters of each industry to popularity and market saturation indices. For instance, the electronics industry, due to rapid technological updates and fluctuating consumer attention, assigns a lower weight to popularity indices (0.35-0.45) and a higher weight to market saturation indices (0.55-0.65). In the fast-moving consumer goods (FMCG) sector, popularity indices are assigned a higher weight (0.65-0.75), while market saturation indices are assigned a lower weight (0.25-0.35). The weighting coefficients are not fixed but dynamically adjusted according to the product's lifecycle stage. During the new product introduction phase, the weight of the popularity indicator is increased, while during the maturity phase, the weight of the market saturation indicator is increased. The weighted average algorithm uses a simple linear combination, but it verifies the effectiveness of different combinations based on historical data to select the optimal fusion strategy. The fusion process also considers the time decay factor, assigning higher weights to recent data to more accurately reflect current market trends. For example, a smartwatch product has a normalized popularity indicator value of 0.65 and a normalized market saturation indicator value of 0.45 in the electronics industry. Based on industry characteristics, weights are set to 0.4 and 0.6, resulting in a market potential assessment indicator of 0.53.

[0074] Through the above technical solutions, this application achieves a scientific integration of product popularity and market saturation, generating potential assessment indicators that objectively reflect the product's market prospects. Normalization resolves the issue of inconsistent dimensions among different indicators, while outlier handling ensures data quality. The industry-specific driven dynamic weight allocation mechanism allows the assessment model to adapt to the characteristics of different industries, avoiding the bias of a "one-size-fits-all" approach. The weighted average fusion method is simple and efficient, considering both product lifecycle stages and time decay factors, ensuring that market potential assessment focuses on both current popularity and long-term development potential.

[0075] This application further proposes a method for constructing a multidimensional product knowledge graph based on product technical feature vectors and using semantic relation extraction algorithms, including:

[0076] Product technical feature vectors are mapped to the node space of the knowledge graph to generate product nodes; the technical similarity between product nodes and existing technical nodes is calculated; when the technical similarity is greater than a preset similarity threshold, a technical association is established between the product node and the technical node; the weight value of the technical association is set according to the magnitude of the technical similarity; product nodes with technical associations are clustered to form technical theme clusters; a multidimensional product knowledge graph is constructed based on the technical theme clusters; the multidimensional product knowledge graph includes product nodes, technical nodes, design element nodes, and patent nodes.

[0077] Specifically, the preprocessed and vectorized product technical feature vectors are mapped to the node space of the knowledge graph. This mapping can be achieved through pre-trained embedding alignment, dimensionality reduction projection, or semantic mapping based on industry ontology to ensure consistency between the vector representation and the semantic entity space of the graph. Then, a similarity search is performed between each product node and existing technology nodes in the graph. The technical similarity between the two nodes is calculated using semantic similarity metrics (e.g., a comprehensive evaluation based on vector cosine similarity, distance metrics, or semantic matching scores). The similarity score is then weighted and corrected by combining textual evidence (such as the frequency of identical terms in product descriptions), structured attribute matching degree, and source credibility. When the obtained technical similarity exceeds a pre-set similarity threshold, a technical association is established between the corresponding product node and technology node. Weights are assigned to this association based on the similarity value, the amount of evidence, and the reliability of the information source. These weights are simultaneously recorded... The database is recorded as auditable metadata for the knowledge graph. For all sets of product nodes and technology nodes with established relationships, an appropriate data-driven clustering method (such as density-based or hierarchical clustering algorithms combined with semantic topic consistency checks) is used to aggregate highly related nodes into technology topic clusters to characterize knowledge clusters in a specific technology field or design direction. When constructing the multidimensional product knowledge graph, product nodes, technology nodes, design element nodes, and patent nodes are used as basic entity types, and various semantic relationships between nodes are clearly defined (such as "implementation-feature", "includes-design element", "depends on-technology module", "limited by-patent claims", etc.). At the same time, a timestamp, confidence level label, and source citation are added to each relationship to support time-series analysis and source tracing. Finally, consistency checks and manual review are performed on the generated graph (including conflict detection, manual labeling of low-confidence relationships, or automatic triggering of secondary extraction).

[0078] The process involves mapping product technical feature vectors to a predefined high-dimensional node space, with each product corresponding to a unique product node. Node attributes include product ID, technical feature vector, and release time. Technical similarity calculation employs an improved cosine similarity algorithm, considering the hierarchical structure and importance differences of technical features, rather than a simple vector dot product. The similarity threshold is not fixed but dynamically adjusted based on the complexity of the technical field; for example, it is set to 0.65 in the electronics field and 0.70 in the mechanical field to accommodate the differences in the expression of technical features across different fields. When the similarity exceeds the threshold, not only is an association established, but the specific technical dimensions of the association are also analyzed, recording whether the similarity is functional, structural, or material technology. Association weights are based not only on similarity values ​​but also on the historical importance and industry influence of the technical nodes, weighted using the PageRank algorithm. The formation of technical topic clusters uses a hierarchical clustering algorithm. First, the similarity matrix between all product nodes is calculated, and then hierarchical merging is performed based on similarity until the average similarity within a cluster reaches a preset threshold (usually 0.75). The knowledge graph is also regularly optimized to identify and merge semantically similar technology nodes, eliminating redundancy. At the same time, the coverage of the knowledge graph is continuously expanded through user feedback and new data. For example, if multiple products are identified as involving "wireless charging" technology, a "wireless charging technology cluster" will be automatically formed, and the relevant product nodes will be assigned to this cluster, while also associating them with sub-technology nodes such as "Qi standard" and "magnetic resonance charging".

[0079] Through the above technical solutions, this application constructs a multi-dimensional product knowledge graph with a clear structure and rich semantics. The relationships between product nodes and technology nodes not only reflect the surface connections between products and technologies but also reveal deeper technological connections through multi-dimensional analysis. Dynamically adjusted similarity thresholds ensure the consistency and accuracy of knowledge graph construction across different technological fields. The formation of technology topic clusters employs a hierarchical clustering algorithm, giving the knowledge graph a clear hierarchical structure that facilitates rapid location of relevant technological fields. The product nodes, technology nodes, design element nodes, and patent nodes contained in the knowledge graph form a complete knowledge system, covering product technical features and intellectual property information.

[0080] This application further proposes that when screening candidate products with market growth potential based on product market potential assessment indicators, the following should be included:

[0081] Set screening thresholds for product market potential assessment indicators; compare the product market potential assessment indicators of all products with the screening thresholds; select products whose product market potential assessment indicators are greater than the screening thresholds as candidate products with market growth potential.

[0082] Specifically, the screening threshold is determined based on the target business scenario and industry characteristics. This threshold can be obtained in various ways—directly set by domain experts based on market experience, determined based on quantiles or baselines obtained from historical data backtesting, or an adaptive threshold based on a sliding time window to respond to market rhythms and seasonal changes. In multi-category or multi-market scenarios, categorized thresholds can be set for different product categories, regions, or sales channels to avoid bias caused by a one-size-fits-all approach. Subsequently, a batch comparison is performed on all product market potential assessment indicators obtained through normalization and fusion calculations. The current indicator of each product is compared with the screening threshold of the corresponding category. If the indicator of a product is greater than the threshold, it is temporarily listed as having market potential. Candidate products with market growth potential; to improve the robustness of the screening, several auxiliary rules can be used in parallel (such as setting a minimum number of days of data coverage, removing abnormal sudden points, and requiring that the popularity index and supply and demand index meet the minimum conditions at the same time) and multi-level threshold strategies (such as three levels of candidates, key candidates, and high-priority candidates). Products on the edge of the threshold are processed by ranking down, Top-K downgrading, or manual review; the screening process should record decision metadata (including the source of the threshold used, timestamp, index version and confidence level) to support traceable audit, and send the candidate list to the subsequent deduplication, infringement risk assessment and compliance filtering stages. Finally, the sorted candidate product list is automatically output and an alarm or manual review is triggered.

[0083] The system employs a dynamic threshold setting method, rather than simply using a fixed threshold. Instead, it determines the threshold based on a comprehensive analysis of product category, market stage, and historical data. The distribution of market potential assessment indicators for successful products in historical data is analyzed to determine the 80th percentile as a baseline threshold. This threshold is then adjusted according to the current market environment: when the overall market is in an upward phase, the threshold is appropriately increased to filter out products with greater potential; when the market is saturated, the threshold is decreased to broaden the screening range. Differences in product categories are also considered; for example, the threshold for electronics is typically set at 0.62, while for household goods it is set at 0.58, reflecting the different market characteristics of each category. The screening process uses a two-stage mechanism: the first stage performs rapid screening to exclude products with market potential assessment indicators significantly below the threshold (e.g., below 0.4); the second stage performs a detailed evaluation of products close to the threshold, considering additional factors such as seasonality and recent market trend changes. The screening results are also visualized, displaying changes in distribution before and after screening, the screening ratio of each product category, and other information to help users understand the screening logic. For example, when screening 1,000 electronic products, with an initial threshold of 0.62, 450 low-potential products were eliminated in the first stage. In the second stage, the remaining products were carefully evaluated, and 220 candidate products with market growth potential were finally selected.

[0084] Through the above technical solutions, this application achieves accurate screening of massive product data, reducing the number of products requiring technical risk assessment. The dynamic threshold setting method allows screening criteria to adapt to changes in the market environment, avoiding screening bias caused by fixed thresholds. The two-stage screening mechanism improves screening efficiency while ensuring screening quality; the first stage quickly eliminates products that clearly do not meet the criteria, and the second stage conducts a detailed evaluation of borderline products. Visual analysis of the screening results provides a transparent screening process, enhancing user trust in the results.

[0085] This application further proposes mapping the technical feature vectors of candidate products to a multi-dimensional product knowledge graph, and performing multi-dimensional comparisons with the technical nodes in the knowledge graph, including:

[0086] Locate the technical node in the multidimensional product knowledge graph that is most similar to the technical feature vector of the candidate product; calculate the technical semantic similarity between the technical feature vector of the candidate product and the located technical node.

[0087] Analyze the extent to which the technical feature vectors of candidate products cover the technical node features, and generate technical feature coverage.

[0088] Search for patent nodes associated with technical nodes and evaluate the matching degree between the technical features of candidate products and patent claims;

[0089] Multi-dimensional comparison results are generated based on technical semantic similarity, technical feature coverage, and matching degree.

[0090] Specifically, candidate product vectors are aligned using the same embedding space or mapping function as graph node vectors to ensure semantic comparability. Then, efficient vector retrieval (e.g., nearest neighbor retrieval or vector indexing acceleration structures) is used to locate the set of technical nodes in the knowledge graph most similar to the candidate product's technical features. The retrieval results simultaneously return the similarity candidate set and retrieval confidence score for each candidate node. For each located technical node, the technical semantic similarity between the candidate product's technical feature vector and each technical node is calculated. This similarity is based not only on distance or similarity measures between vectors but also on textual evidence matching (e.g., key term co-occurrence, term frequency), structured attribute consistency (matching of fields such as function, performance indicators, and materials), and source credibility to weight and correct the similarity score, thus obtaining a more semantically interpretable similarity assessment. Based on this, the coverage of the candidate product's feature vector to the technical node's features is further analyzed, i.e., measuring the proportion of the candidate product containing or implementing the function / feature described by the technical node. The depth and coverage assessment considers the priority of key features (core features have higher weight than secondary features), the completeness of the match, and the impact of missing items, and assigns a confidence level to the coverage. Simultaneously, it retrieves patent nodes associated with these technical nodes, extracts the claims text and specification evidence of relevant patents, and evaluates the semantic and element matching degree between the candidate product's technical features and the patent claims. This matching degree integrates legal semantic clues such as keyword overlap, element correspondence, functional equivalence judgment, and the breadth of the patent subject matter. The retrieved patent legal status and geographical scope are also incorporated as correction factors in the assessment. Finally, the multi-dimensional indicators such as technical semantic similarity, technical feature coverage, and patent matching degree are summarized and normalized according to a preset or adaptive weighting system to generate interpretable multi-dimensional comparison results. These results include both numerical scores and detailed evidence supporting the scores (such as the most matching technical nodes, a list of covered key features, relevant patent citations and matching fragments), confidence levels, and timestamps.

[0091] The process involves using the k-nearest neighbor algorithm to locate the k most similar technical nodes (k is typically set to 5-10) in the knowledge graph to the technical feature vectors of the candidate product, ensuring coverage of major relevant technical fields. Technical semantic similarity calculation considers not only vector similarity but also the hierarchical relationship between technical nodes; for example, if two technical nodes belong to the same technical branch, the similarity score is appropriately increased. Technical feature coverage analysis employs a feature matching algorithm, comparing the technical features of the candidate product with the feature sets of the technical nodes one by one, calculating the proportion of matching features, and weighting them according to their importance. For example, the weight of core functional features is set to 1.0, the weight of auxiliary functional features to 0.7, and the weight of appearance features to 0.5. Patent claim matching evaluation uses a deep learning model to semantically match the technical features of the candidate product with the patent claim text, identifying potential technical feature overlaps. The legal status and geographical scope of the patent are also considered, with reduced matching weights for invalid patents or geographically irrelevant patents. When generating multi-dimensional comparison results, not only are scores for each dimension output, but detailed matching evidence is also provided, such as specific matched technical features, relevant patent numbers, and claim clauses. For example, a Bluetooth headset product has a semantic similarity of 0.85 with an active noise cancellation technology node, a technical feature coverage of 0.78, and a matching degree of 0.82 with 3 valid patents. A detailed comparison report is generated, listing the specific matching technical features and related patent information.

[0092] Through the above technical solutions, this application achieves a comprehensive and accurate evaluation of the technical features of candidate products. The k-nearest neighbor algorithm ensures the comprehensiveness of the technical comparison, avoiding the limitations of focusing only on a single technical node. The calculation of technical semantic similarity combines vector similarity and technical hierarchy relationships, reflecting the degree of technical correlation. The technical feature coverage analysis uses a weighted feature matching algorithm to distinguish the importance of different technical features, making the evaluation results more reasonable. The patent claim matching evaluation uses a deep learning model, which can capture the deep semantic relationship between technical features and patent text, improving matching accuracy.

[0093] This application further proposes that, when generating a patent infringement risk assessment value based on multi-dimensional comparison results and risk quantification, the following should be included:

[0094] The indicators of each dimension in the multi-dimensional comparison results are normalized; the risk weight coefficients of each dimension indicator are set based on legal practice; a weighted summation algorithm is used to calculate the preliminary risk score; the preliminary risk score is corrected according to the legal status and geographical scope to generate a patent infringement risk assessment value.

[0095] Specifically, the various dimensions of the comparison results (such as technical semantic similarity, technical feature coverage, matching degree with patent claims, number of relevant patents retrieved and their confidence levels) undergo unified normalization and quality control processing. This includes strategies for imputing missing values, detection and correction of outliers, and weighted adjustments to confidence levels from different sources to ensure that all indicators are on the same comparable scale and to record metadata of the normalization method. Subsequently, based on legal practice and business strategy, a legal expert group and a data science team jointly formulate risk weight coefficients for each dimension of the indicators. Factors considered in the weight formulation process include: the breadth of the claims (distinction between independent and dependent claims), the coverage and priority period of the patent family, the patentee and its licensing / litigation history, the importance and substitutability of the technology in the target market, and the legal risk threshold that the company can bear. After obtaining the normalized indicators and weights, an interpretable weighted aggregation process is used to generate a preliminary risk score. At the same time, the contribution of each dimension to the final score is retained during the aggregation process to facilitate traceability and manual review. Then, based on the patent's legal status (e.g., granted, under examination, expired, revoked, invalidation result, etc.) and geographical scope (the degree of overlap between the country / region where the patent is valid and the planned product launch location), the initial risk score is modified according to rules. For example, the risk weight is reduced for patents that have expired or have not obtained protection in the target region, and the risk adjustment coefficient is increased for patents that are granted and have recent rights protection records or active enforcement activities by the rights holder. In addition, modifications for temporary factors should also be considered, such as the impact of priority date and publication date on the window of free implementation, the upward / downward adjustment of the judgment result due to the uncertainty of claim interpretation, and the availability of patent invalidation or licensing negotiations. The final output patent infringement risk assessment value is not only a single value, but should also be accompanied by detailed supporting metadata (including the normalized indicators used in the calculation, the source of the weights used, the relevant patent list and legal status, geographical mapping results, contribution decomposition, confidence assessment, and timestamps), and high-risk items exceeding the preset warning threshold will trigger subsequent processing procedures (such as manual legal review, priority generation of design circumvention schemes, or initiation of licensing negotiations).

[0096] The process involves normalizing the indicators in the multi-dimensional comparison results to ensure comparability across different dimensions. During normalization, an industry-standard conversion function is used to convert the original indicator values ​​into risk probability values ​​within the range of 0-1. The risk weighting coefficients are set based on extensive historical case analysis and legal expert opinions: the weight for technical semantic similarity is set at 0.35-0.45, the weight for technical feature coverage is set at 0.25-0.35, and the weight for patent claim matching is set at 0.30-0.40, with specific values ​​fine-tuned according to the technical field and product type. For example, for electronic products, the weight for patent claim matching is slightly higher; for design products, the weight for technical feature coverage is slightly higher. After calculating the initial risk score using a weighted summation algorithm, adjustments are made based on the patent's legal status: valid patents maintain their original score, patents under examination have their score multiplied by 0.8, and invalid patents have a score set to 0. Geographical scope adjustments consider the overlap between the product's sales area and the patent's protection area. For example, if a product is only sold in Europe and America, but the patent is only valid in China, the patent's risk score is multiplied by 0.2. It also implements sensitivity analysis for risk assessment, showing the degree of influence of each dimension indicator on the final risk assessment value, helping users understand the sources of risk. For example, the initial risk score of a Bluetooth speaker product is 0.75. After correction for legal status and geographical scope, the final patent infringement risk assessment value is 0.62. Sensitivity analysis shows that the coverage of technical features is the main source of risk.

[0097] Through the above technical solutions, this application achieves a scientific and quantitative assessment of patent infringement risk. Normalization ensures the comparability of indicators across different dimensions, while the adoption of industry-standard transformation functions gives the assessment results practical legal significance. Risk weighting based on numerous historical cases and legal expert opinions ensures the assessment results conform to legal practice, improving the practicality and credibility of the assessment. The correction mechanism for legal status and geographical scope makes the risk assessment more accurate, avoiding excessive focus on invalid patents or geographically irrelevant patents. Sensitivity analysis in the risk assessment provides transparency into the sources of risk, helping users understand and specifically mitigate risks.

[0098] This application further proposes a method for generating product design avoidance schemes based on technology alternative paths and design variants in a multi-dimensional product knowledge graph, using a graph path analysis algorithm, including:

[0099] Analyze the correlation strength between the technical feature vectors of candidate products and each technical node in the multidimensional product knowledge graph; identify technical nodes whose correlation strength with the technical feature vectors of candidate products is greater than a preset correlation strength threshold; calculate their contribution to the patent infringement risk assessment value; and filter technical nodes whose contribution is greater than a preset contribution threshold based on their contribution to form a set of technical nodes to be avoided.

[0100] Search for adjacent technology nodes in the multidimensional product knowledge graph that are adjacent to each technology node in the set of technology nodes to be avoided; calculate the technology compatibility index between adjacent technology nodes and candidate products; filter technology nodes whose technology compatibility index is greater than the preset compatibility threshold based on the technology compatibility index to form a set of alternative technology nodes.

[0101] Analyze the design variant schemes corresponding to each technology node in the set of alternative technology nodes; calculate the technical effect retention rate of each design variant scheme; evaluate the implementation difficulty coefficient of each design variant scheme; give a comprehensive score to each design variant scheme based on the technical effect retention rate and the implementation difficulty coefficient; identify the design variant schemes with a comprehensive score higher than the preset score threshold as product design avoidance schemes; product design avoidance schemes include feasible design avoidance schemes, technical substitution suggestions and design modification guidance.

[0102] Specifically, a comprehensive evaluation is conducted on the correlation strength between the technical feature vectors of candidate products and each technical node in the knowledge graph. This evaluation considers not only vector similarity but also co-occurrence evidence, functional semantic matching, consistency of implementation components, and source credibility, thereby obtaining a contribution index for each technical node to the candidate product. Next, all technical nodes with correlation strength exceeding a preset threshold and significantly contributing to the patent infringement risk assessment value are identified. These nodes are then grouped into a set of technical nodes to be avoided, and the relative contribution of each node to the overall risk is calculated for quantitative ranking. Subsequently, candidate alternative technical nodes adjacent to or reachable from the nodes to be avoided are searched in the knowledge graph along topological adjacency relationships and semantic substitution paths. For each adjacent node, its technical compatibility with the candidate product in terms of function, performance, interface, materials, and manufacturing processes is automatically evaluated. The compatibility assessment comprehensively considers the impact of substitution on core functions, integration difficulty with existing modules, cost and supply chain feasibility, and potential changes to user experience. For alternative technical nodes that pass the compatibility screening, their corresponding design variants are retrieved, and each variant is analyzed in depth, including estimating the technology. The evaluation process includes: retention rate (the percentage of the alternative solution that retains the original technology's functionality / performance); assessment of implementation difficulty (including design modification workload, trial production verification cycle, compliance and certification requirements, adaptation costs to existing production lines, and external dependency risks); and checking the intellectual property status and feasibility of the variant itself. Based on multi-dimensional indicators such as retention rate and implementation difficulty, each design variant is given an interpretable comprehensive score, and solutions with scores above a preset threshold are listed as preliminary product design avoidance solutions. For the final selected avoidance solutions, structured deliverables are automatically generated—including feasibility statements, a list of key replacement parts / modules, estimated implementation procedures and time, expected performance changes, key points of prototype testing and risk mitigation measures, and a comparison decision table with the original solution. These solutions, along with the evidence chain of the source node (such as supporting text, graph paths, and confidence levels), are saved as auditable metadata for manual review and priority confirmation by the engineering design team and legal reviewers. The entire process supports iterative updates (dynamically refreshing the graph and score based on test feedback, weight adjustments, and new evidence supplementation) and provides a risk-benefit view and multi-solution comparison.

[0103] The analysis focuses on the correlation strength between candidate products and technical nodes in the knowledge graph, employing an improved PageRank algorithm to calculate the contribution of each technical node to the overall infringement risk. The correlation strength threshold and contribution threshold are dynamically set based on product type and industry characteristics; for example, the correlation strength threshold is set to 0.6 and the contribution threshold to 0.4 for electronic products. When searching for technical nodes adjacent to the target technical node in the knowledge graph, not only directly adjacent nodes are considered, but also indirectly related nodes reachable through 2-3 hops, expanding the search scope for alternative solutions. The technical compatibility index calculation comprehensively considers three dimensions: technical principle compatibility, interface compatibility, and performance compatibility, with each dimension assigned different weights based on the technical field. For example, interface compatibility has a higher weight for electronic products, while technical principle compatibility has a higher weight for mechanical products. The technical effect retention rate is predicted through simulation analysis and historical data evaluation to assess the impact of alternative solutions on the core functions of the original product; the implementation difficulty coefficient considers factors such as technological maturity, supply chain feasibility, and R&D costs. The comprehensive scoring uses a weighted scoring method, with the technical effect retention rate weighted at 0.6 and the implementation difficulty coefficient weighted at 0.4, ensuring that the solution is both effective and feasible. For example, the "charging interface sealing structure" of a certain Bluetooth headset was identified as a high-risk node (contribution 0.65). The "magnetic sealing cover" was found as an alternative, with a technical effect retention rate of 0.85, an implementation difficulty coefficient of 0.35, and an overall score of 0.75. It was determined to be a feasible design avoidance solution.

[0104] Through the above technical solutions, this application achieves a complete closed loop from infringement risk identification to avoidance solution generation. The improved PageRank algorithm accurately identifies high-risk technical points and their contribution, avoiding excessive avoidance of low-risk technologies. The multi-hop search mechanism expands the search scope of alternative solutions and increases the probability of finding feasible alternatives. The multi-dimensional evaluation of technical compatibility indicators ensures the coordination between alternative solutions and the overall product design, improving the feasibility of the solution. The comprehensive evaluation of technical effect retention rate and implementation difficulty coefficient balances technical performance and implementation cost, making the avoidance solution both effective and practical.

[0105] This application further proposes to integrate and analyze candidate products, patent infringement risk assessment values, and product design avoidance schemes to generate intelligent product selection decision support information, including:

[0106] Candidate products are categorized and organized according to product type and market potential; risk levels are assigned to candidate products based on patent infringement risk assessment values; product design avoidance solutions are associated with corresponding candidate products; a comprehensive score of market potential and infringement risk is calculated for each candidate product; candidate products are ranked according to the comprehensive score to generate a recommended candidate product list; and the recommended candidate product list, risk labeling information, and product design avoidance solutions are integrated into intelligent product selection decision support information.

[0107] Specifically, the selected candidate products are categorized according to a predefined product classification system (e.g., product line, functional category, target customer group, price range, etc.) and their corresponding product market potential assessment indicators, forming a grouped catalog for horizontal comparison and benchmarking against similar products. Then, each candidate product is labeled with a risk level (e.g., low, medium, high, or tiered color coding) based on its patent infringement risk assessment value, with each level accompanied by a specific risk description and triggering conditions to help reviewers quickly understand the source of the risk. Simultaneously, each candidate product is automatically associated with a corresponding executable product design circumvention scheme, recording the type of circumvention scheme (e.g., alternative technology, structural modification, process adjustment), estimated implementation difficulty, and expected degree of technology retention. Based on this, market potential and infringement risk are comprehensively considered using interpretable rules or strategies determined by historical backtesting (e.g., ...). For example, by prioritizing high-potential, low-risk products, high-potential, high-risk products are included in a candidate pool requiring design avoidance or legal evaluation, while low-potential, low-risk products are listed as low priority. Based on this, a comprehensive evaluation conclusion is generated for each candidate product, clearly stating the recommendation level, the main decision-making driving factors (such as popularity, supply and demand gap, and main patent-related items), and the recommended next steps (such as direct implementation, prior design avoidance followed by implementation, legal review, or termination of implementation). All candidate products are sorted according to the comprehensive evaluation results to generate a candidate product recommendation list, which is output in the form of an interactive report or dashboard. This list includes ranking details, risk decomposition (showing the contribution of each dimension to the final score), avoidance solution links, evidence citations and confidence level annotations, and supports multi-dimensional filtering (by category, region, risk level, potential range, etc.) and export functions.

[0108] The system categorizes candidate products into a structured classification tree based on product category (e.g., electronics, home furnishings, clothing, footwear) and market potential level (high, medium, low). Risk level labeling uses a three-tier classification: low risk (risk assessment value < 0.3), medium risk (0.3 ≤ risk assessment value < 0.6), and high risk (risk assessment value ≥ 0.6), with specific high-risk technical points labeled for high-risk products. Product design mitigation solutions are precisely linked to corresponding candidate products, ensuring each high-risk product has a targeted mitigation solution, and the implementation difficulty and expected effects of the solution are indicated. The comprehensive score calculation uses a weighted scoring model, with market potential assessment index weighted at 0.7 and patent infringement risk assessment value weighted at 0.3. However, the weights can be adjusted according to corporate strategic preferences; for example, innovative companies might increase the market potential weight to 0.8, while conservative companies might increase the risk weight to 0.4. A multi-dimensional sorting function is also implemented, allowing users to sort by market potential, risk level, or comprehensive score, and also set filtering conditions (e.g., displaying only low-risk products). The intelligent product selection decision support information is presented in an interactive dashboard, including modules such as a product list, risk heatmap, and details of avoidance solutions, and supports exporting detailed reports. For example, a smartwatch product has a market potential assessment index of 0.72, a patent infringement risk assessment value of 0.25, and a comprehensive score of 0.43. It is marked as a low-risk product, requiring no avoidance solution, and ranks high in the recommendation list.

[0109] Through the aforementioned technical solutions, this application achieves a comprehensive assessment of market potential and infringement risk, providing a comprehensive and balanced basis for product selection. Multi-dimensional categorization clarifies the decision-making information structure, facilitating quick browsing and comparison of products across different categories and potential levels. Risk level labeling intuitively displays the product's legal risks and provides specific high-risk technical points, helping decision-makers quickly identify problems. The weighted comprehensive score of market potential and infringement risk balances commercial value and legal risk, avoiding either an excessive pursuit of market potential at the expense of legal risks or an overemphasis on risk avoidance leading to missed market opportunities. The interactive dashboard provides flexible decision support; users can adjust weights, filtering criteria, and sorting methods as needed to meet the requirements of different decision-making scenarios.

[0110] In summary, by using distributed web crawlers and platform API interfaces to acquire product description text, user interaction sequences, and market dynamic indicators in parallel, product information across multiple platforms and dimensions can be covered. This enables cross-platform data synchronous mining, avoiding biases caused by data from a single platform and improving data reliability and the comprehensiveness of product selection analysis. Natural language processing technology is used to semantically parse product descriptions, combined with time series analysis to generate product popularity indicators. Market saturation indicators are constructed based on market supply and demand and growth rates. Then, through normalization and weighted fusion, product market potential assessment indicators are formed, achieving digital quantitative analysis from text to market, making potential product identification more scientific and forward-looking. Based on technical feature vectors and semantic relationship extraction technology, a multi-dimensional relationship network between product nodes, technology nodes, design element nodes, and patent nodes is automatically constructed. This intuitively presents technological evolution relationships, design dependencies, and patent protection scope, achieving structured knowledge storage and visualized relational judgment, facilitating subsequent rapid retrieval and automated analysis. By comparing the semantic similarity, feature coverage, and matching degree with patent claims of the technical vectors of candidate products and technical nodes in the knowledge graph, a multi-dimensional technical comparison result is formed. Risk correction is then applied based on legal factors, enabling the quantification and generation of a more objective, transparent, and accurate patent infringement risk assessment value, reducing the uncertainty of traditional manual analysis. Key technical nodes requiring avoidance are identified, and alternative technical nodes are screened through technical compatibility calculations. Corresponding design variants are analyzed, and a comprehensive score is generated based on the retention rate of technical effects and the difficulty of implementation, thus generating the optimal product design avoidance solution. This solution not only avoids patent risks but also maintains product functionality, improving product compliance and innovation capabilities. The market potential, infringement risk level, and avoidance solutions of candidate products are integrated and sorted according to comprehensive scores, automatically generating a product selection recommendation list and decision information. This automates and intelligently integrates the entire product selection process from data acquisition, risk identification, solution generation to final recommendation, improving the platform's product selection efficiency and accuracy while reducing the cost of manual judgment.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for intelligent product selection and infringement early warning based on cross-platform mining and knowledge graphs, characterized in that, include: Based on a distributed crawler system and platform API interfaces, product-related data is obtained from several heterogeneous platforms. The product-related data includes product description text, user interaction behavior sequences, and market dynamic indicators. Semantic analysis of product description text is performed using natural language processing technology to extract product technical feature vectors; user interaction behavior sequences are processed using time series analysis methods to generate product popularity indicators; market saturation indicators are calculated based on the market dynamic indicators; and the product popularity indicators and market saturation indicators are merged to generate product market potential assessment indicators. Based on the product's technical feature vector, a multidimensional product knowledge graph is constructed using a semantic relationship extraction algorithm; candidate products with market growth potential are screened according to the product's market potential assessment indicators. The technical feature vectors of the candidate products are mapped to a multi-dimensional product knowledge graph, and a multi-dimensional comparison is performed with the technical nodes in the knowledge graph. Based on the multi-dimensional comparison results, a patent infringement risk assessment value is generated by risk quantification. When the patent infringement risk assessment value exceeds the preset risk threshold, a product design avoidance scheme is generated based on the technology alternative paths and design variant schemes in the multidimensional product knowledge graph using the graph path analysis algorithm. The candidate products, patent infringement risk assessment values, and product design avoidance schemes are integrated and analyzed to generate intelligent product selection decision support information.

2. The intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph as described in claim 1, characterized in that, When performing semantic parsing on product description text using natural language processing techniques to extract product technical feature vectors, the following are included: The product description text is segmented and part-of-speech tagged to identify core product nouns and technical adjectives; Based on a technical dictionary and domain terminology database, extract key technologies and design elements of the product; The extracted key technologies and design elements of the product are vectorized to generate product technical feature vectors.

3. The intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph as described in claim 2, characterized in that, By combining time series analysis methods to process user interaction behavior sequences, a product popularity index is generated. Based on the aforementioned market dynamics index, a market saturation index is calculated, including: The user interaction behavior sequence is cleaned to remove abnormal clicks and invalid browsing records; the user interaction frequency per unit time is calculated to generate a user interaction frequency sequence; the user interaction frequency sequence is smoothed using a moving average algorithm; the growth rate of the smoothed user interaction frequency sequence is calculated to generate a product popularity index. Obtain market supply and demand data for similar products; calculate the ratio of market supply to market demand to generate a market supply-demand ratio index; calculate the current market growth rate based on historical market data; and non-linearly combine the market supply-demand ratio index with the market growth rate to generate a market saturation index.

4. The intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph as described in claim 3, characterized in that, When combining the product popularity index and the market saturation index to generate a product market potential assessment index, the following are included: The product popularity index and the market saturation index are normalized; weight coefficients for the product popularity index and the market saturation index are set based on industry characteristics; and a weighted average algorithm is used to merge the normalized product popularity index and the market saturation index to generate a product market potential assessment index.

5. The intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph as described in claim 4, characterized in that, When constructing a multidimensional product knowledge graph based on the product's technical feature vector using a semantic relation extraction algorithm, the following steps are included: The product technical feature vector is mapped to the node space of the knowledge graph to generate product nodes; the technical similarity between the product nodes and existing technical nodes is calculated; when the technical similarity is greater than a preset similarity threshold, a technical association is established between the product nodes and the technical nodes; the weight value of the technical association is set according to the magnitude of the technical similarity; product nodes with technical association are clustered to form technical theme clusters; a multidimensional product knowledge graph is constructed based on the technical theme clusters; the multidimensional product knowledge graph includes product nodes, technical nodes, design element nodes, and patent nodes.

6. The intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph as described in claim 5, characterized in that, When screening candidate products with market growth potential based on the aforementioned product market potential assessment indicators, the following are included: Set a screening threshold for the product market potential assessment index; compare the product market potential assessment index of all products with the screening threshold; and select products whose product market potential assessment index is greater than the screening threshold as candidate products with market growth potential.

7. The intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph as described in claim 6, characterized in that, Mapping the technical feature vectors of the candidate products to a multi-dimensional product knowledge graph and performing multi-dimensional comparisons with the technical nodes in the knowledge graph includes: Locate the technical node in the multidimensional product knowledge graph that is most similar to the technical feature vector of the candidate product; calculate the technical semantic similarity between the technical feature vector of the candidate product and the located technical node. Analyze the degree to which the technical feature vectors of the candidate products cover the technical node features, and generate technical feature coverage. Retrieve patent nodes associated with technical nodes and evaluate the matching degree between the technical features of the candidate products and the patent claims; Multi-dimensional comparison results are generated based on the technical semantic similarity, the technical feature coverage, and the matching degree.

8. The intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph as described in claim 7, characterized in that, Based on the multi-dimensional comparison results, when generating a patent infringement risk assessment value through risk quantification, the following are included: The indicators of each dimension in the multi-dimensional comparison results are normalized; risk weight coefficients of each indicator are set based on legal practice; a weighted summation algorithm is used to calculate a preliminary risk score; the preliminary risk score is corrected according to the legal status and geographical scope to generate a patent infringement risk assessment value.

9. The intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph as described in claim 8, characterized in that, When generating product design avoidance solutions based on technology alternative paths and design variants in a multi-dimensional product knowledge graph using graph path analysis algorithms, the following are included: Analyze the correlation strength between the technical feature vector of the candidate product and each technical node in the multidimensional product knowledge graph; identify technical nodes whose correlation strength with the technical feature vector of the candidate product is greater than a preset correlation strength threshold; calculate the contribution to the patent infringement risk assessment value; and filter technical nodes whose contribution is greater than a preset contribution threshold based on the contribution to form a set of technical nodes to be avoided. Search for adjacent technology nodes in the multidimensional product knowledge graph that are adjacent to each technology node in the set of technology nodes to be avoided; calculate the technology compatibility index between the adjacent technology nodes and the candidate products; filter technology nodes whose technology compatibility index is greater than a preset compatibility threshold according to the technology compatibility index to form a set of alternative technology nodes; The design variant schemes corresponding to each technology node in the set of alternative technology nodes are analyzed; the technical effect retention rate of each design variant scheme is calculated; the implementation difficulty coefficient of each design variant scheme is evaluated; a comprehensive score is given to each design variant scheme based on the technical effect retention rate and the implementation difficulty coefficient; design variant schemes with a comprehensive score higher than a preset score threshold are identified as product design avoidance schemes; the product design avoidance schemes include feasible design avoidance schemes, technical substitution suggestions, and design modification guidance.

10. The intelligent product selection and infringement early warning method based on cross-platform mining and knowledge graph as described in claim 9, characterized in that, When integrating and analyzing the candidate products, patent infringement risk assessment values, and product design avoidance schemes to generate intelligent product selection decision support information, the following are included: The candidate products are categorized and organized according to product type and market potential; the candidate products are labeled with risk levels based on the patent infringement risk assessment value; the product design avoidance schemes are associated with the corresponding candidate products; a comprehensive score of market potential and infringement risk for each candidate product is calculated; the candidate products are sorted according to the comprehensive score to generate a candidate product recommendation list; the candidate product recommendation list, risk labeling information, and product design avoidance schemes are integrated into intelligent product selection decision support information.