Intellectual property risk early warning system based on multi-source heterogeneous data fusion

Through the intellectual property risk warning system with multi-source heterogeneous data fusion, multi-dimensional data is automatically processed and analyzed, and a dynamic knowledge graph is built, which solves the problems of low data integration efficiency and lagging risk warning in intellectual property management, and achieves efficient and accurate risk warning and management.

CN120235458AInactive Publication Date: 2025-07-01SHANDONG XINDING MANAGEMENT CONSULTING CO LTD

Patent Information

Application Number
CN202510533869.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the intellectual property management, the existing technology has problems such as low efficiency in the integration of multi-source heterogeneous data, lagging risk warning, and serious data island phenomena, making it difficult to grasp the comprehensive risk situation.

Method used

The intellectual property risk warning system with multi-source heterogeneous data fusion is adopted to automatically extract patented technical features through natural language processing and deep learning models, build dynamic knowledge graphs, conduct efficient data fusion, and perform hierarchical warnings in combination with risk rule databases to realize semantic association and automated cleaning across data sources.

Benefits of technology

It realizes automatic identification and dynamic tracking of intellectual property risks, significantly improves the timeliness and comprehensiveness of infringement discovery, improves the accuracy of infringement judgments, and reduces the company's rights protection costs and management burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235458A_ABST
    Figure CN120235458A_ABST
Patent Text Reader

Abstract

The invention discloses an intellectual property risk early warning system based on multi-source heterogeneous data fusion, which relates to the field of intellectual property management and comprises a multi-source data acquisition module, a data preprocessing and standardization module, a distributed data storage module, a risk analysis engine module, a risk early warning and visualization module and a data communication module. According to the system, heterogeneous data of a patent database, an enterprise bidding and tendering system and a judicial public platform are collected, a natural language processing technology is utilized to extract patent technology feature entities, trademark image similarity is analyzed in combination with a convolutional neural network, and a dynamic knowledge graph related patent technology, competitor layout and infringement litigation cases are constructed; a deep learning model is adopted to fuse multi-dimensional features, risk scores are generated through weighting of a risk rule base, and graded early warning signals are triggered. According to the method, the problems of low multi-source data integration efficiency and risk response lag are solved, the infringement response time is shortened, and enterprises are guided to accurately arrange high-value patents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intellectual property management, and more particularly to an intellectual property risk early warning system based on multi-source heterogeneous data fusion. Background Art

[0002] With the rapid development of big data and artificial intelligence technologies, the field of intellectual property management faces the challenge of integrating multi-source heterogeneous data. The data sources cover patent texts, trademark registration information, judicial litigation records, enterprise bidding data, etc., with diverse formats, large data volumes, and frequent updates. Traditional technologies have significant deficiencies in data fusion, risk early warning, and other aspects: Traditional trademark and patent risk monitoring mainly relies on manual retrieval and regular checks, which has obvious efficiency bottlenecks and recognition blind spots. Taking the patent CN112364034A "Trademark Monitoring Method, Trademark Monitoring Device and Electronic Equipment" as an example, the existing trademark monitoring method mainly relies on users to actively query. Users need to regularly query and analyze trademark information to discover risks or matters. In addition, there is a common data silo phenomenon in existing technical solutions, and there is a lack of an effective correlation analysis mechanism between patent, trademark, domain name, and market sentiment data, making it difficult for enterprises to comprehensively grasp the intellectual property risk situation.

[0003] In view of the above defects, the present invention proposes an intellectual property risk early warning system based on multi-source heterogeneous data fusion. Its technical solution solves the deficiencies of the existing technology through the following innovative points: combining natural language processing and deep learning models, automatically extracting patent technology feature entities, constructing a dynamic knowledge graph, realizing semantic association across data sources, and performing efficient data fusion; automatic cleaning and preprocessing: processing trademark image features through a convolutional neural network for automatic cleaning and preprocessing, significantly reducing the need for manual intervention; based on the model, fusing multi-dimensional features such as patent citation relationships and trademark similarities, and dynamically adjusting weights in combination with a risk rule library to achieve hierarchical early warning and improve prediction accuracy and real-time performance.

[0004] Through the above technical improvements, the present invention realizes the full-process optimization from data collection to risk decision-making in the field of intellectual property management, and solves the bottleneck problems of traditional methods in terms of efficiency, accuracy, and intelligence. Summary of the Invention

[0005] In order to overcome the above defects of the existing technology, the present invention provides an intellectual property risk early warning system based on multi-source heterogeneous data fusion. The technical solution adopted is as follows: An intellectual property risk early warning system based on multi-source heterogeneous data fusion includes a multi-source data collection module, a data preprocessing and standardization module, a distributed data storage module, a risk analysis engine module, a risk early warning and visualization module, and a data communication module; The multi-source data acquisition module collects structured and unstructured data from intellectual property disclosure platforms, enterprise business databases, and Internet public data sources. The data includes patent texts, trademark registration information, enterprise bidding records, product technical parameters, judicial litigation data, and market sentiment texts. The data preprocessing and standardization module cleans, converts the format, and performs semantic annotation on the collected heterogeneous data. Among them, technical feature entities in patent claims are extracted through natural language processing technology, feature vectorization is performed on trademark images, and the technical requirements in bidding documents are associated and mapped with the patent technology efficacy matrix. The distributed data storage module adopts a hybrid architecture of a graph database and a relational database to store knowledge graph associated data and structured business data respectively. The risk analysis engine module performs feature fusion on multi-dimensional data based on a deep learning model to generate an enterprise intellectual property risk score. At the same time, a dynamic knowledge graph is constructed. By associating patent technologies, competitor patent layouts, infringement litigation cases, and policies and regulations, infringement conflicts and technology substitution risks are identified. The risk warning and visualization module triggers a hierarchical warning signal according to the risk score and displays the risk source, associated evidence chain, and countermeasure suggestions through a visualization interface. The data communication module conducts data interaction with the enterprise internal management system to achieve real-time push of warning information.

[0006] Furthermore, the data preprocessing and standardization module further extracts technical keywords, legal status, and rights holder information from unstructured text data through named entity recognition, and uses a convolutional neural network to generate feature codes for trademark image data and compare the similarity with the registered trademark database.

[0007] Furthermore, the training method of the deep learning model in the risk analysis engine module includes fusing patent citation relationships, trademark similarity, bidding technical requirement matching degree, and litigation history record features in the input layer, and generating a risk level probability distribution through the Softmax function in the output layer. The probability distribution results are divided into three risk levels: low, medium, and high. The risk warning and visualization module displays the above risk levels to users through a visualization interface. High risk: The probability value is above 0.75. Medium risk: The probability value is between 0.25 and 0.75. Low risk: The probability value is below 0.25.

[0008] Further, the risk analysis engine module establishes a risk rule library to store preset risk judgment rules, including technical field weights, regional risk coefficients, and infringement history penalty factors, and generates a comprehensive risk score by weighted fusion of the weight parameters in the risk rule library and the risk level probability distribution output by the deep learning model.

[0009] Further, the method for constructing the dynamic knowledge graph includes using the enterprise's core patents as nodes, associating competitor patents, technical standards, and open source agreements through semantic similarity, and predicting the changes in technology substitution risks and market barriers based on the analysis of patent layout trends over time.

[0010] Further, the triggering conditions for the hierarchical warning signals include: High-risk warning: Triggered when the comprehensive risk score ≥ 80 points; Medium-risk warning: Triggered when 60 points ≤ comprehensive risk score < 80 points; Low-risk warning: Triggered when the comprehensive risk score < 60 points.

[0011] Further, the risk analysis engine module classifies the patent layout trend prediction into three types of risk signals: Technology substitution risk: Calculate the remaining life cycle of the original patented technology through a survival analysis model; Market barrier change: Analyze the evolution of market concentration based on the Herfindahl index of the patent regional layout; Technology blank point: Identify technology branches with no patent applications for N consecutive years; The above analysis results are fed back to the dynamic knowledge graph in real time, manifested as changes in the weights of the associated edges between nodes. The node and associated edge data in the dynamic knowledge graph are quantitatively processed, and the quantitative results are mapped into three-level alarms: Red alarm: Node risk score ≥ 80 and associated edge weight > 0.75: Generate a risk analysis report; Yellow alarm: 60 ≤ node risk score < 80: Mark the monitoring list and update the trend analysis regularly; Blue prompt: Node risk score < 60: Archive it in the historical database for subsequent auditing; For the patent layout trend prediction results identified by the risk analysis engine module, the risk warning and visualization module displays them to the user through a visualization interface.

[0012] Further, the risk warning and visualization module generates a risk analysis report containing screenshots of infringement evidence, citations of legal bases, and suggestions for coping strategies according to the warning level, and supports output in PDF, Word, and API data interface formats.

[0013] The technical effects and advantages of the present invention: 1. This system constructs an intelligent intellectual property risk early warning system through multi-dimensional technology integration. Its beneficial effects are mainly reflected in three aspects: First, by adopting multi-source heterogeneous data fusion technology and real-time monitoring algorithms, it realizes the automatic identification and dynamic tracking of intellectual property risks such as trademarks and patents, significantly improving the timeliness and comprehensiveness of infringement discovery; Second, it innovatively combines deep learning models with multi-dimensional weighted evaluation algorithms to establish an intelligent comparison system covering multiple features such as glyphs, pronunciations, graphics, and semantics, significantly improving the accuracy and scientificity of infringement judgment; Finally, through the organic combination of knowledge graph construction and risk rule base, it forms a full-process decision support mechanism including risk rating, evidence chain generation, and countermeasure recommendation, effectively reducing the rights protection cost and management burden of enterprises. These technological innovations together constitute an efficient, accurate, and low-cost intellectual property protection solution, providing a powerful legal risk prevention and control tool for enterprises in complex market competition. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is the main system architecture diagram of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following combines the drawings and preferred embodiments, and details its specific implementation manner, structure, features, and effects as follows. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0017] The following specifically describes the specific solution provided by the present invention with reference to the drawings.

[0018] Reference Figure 1 , which shows that the present invention provides an intellectual property risk early warning system for multi-source heterogeneous data fusion: including: a multi-source data acquisition module, a data preprocessing and standardization module, a distributed data storage module, a risk analysis engine module, a risk early warning and visualization module, and a data communication module; The multi-source data acquisition module collects structured and unstructured data from intellectual property public platforms, enterprise business databases, and Internet public data sources. The data includes patent texts, trademark registration information, enterprise bidding records, product technical parameters, judicial litigation data, and market sentiment texts; The data preprocessing and standardization module cleans, converts the format, and performs semantic annotation on the collected heterogeneous data. Among them, through natural language processing technology, technical feature entities in patent claims are extracted, feature vectorization processing is performed on trademark images, and the technical requirements in bidding documents are associated and mapped with the patent technology efficacy matrix; The distributed data storage module adopts a hybrid architecture of a graph database and a relational database, and stores knowledge graph associated data and structured business data respectively; The risk analysis engine module performs feature fusion on multi-dimensional data based on a deep learning model to generate an enterprise intellectual property risk score. At the same time, a dynamic knowledge graph is constructed. By associating patent technologies, competitor patent layouts, infringement litigation cases, and policies and regulations, infringement conflicts and technology substitution risks are identified; The risk warning and visualization module triggers a hierarchical warning signal according to the risk score, and displays the risk source, associated evidence chain, and countermeasure suggestions through a visualization interface; The data communication module performs data interaction with the enterprise internal management system to realize real-time push of warning information.

[0019] In the first embodiment, when an enterprise monitors trademark warning information through the above system, the multi-source data acquisition module acquires the registration information and usage evidence of the enterprise's own trademarks, including detailed information such as trademark names, vector files of graphic LOGOs, registration categories, and registration numbers. At the same time, parameters such as the monitored geographical scope, category scope, and competitor list are configured. Subsequently, real-time monitoring starts, and the trademark office databases such as the China Trademark Network are automatically scanned daily, as well as common e-commerce platforms and social media, to capture newly applied trademarks or actual usage behaviors that may constitute infringement.

[0020] In terms of graphic trademark comparison, the data preprocessing and standardization module uses a convolutional neural network (CNN) model to extract LOGO feature vectors. The specific process is as follows: First, the LOGO image is preprocessed, including size normalization, background removal, and color space conversion; then, through the multi-layer convolution and pooling operations of the pre-trained CNN model VGG, image features from low-order to high-order are gradually extracted; finally, the feature map is obtained before the fully connected layer of the model and flattened into a feature vector. These feature vectors can effectively capture visual features such as the color distribution, shape contour, and texture pattern of the LOGO.

[0021] When comparing text trademarks, the data preprocessing and standardization module uses natural language processing (NLP) technology to calculate name similarity, converts trademark names into high-dimensional vectors through the word embedding model BERT, and then calculates the cosine similarity between vectors. For the judgment of glyph similarity, the edit distance of the two trademark names is calculated, that is, how many single-character edits, i.e. insertions, deletions or substitutions, are required to change one name to another. At the same time, the data preprocessing and standardization module analyzes category relevance in conjunction with the "Classification of Similar Goods and Services" and establishes a category association matrix, which triggers an early warning when the same or similar trademarks are detected to be registered in related categories.

[0022] In terms of trademark category correlation analysis, the data preprocessing and standardization module constructs a refined correlation matrix based on the "Classification of Similar Goods and Services". This matrix contains three levels of correlation relationships: first, the categories that are officially clearly associated, such as the correlation weight of Class 9 "Computer Software" and Class 42 "Software Development Services" is set to 0.9; second, the industry practice association, such as the association weight of clothing and leather products is 0.8; and finally, the consumption scenario association, such as the association weight of cosmetics and makeup tools is 0.6.

[0023] When enterprises conduct trademark monitoring and early warning, the risk analysis engine module uses multi-dimensional analysis technology to evaluate trademark similarity. For text trademarks or text parts in graphic trademarks, the pinyin similarity, semantic relevance, and text composition are evaluated.

[0024] For the calculation of pinyin similarity, the risk analysis engine module first converts the Chinese trademark into pinyin form. For example, "智服通" is converted to "zhifutong", and then the pinyin edit distance and phoneme matching algorithm are used for analysis. For example, the pinyin edit distance between "智服通" (zhifutong) and "智服达" (zhifuda) is 2, which means that two letters need to be replaced. The risk analysis engine module determines the degree of similarity based on the preset threshold. At the same time, the matching degree of the pinyin first letter is analyzed, such as the difference between "ZT" and "ZF" will be given a higher weight.

[0025] In terms of semantic relevance, the risk analysis engine module calculates the semantic distance between words through the pre-trained word vector model. For example, "dolphin" and "whale" are both marine creatures, and their semantic similarity reaches 0.65. For combined trademarks, the risk analysis engine module analyzes the semantic association of core words, such as "speed" and "lightning" have a strong correlation in the logistics industry.

[0026] In terms of text composition, it is divided into character shape and pronunciation, such as the combination of "康师傅" and "康帅傅" which have highly similar character shapes; in terms of pronunciation, it is analyzed whether the pronunciation of Mandarin and dialect is easy to confuse, such as "飘柔" and "飘油".

[0027] For graphic trademarks, the graphic elements mainly identify the similarity of the composition and coloring of the pattern. When the similarity reaches more than 70%, a warning will be triggered. For cross-class trademarks, the risk analysis engine module first extracts the registered classes of the trademark to be examined, and then queries the association matrix to obtain a list of associated classes. For example, when it is found that a company has registered an identifier similar to its main trademark in Class 25 in Class 18 with an association weight of 0.8, a warning will be issued. Through this multi-level similarity evaluation system, the risk analysis engine module can effectively identify various potential trademark infringement risks, including different forms of infringement such as similar glyphs, similar pronunciations, and cross-class preemptive registration.

[0028] When the risk analysis engine module determines trademark similarity, it makes multi-dimensional comparisons between the trademark to be detected and the trademarks in the database. For text trademarks, in addition to the aforementioned vector similarity and edit distance, pinyin similarity, semantic relevance, and character composition are considered; for graphic trademarks, the Euclidean distance or cosine similarity of the two logo feature vectors is calculated. The risk analysis engine module uses a weighted fusion algorithm to comprehensively calculate the comparison results of multiple dimensions such as glyph similarity, category association degree, and usage scenario to obtain the final similarity score. The risk warning and visualization module will trigger different levels of warnings according to the score.

[0029] The deep learning model training method in the risk analysis engine module includes fusing patent citation relationships, trademark similarity, bidding technical requirement matching degree, and litigation history record features in the input layer, and generating a risk level probability distribution through the Softmax function in the output layer; establishing a risk rule base to store preset risk determination rules, including technical field weights, regional risk coefficients, and infringement history penalty factors, and generating the final risk score by weighted fusion of the weight parameters in the risk rule base and the risk level probability distribution output by the deep learning model.

[0030] Taking the trademark "SunBrite" as an example, its risk levels are divided into low, medium, and high categories. The model fuses the following key features in the input layer: trademark similarity: the similarity with the existing trademark "SunBright" reaches 0.85, litigation history: the applicant has been involved in 3 trademark litigations in the past 5 years, patent citation: the applicant's patent has been cited by other technologies 12 times, and bidding matching degree: the matching degree with the technical requirements of government projects is 0.65. After these features are processed by the neural network, the output layer generates three original numerical logits, which respectively correspond to the initial scores of the three risk levels. For example, the logits output by the model are 1.2 for low risk, 0.5 for medium risk, and 2.3 for high risk.

[0031] To convert logits into a probability distribution, the Softmax function needs to be applied. The steps are as follows: Numerical stability processing: To avoid numerical overflow caused by exponential operations, subtract the maximum value of 2.3 from all logits, and the adjusted logits are [-1.1, -2.8, 0].

[0032] Exponential operation and normalization: Perform exponential operations on the adjusted logits item by item to obtain the exponential results [0.332, 0.061, 1]. Then sum all the exponential values, and the total sum is 1.393. Finally, divide each exponential value by the total sum to complete the normalization.

[0033] Probability generation: After normalization, the low-risk probability is 24%, the medium-risk probability is 4%, and the high-risk probability is 72%.

[0034] If it is necessary to reduce the misjudgment of risks, the probability distribution can be shifted towards low risks by adjusting the weights of positive features such as patent citations in the model. The above risk probabilities will be displayed to users as preliminary analysis results through the risk warning and visualization module on the visualization interface.

[0035] The risk rule library established by the risk analysis engine module will set basic weight coefficients for each evaluation dimension, and these weights will be dynamically adjusted according to trademark types and industry characteristics. Taking the word trademark "Master Kong" as an example, the weights of each dimension may be set as: glyph similarity (40%), pinyin similarity (30%), semantic relevance (20%), and category association (10%). For a graphic trademark such as the mermaid logo of Starbucks, the weight of graphic feature similarity will be increased to 50%, and the proportion of text-related dimensions will be correspondingly reduced.

[0036] In specific calculations, the risk analysis engine module adopts a phased weighting method. In the first stage, calculate the core similarity. For example, to judge the similarity between "Kangshuai Fu" and "Master Kong": First, conduct a glyph comparison, and the edit distance is measured as 1, that is, 1 Chinese character is replaced, and the similarity score is converted to 85 points; the pinyin similarity score is 90 points; the semantic relevance score is 80 points because the core character "Kang" is the same; if it appears in an associated category (such as Class 30 and Class 29), the category association score is 70 points. Multiply the scores of each dimension by their weights and then sum them to obtain the preliminary comprehensive score: 85×0.4 + 90×0.3 + 80×0.2 + 70×0.1 = 84 points.

[0037] In the second stage, introduce the usage scenario correction coefficient. If it is monitored that "Kangshuai Fu" is used on the packaging of convenience foods in the same category and the packaging design is highly similar to that of "Master Kong", the scenario coefficient of 1.2 will be triggered, and the comprehensive score will be corrected to 84×1.2 = 100.8 points, exceeding the preset high-risk threshold of 95 points. For the same trademark used only in a non-associated category (such as Class 15 musical instruments), the scenario coefficient may be only 0.3, and the final score will drop to 25.2 points, belonging to low risk.

[0038] For graphic trademarks, the risk analysis engine module will adopt more complex multi-dimensional matrix calculations. For example, it is detected that a certain beverage uses a graphic similar to the wavy pattern of Coca-Cola: the similarity of graphic features reaches 75%, the color matching degree is 80%, but it is used on clothing in Class 25 which is not related. The risk analysis engine module will assign a higher weight to graphic features (50%), the color similarity is 30%, and the category relevance only accounts for 20%. The calculated score is 75×0.5 + 80×0.3 + 20×0.2 = 71.5 points. If it is also found that this graphic is used for promotional gifts (T-shirts) in the beverage category at the same time, the cross-category usage coefficient of 1.5 will be triggered, and the final score will rise to 107.25 points. The risk warning and visualization module will generate corresponding warning levels according to the score.

[0039] The risk warning and visualization module triggers hierarchical warning signals according to the risk score, and displays the risk source, associated evidence chain and response suggestions through the visualization interface; the triggering conditions of the above hierarchical warning signals include: High-risk warning: Triggered when the comprehensive risk score ≥ 80 points; Medium-risk warning: Triggered when 60 points ≤ comprehensive risk score < 80 points; Low-risk warning: Triggered when the comprehensive risk score < 60 points.

[0040] The risk warning and visualization module can also conduct in-depth analysis, including the detection of usage scenario conflicts and the analysis of historical behavior associations. Arrange the infringement evidence according to the time line and associate relevant legal provisions. According to the risk level, provide differentiated response suggestions: for high-risk infringement acts, a template of the "Trademark Opposition Application" can be generated with one click or directly jump to the complaint page of the e-commerce platform; for medium- and low-risk acts, it is recommended to send a lawyer's letter or continue to monitor. The whole process will greatly shorten the time of traditional manual monitoring and effectively improve the accuracy of graphic comparison, providing an efficient trademark rights protection solution for enterprises.

[0041] The risk warning and visualization module generates a risk analysis report containing screenshots of infringement evidence, citations of legal bases and response strategy suggestions according to the warning level, and supports output in PDF, Word and API data interface formats. The data communication module conducts data interaction with the enterprise internal management system to realize real-time push of warning information.

[0042] In the second embodiment, the risk analysis engine module constructs a dynamic knowledge graph with the enterprise's core patents as nodes, associates competitors' patents, technical standards and open source protocols through semantic similarity, and predicts the changes of technology substitution risks and market barriers based on the time series data analysis of patent layout.

[0043] The realization of patent layout trend analysis by time series data mainly relies on the collaborative calculation of multi-dimensional time series modeling and dynamic knowledge graph. The risk analysis engine module first extracts the historical patent data in the target technology field from the patent database and constructs a patent attribute vector containing a timestamp. The vector includes dynamic indicators such as technical feature word frequency, number of claims, citation growth rate, and regional expansion speed of patents in the same family. The distributed data storage module stores these time-stamped patent data streams through a time series database.

[0044] The risk analysis engine module uses a sliding window mechanism to calculate key indicators of technological evolution on a quarterly basis: for the core patent nodes of an enterprise, it analyzes the trend of its technological influence over time, for example, the frequency of occurrence of keywords in patents in the past three years has increased by 120%; for competitor nodes, it monitors the drift of the technological direction of its patent portfolio, and predicts the technological transformation trend of an enterprise from "rule engine" to "deep learning" through the LSTM model.

[0045] The risk analysis engine module converts patent layout trend prediction into three types of risk signals: Technology substitution risk: Calculate the remaining life cycle of the original patent technology through the survival analysis model, such as predicting the probability that a speech recognition patent will be replaced by a neural network solution within 5 years; Changes in market barriers: Analysis of the evolution of market concentration based on the Herfindahl index of patent regional layout; Technology gaps: Identify technology branches that have not had patent applications for three consecutive years; These time series analysis results are fed back to the dynamic knowledge graph in real time, showing the weight changes of the associated edges between nodes. The node and associated edge data in the dynamic knowledge graph will first undergo the following quantification processing: Node risk value calculation: Each patent / trademark node generates a basic risk score through a weighted formula; Node risk score = technology similarity × 0.6 + market overlap × 0.3 + legal risk coefficient × 0.1; If a competitor's patent has 85% similarity to technology, 70% market overlap, and no litigation history, then the risk score = 0.85×0.6+0.7×0.3+0×0.1=0.81; Risk transmission of associated edges: Risk diffusion calculation is performed through the topological structure of the knowledge graph; The risks of core patent nodes will be transmitted along the "technology derivative" edge with attenuation coefficient 0.8 / jump; the associated nodes of competing enterprises will have superimposed risks through the "market conflict" edge (addition coefficient 1.2); and the dynamic adjustment of time series will impose a 20% emergency weight bonus on nodes whose risk scores have continued to rise for three consecutive months.

[0046] The risk warning and visualization module maps the quantitative results into three levels of alerts: Red Alert: Node risk score ≥ 80 and associated edge weight > 0.75: Generate a risk analysis report; Yellow Alert: 60 ≤ Node risk score < 80: Mark the monitoring list and update the trend analysis regularly; Blue Tip: Node risk score < 60: Archive to the historical database for subsequent auditing; Regarding the infringement conflicts and technology substitution risks identified by the risk analysis engine module, the risk early warning and visualization module will display them to the user through the visualization interface.

[0047] The dynamic weighting algorithm of the present invention fully considers the characteristics of different infringement forms, preventing misjudgment in a single dimension, such as determining infringement only due to the same pinyin, and accurately identifying deliberately evasive infringement behaviors, such as trademark attachment through fine-tuning of glyphs or cross-class registration. The present invention also continuously optimizes the weight parameters through machine learning and automatically adjusts the influence coefficients of each dimension based on the processing results of historical cases. The above technological innovations together constitute an efficient, accurate, and low-cost intellectual property management and protection solution, providing a powerful legal risk prevention and control tool for enterprises in the complex market competition.

[0048] The above content is only an example and illustration of the concept of the present invention. Those skilled in the art of this technology can make various modifications, supplements, or use similar methods to substitute for the specific embodiments described, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, they should fall within the protection scope of the present invention.

Claims

1. An intellectual property risk early warning system based on multi-source heterogeneous data fusion, characterized in that: It includes multi-source data acquisition module, data preprocessing and standardization module, distributed data storage module, risk analysis engine module, risk warning and visualization module and data communication module; The multi-source data collection module collects structured data and unstructured data from the intellectual property disclosure platform, enterprise business database and Internet public data sources. The data includes patent texts, trademark registration information, enterprise bidding records, product technical parameters, judicial litigation data and market public opinion texts; The data preprocessing and standardization module cleans, converts the format and semantically annotates the collected heterogeneous data, extracts the technical feature entities in the patent claims through natural language processing technology, performs feature vectorization on the trademark image, and associates and maps the technical requirements in the bidding documents with the patent technology efficacy matrix; The distributed data storage module adopts a hybrid architecture of graph database and relational database to store knowledge graph associated data and structured business data respectively; The risk analysis engine module generates an enterprise intellectual property risk score by integrating features of multi-dimensional data based on a deep learning model, and constructs a dynamic knowledge graph to identify infringement conflicts and technology substitution risks by associating patent technologies, competitor patent layouts, infringement litigation cases, and policies and regulations; The risk warning and visualization module triggers graded warning signals according to risk scores, and displays risk sources, associated evidence chains, and response suggestions through a visualization interface; The data communication module exchanges data with the enterprise's internal management system to achieve real-time push of early warning information.

2. According to claim 1, the intellectual property risk early warning system based on multi-source heterogeneous data fusion is characterized in that: The data preprocessing and standardization module further extracts technical keywords, legal status and right holder information from unstructured text data through named entity recognition, and uses a convolutional neural network to generate feature encoding for trademark image data, and performs similarity comparison with the registered trademark database.

3. The intellectual property risk early warning system based on multi-source heterogeneous data fusion according to claim 2 is characterized in that: The deep learning model training method in the risk analysis engine module includes an input layer that integrates patent citation relationships, trademark similarity, bidding technology requirements matching, and litigation history record characteristics, and an output layer that generates a risk level probability distribution through a Softmax function. The probability distribution results are divided into three risk levels: low, medium, and high. The risk warning and visualization module displays the above risk levels to users through a visualization interface; High risk: probability value is above 0.75; ‌Medium risk‌: probability value between 0.25 and 0.75; ‌Low Risk‌: Probability value is below 0.

25.

4. The intellectual property risk early warning system based on multi-source heterogeneous data fusion according to claim 3 is characterized in that: The risk analysis engine module establishes a risk rule library to store preset risk determination rules, including technical field weights, regional risk coefficients, and infringement history penalty factors, and generates a comprehensive risk score by weightedly integrating the weight parameters in the risk rule library with the risk level probability distribution output by the deep learning model.

5. The intellectual property risk early warning system based on multi-source heterogeneous data fusion according to claim 4 is characterized in that: The method for constructing the dynamic knowledge graph includes taking the core patents of the enterprise as nodes, associating competitor patents, technical standards and open source protocols through semantic similarity, and analyzing patent layout trends based on time series data to predict technology substitution risks and market barrier changes.

6. The intellectual property risk early warning system based on multi-source heterogeneous data fusion according to claim 5 is characterized in that: The triggering conditions of the graded warning signal include: High-risk warning: triggered when the comprehensive risk score is ≥80 points; Medium risk warning: triggered when 60 points ≤ comprehensive risk score < 80 points; Low risk warning: triggered when the comprehensive risk score is less than 60 points.

7. The intellectual property risk early warning system based on multi-source heterogeneous data fusion according to claim 6 is characterized in that: The risk analysis engine module divides the patent layout trend prediction into three types of risk signals: Technology substitution risk: Calculate the remaining life cycle of the original patent technology through the survival analysis model; Changes in market barriers: Analyze the evolution of market concentration based on the Herfindahl index of patent regional layout; Technology gaps: Identify technology branches that have not had patent applications for N consecutive years; The above analysis results are fed back to the dynamic knowledge graph in real time, which is manifested as the weight change of the associated edges between nodes. The node and associated edge data in the dynamic knowledge graph are quantified and the quantified results are mapped into three levels of alarms: Red alarm: Node risk score ≥ 80 and associated edge weight > 0.75: Generate a risk analysis report; Yellow alarm: 60≤node risk score<80: mark the monitoring list and update the trend analysis regularly; Blue prompt: Node risk score <60: archived to the historical database for subsequent audit; The risk warning and visualization module displays the patent layout trend prediction results identified by the risk analysis engine module to the user through a visual interface.

8. The intellectual property risk early warning system based on multi-source heterogeneous data fusion according to claim 7 is characterized in that: The risk warning and visualization module generates a risk analysis report including screenshots of infringement evidence, citations of legal basis and response strategy recommendations based on the warning level, and supports output in PDF, Word and API data interface formats.

Citation Information

Patent Citations

  • Trademark monitoring method, trademark monitoring device and electronic device

    CN112364034A

  • Knowledge graph-based enterprise risk prediction method and system

    CN108596439A

  • Intellectual property risk management integrated system and working method thereof

    CN113850530A

  • Trademark infringement detection and identification method and system

    CN118365904A

  • Patent risk early warning system driven by artificial intelligence

    CN119515615A

Cited By

  • Intellectual property value evaluation method and system

    CN120524438A

  • Multi-cycle feature fusion analysis method and system based on expired domain name

    CN120915494A

  • Enterprise risk early warning system and method based on intelligent evolution graph

    CN120952538A