Patent data acquisition method based on big data

By expanding the scope of acquisition, introducing artificial intelligence and deep learning models, combining edge computing and blockchain technology, the problems of insufficient coverage and storage efficiency of traditional patent data acquisition methods in emerging technology fields are solved, efficient and accurate patent data evaluation and analysis are achieved, and full-dimensional support is provided for innovative decision-making.

CN120256702APending Publication Date: 2025-07-04YANCHENG XINGUOHUI INTELLIGENT TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510381144.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional patented data acquisition methods are limited to conventional fields and fail to fully cover the field of emerging technology integration, resulting in one-sided value evaluation, and traditional storage architectures are difficult to process massive data, and their correlation analysis capabilities are weak, which cannot meet the needs of innovative decision-making.

Method used

Expand the collection scope to the field of integration of emerging technologies, introduce artificial intelligence and adaptive network crawling technology, combine deep learning models and edge computing, adopt blockchain to ensure data security, use knowledge graphs and visualization platforms to enhance analysis capabilities, and introduce VR/AR technology to accurately extract information by generating adversarial network repair data, and build a real-time quality assessment mechanism.

Benefits of technology

It has achieved a full-dimensional assessment of the patent value in the field of integrated emerging technologies, improved data acquisition efficiency and accuracy, reduced uncertainty in innovation activities, improved data storage and analysis efficiency, and provided forward-looking decision-making support.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to the field of patent data collection, in particular to a patent data collection method based on big data. The patent data collection method based on big data comprises the steps of S1, expanding a collection range, and increasing collection of application possibility data and potential risk data of patents in the emerging technology-technology fusion field on the basis of collection of conventional patent information, value data, influence data and industry dynamics; according to the patent data acquisition method based on the big data, application possibility data and potential risk data of the patent in the field of emerging technology fusion are additionally acquired, and particularly potential application scenes and expected effect data of fusion of the patent and at least three emerging technologies are acquired, so that the patent value can be evaluated from a brand new dimension; the method makes up for the defect that the traditional method only pays attention to the conventional field, resulting in one-sided patent value evaluation, and reduces the uncertainty and potential loss of innovation activities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of patent data collection, and in particular to a patent data collection method based on big data. Background Art

[0002] Patent data refers to the collection of structured and unstructured information related to the entire life cycle of patent application, authorization, protection, operation and expiration. It covers multiple dimensions such as law, technology, economy, and operation, and is the core basis for evaluating patent value, analyzing technology trends, and formulating intellectual property strategies. As science and technology develop at an unprecedented rapid pace, the drawbacks of traditional patent data collection methods have gradually become apparent. In terms of the scope of collection, they are overly limited to conventional fields, and there is almost no exploration of the integration of emerging technologies, resulting in a one-sided assessment of the value of patents, which cannot provide a complete basis for innovative decisions of enterprises and scientific research institutions in the field of emerging technologies. In addition, in terms of storage integration, with the explosive growth of patent data, traditional centralized storage architectures are difficult to meet the needs of rapid storage and processing of massive data. Data is prone to delays or even freezes during storage and reading, affecting the timely application of data. At the same time, traditional storage methods have weak data correlation analysis capabilities and cannot effectively mine the intrinsic connections between data. The visualization platform function is also very simple, and can only perform basic data display and lacks data prediction functions.

[0003] Therefore, it is necessary to provide a new patent data collection method based on big data to solve the above technical problems. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides a patent data collection method based on big data.

[0005] The patent data collection method based on big data provided by the present invention includes: S1: Expand the scope of collection. In addition to collecting conventional patent information, value data, influence data and industry dynamics, add application data and potential risk data of patents in the field of emerging technology integration; S2: Strengthen the collection channels, use artificial intelligence to optimize the interface with the patent office, realize intelligent data screening and push, use adaptive web crawler technology to collect multi-channel data, and use blockchain to ensure the security of industry alliance data sharing; S3: Optimize acquisition technology and introduce deep learning models to improve data classification accuracy; use the IoT to expand the monitoring system to monitor the application performance of technology, and combine VR / AR technology to accurately extract patent drawings and chart information; S4: Deepen cleaning preprocessing, use big data analysis to mine data associations, use generative adversarial networks to repair data, and establish a real-time quality assessment mechanism to ensure data quality; S5: Upgrade storage integration, introduce edge computing to improve the processing efficiency of distributed storage, utilize knowledge graphs to enhance the correlation analysis ability of graph databases, and add a data prediction function to the visualization platform.

[0006] Preferably, in step S1, for the application data in the emerging technology integration field, collect the potential application scenarios and expected effect data of the patent integrated with at least three emerging technologies.

[0007] Preferably, in step S2, the adaptive web crawler technology analyzes websites through machine learning algorithms and automatically adjusts the collection frequency and request method after accessing a certain number of pages.

[0008] Preferably, in step S3, when extracting patent drawings and chart information by combining VR and AR technologies, use spatial perception technology to accurately identify and label the three-dimensional structure information in the patent.

[0009] Preferably, in step S4, the real-time data quality assessment mechanism sets a threshold. When the data quality index is lower than the threshold, it automatically triggers the data repair process and records the quality fluctuation situation.

[0010] Preferably, in step S5, edge computing technology performs preliminary processing and screening of data at the data collection terminal, reduces the data transmission volume, and improves the storage and processing efficiency.

[0011] Preferably, regularly evaluate the data. According to the development of patent technologies, market and policy changes, adjust the collection scope, channels and technologies, and integrate the collected patent data with relevant internal and external enterprise data to form comprehensive data resources to assist decision-making.

[0012] Preferably, use homomorphic encryption technology to encrypt sensitive patent data, so that the data can still perform specific calculations and analyses in the encrypted state, and the calculation results are consistent with the plaintext calculation results after decryption.

[0013] Compared with related technologies, the patent data collection method based on big data provided by the present invention has the following beneficial effects: By increasing the collection of the application possibility data and potential risk data of patents in the emerging technology integration field, especially the potential application scenarios and expected effect data of the patent integrated with at least three emerging technologies, the patent value can be evaluated from a new dimension, which makes up for the defect that the traditional method only focuses on the conventional field and leads to one-sided evaluation of the patent value, and reduces the uncertainty and potential losses of innovation activities; Using artificial intelligence to optimize the interface with the patent office, realizing intelligent data screening and pushing, changing the traditional situation of relying on manual operations or simple batch downloads, which leads to obtaining a large amount of redundant data, achieving rapid and accurate positioning of the required patent data, saving labor and time costs, and improving the efficiency and quality of data acquisition; Introducing a deep learning model to overcome the problems of incorrect classification and category confusion of complex patent data by traditional simple algorithms, being able to accurately identify the weights and associations of different disciplinary and technical elements in patents, improving the accuracy of data classification, and laying a foundation for subsequent accurate analysis of patent technical fields; Introducing edge computing to screen and process data in advance at the data collection terminal, greatly reducing the transmission volume, solving the problems of read-write latency and jamming faced by traditional centralized storage in the face of massive patent data, and being able to efficiently support business scenarios with high real-time requirements such as patent infringement early warning analysis and emergency patent retrieval; With the help of a knowledge graph, strengthening the association analysis of massive patent data by the graph database, quickly sorting out the technical evolution context between patents, the applicant cooperation network, and the market application expansion path, and helping to deeply explore the value of patent data; The visualization platform adds a data prediction function, changing the limitation of only being able to simply display data in the past. Users can intuitively understand the potential laws, development trends, and dynamic relationships between indicators of patent data, providing a forward-looking reference for enterprise strategic planning and scientific research project layout. Detailed implementation manners

[0014] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0015] The following describes the specific implementation of the present invention in detail in conjunction with specific embodiments.

[0016] A patent data collection method based on big data, the patent data collection method based on big data includes the following steps: S1: Expand the collection scope. On the basis of collecting conventional patent information, value data, influence data, and industry trends, add the collection of application data and potential risk data of patents in the field of emerging technology integration. In step S1, for the application data in the field of emerging technology integration, collect the potential application scenarios and expected effect data of the patent integrated with at least three emerging technologies. Through the interface with the patent office database, regularly batch obtain conventional patent information, including basic information such as patent application number, applicant, invention name, patent classification number, application date, and authorization date. Use a professional intellectual property database platform to collect the value data of patents. Through academic databases and patent citation analysis tools, collect influence data such as the number of citations and the rate of other citations of patents. Pay attention to industry information platforms, professional media reports, and information released by industry associations to obtain industry trend data. Establish a special research group for emerging technologies, closely monitor the development trends in emerging technology fields such as artificial intelligence, blockchain, Internet of Things, VR / AR, etc. For the application possibility data of patents in the field of emerging technology integration, collect from the following channels: First, by retrieving professional literature, research reports, technical forums, etc. in the emerging technology field, dig out the potential application scenarios of the patent integrated with at least three emerging technologies. Second, cooperate with scientific research institutions, universities, and innovative enterprises in the industry to carry out joint research and investigations to obtain the expected effect data of the patent in emerging technology integration applications. For potential risk data, analyze the technical risks, legal risks, market risks, etc. that may be faced in the process of emerging technology integration, and collect relevant data through legal consultations, market research, technical expert evaluations, etc.; S2: Strengthen the data collection channels, optimize the interface with the patent office using artificial intelligence to achieve intelligent data screening and pushing. Adopt the adaptive web crawler technology to collect multi-channel data, and ensure the security of data sharing in the industry alliance through blockchain; In step S2, the adaptive web crawler technology analyzes websites through machine learning algorithms and automatically adjusts the collection frequency and request method after accessing a certain number of pages; Develop an intelligent interface system using natural language processing (NLP) and machine learning technologies. First, deeply analyze the data structure and field meanings in the patent office database to construct a semantic understanding model. When a user enters a retrieval requirement, the system uses NLP technology to parse the user's intention, screens out the data that meets the requirements from the massive data in the patent office through machine learning algorithms, sorts them according to indicators such as relevance and importance, and then pushes them. Regularly optimize and upgrade the interface system, adjust the parameters and algorithms of the machine learning model according to user feedback and new data features to improve the accuracy and efficiency of screening and pushing; Develop an adaptive web crawler with machine learning capabilities. When the crawler is initialized, set the basic collection frequency and request method. During the process of accessing websites, use machine learning algorithms to analyze features such as the page structure, data format, and response time of the websites. After accessing a certain number of pages, automatically adjust the collection frequency and request method according to the analysis results. For websites with a strong anti-crawler mechanism, adjust request header information, increase the request interval time and other request methods to ensure that multi-channel data can be collected stably and efficiently. At the same time, establish a data quality monitoring module to conduct real-time quality assessment on the collected data. If data quality problems are found, adjust the crawler strategy in a timely manner; Build a blockchain-based distributed data sharing platform within the industry alliance. Each alliance member encrypts the patent data to be shared and uploads it to the blockchain node. Use the consensus mechanism of blockchain to ensure the consistency and integrity of data among nodes. In terms of data access permission management, define the data access levels and operation permissions of different alliance members through smart contracts. Only authorized members can access and use specific data; S3: Optimize the acquisition technology and introduce deep learning models to improve data classification accuracy; use the Internet of Things to expand the monitoring system to monitor the application performance of the technology, and combine VR / AR technology to accurately extract patent drawings and chart information; in the S3 step, when combining VR and AR technology to extract patent drawings and chart information, use spatial perception technology to accurately identify and annotate the three-dimensional structure information in the patent; collect a large amount of classified patent data as training samples, including the patent's text description, drawings, charts and other information and the corresponding accurate classification labels, and use deep learning architectures such as convolutional neural networks (CNN), recurrent neural networks (RNN) and their variants to build a data classification model. During the model training process, the text data is represented by word vectors, and the drawings and charts are feature extracted and encoded. Through multiple iterative training, the model learns the complex relationship between the characteristics and classification of patent data. After the training is completed, the model is evaluated and optimized using test samples to ensure the accuracy of the model's classification The rate reaches a high level. In actual applications, the newly collected patent data is input into the trained deep learning model, and the model automatically outputs the classification information such as the technical field and application scenario to which the patent belongs, providing an accurate data basis for subsequent patent analysis. For patents that need to monitor the performance of technical applications, IoT devices are deployed in relevant technical application scenarios. IoT devices transmit the collected data to the data processing center through wireless communication technology, and use data analysis algorithms to analyze the data in real time to evaluate the performance of patented technologies in actual applications. At the same time, a historical data repository is established to conduct in-depth mining of long-term accumulated data, analyze the changing trend of technical performance over time, and provide data support for the improvement and optimization of patented technologies. Develop specialized VR / AR patent drawing and chart information extraction software, and integrate a spatial perception technology module into the software. When a user uses the software to open a patent drawing or chart, the software uses spatial perception technology to model and analyze the three-dimensional objects in the drawing, automatically identify key information such as the shape, size, and position relationship of the object, and accurately mark it, providing intuitive and accurate information for the understanding and analysis of patent technology; S4: Deepen the cleaning and preprocessing, use big data analysis to mine data associations, repair data with a generative adversarial network, and establish a real-time quality assessment mechanism to ensure data quality; in step S4, the real-time data quality assessment mechanism sets a threshold. When the data quality index is lower than the threshold, the data repair process is automatically triggered, and the quality fluctuation situation is recorded; use big data analysis tools and algorithms to conduct association analysis on the collected patent data. First, preprocess the patent data, perform word segmentation, stop word removal, etc. on the text data, and standardize the structured data. Then, use association rule mining algorithms to mine the association relationships between patents in terms of technology citation, applicant cooperation, market application, etc., and present the mined association relationships in a visual way, such as constructing a patent association map to facilitate users to intuitively understand the internal connections between patent data; Construct a generative adversarial network (GAN) model for data repair. The model consists of a generator and a discriminator. The role of the generator is to generate possible repair data based on the characteristics of the existing data, and the discriminator is used to judge whether the generated data is real and reasonable. Collect patent data containing problems such as missing values and error values as training samples, and train the GAN model. During the training process, the generator continuously adjusts the parameters of the generated data so that the generated data can pass the judgment of the discriminator, and the discriminator continuously improves its ability to distinguish the generated data. After the training is completed, input the patent data with data problems into the generator, and the generator outputs the repaired data; Establish a real-time data quality assessment system, define a series of data quality indicators, such as data integrity, data accuracy, data consistency, etc., and set reasonable thresholds for each indicator. In each link of data collection, cleaning, and preprocessing, collect real-time information related to data quality, calculate the data quality indicator values. When the data quality indicator is lower than the threshold, the system automatically triggers the data repair process, calls data repair tools such as generative adversarial networks to repair the data. At the same time, the system records the data quality fluctuation situation, such as the change trend of the data quality indicator, the number of times and reasons for triggering the repair process, etc., for subsequent analysis and optimization of data quality; S5: Upgrade storage integration, introduce edge computing to improve the processing efficiency of distributed storage, enhance the association analysis ability of graph databases using knowledge graphs, and add a data prediction function to the visualization platform; In step S5, edge computing technology performs preliminary processing and screening of data at the data collection terminal, reducing the amount of data transmitted and improving storage and processing efficiency; Deploy edge computing devices at the data collection terminal. The edge computing devices have data processing and storage capabilities. After collecting data, they first perform preliminary processing and screening on the data, and then transmit the preliminarily processed and screened data to the distributed storage system, reducing the amount of data transmitted and improving storage and processing efficiency. At the same time, the edge computing devices communicate with the central server and adjust the data processing and screening strategies in a timely manner according to the instructions and updated algorithms of the central server; Construct a patent knowledge graph. First, perform entity recognition and relationship extraction on patent data. Use natural language processing technology to identify entities in patents, such as patents, applicants, inventors, technical fields, application scenarios, etc. By analyzing the logical relationships between patent texts and data, extract the association relationships between entities, and store the identified entities and relationships in a graph database to construct a patent knowledge graph. In practical applications, use knowledge graph analysis tools to perform association analysis on patent data to help deeply explore the value of patent data; Develop a patent data visualization platform and integrate a data prediction module into the platform. The data prediction module uses time series analysis and machine learning prediction algorithms to perform prediction analysis on patent data. Use machine learning prediction algorithms to predict the market value, application prospects, etc. of patents based on factors such as the technical characteristics of patents and market demand data, and display the prediction results in a visual manner on the platform, such as presenting them in the form of trend charts, prediction reports, etc. Users can intuitively understand the potential laws and future development trends of patent data, providing forward-looking references for corporate strategic planning, scientific research project layout, etc.; Regularly evaluate data, adjust the collection scope, channels and technologies according to the development of patent technologies, market and policy changes, integrate the collected patent data with relevant internal and external data of the enterprise to form comprehensive data resources to assist decision-making; construct a monitoring model for patent technology hotspots, gather multi-source data such as patent applications, citations, academic research and market demand, and analyze the monthly increase in the number of patent applications, citation frequency, academic attention and market demand heat in each technical field in real time. When the monthly increase in the number of patent applications in the hot field exceeds 20%, or the weekly increase in the number of mentions on social media exceeds 50%, start the adjustment process. First, evaluate the relevance between the hot field and the existing collection scope. If it belongs to the integration of emerging technologies or has strategic significance, immediately include it in the collection scope. Subsequently, optimize the collection channels and technologies according to the data characteristics and sources of the hot field. For high classification accuracy requirements, optimize the deep learning model and increase the training samples. Comprehensively evaluate the collected data every quarter, and continuously optimize the collection strategy according to the changes in patent technologies, market and policies; build an internal and external data integration platform for the enterprise. When integrating, first clean and standardize the data from different sources to ensure the unity of format and semantics, then use data correlation analysis to establish connections between data. Finally, rely on the integrated comprehensive data to provide all-round support for enterprise decision-making and help formulate new product R & D plans and market competition strategies; Use homomorphic encryption technology to encrypt sensitive patent data, so that specific calculations and analyses can still be performed on the encrypted data, and the calculation results after decryption are the same as those of plaintext calculations; homomorphic encryption and access control are adopted throughout the data processing process. When collecting sensitive patent data, use homomorphic encryption algorithms to encrypt at the collection device end to ensure transmission security. During transmission, combine the SSL / TLS encryption communication protocol to prevent data theft and tampering. In the storage link, store the encrypted data in a database with access control, set different access levels according to user identity and permissions. Senior and core technical personnel can access all sensitive data, while ordinary employees can only access non-sensitive data related to their work. In the data analysis stage, use tools that support homomorphic encryption calculations to perform operations such as statistical analysis and data mining on the encrypted data to ensure that the calculation results after decryption are the same as those of plaintext. Regularly evaluate and update privacy protection measures, and adjust encryption algorithms and access control strategies according to new security threats and technological developments to ensure data privacy and security.

[0017] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A method for collecting patent data based on big data, characterized in that The following steps are involved: S1: Expand the scope of collection. In addition to collecting conventional patent information, value data, influence data and industry dynamics, add application data and potential risk data of patents in the field of emerging technology integration; S2: Strengthen the collection channels, use artificial intelligence to optimize the interface with the patent office, realize intelligent data screening and push, use adaptive web crawler technology to collect multi-channel data, and use blockchain to ensure the security of industry alliance data sharing; S3: Optimize acquisition technology and introduce deep learning models to improve data classification accuracy; Use the IoT to expand the monitoring system to monitor the performance of technology applications, and combine VR / AR technology to accurately extract patent drawings and chart information; S4: Deepen cleaning preprocessing, use big data analysis to mine data associations, use generative adversarial networks to repair data, and establish a real-time quality assessment mechanism to ensure data quality; S5: Upgrade storage integration, introduce edge computing to improve distributed storage processing efficiency, use knowledge graphs to enhance graph database association analysis capabilities, and add data prediction functions to the visualization platform.

2. The method for collecting patent data based on big data according to claim 1, wherein In step S1, for application data in the field of emerging technology integration, data on potential application scenarios and expected effects of the integration of patents with at least three emerging technologies are collected.

3. The method for collecting patent data based on big data according to claim 1, wherein In step S2, the adaptive web crawler technology analyzes the website through machine learning algorithms and automatically adjusts the collection frequency and request method after visiting a certain number of pages.

4. The method for collecting patent data based on big data according to claim 1, wherein In step S3, when VR and AR technologies are combined to extract patent drawings and chart information, spatial perception technology is used to accurately identify and annotate the three-dimensional structure information in the patent.

5. The method for collecting patent data based on big data according to claim 1, wherein In step S4, the real-time data quality assessment mechanism sets a threshold. When the data quality indicator is lower than the threshold, the data repair process is automatically triggered and the quality fluctuation is recorded.

6. The method for collecting patent data based on big data according to claim 1, wherein In step S5, edge computing technology performs preliminary processing and screening of data at the data collection terminal, reducing the amount of data transmission and improving storage and processing efficiency.

7. The method for collecting patent data based on big data according to claim 1, wherein Evaluate data regularly, adjust collection scope, channels and technologies based on patent technology development, market and policy changes, integrate collected patent data with relevant internal and external enterprise data to form a comprehensive data resource to assist decision making.

8. The method for collecting patent data based on big data according to claim 1, wherein Homomorphic encryption technology is used to encrypt sensitive patent data, so that specific calculations and analyses can still be performed on the data in an encrypted state, and the calculation results are consistent with the plaintext calculation results after decryption.

Citation Information

Patent Citations

  • Picture content display method and device

    CN108572772A

  • Intellectual property big data information service platform

    CN111626694A

  • Data monitoring method based on audio and video fusion of smart multimedia management system

    CN118155140A

  • Patent risk early warning system driven by artificial intelligence

    CN119515615A