Patent data acquisition method and device and electronic equipment

By acquiring and analyzing patent data during corporate R&D, extracting demand keywords and using correlation algorithms to screen patents, the problem of companies not considering user needs and R&D deviations is solved, achieving more effective R&D guidance and resource utilization.

CN120705291APending Publication Date: 2025-09-26STATE GRID INFORMATION & TELECOMM BRANCH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510646424.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

During the R&D process, the company fails to fully consider the actual needs of users and the R&D status of other related companies, resulting in waste of R&D resources and deviation from R&D priorities.

Method used

By obtaining the technical requirement data of the target application scenario, extracting the demand keywords, searching the patent database and determining the matching degree using the preset correlation algorithm, we can screen out the patent data that has a high degree of matching with user needs and form a target patent data set.

Benefits of technology

Help enterprises understand user needs and the R&D status of other enterprises, guide R&D direction, improve R&D efficiency and market competitiveness, and avoid waste of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705291A_ABST
    Figure CN120705291A_ABST
Patent Text Reader

Abstract

The invention provides a patent data acquisition method and device and electronic equipment, and the method comprises the steps: obtaining a plurality of pieces of technical demand data in a target application scene, and extracting a demand keyword of the technical demand data for each piece of technical demand data, calling a patent database by utilizing the technical requirement keyword to search to obtain data of a plurality of patents for the target application scene; for any one piece of technical demand data, determining the matching degree between the data of the plurality of patents and the piece of technical demand data by using a preset correlation algorithm to obtain a matching degree peak value; the peak value of the matching degree is the maximum value of the matching degree of the data of the plurality of patents and the technical demand data; determining the technical demand data of which the corresponding matching degree peak value is greater than a first preset threshold value in the plurality of pieces of technical demand data as target technical demand data; and for any target technical demand data, determining a target patent data set from the plurality of patent data by using the patent data corresponding to the matching degree peak value of the target technical demand data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a patent data acquisition method, device, and electronic device. Background Art

[0002] Currently, when conducting R&D, companies usually make iterative updates to existing products or technologies based on the R&D capabilities of their internal R&D personnel.

[0003] However, this R&D approach does not fully consider the actual needs of users and does not understand the R&D status of other related companies on the market, which can easily lead to waste of R&D resources and deviation from R&D priorities. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a patent data acquisition method, device and electronic device.

[0005] According to one aspect of the present disclosure, a method for acquiring patent data is provided, comprising:

[0006] Acquire several pieces of technical requirement data in the target application scenario, extract the requirement keywords of each piece of technical requirement data, and obtain a set of technical requirement keywords corresponding to the target application scenario;

[0007] Using the technical requirement keyword set to search the patent database, several patent data for the target application scenario are obtained;

[0008] For any piece of technical requirement data, a preset correlation algorithm is used to determine the matching degree between the plurality of patent data and the technical requirement data, and obtain a matching degree peak value; the matching degree peak value is the maximum value of the matching degree between the plurality of patent data and the technical requirement data;

[0009] Determining, among the plurality of pieces of technical requirement data, technical requirement data corresponding to a matching degree peak value greater than a first preset threshold as target technical requirement data;

[0010] For any target technology requirement data, the target patent data set is obtained from the plurality of patent data using the patent data corresponding to the peak value of the matching degree.

[0011] According to a second aspect of the present application, a patent data acquisition device is disclosed, comprising:

[0012] An acquisition module is used to acquire a plurality of technical requirement data in a target application scenario, extract the requirement keywords of each technical requirement data, and obtain a set of technical requirement keywords corresponding to the target application scenario;

[0013] A search module is used to use a set of technical requirement keywords to call a patent database to search for a number of patent data for the target application scenario;

[0014] a matching module configured to determine, for any piece of technical requirement data, a degree of matching between the plurality of patent data and the piece of technical requirement data using a preset correlation algorithm, and obtain a peak matching degree; the peak matching degree being the maximum value of the degree of matching between the plurality of patent data and the piece of technical requirement data; and determining, among the plurality of technical requirement data, technical requirement data whose corresponding peak matching degree is greater than a first preset threshold as target technical requirement data;

[0015] The determination module is used to obtain a target patent data set from the plurality of patent data using the patent data corresponding to the peak matching degree of any target technical requirement data.

[0016] According to a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program stored in the memory to implement the steps in the method described in the first aspect above.

[0017] According to a fourth aspect of the present disclosure, a computer program is provided, comprising computer instructions, which implement the steps of the method described in the first aspect when executed by a processor.

[0018] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method described in the first aspect are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A flowchart of a patent data acquisition method provided in accordance with an embodiment of the present disclosure;

[0020] Figure 2 A schematic diagram of a process flow of a patent data acquisition device provided in accordance with an embodiment of the present disclosure;

[0021] Figure 3 A schematic structural diagram of an electronic device provided in one embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] Before introducing the embodiments of the present disclosure, it should be noted that:

[0023] Some embodiments of the present disclosure are described as processing flows. Although the various operation steps of the flow may be given sequential step numbers, the operation steps therein may be implemented in parallel, concurrently, or simultaneously.

[0024] In the embodiments of the present disclosure, the terms "first", "second", etc. may be used to describe various features, but these features should not be limited by these terms. These terms are used only to distinguish one feature from another.

[0025] The term “and / or” may be used in embodiments of the present disclosure. “And / or” includes any and all combinations of one or more of the listed associated features.

[0026] It should be understood that when describing the connection relationship or communication relationship between two components, unless it is explicitly stated that the two components are directly connected or directly communicating, the connection or communication between the two components can be understood as direct connection or communication, or as indirect connection or communication through an intermediate component.

[0027] In order to make the technical solutions and advantages of the embodiments of the present disclosure more clearly understood, the exemplary embodiments of the present disclosure are further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, and are not an exhaustive list of all the embodiments. It should be noted that the embodiments and features in the embodiments of the present disclosure can be combined with each other unless they conflict.

[0028] Currently, when conducting R&D, companies typically rely on the R&D capabilities of their internal R&D personnel to iterate and update existing products or technologies. For example, R&D personnel may optimize certain performance aspects of existing products based on newly learned technologies. However, this R&D approach does not take the actual needs of users into consideration, which can easily lead to newly developed content having no practical application value and wasting R&D resources. In addition, after obtaining technical requirements, R&D activities cannot be effectively carried out without R&D ideas. Therefore, it is possible to obtain certain R&D ideas by understanding the R&D status of other related companies. At the same time, if the R&D status of other related companies is not clear, it is easy to blindly develop new content that lacks market competitiveness.

[0029] In summary, after fully understanding the actual needs of users and the R&D status of other related companies, it is very necessary to conduct research and development based on the understood situation.

[0030] An enterprise's patent data can reflect the enterprise's recent R&D direction and content. Therefore, by obtaining and viewing the patent data of related enterprises, you can understand the R&D status of related enterprises, thereby assisting your own R&D.

[0031] Based on the above problems, this application proposes a patent data acquisition method. This method starts from the actual needs of users and selects patent data that meets the actual needs of users from the patent database, so that R&D personnel can understand the R&D content made by other related companies based on the actual needs of users based on the acquired patent data, and assist their own R&D.

[0032] Specifically, such as Figure 1 As shown, a patent data acquisition method proposed in this application includes:

[0033] S101, obtaining a plurality of technical requirement data in a target application scenario, extracting the requirement keywords of each technical requirement data, and obtaining a set of technical requirement keywords corresponding to the target application scenario;

[0034] S102, using the required keyword set to search the patent database to obtain a number of patent data for the target application scenario;

[0035] S103, for any piece of technical requirement data, using a preset correlation algorithm to determine the matching degree between the plurality of patent data and the piece of technical requirement data, and obtain a matching degree peak; wherein the matching degree peak is the maximum matching degree between the plurality of patent data and the piece of technical requirement data;

[0036] S104: determining, among the plurality of pieces of technical requirement data, technical requirement data whose corresponding matching degree peak value is greater than a first preset threshold as target technical requirement data;

[0037] S105 , for any target technical requirement data, using the patent data corresponding to the matching degree peak, obtaining a target patent data set from the plurality of patent data.

[0038] By adopting the above method, after obtaining the user's technical needs for the target application scenario, the technical needs keywords are extracted from the technical needs to form a technical need keyword set. Based on the technical need keyword set, all patents in the patent database related to the technical needs in the target application scenario can be screened. Furthermore, for each piece of technical need data, a preset correlation algorithm is used to determine the matching degree between each patent data and the technical need data, and a matching degree peak is obtained. The target technical need data is screened based on the matching degree peak. The screened target technical need data is the technical need data that the existing patents are likely to solve. Then, based on the target technical need data, the patent data corresponding to its matching degree peak is used to determine from several patent data that the patent data in the target patent data set are all: patent data that are likely to solve certain actual technical needs of the user. After obtaining the target patent data set, R&D personnel can guide their own R&D content based on the patent data.

[0039] Each step in the patent data acquisition method proposed in this application is described in detail below.

[0040] In S101 above, the target application scenario can be any application scenario faced by the enterprise's products or technologies. When obtaining the technical requirement data for the target application scenario, the technical requirement data that has been previously collected and organized can be directly obtained from a database. In addition, considering that the collection and organization of technical requirement data by corresponding staff is inefficient, in one embodiment, several pieces of technical requirement data for the target application scenario can also be obtained in the following manner.

[0041] Specifically, it can be to obtain several pieces of research data for the target application scenario; wherein the research data can be questionnaires filled out by users, opinions or suggestions put forward by users on public platforms such as forums and official accounts.

[0042] For any survey data, extract the characteristic value corresponding to the characteristic variable in the survey data, input the obtained characteristic value into the pre-trained demand classification model, and obtain the technical demand type corresponding to the survey data; wherein the demand classification model is used to process the input characteristic value to obtain the technical demand type corresponding to the input data; wherein, the demand classification model can be implemented using various classification algorithms, for example, it can be implemented using the random forest algorithm. Random forest is an integrated learning method that can be used for classification or regression tasks. It is a model composed of multiple decision trees. The accuracy and stability of the model are improved by integrating the prediction results of these decision trees. This application does not limit the number and structure of decision trees in the random forest model.

[0043] For example, the survey data is questionnaire data about users' opinions on a certain mobile phone product. The characteristic variables include: age, income, satisfaction, frequency, difficulty of operation, whether it is stuck, charging time, heat dissipation effect, etc. The characteristic values ​​of the characteristic variables include numerical and categorical types. For example, age and income are numerical types, and difficulty of operation and whether it is stuck are categorical types. The characteristic values ​​of the extracted characteristic variables are input into the demand classification model. The technical requirement types output by the demand classification model can include, for example, technical requirement types such as reducing stuck, improving heat dissipation, and improving processing efficiency. In addition, when extracting characteristic variables, data cleaning can be performed, such as directly deleting survey data with a large number of missing values, deleting survey data including obvious outliers, and then performing feature coding and data standardization to finally obtain the extracted characteristic values.

[0044] Using the above method, the demand classification model outputs the technical demand type of each survey data, which can be shown in Table 1.

[0045] Survey data Type of technical requirements 1 Improve heat dissipation 2 Reduce lag ...... ...... N Increase charging speed

[0046] Table 1

[0047] After obtaining all technical requirement types for a number of survey data, the number of requirements corresponding to each technical requirement type is counted. The technical requirement type with a number of requirements greater than a statistical threshold is determined as the target technical requirement type, and the survey data corresponding to the target technical requirement type is used as the target technical requirement data. All target technical requirement data is obtained as several pieces of technical requirement data for the target application scenario.

[0048] For example, the number of demands corresponding to each type of technical demand is shown in Table 2. The number of demands is sorted from top to bottom, and the preset threshold is 500. The target technical demand data is determined to be the survey data corresponding to "reducing lag" and "increasing charging speed".

[0049] Type of technical requirements quantity Reduce lag 1000 Increase charging speed 510 Improve heat dissipation 380 ...... ......

[0050] Table 2

[0051] By using the above method, based on the technical requirement types of statistical survey data, and by analyzing and sorting the survey data, we can select call data that can reflect the actual technical requirements of the majority of users as the technical requirement data for the subsequent determination of the target patent data set. This avoids the need to analyze and screen all survey data, and also avoids the need to manually organize and screen technical requirement data, thereby improving the overall efficiency of patent data screening. It is understandable that the above explanation uses questionnaires as an example, and the above method can be used to process other forms of survey data.

[0052] In the above S101, the requirement keywords of each technical requirement data are extracted respectively. Specifically, for each technical requirement data, the requirement data is segmented to obtain the segmentation result; the word frequency of each segmentation is determined, and the segmentation and its corresponding word frequency are stored to obtain a set of technical requirement keywords.

[0053] For example, for a piece of technical requirement data, "The mobile phone has poor heat dissipation and needs to improve its heat dissipation function. The battery performance is poor," the segmentation results include "mobile phone, heat dissipation, improvement, battery, performance," etc. After segmenting all technical requirement data, the segmentation results for all technical requirement data are obtained. The frequency of each segmentation is then counted, and the segmentation and its corresponding frequency are stored to obtain a set of technical requirement keywords, as shown in Table 3. The segmentations in Table 3 are arranged in reverse order from top to bottom according to the frequency.

[0054] Participle word frequency performance 499 Battery 400 heat dissipation 100 ...... ......

[0055] Table 3

[0056] In S102 above, a search formula can be constructed using the technical requirement keywords in the technical requirement keyword set according to the search rules in the called patent database, thereby retrieving a number of patent data items for the target application scenario from the patent database based on the constructed search formula. Alternatively, a search statement can be constructed using the technical requirement keywords in the technical requirement keyword set, and a number of patent data items for the target application scenario can be retrieved based on the constructed search statement and the semantic search function of the patent database.

[0057] It should be noted here that when using the technical requirement keyword set, all of the technical requirement keywords can be used, or some of the technical requirement keywords can be used. When using some of the technical requirement keywords, a preset number of technical requirement keywords with the largest word frequency can be selected based on the word frequency in the technical requirement keyword set, or technical requirement keywords with a word frequency greater than a preset value can be selected. If all of the technical requirement keywords in the technical requirement keyword set are used, the patent data searched from the patent database can be made more complete. If some of the technical requirement keywords in the technical requirement keyword set are used, the patent data searched from the patent database can be made more in line with the technical requirements. Those skilled in the art can select the corresponding technical solution based on the actual situation.

[0058] After searching and obtaining several patent data for the target application scenario, in the above S103-S104, the target technical requirement data can be determined for any technical requirement data and the searched several patent data.

[0059] It should be noted that users have a variety of needs, some of which can be solved by technologies disclosed in existing patent data, while others cannot. In order to assist R&D personnel in their own research and development, it is necessary to first screen out technical needs that can be solved by technologies disclosed in existing patent data, and then find all the patent data corresponding to these technical needs.

[0060] Therefore, in S103-S104, it is necessary to first select technical requirements that can be solved with a high probability by the technologies disclosed in the existing patent data.

[0061] Among them, the above S103 can be for any technical requirement data, determining at least one morpheme of the technical requirement data; wherein, the morpheme can directly use the keywords extracted in S101, or can be extracted by using other technologies, which is not limited in this application. Calculate the correlation score of each morpheme of the technical requirement data with any patent data; wherein the correlation score is negatively correlated with the frequency of the morpheme appearing in all patent data, and the correlation score is positively correlated with the frequency of the morpheme appearing in the patent data; perform weighted summation on the correlation scores of all morphemes with the patent data to obtain the matching degree between the technical requirement data and the patent data; after obtaining the total matching degree between the technical requirement data and several patent data, determine the maximum value of the matching degree between several patent data and the technical requirement data as the matching degree peak.

[0062] Specifically, the following formula (1) can be used to calculate the matching degree between technology demand data and patent data:

[0063]

[0064] Among them, x is the data of each patent, m is the technical requirements, and m i is the technical requirement morpheme, C x,m It's the degree of matching.

[0065] wi is defined as IDF, specifically as formula (2):

[0066]

[0067] N is a number of patent data for the target application scenario, n(m i ) is the number of patents containing morpheme mi. According to the definition of IDF, for a given patent supply set, the number of patents containing morpheme m is i The more patents a company has, the more m i The more common, i The lower the weight, the i The less important it is when calculating the score.

[0068] m i The similarity score R(m i ,x), refer to formula (3) and formula (4).

[0069]

[0070] Among them, k1, k2, and c are modifiers, and fi represents the required morpheme m in the patent document. i The number of occurrences of mf iis the number of occurrences of each morpheme in the technical requirements, xl is the length of patent document x, and avgxl is the average length of all patent documents. Since morphemes in technical requirements usually appear only once, mf i can be 1, then R(m i ,x) has a right factor of 1, and the formula can be simplified to:

[0071]

[0072] Therefore, the semantic matching degree C between the technical requirement m and any patent data x x,m The formula can be simplified to:

[0073]

[0074] Using the above method, the matching degree between the technical requirement data and the patent data can be calculated. After determining the matching degree between the technical requirement data and each patent data, the maximum matching degree between all the patent data and the technical requirement data can be determined as the matching degree peak value. Then, the technical requirement data with corresponding matching degree peak values ​​greater than the first preset threshold value among several technical requirement data can be determined as the target technical requirement data.

[0075] As shown in Table 4, for example, a total of 4 pieces of technical requirement data are included.

[0076] Technical requirements data Matching peak 1 0.5 2 0.1 3 0.2 4 0.8

[0077] Table 4

[0078] If the peak matching degree of technical requirement data 1 is 0.5, the peak matching degree of technical requirement data 2 is 0.1, the peak matching degree of technical requirement data 3 is 0.2, and the peak matching degree of technical requirement data 4 is 0.8, and if the first preset threshold is 0.4, then technical requirement data 1 and technical requirement data 4 are determined as target technical requirement data. It should be understood that the above description uses four technical requirement data as an example for simplicity. In actual applications, there are usually many technical requirements, and the above method can be used to process them all.

[0079] In this way, based on the morphemes and relevance algorithms of the technical requirement data, target technical requirement data that are likely to be solved by existing disclosed patents can be screened out. In this way, finding patent data corresponding to the target technical requirement data in S105 can truly assist one's own R&D needs.

[0080] The above-mentioned S105 is described below.

[0081] The above S105 may specifically be to obtain, for any target technical requirement data, patent data corresponding to its matching degree peak value as the first patent data;

[0082] Determining the similarity between the first patent data and other patent data in the plurality of patent data based on a preset similarity algorithm;

[0083] Patent data with similarity greater than a preset similarity threshold is determined as a target patent data set of target technology requirement data.

[0084] Since the target technical requirement data is screened out in S103-S104, and the patent data corresponding to its matching degree peak is determined, and the first patent data corresponding to the matching degree peak is the patent data that can solve the target technical requirement with a high probability, therefore, if you want to find other patent data that can solve the target technical requirement data, you can determine the similarity between the first patent data and other patent data in several patent data based on a preset similarity algorithm. If the similarity between some patents and the first patent data is greater than the preset similarity threshold, it means that these patents also have a probability of solving the target technical requirement data, so these patents are used as the target patent data set of the target technical requirement data.

[0085] Furthermore, in order to ensure that the target patent data set is as closely aligned with the target technology requirement data as possible, in the above embodiment, patent data having a similarity greater than a preset similarity threshold and a matching degree with the target technology requirement data greater than a second preset threshold may be selected as the target patent data set; wherein the second preset threshold is not greater than the first preset threshold. In other words, the selected patent data not only has a high similarity with the first patent, but also has a high relevance to the target technology requirement.

[0086] Among them, the preset similarity algorithm can be to obtain all the text data of the first patent data and all the text data of several other patent data, and use text similarity algorithms such as cosine similarity, Jaccard similarity, Word2Ve, etc. to determine the similarity between the first patent data and other patent data in the several patent data.

[0087] In addition, in order to reduce the overall computational complexity of similarity calculation as much as possible, the following method may be used to implement the preset similarity algorithm.

[0088] Obtain the first patent data and the text data of the claims, title, and abstract of the patent to be calculated for similarity; specifically, identify the identifiers of different parts in the patent text data, split the patent data into parts such as the title, abstract, claims, description, and drawings, and then extract the text data of the identified claims, title, and abstract.

[0089] Calculate the similarity between the first patent data and the claim text data and abstract text data of the patent to be similarity calculated using a first similarity algorithm to obtain a first similarity and a second similarity;

[0090] Using a second similarity algorithm to calculate the first patent data and the title data of the patent to be similarity calculated, to obtain a third similarity; wherein the computational complexity of the second similarity algorithm is less than that of the first similarity algorithm;

[0091] The first similarity, the second similarity, and the third similarity are weighted and summed to obtain the final similarity of the first patent data and the patent to be calculated for similarity; the first similarity has the highest weight and the third similarity has the lowest weight.

[0092] Since the title text data is small and contains less semantics, a second similarity algorithm with less computational complexity and less reliance on semantics can be directly used. For example, algorithms such as Jaccard are used to calculate the similarity between the first patent data and the title data of the patent to be similarity calculated to obtain the third similarity.

[0093] The abstract typically includes technical effects and core technical solutions, while the claims include all technical solutions to be protected by the patent. Therefore, when comparing similarities, there is a high reliance on semantics. Therefore, a first similarity algorithm with a large computational load and the ability to take semantics into account can be used to calculate the similarity between the first patent data and the claim text data and abstract text data of the patent to be similarity calculated, thereby obtaining a first similarity and a second similarity. The first similarity algorithm can be implemented using algorithms such as the CDSSM model, the MV-LSTM model, and the ARC-I model, and this application does not limit this.

[0094] When taking the weighted sum of the first, second, and third similarities, the weight of the first similarity is set to the highest because the claims include all the technical solutions to be protected by the patent, and the content and information content in the title are the least, so the weight of the third similarity can be set to the lowest. The final similarity comprehensively considers the similarities of the title, abstract, and claims, and takes the similarity of the claims as the maximum reference value.

[0095] Using the above method, on the one hand, the title, abstract and claims usually include the technical problems, technical effects and complete technical solutions that the patent mainly solves. Therefore, the similarity between patent texts can be effectively compared by using the text data of the above three parts for calculation. At the same time, only the text data of the title, abstract and claims are used to calculate the similarity, which avoids using the full amount of text data in the patent text as data for similarity calculation, thereby reducing the amount of text data required for a large amount of calculation. On the other hand, using a similarity algorithm with a smaller amount of calculation and no need to consider semantics for the title, and using a similarity algorithm with a larger amount of calculation and taking semantics into consideration for the abstract and claims can further reduce the consumption of unnecessary computing resources. The contributions of the above two aspects can reduce resource consumption and improve computing efficiency as a whole.

[0096] In one embodiment, when determining the target technical requirement data, technical requirement data that is likely to be solved by the technology disclosed by the patent can be screened out based on the existing patent matching degree peak. However, if there are many patents that solve a certain technical requirement, it means that the existing technology for this technical requirement is oversaturated. If research and development is still carried out for this technical requirement, it is also a waste of resources. Therefore, although there are corresponding patents, technical requirements that are not yet saturated by the overall technology disclosed by the patent can be further screened out, and the technical requirement data corresponding to such technical requirements can be determined as the target technical requirement data.

[0097] Based on this, this application proposes that any technical requirement data is also marked with a technical requirement type, wherein the technical requirement type can be the one mentioned above and obtained in advance through the demand classification model. For any technical requirement data, the demand intensity of the technical requirement data and the technical supply intensity for the technical requirement data are determined according to its technical requirement type; wherein the technical requirement intensity is obtained according to the number of technical requirement data included in the technical requirement type; and the technical supply intensity is obtained according to the number of patent data whose matching degree with the technical requirement is greater than a first preset threshold.

[0098] In S104 above, the technical requirement data whose corresponding matching degree peak value is greater than the first preset threshold value among several technical requirement data are determined as target technical requirement data. Specifically, the technical requirement data whose demand intensity is greater than the technology supply intensity and whose corresponding matching degree peak value is greater than the first preset threshold value are determined as target technical requirement data.

[0099] Among them, in order to avoid scattered values, the technology demand intensity and technology supply intensity were standardized.

[0100] Technology supply intensity P of technology demand m m The calculation formula is formula (7):

[0101]

[0102] Among them, Nm refers to the total amount of available technology supply to meet technology requirement m. In this application, the number of patent data whose matching degree with technology requirement m is greater than a first preset threshold can be determined as the total amount of available technology supply for technology requirement m. N , Min N is the maximum and minimum value of patent supply for all target technology demand data N = {N1, N2..., Nk}, P m The range is 0 to 1.

[0103] Technical requirement intensity F of technical requirement m m The calculation formula is:

[0104]

[0105] Among them, Md refers to the total number of technical requirements under the technical requirement type corresponding to the technical requirement m. m , Min m It is the maximum and minimum value of the number of technical requirements corresponding to all technical requirement types in the total technical requirement set. The range of F is 0 to 1.

[0106] If F m Greater than P m , indicating that the intensity of technology demand is greater than the intensity of technology supply. If F m Less than P m , indicating that the intensity of technology supply is greater than the intensity of technology demand.

[0107] By selecting technology demand data whose demand intensity is greater than the technology supply intensity and whose corresponding matching degree peak is greater than the first preset threshold as target technology demand data, technology demand data that has corresponding patents but for which the overall disclosed technology of the patents is not yet saturated can be screened out as target technology demand data. Subsequently, the final patent data set can be determined based on this target technology demand data, which can further guide research and development for technology demands that have not yet been oversaturated.

[0108] Based on the same inventive concept, this application also proposes a patented data acquisition device, which is as follows Figure 2 The device includes:

[0109] An acquisition module 210 is configured to acquire a plurality of technical requirement data in a target application scenario, extract the requirement keywords of each technical requirement data, and obtain a set of technical requirement keywords corresponding to the target application scenario;

[0110] Search module 220, configured to use a set of technical requirement keywords to search a patent database to obtain a number of patent data for the target application scenario;

[0111] Matching module 230 is configured to determine, for any piece of technical requirement data, a degree of matching between the plurality of patent data and the piece of technical requirement data using a preset correlation algorithm, and obtain a peak matching degree; the peak matching degree being the maximum value of the degree of matching between the plurality of patent data and the piece of technical requirement data; and determine, among the plurality of technical requirement data, any technical requirement data whose corresponding peak matching degree is greater than a first preset threshold as target technical requirement data;

[0112] The determination module 240 is configured to obtain a target patent data set from the plurality of patent data using the patent data corresponding to the peak matching degree of any target technical requirement data.

[0113] In one embodiment, the acquisition module 210 is specifically configured to acquire a plurality of survey data for the target application scenario;

[0114] For any piece of survey data, extract the characteristic value corresponding to the characteristic variable in the survey data, input the obtained characteristic value into a pre-trained demand classification model to obtain the technical demand type corresponding to the survey data; the demand classification model is used to process the input characteristic value to obtain the technical demand type corresponding to the input data;

[0115] After obtaining all technical requirement types of the plurality of survey data, counting the number of requirements corresponding to each technical requirement type, determining the technical requirement type whose number of requirements is greater than a statistical threshold as the target technical requirement type, and using the survey data corresponding to the target technical requirement type as the target technical requirement data;

[0116] All target technical requirement data are obtained as several pieces of technical requirement data in the target application scenario.

[0117] In one embodiment, the acquisition module 210 is specifically configured to perform word segmentation processing on each piece of technical requirement data to obtain a word segmentation result;

[0118] The word frequency of each word segment is determined, and the word segment and its corresponding word frequency are stored to obtain a technical requirement keyword set.

[0119] In one embodiment, the matching module 230 is specifically configured to determine, for any piece of technical requirement data, at least one morpheme of the technical requirement data;

[0120] Calculate the relevance score between each morpheme of the technical requirement data and any patent data; wherein the relevance score is negatively correlated with the frequency of the morpheme appearing in all patent data, and positively correlated with the frequency of the morpheme appearing in the patent data;

[0121] Perform a weighted summation of the relevance scores of all morphemes and the patent data to obtain the matching degree between the technical requirement data and the patent data;

[0122] After obtaining all the matching degrees between the piece of technical requirement data and the plurality of patent data, the maximum value of the matching degrees between the plurality of patent data and the piece of technical requirement data is determined as the matching degree peak value.

[0123] In one embodiment, the determination module 240 is specifically configured to obtain, for any target technical requirement data, patent data corresponding to its matching degree peak value as the first patent data;

[0124] Determining the similarity between the first patent data and other patent data in the plurality of patent data based on a preset similarity algorithm;

[0125] Patent data with similarity greater than a preset similarity threshold is determined as a target patent data set of the target technology requirement data.

[0126] In one embodiment, the determination module 240 is specifically used to determine patent data whose similarity is greater than a preset similarity threshold and whose matching degree with the target technical requirement data is greater than a second preset threshold as a target patent data set; wherein the second preset threshold is not greater than the first preset threshold.

[0127] In one embodiment, any technical requirement data is marked with a technical requirement type, and the matching module 230 is specifically used to determine the technical requirement data whose demand intensity is greater than the technical supply intensity and whose corresponding matching degree peak is greater than a first preset threshold as the target technical requirement data.

[0128] The solutions in the embodiments of the present application can be implemented using various computer languages, for example, the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0129] In addition, an embodiment of the present invention further provides an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor. The transceiver, the memory, and the processor are respectively connected via a bus. When the computer program is executed by the processor, the various processes of the various embodiments of the above-mentioned patent data acquisition method are implemented, and the same technical effects can be achieved. To avoid repetition, they will not be described here.

[0130] For details, see Figure 3 As shown, the electronic device includes a bus 1110 , a processor 1120 , a transceiver 1130 , a bus interface 1140 , a memory 1150 , and a user interface 1160 .

[0131] In an embodiment of the present invention, the electronic device further includes: a computer program stored in the memory 1150 and executable on the processor 1120, wherein the computer program implements the above-mentioned patented data acquisition method when executed by the processor 1120.

[0132] The transceiver 1130 is configured to receive and send data under the control of the processor 1120 .

[0133] In an embodiment of the present invention, a bus architecture (represented by bus 1110) may include any number of interconnected buses and bridges, and bus 1110 connects various circuits including one or more processors represented by processor 1120 and a memory represented by memory 1150.

[0134] Bus 1110 represents one or more of any of several types of bus structures, including a memory bus and memory controller, a peripheral bus, an Accelerated Graphical Port (AGP), a processor, or a local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA), and a Peripheral Component Interconnect (PCI) bus.

[0135] The processor 1120 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above-mentioned processor includes: a general-purpose processor, a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a programmable logic array (PLA), a microcontroller unit (MCU) or other programmable logic devices, discrete gates, transistor logic devices, discrete hardware components. The various methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. For example, the processor can be a single-core processor or a multi-core processor, and the processor can be integrated into a single chip or located on multiple different chips.

[0136] The processor 1120 can be a microprocessor or any conventional processor. The method steps disclosed in conjunction with the embodiments of the present invention can be directly executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a readable storage medium known in the art, such as a random access memory (RAM), a flash memory (Flash Memory), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), or a register. The readable storage medium is located in a memory, and the processor reads the information in the memory and performs the steps of the above method in conjunction with its hardware.

[0137] The bus 1110 may also connect various other circuits, such as peripheral devices, voltage regulators, or power management circuits. The bus interface 1140 provides an interface between the bus 1110 and the transceiver 1130. These are all well known in the art and are therefore not further described in this embodiment of the present invention.

[0138] The transceiver 1130 can be a single component or multiple components, such as multiple receivers and transmitters, providing a means for communicating with various other devices over a transmission medium. For example, the transceiver 1130 receives external data from other devices and transmits data processed by the processor 1120 to other devices. Depending on the nature of the computer system, a user interface 1160 may also be provided, such as a touch screen, physical keyboard, display, mouse, speaker, microphone, trackball, joystick, or stylus.

[0139] It should be understood that in an embodiment of the present invention, the memory 1150 may further include a memory remotely located relative to the processor 1120, and these remotely located memories may be connected to a server via a network. One or more parts of the aforementioned network may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a wireless wide area network (WWAN), a metropolitan area network (MAN), the Internet, a public switched telephone network (PSTN), a plain old telephone service network (POTS), a cellular telephone network, a wireless network, a wireless fidelity (Wi-Fi) network, or a combination of two or more of the aforementioned networks. For example, the cellular telephone network and the wireless network can be a Global System for Mobile Communications (GSM) system, a Code Division Multiple Access (CDMA) system, a Worldwide Interoperability for Microwave Access (WiMAX) system, a General Packet Radio Service (GPRS) system, a Wideband Code Division Multiple Access (WCDMA) system, a Long Term Evolution (LTE) system, an LTE Frequency Division Duplex (FDD) system, an LTE Time Division Duplex (TDD) system, an Advanced Long Term Evolution (LTE-A) system, a Universal Mobile Telecommunications (UMTS) system, an Enhanced Mobile Broadband (eMBB) system, a Massive Machine Type of Communication (mMTC) system, an Ultra Reliable Low Latency Communications (uRLLC) system, and the like.

[0140] It should be understood that the memory 1150 in the embodiment of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Non-volatile memories include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory.

[0141] Volatile memory includes random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1150 of the electronic device described in the embodiments of the present invention includes, but is not limited to, the above and any other suitable types of memory.

[0142] In the embodiment of the present invention, the memory 1150 stores the following elements of the operating system 1151 and the application 1152: executable modules, data structures, or subsets thereof, or extended sets thereof.

[0143] Specifically, the operating system 1151 includes various system programs, such as a framework layer, a core library layer, and a driver layer, which are used to implement various basic services and process hardware-based tasks. The application 1152 includes various application programs, such as a media player and a browser, which are used to implement various application services. The program that implements the method of the embodiment of the present invention may be included in the application 1152. The application 1152 includes applets, objects, components, logic, data structures, and other computer system executable instructions that perform specific tasks or implement specific abstract data types.

[0144] In addition, an embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the various processes of the various embodiments of the above-mentioned patent data acquisition method are implemented, and the same technical effects can be achieved. To avoid repetition, they will not be described here.

[0145] An embodiment of the present invention also provides a computer program comprising one or more computer instructions, which, when executed by a processor, implement the steps in the above-mentioned patent data acquisition method and generate, in whole or in part, the process or function described in the embodiment of the present application.

[0146] When the computer instructions are loaded and executed on the processor, the process or function described in the embodiments of the present application is generated in whole or in part.

[0147] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0148] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A patent data acquisition method, characterized in that: include: Acquire several pieces of technical requirement data in the target application scenario, extract the requirement keywords of each piece of technical requirement data, and obtain a set of technical requirement keywords corresponding to the target application scenario; Using the technical requirement keyword set to search the patent database, several patent data for the target application scenario are obtained; For any piece of technical requirement data, a preset correlation algorithm is used to determine the matching degree between the plurality of patent data and the technical requirement data, and obtain a matching degree peak value; the matching degree peak value is the maximum value of the matching degree between the plurality of patent data and the technical requirement data; Determining, among the plurality of pieces of technical requirement data, technical requirement data corresponding to a matching degree peak value greater than a first preset threshold as target technical requirement data; For any target technology requirement data, the target patent data set is obtained from the plurality of patent data using the patent data corresponding to the peak value of the matching degree.

2. The method according to claim 1, characterized in that The acquisition of several technical requirement data in the target application scenario includes: Obtaining several pieces of research data for the target application scenario; For any piece of survey data, extract the characteristic value corresponding to the characteristic variable in the survey data, input the obtained characteristic value into a pre-trained demand classification model to obtain the technical demand type corresponding to the survey data; the demand classification model is used to process the input characteristic value to obtain the technical demand type corresponding to the input data; After obtaining all technical requirement types of the plurality of survey data, counting the number of requirements corresponding to each technical requirement type, determining the technical requirement type whose number of requirements is greater than a statistical threshold as the target technical requirement type, and using the survey data corresponding to the target technical requirement type as the target technical requirement data; All target technical requirement data are obtained as several pieces of technical requirement data in the target application scenario.

3. The method according to claim 1, characterized in that The step of extracting the requirement keywords of each piece of technical requirement data includes: For each piece of technical requirement data, perform word segmentation processing on the requirement data to obtain a word segmentation result; The word frequency of each word segment is determined, and the word segment and its corresponding word frequency are stored to obtain a technical requirement keyword set.

4. The method according to claim 1, wherein For any piece of technical requirement data, a preset correlation algorithm is used to determine the matching degree between the plurality of patent data and the technical requirement data, and obtain a matching degree peak, including: For any piece of technical requirement data, determining at least one morpheme of the technical requirement data; Calculate the relevance score between each morpheme of the technical requirement data and any patent data; wherein the relevance score is negatively correlated with the frequency of the morpheme appearing in all patent data, and positively correlated with the frequency of the morpheme appearing in the patent data; Perform a weighted summation of the relevance scores of all morphemes and the patent data to obtain the matching degree between the technical requirement data and the patent data; After obtaining all the matching degrees between the piece of technical requirement data and the plurality of patent data, the maximum value of the matching degrees between the plurality of patent data and the piece of technical requirement data is determined as the matching degree peak value.

5. The method according to claim 1, wherein The method of determining a target patent data set from the plurality of patent data by using the patent data corresponding to the peak value of the matching degree for any target technology requirement data includes: For any target technical requirement data, obtain the patent data corresponding to its matching degree peak as the first patent data; Determining the similarity between the first patent data and other patent data in the plurality of patent data based on a preset similarity algorithm; Patent data with similarity greater than a preset similarity threshold is determined as a target patent data set of the target technology requirement data.

6. The method according to claim 5, characterized in that The step of determining the patent data having a similarity greater than a preset similarity threshold as the target patent data set includes: Patent data whose similarity is greater than a preset similarity threshold and whose matching degree with the target technical requirement data is greater than a second preset threshold is determined as a target patent data set; wherein the second preset threshold is not greater than the first preset threshold.

7. The method according to claim 5, characterized in that The preset similarity algorithm includes: Obtaining the first patent data and text data of claims, titles, and abstracts of the patents to be calculated for similarity; Calculating similarity between the first patent data and the claim text data and abstract text data of the patent data to be similarity calculated using a first similarity algorithm to obtain a first similarity and a second similarity; A second similarity algorithm is used to calculate the title data of the first patent data and the patent data to be similarity calculated, to obtain a third similarity; the computational complexity of the second similarity algorithm is less than the computational complexity of the first similarity algorithm; The first similarity, the second similarity, and the third similarity are weighted and summed to obtain the final similarity between the first patent data and the patent data to be calculated for similarity; the first similarity has the highest weight, and the third similarity has the lowest weight.

8. The method according to claim 1, characterized in that Any technical requirement data is marked with a technical requirement type, and the method further includes: For any technical requirement data, determine the demand intensity of the technical requirement data and the technical supply intensity for the technical requirement data according to its technical requirement type; wherein the technical requirement intensity is obtained based on the number of technical requirement data included in the technical requirement type; and the technical supply intensity is obtained based on the number of patent data whose matching degree with the technical requirement is greater than a first preset threshold; The step of determining the technical requirement data whose corresponding matching degree peak value is greater than a first preset threshold among the plurality of technical requirement data as target technical requirement data includes: The technical demand data whose demand intensity is greater than the technical supply intensity and whose corresponding matching degree peak is greater than the first preset threshold is determined as the target technical demand data.

9. A patent data acquisition device, characterized in that: The device includes: An acquisition module is used to acquire a plurality of technical requirement data in a target application scenario, extract the requirement keywords of each technical requirement data, and obtain a set of technical requirement keywords corresponding to the target application scenario; A search module is used to use a set of technical requirement keywords to call a patent database to search for a number of patent data for the target application scenario; a matching module configured to determine, for any piece of technical requirement data, a degree of matching between the plurality of patent data and the piece of technical requirement data using a preset correlation algorithm, and obtain a peak matching degree; the peak matching degree being the maximum value of the degree of matching between the plurality of patent data and the piece of technical requirement data; and determining, among the plurality of technical requirement data, technical requirement data whose corresponding peak matching degree is greater than a first preset threshold as target technical requirement data; The determination module is used to obtain a target patent data set from the plurality of patent data using the patent data corresponding to the peak matching degree of any target technical requirement data.

10. An electronic device comprising a processor and a memory, wherein the memory stores a computer program, wherein: The processor executes the computer program stored in the memory to implement the steps in the method according to any one of claims 1 to 8.

11. A computer program comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 8 are implemented.