Intellectual property supply and demand matching system based on natural language processing and multi-dimensional vectors and use method of intellectual property supply and demand matching system

Through the intellectual property supply and demand matching system of natural language processing and multi-dimensional vector matching, the problem of irrelevant information in the search results in the prior art is solved, and the intellectual property search with high accuracy and efficiency is achieved, and the technical threshold is lowered.

CN120144734APending Publication Date: 2025-06-13ADVANCED INST OF INFORMATION TECH (AIIT) PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311690727.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-13

Smart Images

  • Figure CN120144734A_ABST
    Figure CN120144734A_ABST
Patent Text Reader

Abstract

The invention provides an intellectual property supply and demand matching system based on natural language processing and multi-dimensional vectors and a use method. A supply document vectorization component C1, a demand document vectorization component C2, an intellectual property library Db, an intellectual property vector library Db1, a deep learning training component M, a vector matching component M0, a vector reverse query intellectual property component C3 and a model positive feedback component C4 are combined to obtain a supply and demand matching device; deploying the combined device as a web service of a B-S framework through an application server; the system is wide in applicability and high in accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of Internet application systems, and particularly relates to an intellectual property supply and demand matching system and a usage method based on natural language processing and multi-dimensional vectors. Background Art

[0002] For new models, new processes, new methods, new materials, or solution, software systems, etc. required by an enterprise for technological upgrading and transformation, there may be similar cases and achievements in other enterprises. Therefore, the enterprise needs to query and retrieve these existing achievements, thus giving rise to a considerable number of intellectual property retrieval and matching requirements.

[0003] The existing intellectual property retrieval is realized through a search engine. The retrieval system segments the intellectual property documents to generate an inverted index table and then archives it. When a user uses it, one or more retrieval keywords are input. The retrieval system finds these keywords in the archived inverted index table through a hash algorithm, and then returns the associated intellectual property documents to the user.

[0004] This retrieval method is time-consuming in establishing the inverted index table in the early stage, but has a high retrieval efficiency during actual user use. The disadvantage is that the accuracy of the retrieval results depends on the specificity of the keywords selected by the user. If the keyword specificity is relatively high, the probability of retrieving accurate results is large; if the keywords selected by the user are relatively common, the retrieved results often contain a lot of irrelevant information. This poses requirements on the information technology level of the relevant personnel in the enterprise. However, the information technology level of most employees in traditional enterprises often fails to meet such requirements.

[0005] Therefore, designing a general and highly accurate intellectual property supply and demand matching method and device is an inevitable requirement for traditional enterprises to carry out technological upgrading and transformation. Summary of the Invention

[0006] This patent provides a general and highly accurate intellectual property supply and demand matching method and a supporting usage device to provide support for the technological upgrading and transformation of traditional enterprises. The following is included: An intellectual property supply and demand matching system based on natural language processing and multi-dimensional vectors. The construction of the intellectual property supply and demand matching system includes the following steps:

[0007] Step 1: Collect n supply description documents, that is, intellectual property documents, and demand description documents, that is, natural language description samples of the intellectual property needs of traditional enterprises with retrieval needs, and make relevant annotations. If supply description document i and demand description document j are judged to be associated, they are marked as "associated", where n≥1000;

[0008] The supply description document is divided into three parts: title, content, and usage scenario, and processed into three vectors through NLP technology. The three vectors are combined into a multi-dimensional vector V. This process is encapsulated as component C1. The input of C1 is the supply description document D, and the output is the multi-dimensional vector V corresponding to D;

[0009] The demand description document is divided into three parts: title, brief description, and usage scenario, and processed into three vectors through NLP technology. The three vectors are combined into a multi-dimensional vector v. The processing process is encapsulated as component C2. The input of C2 is the demand description document d, and the output is the multi-dimensional vector v corresponding to d;

[0010] Step 2: Through deep learning model training, a calculation model M0 that can achieve pairwise matching of (V1, v1), (V2, v2),..., (Vn, vn) is trained;

[0011] Step 3: Establish an intellectual property database Db, perform NLP processing on all intellectual property documents in Db, and store the generated multi-dimensional vectors as database Db1. Db1 is unidirectionally associated with Db;

[0012] Step 4: Deploy the deep learning model training result M0 to the application server, combine it with the intellectual property database to provide an intellectual property retrieval system, and deploy it as a B-S architecture web service for enterprise use.

[0013] Preferably, step 2 of the intellectual property supply and demand matching system based on natural language processing and multi-dimensional vectors includes the following process:

[0014] (1) Select n (n≥1000) demand description documents (d1, d2,..., dn) as the sample set, and use component C2 to process them respectively to obtain n multi-dimensional vectors v1, v2,..., vn;

[0015] (2) Use component C1 to process the corresponding n supply description documents (D1, D2,..., Dn) of the demand description documents respectively to obtain n multi-dimensional vectors V1, V2,..., Vn;

[0016] (3) Construct a deep learning component M, and through training, obtain a vector matching component M0 such that M0(v1) = (V1,...), M0(v2) = (V2,...),..., M0(vn) = (Vn...), that is, when M0 acts on vx, the result is a multi-dimensional vector set containing Vx, and Vx is ranked first; that is, it is required that M0 can match successfully for all demand documents and supply documents in the sample set. If the matching fails, adjust the parameters of M0 and retrain until the matching is successful.

[0017] Preferably, step four of the intellectual property supply-demand matching system based on natural language processing and multi-dimensional vectors includes the following process. A matching interface and a feedback interface are set on the supply-demand matching system, and a connection firewall is set in the application server, connecting to the network and the browser port.

[0018] A method for using an intellectual property supply-demand matching system based on natural language processing and multi-dimensional vectors, for a demand description document di input by any user, includes the following processing steps:

[0019] (1). Calculate the corresponding multi-dimensional vector vi through component C2;

[0020] (2). Through the deep learning model M0, combine with the vector database Db1 to calculate M0(vi), and obtain (V’1, V’2, …, V’m);

[0021] (3). According to the association relationship between Db1 and Db, match the intellectual property document set (D’1, D’2, …, D’m), and encapsulate the matching process as component C3;

[0022] (4). The user selects the most matching document Di from the intellectual property document set (D’1, D’2, …, D’m);

[0023] (5). Add the supply-demand matching relationship (di, Di) to the sample set, re-execute the process in step two of claim 2, and encapsulate this process as the model positive feedback component C4.

[0024] Advantages or positive effects of the present invention:

[0025] The present invention performs supply-demand matching of intellectual property through natural language processing and multi-dimensional vector matching, and has the following advantages:

[0026] 1. The demand description is presented in natural language, reducing the usage threshold for users;

[0027] 2. The matching model obtained through deep learning of multi-dimensional vectors has higher matching accuracy;

[0028] 3. The device is equipped with a feedback module, which can continuously adaptively optimize, and continuously improve the matching accuracy as the intellectual property library expands and the number of usage times increases. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a schematic diagram of the process structure of the system involved in this patent.

[0031] Figure 2 It is a schematic diagram of the device completed by this patent portfolio being deployed as a B-S architecture web service through an application server. Specific implementation manners

[0032] The present invention will be further described in detail below in conjunction with embodiments. The following embodiments are explanations of the present invention and the present invention is not limited to the following embodiments.

[0033] Embodiment:

[0034] The specific implementation of the present invention includes the following four steps

[0035] 1. Processing of supply description documents

[0036] 1.1 Divide the supply description document into three parts: title, content, and usage scenario;

[0037] 1.2 Process the three parts respectively using the word2vec technology of NLP to obtain three vectors, and combine the three vectors into a multi-dimensional vector V;

[0038] 1.3 Package the processing processes of 1.1 and 1.2 into component C1. The input of C1 is the supply description document D, and the output is the multi-dimensional vector V corresponding to D.

[0039] 2. Processing of requirement description documents

[0040] 2.1 Divide the requirement description document into three parts: title, brief description, and usage scenario;

[0041] 2.2 Process the three parts respectively using the word2vec technology of NLP to obtain three vectors, and combine the three vectors into a multi-dimensional vector v;

[0042] 2.3 Package the processing processes of 2.1 and 2.2 into component C2. The input of C2 is the requirement description document d, and the output is the multi-dimensional vector v corresponding to d.

[0043] 3. Processing of supply-demand matching

[0044] 3.1 Select n (n≥1000) requirement description documents (d1, d2,..., dn) as a sample set, and use component C2 to process them respectively to obtain n multi-dimensional vectors v1, v2,..., vn;

[0045] 3.2 Select n supply description documents (D1, D2, …, Dn) corresponding to the requirement description document in 3.1, and process them separately using component C1 to obtain n multi-dimensional vectors V1, V2, …, Vn;

[0046] 3.3 Construct a deep learning component M, and obtain a vector matching component M0 through training, such that M0(v1) = (V1, …), M0(v2) = (V2, …), …, M0(vn) = (Vn …), that is, when M0 acts on vx, the result is a multi-dimensional vector set containing Vx, and Vx is ranked first; that is, it is required that M0 can successfully match all requirement documents and supply documents in the sample set. If the match fails, adjust the parameters of M0 and retrain until the match is successful;

[0047] The training of M0 is based on the following steps:

[0048] (1) Model construction

[0049] Use the NLP model in 2.2 as the embedding layer to construct a neural network architecture including a word sequence modeling layer and an output layer, and introduce an attention mechanism.

[0050] The word embedding layer uses the word embedding method (word2vec) to convert words into vectors to capture the semantic similarity between words. The sequence modeling layer uses neural network structures such as recurrent neural network (RNN), long short-term memory network (LSTM), or gated recurrent unit (GRU) to process sequence data and capture the temporal dependencies in the text. The output layer uses a fully connected layer to map the output of the model to a multi-dimensional vector. At the same time, an attention mechanism is introduced to improve the performance and generalization ability of the model by calculating the correlations between different positions in the sequence.

[0051] (2) Model training

[0052] Assign weight parameters to each connection point on the embedding layer, sequence modeling layer, and output layer as a set of parameters for adjusting the training results and optimizing the model. At the same time, introduce control parameters such as learning rate, batch size, and number of training epochs to better optimize the model.

[0053] Based on the above neural network architecture, perform cyclic training in combination with appropriate loss functions (such as cross-entropy loss function) and appropriate optimizers (such as stochastic gradient descent for updating model parameters) and other control logics.

[0054] (3) Model evaluation

[0055] Divide the n requirement description documents in 3.1 into three sets: a training set, a validation set, and a test set. Perform model training on the training set, introduce appropriate evaluation metrics (such as accuracy, recall, etc.), verify and evaluate the model training results on the validation set, and monitor the performance of the validation set during the training process. When the performance no longer improves, stop training early to prevent overfitting. After the training and validation are completed, the model is tested and evaluated through the test set, and only the models that pass the evaluation are formally deployed.

[0056] 3.4 Establish an intellectual property database Db, process the intellectual property documents in Db through component C1, and save the processing results to the vector database Db1;

[0057] 3.5 For any requirement description document di input by a user, perform supply-demand matching through the following steps:

[0058] 3.5.1 Calculate the corresponding multi-dimensional vector vi through component C2;

[0059] 3.5.2 Through the deep learning model M0, combine with the vector database Db1 to calculate M0(vi), and obtain (V’1, V’2, …, V’m);

[0060] 3.5.3 According to the association relationship between Db1 and Db, match the intellectual property document set (D’1, D’2, …, D’m), and encapsulate the matching process as component C3;

[0061] 3.5.4 The user selects the most matching document Di from the intellectual property document set (D’1, D’2, …, D’m);

[0062] 3.5.5 Add the supply-demand matching relationship (di, Di) to the sample set, re-execute 3.1 - 3.3, and encapsulate this process as the model positive feedback component C4.

[0063] 4. Assembly and Deployment of the Device

[0064] Combine the above components (supply document vectorization component C1, demand document vectorization component C2, intellectual property database Db, intellectual property vector database Db1, deep learning training component M, vector matching component M0, vector reverse lookup intellectual property component C3, model positive feedback component C4) to obtain a supply-demand matching device, and the combination method is as Figure 1 shown.

[0065] The combined device is deployed as a B - S architecture web service through an application server, and the deployment method is as Figure 2 shown.

[0066] In addition, it should be noted that for the specific embodiments described in this specification, the names of each system step and the like can be different. Any equivalent or simple changes made according to the principles described in the inventive concept of this invention patent are included within the protection scope of this invention patent. Those skilled in the art to which this invention pertains can make various modifications, supplements, or use similar ways of substitution to the specific embodiments described, as long as they do not deviate from the structure of this invention or exceed the scope defined by this claims, they should all fall within the protection scope of this invention.

Claims

1. An intellectual property supply-demand matching system based on natural language processing and multi-dimensional vectors, characterized in that the construction of the intellectual property supply-demand matching system includes the following steps: Step 1: Collect n supply description documents, namely intellectual property documents, and demand description documents, namely natural language description samples of the intellectual property needs of traditional enterprises with retrieval needs, and make relevant annotations. If the supply description document i and the demand description document j are judged to be related, they are marked as "related", n≥1000; The supply description document is divided into three parts: title, content, and usage scenario, and processed into three vectors through NLP technology. The three vectors are combined into a multi-dimensional vector V. This process is encapsulated as component C1. The input of C1 is the supply description document D, and the output is the multi-dimensional vector V corresponding to D; The demand description document is divided into three parts: title, brief description, and usage scenario, and processed into three vectors through NLP technology. The three vectors are combined into a multi-dimensional vector v. The processing process is encapsulated as component C2. The input of C2 is the demand description document d, and the output is the multi-dimensional vector v corresponding to d; Step 2: Through deep learning model training, train a calculation model M0 that can achieve pairwise matching of (V1, v1), (V2, v2),..., (Vn, vn); Step 3: Establish an intellectual property database Db, perform NLP processing on all the intellectual property documents in Db, and store the generated multi-dimensional vectors as database Db1. Db1 is unidirectionally associated with Db; Step 4: Deploy the deep learning model training result M0 to the application server, combine it with the intellectual property database to provide an intellectual property retrieval system, and deploy it as a B-S architecture web service for enterprises to use.

2. The intellectual property supply-demand matching system based on natural language processing and multi-dimensional vectors according to claim 1, characterized in that the said Step 2 includes the following process: (1) Select n (n≥1000) demand description documents (d1, d2,..., dn) as a sample set, and use component C2 to process them respectively to obtain n multi-dimensional vectors v1, v2,..., vn; (2) Use component C1 to process the n supply description documents (D1, D2,..., Dn) corresponding to the demand description documents respectively to obtain n multi-dimensional vectors V1, V2,..., Vn; (3) Construct a deep learning component M, and through training, obtain a vector matching component M0, such that M0(v1)=(V1,...), M0(v2)=(V2,...),..., M0(vn)=(Vn...), that is, when M0 acts on vx, the result obtained is a multi-dimensional vector set containing Vx, and Vx is ranked first; that is, it is required that M0 can match successfully for all demand documents and supply documents in the sample set. If it cannot match successfully, adjust the parameters of M0 and retrain until it matches successfully.

3. The intellectual property supply-demand matching system based on natural language processing and multi-dimensional vectors according to claim 1, characterized in that the said Step 4 includes the following process. A matching interface and a feedback interface are set on the supply-demand matching system, and a connection firewall is set in the application server, connecting to the network and browser ports.

4. Method for using an intellectual property supply and demand matching system based on natural language processing and multi-dimensional vectors, characterized in that: For any demand description document di input by a user, the following processing steps are included: (1). Calculate the corresponding multi-dimensional vector vi through component C2; (2). Through the deep learning model M0, combine with the vector database Db1 to calculate M0(vi), and obtain (V’1, V’2, …, V’m); (3). According to the association relationship between Db1 and Db, match the intellectual property document set (D’1, D’2, …, D’m), and encapsulate the matching process as component C3; (4). The user selects the most matching document Di from the intellectual property document set (D’1, D’2, …, D’m); (5). Add the supply and demand matching relationship (di, Di) to the sample set, re-execute the process in step two of claim 2, and encapsulate this process as the model positive feedback component C4.