Method, apparatus, electronic device, and storage medium for identifying technical competitors
By employing a fine-tuned model to generate semantic vectors from multilingual technical texts, the method addresses the limitation of single-language data in competitor identification, enabling efficient and cost-effective global competitor recognition.
Patent Information
- Application Number
- CN202411674057.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The existing technology cannot effectively transcend language and country recognition technology competitors, and cannot adapt to the global technology competition environment.
By constructing a total text set, retraining the first embedded model of multilinguals using a contrast learning mechanism, a first model that can span the competitors of language and country recognition technology is generated, and the text of the first object is converted into a semantic vector, and the competitors are identified based on the similarity of semantic vectors.
It realizes the recognition of technology competitors across languages and countries without relying on manual participation, saves human resources costs, and provides global industrial analysis and technical layout support.
Smart Images

Figure CN119557654B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data mining. Specifically, the present application relates to a method, an apparatus, an electronic device, and a storage medium for identifying technical competitors Background Art
[0002] The identification of technical competitors refers to identifying enterprises, universities, research institutions, or individuals with a technical competition relationship in a specific technical field through systematic analysis and evaluation. These technical competitors have direct or indirect competition relationships with the competing entity in aspects such as technological innovation, patent applications, and market development. Identifying technical competitors is of great significance for an enterprise's strategic planning, adjustment of technological R & D directions, and improvement of market competitiveness
[0003] In the prior art, the identification of technical competitors is usually based on industrial data (including patent data, product data, user review data, web page data, etc.) and artificial intelligence algorithms are used. Currently, no matter the identification method implemented from the perspective of enterprises, customers, a mixed perspective of enterprises and customers, or the Internet perspective, it is based on monolingual data and is difficult to adapt to the current trend of technological globalization, and it is impossible to identify technical competitors across languages and countries Summary of the Invention
[0004] Embodiments of the present disclosure provide a method, an apparatus, an electronic device, and a storage medium for identifying technical competitors, which are used to solve the technical problem that the existing method for identifying technical competitors cannot identify technical competitors across languages and countries
[0005] According to one aspect of the embodiments of the present disclosure, a method for identifying technical competitors is provided, including:
[0006] Obtain a text set, where the text set includes a plurality of first texts, and the first text is a technical text of a first object in the target technical field
[0007] According to a pre-trained first model, obtain a first semantic vector of each first text, and the first model is obtained by retraining a pre-trained multilingual first embedding model with a contrastive learning mechanism according to a total text set, and the total text set includes multilingual technical texts in the target technical field
[0008] Calculate the similarity between each first semantic vector and each second semantic vector, and obtain a target semantic vector from each second semantic vector according to the similarity, where the second semantic vector is the semantic vector of the technical text of a second object in the target technical field
[0009] For each second object, according to the similarities corresponding to the respective target semantic vectors of the second object, obtain the competition intensity between the second object and the first object, where the competition intensity characterizes the similarity degree of the technologies between the first object and the second object;
[0010] According to the competition intensity of the second object, determine the technological competitors of the first object from among the various second objects.
[0011] According to another aspect of the embodiments of the present disclosure, there is provided an apparatus for identifying technological competitors, including:
[0012] A technical text acquisition module, configured to acquire a text set, the text set including a plurality of first texts, and the first text being the technical text of the first object in the target technical field;
[0013] A semantic vector acquisition module, configured to obtain the first semantic vector of each first text according to a pre-trained first model, where the first model is obtained by re-training a pre-trained multilingual first embedding model according to a contrastive learning mechanism with respect to a total text set, and the total text set includes multilingual technical texts in the target technical field;
[0014] A similarity comparison module, configured to calculate the similarity between each first semantic vector and each second semantic vector, and obtain a target semantic vector from among the second semantic vectors according to the similarity, where the second semantic vector is the semantic vector of the technical text of the second object in the target technical field;
[0015] A competition intensity acquisition module, configured to, for each second object, obtain the competition intensity between the second object and the first object according to the similarities corresponding to the respective target semantic vectors of the second object, where the competition intensity characterizes the similarity degree of the technologies between the first object and the second object;
[0016] A competitor determination module, configured to determine the technological competitors of the first object from among the various second objects according to the competition intensity of the second object.
[0017] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method provided in any one of the above embodiments.
[0018] According to still another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method provided in any one of the above embodiments are implemented.
[0019] According to one aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method provided in any one of the above embodiments are implemented.
[0020] The beneficial effects brought by the technical solutions provided in the embodiments of the present disclosure are as follows:
[0021] The technical solutions provided in the embodiments of the present disclosure can construct a total text set based on a large number of multi - language technical texts in the target technical field. By retraining the multi - language first embedding model through a contrastive learning mechanism to obtain a first model, it can finely tune and optimize the ability of the first embedding model to express text vectors in multiple languages in the target technical field. Using the trained first model, without relying on any manual participation, for a determined first object, each first text in the text set of the first object can be converted into a first semantic vector, and according to the similarity between the first semantic vector and each second semantic vector in the target technical field, a target semantic vector is determined. The second object is determined based on the target semantic vector, and the competition intensity of the second object is calculated. The technical competitors of the first object are determined from each second object based on the competition intensity, enabling the identification of technical competitors to transcend language and national boundaries, thus greatly saving the human resource cost for judging technical competitors and providing strong support for analysis scenarios such as global industrial analysis and technology layout analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description in the embodiments of the present disclosure.
[0023] Figure 1 It is a schematic flow chart of a method for identifying technical competitors provided in an embodiment of the present disclosure;
[0024] Figure 2 It is a schematic flow chart of an unsupervised SimCSE training method provided in an embodiment of the present disclosure;
[0025] Figure 3 It is a schematic flow chart of a supervised SimCSE training method provided in an embodiment of the present disclosure;
[0026] Figure 4 It is a schematic flow chart of a method for identifying technical competitors of patent holders provided in an embodiment of the present disclosure;
[0027] Figure 5 It is a schematic structural diagram of an apparatus for identifying technical competitors provided in an embodiment of the present disclosure;
[0028] Figure 6 It is a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] Embodiments of the present disclosure will be described below with reference to the accompanying drawings in the present disclosure. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present disclosure, and do not constitute limitations on the technical solutions of the embodiments of the present disclosure.
[0030] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0031] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.
[0032] The technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0033] The terms and related technologies involved in the present application will be described below:
[0034] In the prior art, the identification of technology competitors is usually based on industrial data (including patent data, product data, user review data, web data, etc.) and artificial intelligence algorithms are used.
[0035] The current related technologies mainly include:
[0036] Related Technology 1: A method for identifying competitors from the perspective of an enterprise, which identifies competitors through the basic characteristics of the enterprise itself, such as enterprise technology, products, and strategies. The method for identifying competitors from the perspective of an enterprise is mainly represented by the identification standard method, the patent analysis method, the strategic group method, and the product information clustering method, etc.
[0037] Related Technology 2: A method for identifying competitors from the customer perspective. Taking customers as the main body, it determines whether there is a competitive relationship between enterprises by judging the customer overlap. Common methods include questionnaire surveys and brand switching analysis. With the development of Internet information, the research on text mining methods of online reviews has received increasing attention.
[0038] Related Technology 3: A method for identifying competitors from the mixed perspective of enterprises and customers. It usually combines structured and unstructured data such as the enterprise's product data and customer review data, which is more comprehensive and accurate than the identification from a single perspective. For example, based on the customer value leadership strategy, the enterprise financial statements and consumer reviews on e-commerce platforms are selected as information sources, and relying on the Back Propagation (BP) neural network, competitor evaluation models based on financial characteristics and comprehensive characteristics are established respectively.
[0039] Related Technology 4: A method for identifying competitors from the Internet perspective. This type of method emerged with the development of the Internet. The methods for identifying competitors from the Internet perspective can be roughly divided into two ways: based on web content and based on web link structure.
[0040] Among them, for the method based on web content, the Support Vector Machine (SVM) can be used to identify the competitive intention in news articles and can effectively identify documents with competitive intention. For the method based on web link structure, the target industry can be selected as the research object, and the Uniform Resource Locator (URL) co-occurrence analysis method can be used to display the competitive pattern of the target industry through a Multi Dimensional Scaling (MDS) graph, and the competitive level of competitors can be divided from the perspective of relevant web pages in the target industry to analyze and identify the main competitors.
[0041] Currently, whether it is the method for identifying competitors based on the enterprise perspective, customer perspective, the mixed perspective of enterprises and customers, or the Internet perspective, it is all based on monolingual data. Currently, there is little research on identifying technical competitors across languages and countries based on multilingual data, which is difficult to adapt to the current trend of technological globalization and cannot identify technical competitors across languages and countries.
[0042] Based on this, the embodiments of the present disclosure provide an identification of technical competitors to solve, to a certain extent, the problem in the above related technologies that it is impossible to identify technical competitors across languages and countries.
[0043] The technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referenced, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be repeatedly described.
[0044] It can be understood that in the method for identifying technical competitors provided by the embodiments of the present disclosure, any method step can be executed by an electronic device and / or a server. All the steps in the method can be independently executed by the electronic device or the server, or jointly executed by the electronic device and the server.
[0045] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The electronic device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device (such as a smart speaker), a wearable electronic device (such as a smart watch), a vehicle-mounted terminal, a smart home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto.
[0046] Subsequently, the embodiments of the present disclosure will be introduced with the server as the execution subject. However, this does not constitute a limitation on the embodiments of the present disclosure. The method provided by the embodiments of the present disclosure can be applied to the calculation and comparison of the text similarity corresponding to each object in other scenarios in addition to identifying technical competitors.
[0047] Figure 1 It is a schematic flowchart of a method for identifying technical competitors provided by the embodiments of the present disclosure. As Figure 1 shown, the technical solutions provided by the embodiments of the present disclosure include the following steps:
[0048] Step S101, obtain a text set, the text set includes multiple first texts, and the first text is a technical text of a first object in the target technical field.
[0049] Among them, the object refers to an enterprise, a university, a research institution, an organization, an individual, etc.
[0050] The technical text refers to a text with technical expression ability, such as a paper text, a patent text, a fund project text, a book text, a research report text, etc. In the actual application process, the type and quantity of the technical text can be determined according to actual needs.
[0051] In addition, considering the length of the technical text and the amount of information that each part of the content can represent, the input (and training samples) of each model in the embodiments of the present disclosure can be the technical text of the complete length or the pre-processed technical text (or pre-processed by the model). The pre-processed technical text can be the text or text representation corresponding to one or more of the title, abstract, keywords, conclusion, etc. of the technical text of the complete length.
[0052] Specifically, taking the identification of the technical competitors of any first object as an example, the method for identifying the technical competitors provided in the embodiments of the present disclosure will be described in detail.
[0053] It can be understood that there is a corresponding relationship between the object and the technical field it involves. For example, when the object is an enterprise, the technical field can be the technical field related to the industry involved by the enterprise. When the object is a university or a research institution, the technical field can be the technical field of the research topic studied by the university or the research institution, etc.
[0054] When determining the first object, the target technical field corresponding to the first object can be determined synchronously. The target technical field can refer to multiple technical fields involved by the first object, or can be set as part or a single technical field involved by the first object according to actual needs.
[0055] For example, when determining that the first object is an enterprise with cross-field industries, the first object involves multiple technical fields at the same time. If the goal is to identify the technical competitors of the first object in the field of communication technology, then the communication field is determined as the target technical field.
[0056] After determining the first object (i.e., the target research object), in step S101, among the technical texts in the target field, obtain the technical texts belonging to the first object as the first texts, and construct a text set from each first text. The number of first texts in the first text set is determined according to the actual situation of the first object, and the embodiments of the present disclosure do not limit this.
[0057] It can be understood that the first text is the technical text of the first object in the target technical field. Considering the influence of the data volume of the technical texts in the target field on the identification, the technical texts of the first object can be screened when determining the first text. For example, for patent texts, authorized patent texts within a certain number of years can be screened as the first texts. For papers, texts within a certain number of years and with an impact factor meeting the conditions can be screened as the first texts, etc. The determination method of the first text can be adjusted according to actual needs.
[0058] Step S102: Obtain the first semantic vectors of each first text according to the pre-trained first model. The first model is obtained by retraining the pre-trained multilingual first embedding model with a contrastive learning mechanism based on the total text set, where the total text set includes multilingual technical texts in the target technical field.
[0059] Specifically, the first model is an embedding model that converts text into semantic vectors. Before applying the first model, it is necessary to pre-train the first model. The first model is obtained by retraining the pre-trained multilingual first embedding model with a contrastive learning mechanism based on the total text set. The goal of using contrastive learning is to enable the first embedding model to distinguish positive examples (similar samples) and negative examples (dissimilar samples), and learn more meaningful feature representations from them, thereby improving the first embedding model's ability to grasp the semantics of multilingual texts.
[0060] Among them, the multilingual first embedding model is a pre-trained model (large model) with the ability to represent multilingual texts (i.e., convert text into semantic vectors). Multilingual means at least two languages. The specific types and quantities of the languages involved, as well as the specific type of the multilingual first embedding model, can be determined according to actual needs. For example, when the languages are Chinese and English, the specific type of the first embedding model can be OpenAI, BGE-Large, BCE, GTE-Large, etc.
[0061] It can be understood that the total text set includes multilingual technical texts in the target technical field. When obtaining the total text set, it is also possible to screen the multilingual technical texts in the target technical field to limit the size of the total text set. The specific screening method can be determined according to actual needs, which is the same or similar to the method of screening the technical texts of the first object.
[0062] In step S102, input each first text into the pre-trained first model, and the first model converts the text into a semantic vector to obtain the first semantic vector of each first text output by the first model.
[0063] It can be understood that in the embodiments of the present disclosure, the semantic vectors converted from text can be low-dimensional semantic vectors or high-dimensional semantic vectors. Considering that the embodiments of the present disclosure are applied to large-scale and complex text data, converting each text into high-dimensional semantic vectors can better capture more semantic details, capture the subtle differences and complex semantics in the text, and improve the accuracy when calculating the similarity between texts.
[0064] Step S103: Calculate the similarity between each first semantic vector and each second semantic vector, and obtain the target semantic vector from each second semantic vector according to the similarity. The second semantic vector is the semantic vector of the technical text of the second object in the target technical field.
[0065] Specifically, the second semantic vector is the semantic vector of the technical text of the second object in the target technical field. The second object is different from the first object, and the second semantic vector is the semantic vector of the technical text that is not the first object in the target technical field.
[0066] After obtaining the first semantic vector, in step S103, calculate the similarity between each first semantic vector and each second semantic vector, and obtain the target semantic vector from each second semantic vector according to the similarity between each first semantic vector and the second semantic vector. Among them, the similarity between semantic vectors measures the degree of proximity of two vectors in the semantic space, and the similarity metric can be determined according to actual needs, including but not limited to cosine similarity (Cosine Similarity), Pearson correlation coefficient (Pearson Correlation Coefficient), etc.
[0067] It can be understood that considering that the goal of calculating the semantic vector similarity in the embodiments of the present disclosure is to identify the technical competitors of the first object, that is, the target requirement is to identify the technical text with a relatively high similarity to the first text of the first object.
[0068] Based on this, in the step of obtaining the target semantic vector from each second semantic vector according to the similarity, a preset screening rule can be set to determine the target semantic vector with a relatively high similarity from each second semantic vector, and the screening rule can be determined according to actual needs.
[0069] For example, based on the technical texts in multiple languages in all target technical fields in the total text set, obtain the semantic vector corresponding to each technical text according to the preset first model, and use the semantic vector of the technical text that is not the first object as the second semantic vector. For each first semantic vector of the first object, calculate the similarity between the first semantic vector and each second semantic vector.
[0070] When obtaining the target semantic vector, a preset screening rule can be set. For each first semantic vector, select the top A second semantic vectors with the highest similarity to the first semantic vector as the target semantic vector, or sort all the similarities and select the top B similarities with the highest similarity, and use the second semantic vectors corresponding to the selected similarities as the target semantic vectors.
[0071] It should be noted that the above examples are only used as a specific example to assist in explaining the method steps of the embodiments of the present disclosure, and do not limit the embodiments of the present disclosure.
[0072] Step S104: For each second object, obtain the competition intensity between the second object and the first object according to the similarities of the respective target semantic vectors corresponding to the second object. The competition intensity represents the similarity degree of the technologies between the first object and the second object.
[0073] Specifically, considering that there may be multiple technical texts belonging to the same second object that are similar to the first text of the first object, in order to determine whether the second object belongs to the technical competitor of the first object and accurately measure the intensity of competition between the second object and the first object in terms of technology, the embodiments of the present disclosure set the competition intensity as a measurement index, which represents the similarity of the technologies between the first object and the second object and is used to measure the intensity of competition between the first object and the second object in terms of technology.
[0074] For each target semantic vector in step S103, there is a corresponding second object. The second objects corresponding to the respective target semantic vectors may be the same or different. It is necessary to take the second object as the processing perspective and uniformly process the target semantic vectors belonging to the same second object.
[0075] It can be understood that there may be a situation where a certain technical text is shared by two or more objects. When processing, the technical text shared by multiple objects can be recorded as belonging to the object with the first signature order, or it can be recorded as being shared by multiple objects. The processing method can be determined according to actual needs.
[0076] In step S104, for each determined second object, obtain at least one target semantic vector corresponding to the second object, and obtain the competition intensity between the second object and the first object according to the similarities of the respective target semantic vectors corresponding to the second object.
[0077] It can be understood that the competition intensity represents the similarity degree of the technologies between the first object and the second object, and the competition intensity calculation formula can be determined according to actual needs.
[0078] For example, for each second object, the competition intensity can be the average of the similarities of the respective target semantic vectors corresponding to the second object. It can also be considered to add corresponding weights to the similarities of different target semantic vectors, and the competition intensity is determined by weighted fusion of the similarities of the respective target semantic vectors corresponding to the second object. It can be understood that when determining the weights, the information referred to can include but is not limited to the publication time of the technical text corresponding to the target semantic vector, the market share of the products related to the technical text, whether the technical text is shared by multiple objects, and the number of shared objects, etc.
[0079] Step S105: Determine the technical competitors of the first object from each second object according to the competition intensity of the second object.
[0080] Specifically, the technical competitor of the first object can be understood as the second object that has a relatively high similarity with the first object from a technical perspective. Therefore, it is necessary to screen the second objects according to their competition intensity and select the technical competitors of the first object from each second object.
[0081] In step S105, according to the competition intensity of each second object determined in step S104, screen the second objects and determine the technical competitors of the first object from each second object.
[0082] It can be understood that the purpose of screening the second objects is to select the second objects with a relatively high competition intensity as technical competitors from each second object. The specific screening method can be determined according to actual needs.
[0083] For example, all competition intensities can be sorted, and the top C or top D% of the competition intensities with the highest values can be selected, and the second objects corresponding to the selected competition intensities are used as technical competitors.
[0084] The technical solution provided by the embodiments of the present disclosure can construct a total text set based on a large number of multi - language technical texts in the target technical field, retrain the multi - language first embedding model through a contrastive learning mechanism to obtain a first model, which can fine - tune and optimize the ability of the first embedding model to express text vectors in multiple languages in the target technical field. By using the trained first model, without relying on any manual participation, for a determined first object, each first text in the text set of the first object can be converted into a first semantic vector, and according to the similarity between the first semantic vector and each second semantic vector in the target technical field, a target semantic vector is determined. Based on the target semantic vector, a second object is determined and the competition intensity of the second object is calculated, and the technical competitors of the first object are determined from each second object according to the competition intensity, enabling the identification of technical competitors to cross language and country limitations, thus greatly saving the human resource cost for judging technical competitors and providing strong support for analysis scenarios such as global industrial analysis and technology layout analysis.
[0085] In the application scenario of multi - language text processing, the existing technology usually uses a multi - language machine translation (dictionary - based translation or statistic - based translation) engine to introduce multi - language technical texts into the machine translation engine and translate them into single - language technical texts (for example, translate non - English technical texts such as English, German, Japanese, Korean, Russian, etc. into English technical texts), and then convert the obtained single - language technical texts into semantic vectors in the form of a single - language embedding model. This method relies on the direct text conversion of the machine translation engine and may not be able to accurately capture complex semantic relationships and deep - level similarities between texts.
[0086] In the embodiments of the present disclosure, the first embedding model for multiple languages is retrained through contrastive learning to obtain the first model, enabling the first model to learn deeper and more fine-grained language features and semantic information, and better understand and represent the text content between different languages. Compared with the above existing technologies, in the manner of directly converting the technical text in multiple languages into semantic vectors by the first model, it does not rely on machine translation engines with uneven translation quality, avoids the possible translation errors of machine translation engines, maintains the accuracy and professionalism of the original text during the conversion process, avoids the problem of inconsistent terms caused by translation between different languages, and simplifies the processing flow of obtaining semantic vectors from technical texts during application, reducing the translation and processing costs to a certain extent and improving the processing efficiency.
[0087] In a possible implementation manner, the first model is generated through the following steps:
[0088] Construct multiple batches of first sample sets according to the total text set;
[0089] Perform unsupervised training on the first embedding model batch by batch according to each batch of first sample sets, with the goal of minimizing the contrastive loss function, fine-tune the parameters of the first embedding model, and obtain the first model;
[0090] Among them, any batch of first sample sets is constructed through the following steps:
[0091] Determine the first technical text corresponding to the batch from the total text set, and process the first technical text with data augmentation technology to obtain a number of second technical texts;
[0092] Obtain a number of technical texts different from the first technical text from the total text set as the third technical text;
[0093] According to the first embedding model, obtain the semantic vectors corresponding to the first technical text, the second technical text, and the third technical text respectively;
[0094] Use the semantic vector of the first technical text as the original sample, the semantic vector of the second technical text as the positive sample, and the semantic vector of the third technical text as the negative sample to obtain the first sample set.
[0095] Specifically, when training the first model, the contrastive mechanism adopted is the unsupervised SimCSE (Simple Contrastive Learning of Sentence Embeddings) method, and a data augmentation strategy is adopted. In a batch training manner, the contrastive loss function is used to optimize the model.
[0096] Specifically, before training, it is necessary to determine the training samples used in each batch of model training. The first sample set of any batch is constructed as follows:
[0097] For any first technical text in the total text set, the first technical text is processed by data augmentation techniques, different sentence representations are generated based on the first technical text, and a number of second technical texts corresponding to the first text are obtained. The number of second technical texts can be determined according to the number requirement of positive samples in each batch.
[0098] A number of technical texts different from the first technical text are obtained from the total text set as the third technical text. The number of third technical texts can be determined according to the number requirement of negative samples in each batch.
[0099] According to the first embedding model, the determined technical texts are converted into semantic vectors, and the semantic vectors corresponding to the first technical text, the second technical text, and the third technical text are obtained respectively.
[0100] The semantic vector corresponding to the obtained first technical text is used as the original sample, the semantic vector of the second technical text is used as the positive sample, and the semantic vector of the third technical text is used as the negative sample to obtain the first sample set.
[0101] In the above way, for the total text set, multiple batches of the first sample set can be constructed. According to the first sample set of each batch, the first embedding model is trained without supervision batch by batch.
[0102] During each batch of training, the original sample and the positive sample in the first sample set form a positive sample pair, and the original sample and the negative sample form a negative sample pair. With minimizing the contrast loss function as the optimization goal, the contrast loss is backpropagated and the parameters of the first embedding model are fine-tuned (update the model parameters), and the updated model is saved until the fine-tuning ends to obtain the first model.
[0103] It can be understood that the contrast loss function is a loss function used to measure the similarity between two input samples in the embedding space. Taking minimizing the contrast loss function as the optimization goal means that the first embedding model learns to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs in the embedding space to learn more effective feature representations.
[0104] The contrast loss function adopted can include but is not limited to Information Noise-Contrastive Estimation loss (InfoNCE for short), the normalized temperature-scaled cross entropy loss (NT-Xent for short), etc.
[0105] For example, the first model obtained after the BCE model is fine-tuned by unsupervised SimCSE is denoted as the BCE-UF-SimCSE model. The "U" in the model represents unsupervised, and the "F" represents fine-tune.
[0106] Figure 2 As shown in the flowchart of an unsupervised SimCSE training method provided by an embodiment of the present disclosure, for each training batch using the total text set, based on the technical text corresponding to the original sample, positive samples are obtained using a data augmentation strategy, and negative samples are obtained from the technical text corresponding to non-original samples in the total text set to obtain a first sample set. The first embedding model is fine-tuned by unsupervised SimCSE using multiple first sample sets to obtain the BCE-UF-SimCSE model. Here, the methods for obtaining positive and negative samples are not elaborated. Figure 2 As shown, for each training batch using the total text set, based on the technical text corresponding to the original sample, positive samples are obtained using a data augmentation strategy, and negative samples are obtained from the technical text corresponding to non-original samples in the total text set to obtain a first sample set. The first embedding model is fine-tuned by unsupervised SimCSE using multiple first sample sets to obtain the BCE-UF-SimCSE model. Here, the methods for obtaining positive and negative samples are not elaborated.
[0107] The contrast loss function formula is as follows:
[0108]
[0109] In the formula, represents the original sample, represents the positive sample, represents the negative sample, is the temperature coefficient, is the cosine similarity.
[0110] The technical solution provided by the embodiment of the present disclosure can obtain a large number of positive samples using a data augmentation strategy, without the need for complex sample annotation work, greatly reducing the cost in the training sample preparation stage. Unsupervised learning in a batch training manner can improve the generalization ability of the model. Using the contrast loss function to fine-tune the parameters of the first embedding model to obtain the first model can enhance the representation ability of the first model for texts in different languages, making similar texts in different languages closer in the embedding space, which helps to achieve alignment between different languages when processing cross-lingual texts and improve the performance of the first model.
[0111] In a possible implementation, the first model is generated in the following manner:
[0112] Construct multiple batches of second sample sets according to the total text set;
[0113] According to the second sample sets of each batch, perform supervised training on the first embedding model in batches, with minimizing the contrast loss function as the optimization objective, and fine-tune the parameters of the first embedding model to obtain the first model;
[0114] Among them, the second sample set of any batch is constructed in the following manner:
[0115] Determine the fourth technical text in the first preset language corresponding to the batch from the total text set, and the fifth technical text having a citation relationship with the fourth technical text;
[0116] For the fourth technical text and each fifth technical text, obtain the translation of the fourth technical text and the translation of the fifth text according to the second preset language;
[0117] According to the first embedding model, obtain the semantic vectors corresponding to the fourth technical text, the fifth technical text, the translation of the fourth technical text, and the translation of the fifth technical text respectively;
[0118] Calculate the similarity between the semantic vector of the fourth technical text and the semantic vectors of each sixth technical text in the target technical field, and determine the semantic vector of the target sixth technical text that meets the preset similarity threshold from each sixth technical text. The sixth technical text corresponds to the first language, and the semantic vector of the sixth technical text is obtained by processing the sixth technical text by the second model. The second model is a single-language embedding model for processing the target language;
[0119] Calculate the similarity between the semantic vector of the translation of the fourth technical text and the semantic vectors of each seventh technical text in the target technical field, and determine the semantic vector of the target seventh technical text that meets the preset similarity threshold from each seventh technical text. The seventh technical text corresponds to the second language, and the semantic vector of the seventh technical text is obtained by processing the seventh technical text by the third model. The third model is a single-language embedding model for processing the second language;
[0120] According to the first embedding model, obtain the semantic vectors corresponding to the target sixth technical text and the target seventh technical text respectively;
[0121] Take the semantic vector of the fourth technical text as the original sample, take the semantic vectors corresponding to the translation of the fourth technical text, the fifth technical text, and the translation of the fifth technical text as positive samples, and take the semantic vectors corresponding to the target sixth technical text and the target seventh technical text as negative samples to obtain the second sample set.
[0122] Specifically, in order to enable the first model to further improve its performance in processing technical texts and more accurately capture the language features and semantic relationships in technical texts, in the training of the first model in the embodiments of the present disclosure, the adopted contrast mechanism is the supervised SimCSE algorithm, positive samples are labeled based on the citation relationship between technical texts, and negative samples are labeled based on the similarity between technical texts.
[0123] Specifically, before training, it is necessary to determine the training samples used in each batch of model training. In the embodiments of the present disclosure, the idea of semantic alignment is adopted to determine the training samples. The second sample set of any batch is constructed in the following manner:
[0124] Pre-set the first language and the second language as the basis for semantic alignment. Both the first language and the second language are the language types that the first model aims to process. Determine the second model for processing the first language and the third model for processing the second language. Both the second model and the third model are single-language embedding models.
[0125] Based on the language types corresponding to each technical text in the total text set, determine any fourth technical text from the technical texts in the total text set whose language is the first language, and use the fourth text as the technical text corresponding to the original sample of this batch.
[0126] For the fourth technical text, according to the citation relationship existing between technical texts, determine the fifth technical text that has a citation relationship with the fourth technical text. It can be understood that the citation relationship includes citation and being cited, and can be determined according to actual needs.
[0127] For the obtained fourth technical text and each fifth technical text, obtain the translation of the fourth technical text and the translation of the fifth technical text according to the pre-set first language and second language.
[0128] It can be understood that the fourth technical text is in the first language, and its corresponding translation is in the second language. The language types involved in each fifth technical text can be various. When obtaining the translation, if the fifth technical text is in the first language, obtain the corresponding translation in the second language. If the fifth technical text is in the second language, obtain the corresponding translation in the first language. If the fifth technical text is in a language other than the first language and the second language, obtain the corresponding translations in the first language and the second language of the fifth technical text at the same time. By obtaining the translations in this way, multiple pairs of aligned corpus with the same semantics in the first language and the second language can be obtained.
[0129] For example, when the language types that the first model aims to process include Chinese, English, and German, set the first language and the second language as Chinese and English respectively. The fourth technical text is a technical text in Chinese. Determine the fifth technical text that has a citation relationship with the fourth technical text. The language types corresponding to each fifth technical text include Chinese, English, and German. Obtain the Chinese translation of the fifth technical text in English, obtain the English translation of the fifth technical text in Chinese, and obtain the Chinese and English translations of the fifth technical text in German. It can be understood that according to the requirement for the number of samples, the order of the first language and the second language can also be replaced. For the case where the fourth technical text is a technical text in English, generate multiple pairs of aligned corpus with the same semantics in the same way.
[0130] According to the first embedding model, convert each determined technical text into a semantic vector, and obtain the semantic vectors corresponding to the fourth technical text, the translation of the fourth technical text, and the translation of the fifth technical text respectively.
[0131] In the total text set, except for the fourth technical text, the technical texts corresponding to the first language are denoted as the sixth technical text, and based on the second model (i.e., the monolingual embedding model for processing the first language), obtain the semantic vector of each sixth technical text.
[0132] In the total text set, except for the fourth technical text, the technical texts corresponding to the second language are denoted as the seventh technical text, and based on the third model (i.e., the monolingual embedding model for processing the second language), obtain the semantic vector of each seventh technical text.
[0133] Calculate the similarity between the semantic vector of the fourth technical text and the semantic vectors of each sixth technical text, determine the target sixth technical text that meets the pre-set similarity threshold from each sixth technical text, and determine the semantic vector of the target sixth technical text.
[0134] Calculate the similarity between the semantic vector of the translation of the fourth technical text and the semantic vectors of each seventh technical text, determine the target seventh technical text that meets the pre-set similarity threshold from each seventh technical text, and determine the semantic vector of the target seventh technical text.
[0135] It should be noted that when determining the target sixth semantic vector and the seventh semantic vector, in order to improve processing efficiency, a vector database corresponding to each technical text in the first language and the second language can be constructed based on the total text set. For example, the vector database of the first language can include the semantic vectors of all technical texts corresponding to the first language in the total text set (that is, for all non-first-language technical texts in the total text set, obtain their translations in the first language, and do not process the first-language technical texts), and the vector database of the second language can be obtained in the same way.
[0136] Based on the vector databases corresponding to the first language and the second language, the target sixth technical text corresponding to the fourth technical text and the target seventh technical text corresponding to the translation of the fourth technical text can be directly determined from the two vector databases according to the retrieval algorithm.
[0137] It is understandable that the similarity threshold is used to screen out the technical texts with different semantics from the fourth technical text from each sixth technical text and each seventh technical text based on the respective similarities. That is, the similarity threshold indicates the tail interval of the similarity value. The similarity thresholds used to screen the sixth technical text or the seventh technical text can be the same or different. The specific value of the similarity threshold can be the direct similarity value or other values reflecting the similarity magnitude (such as the normal distribution value interval of similarity, the sorting value interval), etc., which can be determined according to actual needs.
[0138] After obtaining the target sixth technical text and the target seventh technical text, the corresponding semantic vectors of each technical text are obtained according to the first embedding model.
[0139] Taking the semantic vector of the obtained fourth technical text as the original sample, taking the semantic vectors corresponding to the translation of the fourth technical text, the fifth technical text, and the translation of the fifth technical text as positive samples, and taking the semantic vectors of the target sixth technical text and the target seventh technical text as negative samples, a second sample set is obtained.
[0140] It is understandable that in machine learning and natural language processing (Natural Language Processing, abbreviated as NLP) tasks, a negative sample is a data sample that does not match the target or the correct answer, and is usually used to train the model to learn to distinguish between correct and incorrect situations. In frameworks such as contrastive learning (such as SimCSE), negative samples can help the model better identify dissimilar samples to improve the performance of the model, and their role is particularly important.
[0141] For example, the first model obtained after the BCE model is fine-tuned by supervised SimCSE is denoted as the BCE-SF-SimCSE model, where "S" in the model represents supervised, and "F" represents fine-tune.
[0142] Figure 3 It is a schematic flowchart of a supervised SimCSE training method provided by an embodiment of the present disclosure. As Figure 3 shown, for each training batch of the total text set, based on the technical text corresponding to the respective original samples, positive samples are determined based on the citation relationship between the technical texts, and negative samples are determined by combining the semantic vectors generated by the single-language embedding model and combining similarity calculation and screening methods to obtain a second sample set. The first embedding model is fine-tuned by supervised SimCSE using multiple second sample sets to obtain the BCE-SF-SimCSE model. Among them, the acquisition methods of positive samples and negative samples are not elaborated here.
[0143] The contrast loss function formula is as follows:
[0144]
[0145] In the formula, represents the original sample, represents the positive sample, represents the negative sample, is the temperature coefficient, is the cosine similarity.
[0146] Based on the second model and the third model for processing a single language in the embodiments of the present disclosure, multiple groups of aligned corpora with the same semantics are obtained. Based on the citation relationship between technical texts, positive samples are constructed, so that on the basis of semantic similarity, the first embedding model can also learn the aligned semantics. By combining the screening of the similarity between semantic vectors, the semantic vectors of the technical texts with different semantics from the corresponding technical texts of the original samples are screened out from the technical texts in the first language and the second language as negative samples, which can make the obtained negative samples more accurate in single-language expression, enable the multi-language embedding model to better capture language features and semantic relationships during training, and improve the performance of the model when processing technical texts.
[0147] In the above manner, for the total text set, multiple batches of second sample sets can be constructed, and the first embedding model is supervised and trained batch by batch according to each batch of second sample sets.
[0148] During the training of each batch, the original samples and the positive samples in the second sample set form positive sample pairs, and the original samples and the negative samples form negative sample pairs. With the minimization of the contrast loss function as the optimization goal, the contrast loss is backpropagated and the parameters of the first embedding model are fine-tuned (update the model parameters), and the updated model is saved until the fine-tuning is completed to obtain the first model.
[0149] The technical solution provided by the embodiments of the present disclosure adopts a supervised contrast learning mechanism, considers the relevance between texts, annotates positive samples based on the citation relationship between technical texts, enables the positive samples to help the model capture complex technical concepts, and determines negative samples by combining the semantic vectors generated by the single-language embedding model and combining similarity calculation and screening. On the basis of ensuring that the negative samples have sufficient differences in semantics from the original samples, it can more accurately capture the language features and semantic relationships between different languages, helps the model learn more refined discrimination capabilities, and can effectively improve the performance of the multi-first model when processing technical texts, meeting the requirements for understanding complex technical content and cross-language semantic relationships.
[0150] In a possible implementation manner, the second model is generated in the following way:
[0151] Determine the sub-text set corresponding to the first language from the total text set in the target technical field;
[0152] Retrain the pre-trained monolingual embedding model using a subset of texts, and repeat the following first operation until the training stop condition is met to obtain a second model;
[0153] Among them, the first operation includes:
[0154] Perform a masking operation on the target segment in the eighth technical text in the subset of texts to obtain a masked text;
[0155] Input the masked text into the monolingual embedding model to obtain the semantic vector of the masked text;
[0156] Obtain an output result through a decoding operation on the semantic vector of the masked text, and the output result is used to indicate the prediction result of the target segment;
[0157] Update the parameters of the monolingual embedding model in this round of iteration according to the loss between the target segment and the prediction result to obtain the monolingual embedding model in the next round of iteration.
[0158] Specifically, before applying the second model, a dual structure of an encoder and a decoder can also be introduced to perform secondary pre-training on the embedding model to improve the text representation ability. The architectures adopted include but are not limited to Masked Autoencoders and RetroMAE (a retrieval-oriented pre-training paradigm based on MAE), etc.
[0159] Determine the technical texts corresponding to the target language from the total text set in the target technical field, and construct a subset of texts. The subset of texts includes multiple technical texts in the first language, denoted here as the eighth technical text. The eighth technical text can be obtained by screening or translating each technical text in the total text set. Retrain the pre-trained monolingual embedding model using the subset of texts, and repeat the following first operation:
[0160] Retrain the pre-trained monolingual embedding model using the subset of texts. In each training process, determine the corresponding eighth technical text. It can be understood that the number of eighth technical texts corresponding to one training can be determined according to actual needs.
[0161] For each eighth technical text, encode the eighth technical text in the subset of texts through an encoder, perform a masking operation on the target segment in the eighth technical text to obtain a masked text. Input the masked text into the monolingual embedding model that processes the target language text to obtain the semantic vector of the masked text.
[0162] Through the decoder, perform a decoding operation on the semantic vector of the masked text, predict the target segment of the mask in the masked text, obtain an output result, and the output result is used to indicate the prediction result of the target segment.
[0163] For the loss between the target segment corresponding to each eighth technical text and the prediction result, the total loss of the current round of iterative training can be determined, and the parameters of the encoder and decoder are updated based on the total loss. Based on the updated parameters of the encoder and decoder, the parameters of the monolingual embedding model of the current round of iteration are updated to obtain the monolingual embedding model of the next round of iteration.
[0164] Repeat the above first operation to perform iterative training on the monolingual embedding model until the training stop condition is met to obtain the second model, and the training stop condition can be determined according to actual needs.
[0165] RetroMAE (Retrospective Masked Autoencoder) performs secondary pre-training on the embedding model. The purpose of the secondary pre-training is to further improve the performance of the model when dealing with technical texts in the target language of the target domain, so that it can more accurately capture the language features and semantic relationships in the technical texts.
[0166] It can be understood that the secondary pre-training method of the third model is the same as that of the second model, and will not be elaborated here.
[0167] The technical solution provided by the embodiments of the present disclosure performs secondary pre-training on the monolingual embedding model by introducing the dual structure of the encoder and decoder to further improve the performance of the second model when dealing with technical texts in the target language of the target technical field, so that it can more accurately capture the language features and semantic relationships in the technical texts of the target technical field. In this way, the quality of the negative samples during the supervised training of the first embedding model can be further improved.
[0168] In a possible implementation, calculate the similarity between each first semantic vector and each second semantic vector, and based on the similarity, obtain the target semantic vector from each second semantic vector, including:
[0169] For each first semantic vector, use the search algorithm corresponding to the index structure of the pre-constructed vector database to determine a preset number of second semantic vectors with the largest similarity from the second semantic vectors stored in the vector database as the target semantic vectors;
[0170] Among them, the vector database stores the semantic vectors of each technical text in the total text set.
[0171] Specifically, after converting each technical text into a semantic vector, directly calculating the similarity matrix between semantic vectors will consume huge storage space and computing resources, and the increase in the vector dimension will lead to an exponential growth in the computational complexity. The embodiments of the present disclosure combine the index structure of the vector database to achieve efficient data management, and through the corresponding search algorithm, improve the similarity between each semantic vector, and improve the efficiency of obtaining the target semantic vector from each second semantic vector, and shorten the retrieval time.
[0172] For the semantic vectors of each technical text in the obtained total text set, each semantic vector is stored in the vector database, and the index of the vector database is constructed. The type of the vector database can be determined according to actual needs, including but not limited to Faiss, Elasticsearch, and Milvus, etc.
[0173] When determining the target semantic vector corresponding to the first semantic vector, for each first semantic vector, the semantic vectors from different objects from the first semantic vector are denoted as second semantic vectors. That is, the semantic vectors belonging to the first object in the vector database are the first semantic vectors, and the semantic vectors belonging to non-first objects (each second object) are the second semantic vectors.
[0174] Using the index structure, execute the search algorithm corresponding to the pre-constructed index structure of the vector database, search for a preset number of second semantic vectors with the largest similarity to the first semantic vector from the second semantic vectors stored in the vector database, and use the searched second semantic vectors as the target semantic vectors.
[0175] It can be understood that the index structure and the corresponding search algorithm can be determined according to actual needs. Among them, the type of the index structure can include but not limited to Inverted File Product Quantization Index (IVF_PQ for short), Inverted File Flat Index (IVF_FLAT for short), and Hierarchical Navigable Small World (HNSW for short), etc. The search algorithm can be determined according to the corresponding index structure, and the present embodiment does not make any limitations in this regard.
[0176] It should be noted that in addition to the steps of determining the target semantic vector, the vector database can also be applied in the relevant steps of determining negative samples based on the similarity between semantic vectors. The specific application method is the same as the principle described above, and will not be elaborated here.
[0177] The technical solution provided by the embodiments of the present disclosure uses a vector database in combination with an index structure and a search algorithm for similarity retrieval, which can efficiently find similar semantic vectors in a large-scale vector space, improve the efficiency of vector similarity calculation, shorten the retrieval time, and save computing resources.
[0178] In a possible implementation, the index of the vector database is constructed in the following manner:
[0179] According to the first model, obtain the semantic vectors of each technical text in the total text set;
[0180] Construct the graph index of the vector database according to the hierarchical navigable small world graph index structure;
[0181] Among them, the graph index includes multiple layers of navigable small world graphs arranged from top to bottom and connected directionally. Each layer of the graph includes multiple nodes. Each node in the bottom layer graph is respectively used to represent the semantic vector of each technical text in the total text set; for any non-bottom layer graph, each node in this graph is a representative node among at least one node in the next layer graph; in each layer of the graph, the edge between any node pair indicates the similarity between the two semantic vectors represented by the node pair.
[0182] Specifically, the embodiments of the present disclosure apply the hierarchical navigable small world graph as the index structure. After obtaining the semantic vectors of each technical text in the total text set according to the first model, construct the graph index of the vector database (i.e., the hierarchical navigable small world graph index) according to the hierarchical navigable small world graph index structure.
[0183] The graph index is composed of multiple layers of navigable small world graphs, and each layer of navigable small world graph is arranged from top to bottom and connected directionally. In each layer of the graph, there is at least one node, and the edge between any node pair indicates the similarity between the two semantic vectors represented by the node pair.
[0184] In different graphs, the information granularity reflected by the nodes is different. In the bottom layer graph, each node directly corresponds to the semantic vector of a corresponding technical text in the total text set. For any non-bottom layer graph, for any non-bottom layer graph, each node in this graph is a representative node among at least one node in the next layer graph, that is, a node in the upper layer graph can represent a node group composed of multiple nodes with similarity in the semantic space in the next layer graph.
[0185] In the search stage, by utilizing the sparse connection characteristics of the hierarchical navigable small world graph, the memory usage is significantly reduced, which is especially suitable for the processing of large-scale data sets. The hierarchical navigable small world graph can effectively balance the query speed and accuracy through its multi-level graph structure.
[0186] For example, after constructing the index structure corresponding to the vector database, the first semantic vector can be input into the hierarchical navigable small world graph index through a graph-based approximate nearest neighbor search algorithm, and retrieved in each layer of the navigable small world graph in sequence from top to bottom until multiple neighboring nodes of the first semantic vector are determined from the multiple nodes included in the bottommost navigable small world graph. Nodes corresponding to semantic vectors that belong to the first object together with the first semantic vector are removed from the multiple neighboring nodes, and a preset number of neighboring nodes with the largest similarity are determined from each neighboring node according to the magnitude of the similarity, and the corresponding target semantic vectors are determined according to the obtained neighboring nodes.
[0187] The technical solution provided by the embodiments of the present disclosure uses a hierarchical navigable small world graph as the index structure. When performing a similarity search, the number of node connections in each layer of the graph is restricted within a small range, greatly reducing the search path length and computational complexity, significantly improving the search efficiency, enabling similar vectors to be found efficiently in a large-scale vector space, and significantly improving the retrieval speed and accuracy of similar texts.
[0188] In a possible implementation manner, for each second object, the competition intensity between the second object and the first object is obtained according to the similarities corresponding to the respective target semantic vectors of the second object, including:
[0189] For each second object, the similarities corresponding to the respective target semantic vectors of the second object are weighted and fused to obtain the total similarity;
[0190] The total similarity of each second object is normalized, and the normalized total similarity is used as the competition intensity between the second object and the first object.
[0191] Specifically, for each determined second object, at least one target semantic vector corresponding to the second object is obtained, and the similarities corresponding to the respective target semantic vectors of the second object are weighted and fused to obtain the total similarity, which is the total similarity between the second object and the first object. The weights applied during the weighted fusion can be set according to actual requirements.
[0192] After obtaining the total similarity of each second object, the total similarity of each second object is normalized based on the maximum total similarity value, and the normalized total similarity is used as the competition intensity between the second object and the first object.
[0193] The technical solution provided by the embodiments of the present disclosure calculates the total similarity between the second object and the first object, and through normalization, converts the total similarity into a quantifiable value with a consistent measurement standard as the competition intensity, which can more objectively evaluate the proximity of different second objects to the first object at the technical level, making the competition intensity between different second objects comparable, so as to accurately screen out the most likely technical competitors.
[0194] Considering that patent data is a direct manifestation of technological innovation and R & D activities, containing a large amount of information about technological development trends, R & D achievements, and technological layouts. The following combines a specific application example, taking patent data as the source of technical text and multiple languages as Chinese and English as an example, to illustrate the method for identifying technical competitors provided by the embodiments of the present disclosure:
[0195] The embodiments of the present disclosure evaluate existing multiple multilingual embedding models (including OpenAI, BGE-Large, BCE, and GTE-Large models), and construct a bilingual and cross-lingual evaluation dataset covering multiple fields such as computer science, physics, biology, economics, mathematics, and quantitative finance. The evaluation uses two key indicators: hit rate and mean reciprocal rank. The hit rate reflects the ability of the model to find the correct answer in the previous few retrievals, while the mean reciprocal rank reflects the average accuracy of the model in locating relevant documents in all queries. The larger these two indicators are, the better.
[0196] Based on the evaluation results, comprehensively considering the hit rate and mean reciprocal rank, the embodiments of the present disclosure select the BCE model as the first embedding model. The BCE model makes full use of the advantages of the Youdao translation engine, has excellent bilingual and cross-lingual capabilities, and eliminates the differences between Chinese and English languages in semantic retrieval, thereby achieving a powerful bilingual and cross-lingual semantic representation ability.
[0197] The embodiments of the present disclosure are based on the Chinese and English patent text data of patents in the new material field, and select the "title + abstract" text obtained by combining the title and abstract of each patent document as the technical text to construct the total text set. It can be understood that in the total text set, for each patent document, based on the translation of the technical text, the Chinese version and the English version of the technical text can be obtained, effectively ensuring the consistency and alignment of the data in different languages, ensuring the content consistency of the cited patents in a multilingual environment, and helping to achieve accurate matching in subsequent cross-lingual similarity calculations.
[0198] Based on the total text set, for each technical text, according to the citation relationship of the corresponding patent document, retrieve the corresponding cited patent document from the patent database, and construct positive samples according to the technical text corresponding to the cited patent document.
[0199] It is understandable that in actual operation, there may be cases of missing data in the cited patent fields. In addition, there may be differences in the integrity of Chinese and foreign patent data. Moreover, the number of patents cited in each target patent document may vary, resulting in inconsistent numbers of cited patent texts finally generated for each target patent document. Therefore, in the actual data processing process, the impact of the inconsistency in quantity on model training can be considered to ensure the accuracy and reliability of the mining results.
[0200] For the annotation of negative samples, select chinese-bert-wwm-ext as the Chinese embedding model and bert-for-patents as the English embedding model. Represent each technical text as a vector and store it in the vector database. By constructing a hierarchical navigation small-world graph as the index structure and using cosine similarity as the similarity metric, for each technical text, select 6 to construct negative samples from the pre-set similarity threshold range according to the similarity between semantic vectors.
[0201] Based on the negative sample and positive sample annotation methods described above, construct the corresponding second sample set for each batch (Batch).
[0202] For example, when the original sample comes from Chinese technical texts, the positive samples come from the English translation of the Chinese technical text + the Chinese-English technical text pairs of 3 cited patents, and the negative samples are 3 Chinese technical texts mined by the chinese-bert-wwm-ext model pre-trained twice by RetroMAE + 3 English technical texts mined by the bert-for-patents model pre-trained twice by RetroMAE.
[0203] Correspondingly, when the original sample comes from English technical texts, the positive samples come from the Chinese translation of the English technical text + the Chinese-English technical text pairs of 3 cited patents, and the negative samples are 3 Chinese technical texts mined by the chinese-bert-wwm-ext model pre-trained twice by RetroMAE + 3 English technical texts mined by the bert-for-patents model pre-trained twice by RetroMAE.
[0204] After obtaining the corresponding technical texts, process the technical texts with the BCE model to obtain the corresponding semantic vectors, and obtain the original texts, negative samples, and positive samples included in each second sample set.
[0205] Based on the obtained multiple second sample sets, fine-tune the BCE model using the supervised SimCSE fine-tuning method to obtain the BCE-SF-SimCSE model.
[0206] For example,Figure 4 The figure is a schematic flow chart of a method for identifying a technical competitor of a patentee provided by an embodiment of the present disclosure. As Figure 4 shown, the BCE-SF-SimCSE model is used to obtain the semantic vector corresponding to each technical text in the total text set, and after indexing and storing in the vector database, patent similarity calculation is performed, and based on this, the technical competitor is identified.
[0207] For the patentee A (the first object), obtain the patent list M of the patentee A, and the patent list M includes multiple patents m.
[0208] For each patent m, find N similar patents similar to the patent m from the vector database, obtain the similar patent set Nm, the similarity set Sm, and the patentee (the second object) set Pm. It can be understood that there is a mapping relationship between the data in the above three sets.
[0209] For each patentee in the patentee set Pm, determine the similarity data related to it from the similarity set Sm, and perform weighted summation to calculate the total similarity corresponding to each patentee.
[0210] Perform normalization processing on the total similarity corresponding to each patentee, and calculate the competition intensity corresponding to each patentee.
[0211] Output the patentee set Pm corresponding to the patentee A, and mark the corresponding competition intensity, and select the top 5 patentees with the competition intensity as technical competitors.
[0212] Identifying technical competitors using patent data has the following advantages:
[0213] Comprehensiveness and systematicness: Patent data covers technical innovation information worldwide, and patent data is systematic and standardized. By analyzing patent data, the competition situation in a specific technical field can be comprehensively understood, and technical competitors can be comprehensively identified.
[0214] Timeliness and dynamics: Patent data is constantly updated, which can reflect the latest technological developments and competition dynamics. By analyzing the latest patent data, emerging technical competitors can be identified in a timely manner, and technological trends can be understood, helping enterprises to maintain a leading position in technological competition.
[0215] Technical details and innovation points: Patent data details the technical innovation details and core innovation points. By analyzing patent data, the technical strength and innovation ability of competitors can be deeply understood. This information has important reference value for evaluating the technical level of competitors and formulating technology R & D strategies.
[0216] Legal protection and competitive barriers: Patents have legal protection functions. By analyzing patent data, one can understand the protection strategies and competitive barriers of competitors in the technical field. This helps enterprises avoid infringement risks in technology R & D and market promotion, and build their own technical protection systems by applying for patents.
[0217] Technical cooperation and merger & acquisition opportunities: Patent data can also reveal potential technical cooperation partners or merger & acquisition targets. By analyzing the patent layout of competitors, one can discover opportunities for technological complementarity or cooperation, thereby promoting technical cooperation and resource integration and enhancing the competitiveness of enterprises.
[0218] In summary, using patent data to identify technical competitors can not only comprehensively understand the technical layout and innovation capabilities of competitors, but also provide an important basis for enterprises' technology R & D, market expansion and strategic planning.
[0219] The embodiments of the present disclosure use a massive multilingual patent dataset as the basis of technical texts. By designing an artificial intelligence algorithm to construct a first model for converting technical texts into semantic vectors, it can enable the identification of technical competitors to cross language and country limitations without relying on any manual participation, thus greatly saving the human resource cost for judging technical competitors and providing strong support for analysis scenarios such as global industrial analysis and technical layout analysis.
[0220] It can be understood that the embodiments of the present disclosure are only used to illustrate the present invention in detail as a specific example, and the described embodiments are only intended to facilitate the understanding of the present invention and do not impose any limitations on it.
[0221] Figure 5 The following is a schematic structural diagram of an apparatus for identifying technical competitors provided by an embodiment of the present disclosure. As Figure 5 shown, the apparatus 500 for identifying technical competitors includes:
[0222] A technical text acquisition module 501, configured to acquire a text set, where the text set includes multiple first texts, and the first text is a technical text of a first object in the target technical field;
[0223] A semantic vector acquisition module 502, configured to obtain a first semantic vector of each first text according to a pre-trained first model, and the first model is obtained by re-training a pre-trained multilingual first embedding model with a contrastive learning mechanism based on a total text set, and the total text set includes multilingual technical texts in the target technical field;
[0224] The similarity comparison module 503 is configured to calculate the similarity between each first semantic vector and each second semantic vector, and obtain a target semantic vector from each second semantic vector according to the similarity. The second semantic vector is the semantic vector of the technical text of the second object in the target technical field;
[0225] The competition intensity obtaining module 504 is configured to, for each second object, obtain the competition intensity between the second object and the first object according to the similarity corresponding to each target semantic vector corresponding to the second object. The competition intensity characterizes the similarity degree of the technologies between the first object and the second object;
[0226] The competitor determination module 505 determines the technical competitors of the first object from each second object according to the competition intensity of the second object.
[0227] In a possible implementation manner, the technical competitor identification device further includes: a model generation module; the model generation module is configured to generate a first model;
[0228] The first model is generated in the following manner:
[0229] According to the total text set, construct multiple batches of first sample sets;
[0230] According to each batch of first sample sets, perform unsupervised training on the first embedding model in batches, and take minimizing the contrast loss function as the optimization objective to fine-tune the parameters of the first embedding model to obtain the first model;
[0231] Wherein, any batch of first sample sets is constructed in the following manner:
[0232] Determine the first technical text corresponding to the batch from the total text set, and process the first technical text with data augmentation technology to obtain a number of second technical texts;
[0233] Obtain a number of technical texts different from the first technical text from the total text set as the third technical text;
[0234] According to the first embedding model, obtain the semantic vectors corresponding to the first technical text, the second technical text, and the third technical text respectively;
[0235] Take the semantic vector of the first technical text as the original sample, the semantic vector of the second technical text as the positive sample, and the semantic vector of the third technical text as the negative sample to obtain the first sample set.
[0236] In a possible implementation manner, the technical competitor identification device further includes: a model generation module; the model generation module is configured to generate a first model;
[0237] The first model is generated in the following manner:
[0238] Determine the fourth technical text in the first preset language corresponding to the batch from the total text set, and the fifth technical text having a citation relationship with the fourth technical text;
[0239] For the fourth technical text and each fifth technical text, obtain the translation of the fourth technical text and the translation of the fifth text according to the second preset language;
[0240] According to the first embedding model, obtain the semantic vectors corresponding to the fourth technical text, the fifth technical text, the translation of the fourth technical text, and the translation of the fifth technical text respectively;
[0241] Calculate the similarity between the semantic vector of the fourth technical text and the semantic vectors of each sixth technical text in the target technical field, and determine the semantic vector of the target sixth technical text that meets the preset similarity threshold from each sixth technical text. The sixth technical text corresponds to the first language, and the semantic vector of the sixth technical text is obtained by the second model processing the sixth technical text. The second model is a single-language embedding model for processing the target language;
[0242] Calculate the similarity between the semantic vector of the translation of the fourth technical text and the semantic vectors of each seventh technical text in the target technical field, and determine the semantic vector of the target seventh technical text that meets the preset similarity threshold from each seventh technical text. The seventh technical text corresponds to the second language, and the semantic vector of the seventh technical text is obtained by the third model processing the seventh technical text. The third model is a single-language embedding model for processing the second language;
[0243] According to the first embedding model, obtain the semantic vectors corresponding to the target sixth technical text and the target seventh technical text respectively;
[0244] Take the semantic vector of the fourth technical text as the original sample, take the semantic vectors corresponding to the translation of the fourth technical text, the fifth technical text, and the translation of the fifth technical text as positive samples, and take the semantic vectors corresponding to the target sixth technical text and the target seventh technical text as negative samples to obtain the second sample set.
[0245] In a possible implementation manner, the model generation module is used to generate the second model;
[0246] The second model is generated in the following manner:
[0247] Determine the subset of texts corresponding to the first language from the total text set in the target technical field;
[0248] Use the subset of texts to retrain the pre-trained single-language embedding model, and repeatedly execute the following first operation until the training stop condition is met to obtain the second model;
[0249] Wherein, the first operation includes:
[0250] Perform a masking operation on the target segment in the eighth technical text in the pair of text sets to obtain a masked text;
[0251] Input the masked text into a monolingual embedding model to obtain the semantic vector of the masked text;
[0252] Obtain an output result through a decoding operation on the semantic vector of the masked text, and the output result is used to indicate the prediction result of the target segment;
[0253] Update the parameters of the monolingual embedding model in this round of iteration according to the loss between the target segment and the prediction result to obtain the monolingual embedding model for the next round of iteration.
[0254] In a possible implementation, the similarity comparison module is used for:
[0255] For each first semantic vector, use the search algorithm corresponding to the index structure of the pre-constructed vector database to determine a preset number of second semantic vectors with the highest similarity from the second semantic vectors stored in the vector database as the target semantic vectors;
[0256] Among them, the vector database stores the semantic vectors of each technical text in the total text set.
[0257] In a possible implementation, the technical competitor identification device further includes: a database processing module; the database processing module is used to construct an index of the vector database;
[0258] The index of the vector database is constructed in the following manner:
[0259] Obtain the semantic vectors of each technical text in the total text set according to the first model;
[0260] Construct a graph index of the vector database according to the hierarchical navigable small world graph index structure;
[0261] Among them, the graph index includes multiple layers of navigable small world graphs arranged from top to bottom and connected in a directed manner. Each layer of the graph includes multiple nodes. Each node in the bottom layer graph is used to represent the semantic vector of each technical text in the total text set; for any non-bottom layer graph, each node in this graph is a representative node of at least one node in the next layer graph; in each layer of the graph, the edge between any node pair indicates the similarity between the two semantic vectors represented by the node pair.
[0262] In a possible implementation, the competition intensity acquisition module is used for:
[0263] For each second object, perform weighted fusion on the similarities corresponding to the respective target semantic vectors of the second object to obtain the total similarity;
[0264] Normalize the total similarity of each second object, and use the normalized total similarity as the competition intensity between the second object and the first object.
[0265] The device according to the embodiments of the present disclosure can execute the method provided by the embodiments of the present disclosure, and the implementation principle is similar. The actions performed by each module in the device according to the embodiments of the present disclosure correspond to the steps in the methods according to the embodiments of the present disclosure. For the detailed function description of each module of the device, reference can be specifically made to the description in the corresponding method shown above, and details are not described herein again.
[0266] In addition, in the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit including the function of the module or unit.
[0267] An electronic device (computer device / equipment / system) is provided in the embodiments of the present disclosure, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the method provided by any optional embodiment of the present disclosure, and achieve the corresponding technical effects.
[0268] In an optional embodiment, an electronic device is provided. Figure 6 As a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure, as Figure 6 shown, the electronic device 600 includes: a processor 601 and a memory 603. Among them, the processor 601 and the memory 603 are connected, such as connected through a bus 602. Optionally, the electronic device 600 may further include a transceiver 604, and the transceiver 604 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation to the embodiments of the present disclosure.
[0269] The processor 601 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the present disclosure. The processor 601 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0270] The bus 602 may include a path for transmitting information between the above components. The bus 602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 602 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0271] The memory 603 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.
[0272] The memory 603 is used to store the computer program for implementing the embodiments of the present disclosure, and is controlled by the processor 601 to execute. The processor 601 is used to execute the computer program stored in the memory 603 to implement the steps shown in the foregoing method embodiments.
[0273] The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0274] The embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding content shown in the foregoing method embodiments can be implemented.
[0275] The embodiments of the present disclosure also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding content shown in the foregoing method embodiments can be implemented.
[0276] It should be noted that the computer-readable storage medium in the present disclosure above may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0277] In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination of the foregoing.
[0278] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include but are not limited to object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0279] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description, claims, and drawings of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than that shown or described in words.
[0280] It should be understood that although the flowcharts in the embodiments of the present disclosure indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart may be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenarios. Some or all of these sub-steps or stages may be executed at the same time, and each sub-step or stage among these sub-steps or stages may also be executed at different times. In scenarios where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present disclosure do not limit this.
[0281] The above are only optional implementation manners of some implementation scenarios of the present disclosure. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the technical concept of the solution of the present disclosure, adopting other similar implementation means based on the technical idea of the present disclosure also belongs to the protection scope of the embodiments of the present disclosure.
Claims
1. A method for identifying technical competitors, characterized in that Including: Obtain a text set, where the text set includes multiple first texts, and the first texts are technical texts of a first object in the target technical field; According to a pre-trained first model, obtain first semantic vectors of each first text, where the first model is obtained by retraining a pre-trained multi-lingual first embedding model with a contrastive learning mechanism based on a total text set, and the total text set includes multi-lingual technical texts in the target technical field; Calculate the similarity between each first semantic vector and each second semantic vector, and based on the similarity, obtain a target semantic vector from each second semantic vector, where the second semantic vector is the semantic vector of the technical text of a second object in the target technical field; For each second object, obtain the competition intensity between the second object and the first object according to the similarity corresponding to each target semantic vector corresponding to the second object, where the competition intensity represents the similarity degree of the technology between the first object and the second object; Determine the technical competitors of the first object from each second object according to the competition intensity of the second object; Wherein, the contrastive learning mechanism includes an unsupervised SimCSE algorithm or a supervised SimCSE algorithm; When the unsupervised SimCSE algorithm is adopted, positive samples are augmented with a data augmentation strategy, and the first embedding model is retrained in a batch training manner using a contrastive loss function to obtain the first model; When the supervised SimCSE algorithm is adopted, positive samples are labeled based on the citation relationship between technical texts, negative samples are labeled based on the similarity between technical texts, and the first embedding model is retrained in a batch training manner using a contrastive loss function to obtain the first model.
2. The method for identifying a technological competitor according to claim 1, wherein The first model is generated in the following manner: Construct multiple batches of first sample sets according to the total text set; According to each batch of first sample sets, perform unsupervised training on the first embedding model in batches, taking minimizing the contrastive loss function as the optimization target, fine-tuning the parameters of the first embedding model, and obtaining the first model; Wherein, any batch of first sample sets is constructed in the following manner: Determine the first technical text corresponding to the batch from the total text set, and process the first technical text with a data augmentation technique to obtain several second technical texts; Obtain several technical texts different from the first technical text from the total text set as third technical texts; According to the first embedding model, obtain the semantic vectors corresponding to the first technical text, the second technical text, and the third technical text respectively; Use the semantic vector of the first technical text as the original sample, the semantic vector of the second technical text as the positive sample, and the semantic vector of the third technical text as the negative sample to obtain the first sample set.
3. The method for identifying a technological competitor according to claim 1, wherein The first model is generated in the following manner: Construct multiple batches of second sample sets according to the total text set; Supervise and train the first embedding model batch by batch according to the second sample sets of each batch, aiming to minimize the contrast loss function, and fine-tune the parameters of the first embedding model to obtain the first model; Among them, the second sample set of any batch is constructed in the following way: Determine the fourth technical text in the first preset language corresponding to the batch and the fifth technical text having a citation relationship with the fourth technical text from the total text set; For the fourth technical text and each fifth technical text, obtain the fourth technical text translation and the fifth text translation according to the preset second language; According to the first embedding model, obtain the semantic vectors corresponding to the fourth technical text, the fifth technical text, the fourth technical text translation, and the fifth technical text translation respectively; Calculate the similarity between the semantic vector of the fourth technical text and the semantic vectors of each sixth technical text in the target technical field, and determine the semantic vector of the target sixth technical text that meets the preset similarity threshold from each sixth technical text. The sixth technical text corresponds to the first language, and the semantic vector of the sixth technical text is obtained by the second model processing the sixth technical text. The second model is a single-language embedding model for processing the target language; Calculate the similarity between the semantic vector of the fourth technical text translation and the semantic vectors of each seventh technical text in the target technical field, and determine the semantic vector of the target seventh technical text that meets the preset similarity threshold from each seventh technical text. The seventh technical text corresponds to the second language, and the semantic vector of the seventh technical text is obtained by the third model processing the seventh technical text. The third model is a single-language embedding model for processing the second language; According to the first embedding model, obtain the semantic vectors corresponding to the target sixth technical text and the target seventh technical text respectively; Use the semantic vector of the fourth technical text as the original sample, use the semantic vectors corresponding to the fourth technical text translation, the fifth technical text, and the fifth technical text translation as positive samples, and use the semantic vectors corresponding to the target sixth technical text and the target seventh technical text as negative samples to obtain the second sample set.
4. The method for identifying a technological competitor according to claim 3, wherein The second model is generated in the following way: Determine the sub-text set corresponding to the first language from the total text set in the target technical field; Retrain the pre-trained single-language embedding model through the sub-text set, and repeatedly execute the following first operation until the training stop condition is met to obtain the second model; Among them, the first operation includes: Perform a masking operation on the target segment in the eighth technical text in the sub-text set to obtain a masked text; Input the masked text into the single-language embedding model to obtain the semantic vector of the masked text; Through the decoding operation of the semantic vector of the masked text, obtain the output result, and the output result is used to indicate the prediction result of the target segment; Update the parameters of the monolingual embedding model in this round of iteration according to the loss between the target segment and the prediction result to obtain the monolingual embedding model in the next round of iteration.
5. The method for identifying a technical competitor according to any one of claims 1-4, characterized in that, The calculating the similarity between each first semantic vector and each second semantic vector, and obtaining a target semantic vector from each second semantic vector according to the similarity includes: For each first semantic vector, use the search algorithm corresponding to the index structure of the pre-constructed vector database to determine a preset number of second semantic vectors with the greatest similarity from the second semantic vectors stored in the vector database as the target semantic vectors; Wherein, the vector database stores the semantic vectors of each technical text in the total text set.
6. The method for identifying a technological competitor according to claim 5, characterized in that The index of the vector database is constructed in the following manner: Obtain the semantic vectors of each technical text in the total text set according to the first model; Construct a graph index of the vector database according to the hierarchical navigable small world graph index structure; Wherein, the graph index includes multiple hierarchical navigable small world graphs arranged from top to bottom and connected in a directed manner. Each layer of the graph includes multiple nodes. Each node in the bottom layer graph is respectively used to represent the semantic vectors of each technical text in the total text set; for any non-bottom layer graph, each node in this graph is a representative node of at least one node in the next layer graph; in each layer of the graph, the edge between any node pair indicates the similarity between the two semantic vectors represented by the node pair.
7. The method for identifying a technological competitor according to any one of claims 1 to 4, characterized in that For each second object, obtain the competition intensity between the second object and the first object according to the similarity corresponding to each target semantic vector corresponding to the second object, including: For each second object, perform weighted fusion on the similarities corresponding to each target semantic vector corresponding to the second object to obtain the total similarity; Perform normalization processing on the total similarity of each second object, and use the normalized total similarity as the competition intensity between the second object and the first object.
8. An identification device for technological competitors, characterized in that, Including: A technical text acquisition module, configured to acquire a text set, where the text set includes multiple first texts, and the first texts are technical texts of a first object in a target technical field; A semantic vector acquisition module, configured to obtain a first semantic vector of each first text according to a pre-trained first model, where the first model is obtained by retraining a pre-trained multilingual first embedding model according to a contrast learning mechanism based on a total text set, and the total text set includes multilingual technical texts in the target technical field; A similarity comparison module, configured to calculate the similarity between each first semantic vector and each second semantic vector, and obtain a target semantic vector from each second semantic vector according to the similarity, where the second semantic vector is the semantic vector of the technical text of a second object in the target technical field; A competition intensity acquisition module, configured to, for each second object, obtain the competition intensity between the second object and the first object according to the similarity corresponding to each target semantic vector corresponding to the second object, where the competition intensity characterizes the similarity degree of the technology between the first object and the second object; A competitor determination module determines the technical competitors of the first object from each second object according to the competition intensity of the second object; Among them, the contrastive learning mechanism includes an unsupervised SimCSE algorithm or a supervised SimCSE algorithm; When the unsupervised SimCSE algorithm is adopted, positive samples are augmented with a data augmentation strategy, and the first embedding model is retrained with a contrastive loss function in a batch training manner to obtain the first model; When the supervised SimCSE algorithm is adopted, positive samples are labeled based on the citation relationship between technical texts, and negative samples are labeled based on the similarity between technical texts. The first embedding model is retrained with a contrastive loss function in a batch training manner to obtain the first model.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Competitor mining method fused with multi-algorithm model
CN116823306A
Pre-training language model-based summarization generation method
US20230418856A1