Methods, devices, equipment, media and products for processing the intensity of regional technological competition

By obtaining a set of patent vectors in the target region and using a fine-tuned text embedding model to calculate the similarity of patent text features, the problem of low evaluation accuracy in the existing technology is solved, and a more accurate regional technology competition assessment is achieved.

CN120031683BActive Publication Date: 2025-09-19INST OF SCI & TECHN INFORMATION OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510164313.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-09-19
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In the existing technology, the regional technology competition assessment method based on patent text fails to effectively utilize patent semantic information, resulting in low assessment accuracy.

Method used

By obtaining the patent vector sets of the two target areas, using the fine-tuned text embedding model to extract patent text features, calculating the similarity between patent vectors, determining the intensity of technological competition based on the similarity, and performing weighted fusion to improve evaluation accuracy.

Benefits of technology

The accuracy of the assessment of technological competition between two target areas has been improved, and the precision of the assessment has been enhanced by utilizing patent semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031683B_ABST
    Figure CN120031683B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a method, apparatus, device, medium and product for processing the technical competition intensity of a region, involving fields such as regional competition, and application scenarios include but are not limited to regional technical competition scenarios. The method comprises: obtaining patent vector sets of two target regions, the patent vector set of each target region includes a subset of patent vectors corresponding to at least one technical category, the subset of patent vectors corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; for each technical category, based on the similarity between the patent vectors of the technical category of the two target regions, determining the technical competition intensity between the two target regions corresponding to the technical category; based on the technical competition intensity of the two target regions corresponding to various technical categories, determining the technical competition intensity between the two target regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and more specifically, to a method, apparatus, device, medium, and product for processing the technical competition intensity of a region. Background Art

[0002] In the existing technology, the technological competition between two regions is evaluated based on patent texts. The feature representation method of patent texts is implemented through traditional keyword extraction-based TF-IDF (Term Frequency–Inverse Document Frequency) or Word2Vec (word embedding model), which does not make good use of patent semantic information. The existing technology has conducted few studies on the technological competition between two regions from the perspective of patent texts. The technological competition between two regions is often more focused on the construction of indicators, resulting in low accuracy in evaluating the technological competition between the two regions. Summary of the Invention

[0003] In response to the shortcomings of existing methods, the present disclosure proposes a method, device, equipment, computer-readable storage medium and computer program product for processing the technological competition intensity of a region, which are used to solve the problem of how to improve the accuracy of evaluating the technological competition between two regions.

[0004] In a first aspect, the present disclosure provides a method for processing regional technological competition intensity, comprising:

[0005] Obtain patent vector sets for two target regions, where each patent vector set for the target region includes a patent vector subset corresponding to at least one technical category, and each patent vector subset corresponding to each technical category includes multiple patent vectors for that technical category. A patent vector for each technical category is a feature representation of a patent text for that technical category.

[0006] For each technology category, based on the similarity between the patent vectors of the technology category in the two target regions, determine the technology competition intensity corresponding to the technology category between the two target regions;

[0007] The technological competition intensity between the two target areas is determined based on the technological competition intensity of the two target areas corresponding to various technological categories.

[0008] In one embodiment, for each target region, the patent vector set of the target region is obtained by:

[0009] Acquire a plurality of patent texts corresponding to each technical category of at least one technical category in the target region;

[0010] For each patent text, input each patent text into the fine-tuned text embedding model for feature extraction to obtain the patent vector corresponding to each patent text;

[0011] Based on the patent vectors of each patent text in each technology category, a patent vector set for the target area is constructed.

[0012] In one embodiment, the fine-tuned text embedding model is trained by:

[0013] Obtain multiple patent training texts;

[0014] Input each patent training text into the pre-trained text embedding model to perform feature extraction and obtain the patent vector corresponding to each patent training text;

[0015] Determine the first similarity between patent vectors corresponding to each pair of patent training texts;

[0016] Obtaining the standard similarity between two patent training texts obtained through expert scoring;

[0017] Based on the first similarity and the corresponding standard similarity between each pair of patent training texts, the plurality of patent training texts are divided into positive samples and negative samples, and each positive sample is used as a positive sample set;

[0018] Each negative sample and the prompt word are input into the trained large language model for text enhancement to obtain patent text similar to each negative sample. The prompt word is used to instruct the trained large language model to output patent text similar to each negative sample.

[0019] Each negative sample and patent text similar to each negative sample are regarded as a negative sample set;

[0020] The positive sample set and the negative sample set are used as fine-tuning datasets. Based on the fine-tuning datasets, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model.

[0021] In one embodiment, based on the first similarity and the corresponding standard similarity between each pair of patent training texts, a plurality of patent training texts are divided into positive samples and negative samples, including:

[0022] For each patent training text, if the difference between each first similarity corresponding to the patent training text and the corresponding standard similarity is greater than a preset threshold, the patent training text is determined as a negative sample;

[0023] If the differences between each first similarity corresponding to the patent training text and the corresponding standard similarity are all less than or equal to a preset threshold, the patent training text is determined as a positive sample.

[0024] In one embodiment, for each technology category, the patent vector subsets of the technology category in the two target regions are used as the first vector set and the second vector set, respectively, to determine the technology competition intensity corresponding to the technology category between the two target regions, including:

[0025] Determining a second similarity between each patent vector in the first vector set and each patent vector in the second vector set;

[0026] According to the order of the second similarities from largest to smallest, a preset number of second similarities ranked first among the second similarities are determined as a similarity set of the two target areas corresponding to the technology category;

[0027] Based on the similarity set, the technical competition intensity of the two target areas corresponding to the technology category is determined.

[0028] In one embodiment, determining the technical competition intensity of two target regions corresponding to the technology category based on the similarity set includes:

[0029] A preset number of second similarities in the similarity set are fused, and a fusion result is determined as the technical competition intensity of the two target areas corresponding to the technical category.

[0030] In one embodiment, determining the technological competition intensity between the two target regions based on the technological competition intensity of the two target regions corresponding to various technological categories includes:

[0031] Get the weight corresponding to each technology category;

[0032] Based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weightedly integrated to obtain the technology competition intensity of the two target areas.

[0033] In one embodiment, obtaining a plurality of patent texts corresponding to each of at least one technical category in the target region includes:

[0034] Acquire regional information of the target area, where the regional information includes specific technical fields of the target area;

[0035] Obtain multiple patent texts for each technology category in at least one technology category related to a specific technology field.

[0036] In a second aspect, the present disclosure provides a device for processing regional technical competition intensity, comprising:

[0037] A first processing module is configured to obtain patent vector sets for two target regions, wherein the patent vector set for each target region includes a subset of patent vectors corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of that technical category, wherein a patent vector for each technical category is a feature representation of a patent text of that technical category;

[0038] A second processing module is configured to determine, for each technology category, the technology competition intensity corresponding to the technology category between the two target regions based on the similarity between the patent vectors of the technology category in the two target regions;

[0039] The third processing module is configured to determine the technological competition intensity between the two target areas based on the technological competition intensity of the two target areas corresponding to various technological categories.

[0040] In a third aspect, the present disclosure provides an electronic device, comprising: a processor, a memory, and a bus;

[0041] Bus, used to connect the processor and memory;

[0042] a memory for storing operation instructions;

[0043] The processor is configured to execute the method for processing the technical competition intensity of the region according to the first aspect of the present disclosure by calling an operation instruction.

[0044] In a fourth aspect, the present disclosure provides a computer-readable storage medium storing a computer program, which is used to execute the method for processing the technical competition intensity of a region according to the first aspect of the present disclosure.

[0045] In a fifth aspect, the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method for processing the technical competition intensity of a region in the first aspect of the present disclosure.

[0046] The technical solutions provided by the embodiments of the present disclosure have at least the following beneficial effects:

[0047] A patent vector set of two target areas is obtained, wherein the patent vector set of each target area includes a subset of patent vectors corresponding to at least one technical category, the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; for each technical category, based on the similarity between the patent vectors of the technical category in the two target areas, the technical competition intensity corresponding to the technical category between the two target areas is determined; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined; in this way, based on the patent vector (feature representation of the patent text, i.e., the patent semantic information in the patent text), the technical competition intensity between the two target areas corresponding to the technical category is calculated; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined, i.e., the technical competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technical competition between the two target areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for describing the embodiments of the present disclosure.

[0049] Figure 1 A schematic diagram of the architecture of a system for processing regional technical competition intensity provided by an embodiment of the present disclosure;

[0050] Figure 2 A flowchart of a method for processing regional technical competition intensity provided by an embodiment of the present disclosure;

[0051] Figure 3 A schematic diagram of processing the technical competition intensity of a region provided by an embodiment of the present disclosure;

[0052] Figure 4 A flowchart of a method for processing regional technical competition intensity provided by an embodiment of the present disclosure;

[0053] Figure 5 A schematic diagram of the structure of a device for processing the technical competition intensity of a region provided by an embodiment of the present disclosure;

[0054] Figure 6 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0055] The following describes embodiments of the present disclosure in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions of the embodiments of the present disclosure.

[0056] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to the connection relationship between the element and the other element established through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" indicates implementation as "A", or implementation as "B", or implementation as "A and B".

[0057] It is understandable that in the specific implementation of the present disclosure, when the above embodiments of the present disclosure are applied to specific products or technologies, data related to the processing of regional technological competition intensity is required, and the user's permission or consent is required, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0058] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.

[0059] The embodiment of the present disclosure is a method for processing regional technical competition intensity provided by a regional technical competition intensity processing system. The method for processing regional technical competition intensity involves fields such as regional competition.

[0060] In order to better understand and illustrate the solutions of the embodiments of the present disclosure, some technical terms involved in the embodiments of the present disclosure are briefly explained below.

[0061] BGE model: The BGE model is a series of text embedding models designed to convert text into low-dimensional dense vectors for efficient calculation and analysis.

[0062] Large language models: Large language models have demonstrated significant capabilities in text enhancement, which is mainly reflected in text understanding, generation, and processing of multimodal content.

[0063] A vector database is a database system specifically designed to store and manage high-dimensional vector data. It is widely used in fields such as machine learning, image processing, and natural language processing. Currently, mainstream vector databases include Faiss, Elasticsearch, and Milvus.

[0064] Faiss is an efficient similarity search library primarily used for similarity searches on dense vectors. It provides a variety of index structures, including flat indexes, quantized indexes, and hierarchical clustering indexes, enabling efficient approximate nearest neighbor searches on large datasets. Faiss's core strength lies in its highly optimized implementation and support for GPU (Graphics Processing Unit) acceleration, enabling it to excel in large-scale vector retrieval tasks. Faiss is suitable for scenarios requiring high precision and has been widely used in image retrieval and recommendation systems.

[0065] Elasticsearch is an open-source, full-text search engine with powerful full-text retrieval and analysis capabilities. In recent years, Elasticsearch has introduced vector search capabilities, enabling it to handle high-dimensional vector data. By integrating a vector similarity calculation plugin, Elasticsearch can perform approximate nearest neighbor searches based on vectors. Elasticsearch's advantage lies in its seamless integration with full-text search, making it suitable for hybrid query scenarios, such as applications requiring both text and vector search. However, because its vector search functionality is implemented using a plugin, its performance and scalability may not be as good as those of dedicated vector databases when handling very large amounts of vector data.

[0066] Milvus is an open-source vector database designed to provide a high-performance, highly scalable vector retrieval solution. Designed specifically to handle large-scale, high-dimensional vector data, Milvus is particularly well-suited for applications requiring real-time response and efficient retrieval. As a specialized vector database, Milvus demonstrates significant advantages in handling high-dimensional vector data. When using Milvus, vector indexing is a key element that requires special attention. This indexing mechanism, based on a specific mathematical model, constructs a data structure that is superior in both time and space efficiency. This type of indexing enables efficient retrieval of multiple vectors that are highly similar to a target vector. Commonly used vector indexing methods fall under the category of approximate nearest neighbor search. The core concept of this method is not to pursue absolute accuracy, but rather to significantly improve retrieval efficiency by exploring possible nearest neighbors, while moderately sacrificing accuracy within an acceptable range. Table 1 shows the indexing methods used in the Milvus vector database. In addition to indexing, the method for calculating distances between vectors is also a crucial aspect of vector databases. The distance calculation method directly affects the accuracy and efficiency of vector search. Currently, available methods for calculating the distance between vectors (vector distance calculation methods) include Euclidean distance, dot product (inner product), and cosine similarity. Table 2 shows the vector distance calculation methods.

[0067] Table 1: Milvus vector database indexing method

[0068]

[0069] Table 2: Vector distance calculation method

[0070]

[0071] The solutions provided by the embodiments of this disclosure involve technologies for processing regional technical competition intensity. The technical solutions of this disclosure are described in detail below using specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of this disclosure will be described below in conjunction with the accompanying drawings.

[0072] In order to better understand the solution provided by the embodiment of the present disclosure, the solution is described below in conjunction with a specific application scenario.

[0073] In one embodiment, Figure 1 FIG. 1 shows a schematic diagram of the architecture of a system for processing the technical competition intensity of a region applicable to an embodiment of the present disclosure. It can be understood that the method for processing the technical competition intensity of a region provided by the embodiment of the present disclosure can be applied to, but not limited to, Figure 1 In the application scenario shown.

[0074] In this example, Figure 1 As shown, the architecture of the regional technical competitiveness intensity processing system in this example may include but is not limited to a server 10, a terminal 20, and a database 30. The server 10, the terminal 20, and the database 30 may interact with each other through a network 40.

[0075] The server 10 obtains patent vector sets for two target areas, where the patent vector set for each target area includes a subset of patent vectors corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of that technical category, and a patent vector for each technical category is a feature representation of a patent text of that technical category. For each technical category, the server 10 determines the technical competition intensity between the two target areas corresponding to that technical category based on the similarity between the patent vectors of that technical category in the two target areas. The server 10 determines the technical competition intensity between the two target areas based on the technical competition intensity of the two target areas corresponding to various technical categories. The server 10 sends the technical competition intensity of the two target areas corresponding to various technical categories and the technical competition intensity between the two target areas to the terminal 20 for display, and the server 10 sends the technical competition intensity of the two target areas corresponding to various technical categories and the technical competition intensity between the two target areas to the database 30 for storage.

[0076] It is understood that the above is only an example and is not limited to this embodiment.

[0077] Among them, terminals include but are not limited to smartphones (such as Android phones, iOS phones, etc.), mobile phone simulators, tablet computers, laptops, digital broadcast receivers, MIDs (Mobile Internet Devices), PDAs (Personal Digital Assistants), intelligent voice interaction devices, smart home appliances, and vehicle-mounted terminals.

[0078] A server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.

[0079] The aforementioned networks may include, but are not limited to, wired networks and wireless networks. Wired networks include local area networks, metropolitan area networks, and wide area networks, and wireless networks include Bluetooth, Wi-Fi, and other wireless communication networks. The specific network type may be determined based on actual application scenarios and is not limited here.

[0080] See also Figure 2 , Figure 2 The flowchart of a method for processing the technical competition intensity of a region provided by an embodiment of the present disclosure is shown, wherein the method can be executed by any electronic device, such as a server, etc. As an optional implementation, the method can be executed by a server. For the convenience of description, in the description of some optional embodiments below, the method will be executed by a server as an example of the execution subject. Figure 2 As shown, the method for processing the technical competition intensity of a region provided by the embodiment of the present disclosure includes the following steps:

[0081] S201, obtain patent vector sets for two target areas, wherein the patent vector set for each target area includes a patent vector subset corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector for each technical category is a feature representation of a patent text of the technical category.

[0082] Specifically, a target area is, for example, Beijing, and multiple target areas are, for example, Beijing, Shanghai, Wuhan, etc. For example, a patent vector set for a target area includes patent vector subsets corresponding to multiple technology categories, such as a technology category such as brain-like chips, and multiple technology categories such as GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), ASIC (Application-Specific Integrated Circuit), brain-like chips, NPU (Neural Processing Unit), etc. For example, a patent vector for a technology category is a feature representation of a patent text for the technology category, and the feature representation of the patent text corresponds to the patent semantic information in the patent text.

[0083] S202: For each technology category, based on the similarity between the patent vectors of the technology category in the two target regions, determine the technology competition intensity corresponding to the technology category between the two target regions.

[0084] Specifically, for example, the greater the intensity of technological competition between two target areas corresponding to a certain technology category, the closer the technological strength between the two target areas for this technology category, and the two target areas are each other's main competitors.

[0085] For example, for a technology category, the two target areas are target area A and target area B, the patent vector set of target area A includes patent vector subset 1 corresponding to the technology category, and the patent vector set of target area B includes patent vector subset 2 corresponding to the technology category, patent vector subset 1 includes multiple patent vectors, and patent vector subset 2 includes multiple patent vectors; determine the second similarity between a patent vector in vector set 1 and each patent vector in vector set 2, that is, obtain multiple second similarities corresponding to the patent vector in vector set 1, sort the multiple second similarities from large to small, and determine the top n second similarities among the multiple second similarities as the similarity set corresponding to the patent vector in vector set 1, where n is a positive integer; based on the similarity set corresponding to each patent vector in patent vector subset 1, determine the technical competition intensity between target area A and target area B corresponding to the technology category.

[0086] S203 : Determine the technological competition intensity between the two target areas based on the technological competition intensity of the two target areas corresponding to various technological categories.

[0087] Specifically, for example, the greater the intensity of technological competition between two target areas, the closer the overall technological strength between the two target areas, and the two target areas become each other's main competitors.

[0088] For example, a specific technology field includes various technology categories. The greater the intensity of technological competition between two target areas, the closer the overall technological strength between the two target areas for this specific technology field; among them, specific technology fields include artificial intelligence hardware platforms, and various technology categories include GPU, FPGA, ASIC, brain-like chips, NPU, etc.

[0089] For example, the weight corresponding to each technology category is obtained; based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weightedly integrated to obtain the technology competition intensity of the two target areas.

[0090] In an embodiment of the present disclosure, patent vector sets of two target areas are obtained, and the patent vector set of each target area includes a subset of patent vectors corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; for each technical category, based on the similarity between the patent vectors of the technical category in the two target areas, the technical competition intensity corresponding to the technical category between the two target areas is determined; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined; in this way, based on the patent vectors (feature representation of the patent text, i.e., the patent semantic information in the patent text), the technical competition intensity between the two target areas corresponding to the technical category is calculated, and based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined, that is, the technical competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technical competition between the two target areas.

[0091] In one embodiment, for each target region, the patent vector set of the target region is obtained by:

[0092] Acquire a plurality of patent texts corresponding to each technical category of at least one technical category in the target region;

[0093] For each patent text, input each patent text into the fine-tuned text embedding model for feature extraction to obtain the patent vector corresponding to each patent text;

[0094] Based on the patent vectors of each patent text in each technology category, a patent vector set for the target area is constructed.

[0095] Specifically, for example, multiple patent texts corresponding to each of multiple technology categories in a certain target area are obtained; multiple technology categories such as GPU, FPGA, ASIC, brain-like chip, NPU, etc.

[0096] A fine-tuned text embedding model, such as a fine-tuned BGE model, can be used. A BGE model, such as the BGE-large-zh-1.5 model, can be used. For example, for each patent document, each patent document is input into the fine-tuned BGE-large-zh-1.5 model for feature extraction, resulting in a patent vector corresponding to each patent document. The patent vector corresponding to each patent document has 768 dimensions. The patent vectors are stored in a vector database, such as Milvus, using a graph-based indexing method such as HNSW. Patent vector similarity calculation (vector distance calculation method) uses cosine similarity.

[0097] It should be noted that, based on the fine-tuning dataset, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model; a more accurate patent vector is output by the fine-tuned text embedding model, that is, the feature representation of the patent text is optimized.

[0098] In one embodiment, the fine-tuned text embedding model is trained by:

[0099] Obtain multiple patent training texts;

[0100] Input each patent training text into the pre-trained text embedding model to perform feature extraction and obtain the patent vector corresponding to each patent training text;

[0101] Determine the first similarity between patent vectors corresponding to each pair of patent training texts;

[0102] Obtaining the standard similarity between two patent training texts obtained through expert scoring;

[0103] Based on the first similarity and the corresponding standard similarity between each pair of patent training texts, the plurality of patent training texts are divided into positive samples and negative samples, and each positive sample is used as a positive sample set;

[0104] Each negative sample and the prompt word are input into the trained large language model for text enhancement to obtain patent text similar to each negative sample. The prompt word is used to instruct the trained large language model to output patent text similar to each negative sample.

[0105] Each negative sample and patent text similar to each negative sample are regarded as a negative sample set;

[0106] The positive sample set and the negative sample set are used as fine-tuning datasets. Based on the fine-tuning datasets, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model.

[0107] Specifically, for example, the feature representation optimization method of patent text based on expert feedback fine-tuning is as follows Figure 3As shown, (1) the model domain continued training and expert evaluation dataset construction include: based on a large-scale dataset, the initial pre-trained text embedding model is trained in the model domain to obtain a pre-trained text embedding model (domain-trained model); multiple patent training texts are obtained; each patent training text is input into the pre-trained text embedding model, and feature extraction is performed to obtain the patent vector corresponding to each patent training text (feature representation of the patent text); the first similarity between the patent vectors corresponding to each two patent training texts is determined, that is, the patent similarity calculation; (2) the expert evaluation dataset construction includes: each first similarity and multiple patent training texts are stored in the database as similar patent data; the standard similarity between the two patent training texts is obtained through expert scoring (expert scoring and evaluation) (expert evaluation data); (3) the fine-tuning dataset construction includes: based on the first similarity between the two patent training texts and the corresponding standard similarity (expert evaluation data), the multiple patent training texts are divided into positive samples and negative samples, each positive sample is used as a positive sample set (positive example dataset), and each negative sample is used as a negative example dataset; each negative sample and the prompt word (Prompt) are input into the trained The large language model is used to perform text enhancement processing to obtain patent texts similar to each negative sample, wherein the prompt word is used to instruct the trained large language model to output patent texts similar to each negative sample, such as BLOOM, Baichuan, atom, etc.; (4) Model fine-tuning includes: taking each negative sample and patent texts similar to each negative sample as a negative sample set, wherein the patent texts similar to each negative sample are negative example data sets generated by the trained large language model, for example, negative example data sets are generated based on different trained large language models; the positive sample set and the negative sample set are combined. This set is used as a fine-tuning dataset with balanced positive and negative examples (the ratio of positive and negative data is determined). In this way, fine-tuning datasets based on different pre-trained text embedding models can be obtained; pre-trained text embedding models are selected; based on the fine-tuning dataset and fine-tuning strategy, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model. Different pre-trained text embedding models correspond to different fine-tuning strategies, that is, fine-tuning strategies based on different pre-trained text embedding models; the performance of the fine-tuned text embedding model is evaluated based on the validation set, that is, the performance of the fine-tuned text embedding model is evaluated on the validation set.

[0108] It should be noted that, based on the fine-tuning dataset, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model; the fine-tuned text embedding model outputs a more accurate patent vector, that is, the feature representation of the patent text is optimized; the method provided by the embodiment of the present disclosure is intended to solve the limitations of the traditional patent text similarity calculation method when considering the professional patent text context, that is, the patent text similarity calculation in a specific technical field (specific technical field such as artificial intelligence hardware platform) based on the pre-trained model (pre-trained text embedding model) obtained by corpus training has the limitation of fine-tuning data construction; (1) In the method provided by the embodiment of the present disclosure, the initial pre-trained model is continuously trained on a large scale to learn relevant field knowledge and obtain a pre-trained model; after obtaining the feature representation of the patent text through the pre-trained model, the similarity between patents is calculated based on the vector database, and experts score the similarity of similar patent pairs on this data and give reasons for the score, and a fine-tuning dataset is constructed through the feedback of experts on the similarity calculation results; because in the process of model fine-tuning, attention is paid to whether the similarity calculated by the pre-trained model can be The similarity (first similarity) of the patent is closer to the similarity (standard similarity) given by the expert, so the focus is on the changes in the fine-tuning of patents with a large gap between the similarity (first similarity) calculated by the pre-training model and the similarity (standard similarity) given by the expert; (2) by means of data enhancement, the data (patent text) with a large gap between the similarity given by the pre-training model and the similarity given by the expert is used as negative examples (negative samples), and the data with a small gap is used as positive examples. The amount of positive examples and negative examples is unbalanced, so the negative examples are expanded. On the basis of maintaining the original semantic information, the evaluation perspective of the expert is further integrated to obtain a richer and more accurate semantic representation and a fine-tuning dataset with balanced positive and negative examples; (3) one or more pre-training models are fine-tuned on the constructed fine-tuning dataset with balanced positive and negative examples. Specifically, the pre-training model is supervised training using the positive and negative example dataset that has been text-enhanced by the large language model; usually a small learning rate and an appropriate number of training steps are used for fine-tuning to retain the semantic representation ability of the pre-training model; the base selected in the method provided in the embodiment of the present disclosure is the BGE model.

[0109] In one embodiment, based on the first similarity and the corresponding standard similarity between each pair of patent training texts, a plurality of patent training texts are divided into positive samples and negative samples, including:

[0110] For each patent training text, if the difference between each first similarity corresponding to the patent training text and the corresponding standard similarity is greater than a preset threshold, the patent training text is determined as a negative sample;

[0111] If the differences between each first similarity corresponding to the patent training text and the corresponding standard similarity are all less than or equal to a preset threshold, the patent training text is determined as a positive sample.

[0112] Specifically, for example, there are 10 patent training texts, and there are 9 first similarities between patent training text 1 and the other 9 patent training texts in the 10 patent training texts. Each of the 9 first similarities corresponds to a standard similarity. The difference between each first similarity and the corresponding standard similarity in the 9 first similarities is calculated to obtain 9 differences. If there is a difference greater than a preset threshold among the 9 differences, that is, there is a large gap between the first similarities and the corresponding standard similarities among the 9 first similarities, then patent training text 1 is determined as a negative sample; if there is no difference greater than the preset threshold among the 9 differences (all 9 differences are less than or equal to the preset threshold), that is, the gap between each first similarity and the corresponding standard similarity in the 9 first similarities is small, then patent training text 1 is determined as a positive sample.

[0113] It should be noted that the patent training texts with a large gap between the first similarity and the standard similarity given by the expert are taken as negative samples, and the patent training texts with a small gap between the first similarity and the standard similarity given by the expert are taken as positive samples.

[0114] In one embodiment, for each technology category, the patent vector subsets of the technology category in the two target regions are used as the first vector set and the second vector set, respectively, to determine the technology competition intensity corresponding to the technology category between the two target regions, including:

[0115] Determining a second similarity between each patent vector in the first vector set and each patent vector in the second vector set;

[0116] According to the order of the second similarities from largest to smallest, a preset number of second similarities ranked first among the second similarities are determined as a similarity set of the two target areas corresponding to the technology category;

[0117] Based on the similarity set, the technical competition intensity of the two target areas corresponding to the technology category is determined.

[0118] Specifically, for example, the first vector set includes patent vector 1 and patent vector 2, and the second vector set includes patent vector 6, patent vector 7, patent vector 8, patent vector 9 and patent vector 10; determine the second similarity between each patent vector in the first vector set and each patent vector in the second vector set; the second similarities between patent vector 1 and patent vector 6, patent vector 7, patent vector 8, patent vector 9 and patent vector 10 are second similarity 1, second similarity 2, second similarity 3, second similarity 4 and second similarity 5 respectively, sort the second similarity 1, second similarity 2, second similarity 3, second similarity 4 and second similarity 5 from large to small, and select the first two second similarities, for example, the first two second similarities 2 and second Similarity 4, the preset number is 2; the second similarity 2 and the second similarity 4 are set in the similarity set of the two target areas corresponding to the technology category; the second similarities between patent vector 2 and patent vector 6, patent vector 7, patent vector 8, patent vector 9 and patent vector 10 are second similarity 6, second similarity 7, second similarity 8, second similarity 9 and second similarity 10 respectively, and the second similarity 6, second similarity 7, second similarity 8, second similarity 9 and second similarity 10 are sorted from large to small, and the first two second similarities are selected, and the first two second similarities are, for example, second similarity 6 and second similarity 9, the preset number is 2; the second similarity 6 and second similarity 9 are set in the similarity set of the two target areas corresponding to the technology category.

[0119] For example, determine the intensity of technological competition between two target regions, including:

[0120] (1) Data preprocessing.

[0121] Label patent data in a certain field by year. For example, for each technology category t, collect patent data (patent text) in a certain field related to technology category t, and label each patent (patent text) by region (target region) to obtain a regional label, such as Beijing.

[0122] For example, in patent information data processing, patents are classified into corresponding technology categories according to the "artificial intelligence classification name" field in the patent data and the artificial intelligence classification framework.

[0123] For example, in the data processing of patent holder information, since the patent holder information given for each patent is inconsistent, there are errors in some patent holder information data, especially in the fields of cities and districts and counties. Therefore, it is necessary to verify the data of the patent holder information; in the information of cities, districts and counties to which the patent holder belongs, the area with the most appearances in all patents corresponding to the patent holder is selected as the final patent holder area; then the data of the patent holders in the cities, districts and counties that need to be analyzed are extracted and merged to facilitate subsequent retrieval.

[0124] (2) Calculation of the intensity of technological competition between regions (target regions) by technology category (the intensity of technological competition between two target regions corresponding to technology category t), including steps 1 to 3:

[0125] Step 1: Extract the patent sets of region A (target region A) and region B (target region B) in technology category t, denoted as and ;

[0126] Step 2: For each patent in region A, obtain the top n most similar patents in region B in technology category t, and store the cosine similarity (the top n second similarities), where n is a positive integer.

[0127] Step 3: Repeat step 2 for all patents in region A in technology category t to obtain all patents in region B related to region A and their cosine similarities; sum each cosine similarity (each second similarity in the similarity set) to obtain the technological competition intensity between region A and region B. , that is, the technological competition intensity of the two target areas corresponding to technology category t ,calculate The formula (1) is as follows:

[0128] Formula (1)

[0129] in, and represent the patent vector of patent p and the patent vector of patent q respectively; : the patent set of region A in technology category t; p: each patent in region A; : The set of patents similar to patent p in region B; cos(Vec(p),Vec(q)): The cosine similarity between the patent vector of patent p and the patent vector of patent q, indicating the degree of similarity between the patent vector of patent p and the patent vector of patent q in the vector space; : Find the top n patents in region B that are most similar to patent p.

[0130] (3) Regional (target region) technological competition intensity vector

[0131] For any region, the intensity of technological competition with a region B (assuming there are m regions in total) in all technological categories (assuming there are n technological categories in total) will form a technological competition intensity vector , As shown in formula (2):

[0132] Formula (2)

[0133] in, Represents the region index, Indicates the technology category index, represents the intensity of technological competition between region i and region B in the jth technological category; there are n technological categories in total. The dimension is n.

[0134] (4) Calculation of inter-regional technological competition intensity that incorporates regional technological characteristics, that is, determining the technological competition intensity between two target regions.

[0135] Based on formula (1) and formula (2), determine the technical competition intensity vector of any region B with a certain region B in n technology categories: ; In the technology competition intensity vector In the equation, each component expresses the technical competition intensity between any region i and a region B in the jth technology category; the technical competition intensity between any region and a region B is calculated as , a simple calculation method is to vector The components of are summed up, as shown in formula (3):

[0136] Formula (3)

[0137] Among them, in the calculation of the technological competition intensity between any region and a region B, the technological competition intensity vector The amount are equally weighted.

[0138] In fact, the regional technological characteristics of any region in each technology category are not effectively expressed. For example, the distribution of patents in any region A across n technology categories is not uniform. Region A may be strong in some technology categories and have a large number of patents, while in some technology categories, it may have almost no patents. In order to express this difference, the regional technological characteristics factor is introduced. , The calculation method is shown in formula (4):

[0139] Formula (4)

[0140] Therefore, when the regional technology characteristic factor is introduced, the technological competition intensity between any region and a region B is The calculation method is changed to formula (5):

[0141] Formula (5)

[0142] in, represents the intensity of technological competition between region i and region B in the jth technology category, Represents the regional technology characteristic factor of region i in the jth technology category.

[0143] It should be noted that the intensity of technological competition between the two target areas corresponding to the technology categories is calculated based on the patent vector, and the intensity of technological competition between the two target areas corresponding to various technology categories is determined based on the intensity of technological competition between the two target areas, that is, the technological competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technological competition between the two target areas.

[0144] In one embodiment, determining the technical competition intensity of two target regions corresponding to the technology category based on the similarity set includes:

[0145] A preset number of second similarities in the similarity set are fused, and a fusion result is determined as the technical competition intensity of the two target areas corresponding to the technical category.

[0146] Specifically, for example, the preset number is 50, and the 50 second similarities are summed to obtain a sum result, which is then determined as the technical competition intensity of the two target areas corresponding to the technical category.

[0147] For example, as shown in formula (1), the sum of each cosine similarity (each second similarity in the similarity set) is calculated to obtain the technical competition intensity of region A and region B (two target regions) corresponding to technology category t: .

[0148] In one embodiment, determining the technological competition intensity between the two target regions based on the technological competition intensity of the two target regions corresponding to various technological categories includes:

[0149] Get the weight corresponding to each technology category;

[0150] Based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weightedly integrated to obtain the technology competition intensity of the two target areas.

[0151] Specifically, the weight corresponding to each technology category is the regional technology characteristic factor shown in formula (4): For example, as shown in formula (5), based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weighted and integrated to obtain the technology competition intensity of the two target areas (area i and area B). .

[0152] It should be noted that the intensity of technological competition between the two target areas corresponding to the technology categories is calculated based on the patent vector, and the intensity of technological competition between the two target areas corresponding to various technology categories is determined based on the intensity of technological competition between the two target areas, that is, the technological competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technological competition between the two target areas.

[0153] In one embodiment, obtaining a plurality of patent texts corresponding to each of at least one technical category in the target region includes:

[0154] Acquire regional information of the target area, where the regional information includes specific technical fields of the target area;

[0155] Obtain multiple patent texts for each technology category in at least one technology category related to a specific technology field.

[0156] Specifically, for example, a specific technology field includes various technology categories. The greater the intensity of technological competition between two target areas, the closer the overall technological strength between the two target areas for this specific technology field; among them, specific technology fields include artificial intelligence hardware platforms, and various technology categories include GPU, FPGA, ASIC, brain-like chips, NPU, etc.

[0157] The application of the embodiments of the present disclosure has at least the following beneficial effects:

[0158] Based on the fine-tuning dataset, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model; a more accurate patent vector is output through the fine-tuned text embedding model, that is, the feature representation of the patent text is optimized; based on the patent vector (the feature representation of the patent text, that is, the patent semantic information in the patent text), the technical competition intensity corresponding to the technical categories between the two target areas is calculated, and based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined, that is, the technical competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technical competition between the two target areas.

[0159] In order to better understand the method provided by the embodiment of the present disclosure, the solution of the embodiment of the present disclosure is further described below with reference to examples of specific application scenarios.

[0160] In one embodiment, for example, Table 3 shows that "artificial intelligence hardware platform" includes several technology categories, such as GPU, FPGA, ASIC, brain-like chip, NPU, etc.; through the method provided by the embodiment of the present disclosure, the technology competition city and technology competition intensity of "Beijing" are obtained, as shown in Table 3:

[0161] Table 3: Technology competition cities and technology competition intensity of “Beijing”

[0162]

[0163] In a specific application scenario, such as a regional technology competition scenario, see Figure 4 , shows the processing flow of a method for processing the technical competition intensity of a region, such as Figure 4 As shown, the processing flow of the method for processing the technical competition intensity of a region provided by the embodiment of the present disclosure includes the following steps:

[0164] S401: The server obtains a fine-tuning dataset.

[0165] Specifically, for example, Figure 3 As shown, based on the first similarity and the corresponding standard similarity (expert evaluation data) between each pair of patent training texts, multiple patent training texts are divided into positive samples and negative samples, each positive sample is used as a positive sample set (positive data set), and each negative sample is used as a negative data set; each negative sample and a prompt word (Prompt) are input into the trained large language model for text enhancement processing to obtain a patent text similar to each negative sample, wherein the prompt word is used to instruct the trained large language model to output a patent text similar to each negative sample; each negative sample and the patent text similar to each negative sample are used as a negative sample set, and the positive sample set and the negative sample set are used as fine-tuning data sets.

[0166] S402: The server fine-tunes the pre-trained text embedding model based on the fine-tuning dataset to obtain a fine-tuned text embedding model.

[0167] Specifically, for example, a smaller learning rate and an appropriate number of training steps are used to fine-tune the pre-trained text embedding model to retain the semantic representation ability of the pre-trained text embedding model.

[0168] S403: The server obtains a plurality of patent texts corresponding to each of the plurality of technical categories in each of the two target areas.

[0169] Specifically, for example, regional information of the target area is obtained, where the regional information includes a specific technical field of the target area; and multiple patent texts of each technical category in multiple technical categories related to the specific technical field are obtained.

[0170] In step S404, the server performs feature extraction based on multiple patent texts corresponding to various technical categories in each target area using a fine-tuned text embedding model to obtain a patent vector set for the target area.

[0171] Specifically, for example, for each patent text, each patent text is input into the fine-tuned text embedding model for feature extraction to obtain the patent vector corresponding to each patent text; based on the patent vectors of each patent text in each technical category, a patent vector set for the target area is constructed; and the patent vector set for the target area is stored in a vector database.

[0172] S405 , the server determines, for each technology category, the technology competition intensity corresponding to the technology category between the two target regions based on the similarity between the patent vectors of the technology category in the two target regions.

[0173] Specifically, for example, as shown in formula (1), the sum of each cosine similarity (each second similarity in the similarity set) is calculated to obtain the technical competition intensity of region A and region B (two target regions) corresponding to technology category t: .

[0174] S406: The server performs weighted fusion of the technical competition intensities of the two target areas corresponding to the various technical categories based on the weights corresponding to the various technical categories, to obtain the technical competition intensities of the two target areas.

[0175] Specifically, the weight corresponding to each technology category is the regional technology characteristic factor shown in formula (4): For example, as shown in formula (5), based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weighted and integrated to obtain the technology competition intensity of the two target areas (area i and area B). .

[0176] The application of the embodiments of the present disclosure has at least the following beneficial effects:

[0177] From a methodological and technical perspective, the precise identification of regional technological competitors is explored based on patent data. The basis for calculating regional technological competitors is a feature representation optimization method of patent texts based on expert feedback. Similar patents to any patent are obtained through patent text similarity, and then the intensity of technological competition between regions is obtained based on the regional labels of similar patents. From the practical application perspective of regional technological competitor identification, the provided patent recommendation and analysis methods have brought more accurate Chinese patent analysis tools to enterprises, scientific research institutions and science and technology management departments. In the domestic market, patent recommendations achieved through model optimization can help enterprises and scientific research institutions to more comprehensively understand the technological layout and innovation direction of regional competitors and clarify their relative technological strength in the industry. At the national level, this method provides science and technology management departments with a global perspective, facilitating the identification of major competitors and potential challengers in key technology fields, thereby supporting the formulation of science and technology policies and resource allocation. By accurately identifying technological competitors at the regional level, enterprises and countries can better allocate R&D resources, avoid duplicate investment, improve innovation efficiency, and promote the overall improvement of domestic technological strength.

[0178] The embodiment of the present disclosure also provides a device for processing the technical competition intensity of a region. The structural diagram of the device for processing the technical competition intensity of a region is as follows: Figure 5 As shown, the regional technology competition intensity processing device 50 includes a first processing module 501 , a second processing module 502 and a third processing module 503 .

[0179] The first processing module 501 is configured to obtain patent vector sets for two target regions, where each patent vector set for the target region includes a subset of patent vectors corresponding to at least one technology category, and each patent vector subset corresponding to each technology category includes multiple patent vectors for that technology category. A patent vector for each technology category is a feature representation of a patent text for that technology category.

[0180] The second processing module 502 is configured to determine, for each technology category, the technology competition intensity corresponding to the technology category between the two target regions based on the similarity between the patent vectors of the technology category in the two target regions;

[0181] The third processing module 503 is configured to determine the technological competition intensity between the two target areas based on the technological competition intensity of the two target areas corresponding to various technological categories.

[0182] In one embodiment, for each target region, the patent vector set of the target region is obtained by the first processing module 501 in the following manner:

[0183] Acquire a plurality of patent texts corresponding to each technical category of at least one technical category in the target region;

[0184] For each patent text, input each patent text into the fine-tuned text embedding model for feature extraction to obtain the patent vector corresponding to each patent text;

[0185] Based on the patent vectors of each patent text in each technology category, a patent vector set for the target area is constructed.

[0186] In one embodiment, the fine-tuned text embedding model is trained by the first processing module 501 in the following manner:

[0187] Obtain multiple patent training texts;

[0188] Input each patent training text into the pre-trained text embedding model to perform feature extraction and obtain the patent vector corresponding to each patent training text;

[0189] Determine the first similarity between patent vectors corresponding to each pair of patent training texts;

[0190] Obtaining the standard similarity between two patent training texts obtained through expert scoring;

[0191] Based on the first similarity and the corresponding standard similarity between each pair of patent training texts, the plurality of patent training texts are divided into positive samples and negative samples, and each positive sample is used as a positive sample set;

[0192] Each negative sample and the prompt word are input into the trained large language model for text enhancement to obtain patent text similar to each negative sample. The prompt word is used to instruct the trained large language model to output patent text similar to each negative sample.

[0193] Each negative sample and patent text similar to each negative sample are regarded as a negative sample set;

[0194] The positive sample set and the negative sample set are used as fine-tuning datasets. Based on the fine-tuning datasets, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model.

[0195] In one embodiment, the first processing module 501 is specifically configured to:

[0196] For each patent training text, if the difference between each first similarity corresponding to the patent training text and the corresponding standard similarity is greater than a preset threshold, the patent training text is determined as a negative sample;

[0197] If the differences between each first similarity corresponding to the patent training text and the corresponding standard similarity are all less than or equal to a preset threshold, the patent training text is determined as a positive sample.

[0198] In one embodiment, for each technology category, subsets of patent vectors of the technology category in the two target regions are used as a first vector set and a second vector set, respectively. The second processing module 502 is specifically configured to:

[0199] Determining a second similarity between each patent vector in the first vector set and each patent vector in the second vector set;

[0200] According to the order of the second similarities from largest to smallest, a preset number of second similarities ranked first among the second similarities are determined as a similarity set of the two target areas corresponding to the technology category;

[0201] Based on the similarity set, the technical competition intensity of the two target areas corresponding to the technology category is determined.

[0202] In one embodiment, the second processing module 502 is specifically configured to:

[0203] A preset number of second similarities in the similarity set are fused, and a fusion result is determined as the technical competition intensity of the two target areas corresponding to the technical category.

[0204] In one embodiment, the third processing module 503 is specifically configured to:

[0205] Get the weight corresponding to each technology category;

[0206] Based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weightedly integrated to obtain the technology competition intensity of the two target areas.

[0207] In one embodiment, the first processing module 501 is specifically configured to:

[0208] Acquire regional information of the target area, where the regional information includes specific technical fields of the target area;

[0209] Obtain multiple patent texts for each technology category in at least one technology category related to a specific technology field.

[0210] The application of the embodiments of the present disclosure has at least the following beneficial effects:

[0211] A patent vector set of two target areas is obtained, wherein the patent vector set of each target area includes a subset of patent vectors corresponding to at least one technical category, the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; for each technical category, based on the similarity between the patent vectors of the technical category in the two target areas, the technical competition intensity corresponding to the technical category between the two target areas is determined; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined; in this way, based on the patent vector (feature representation of the patent text, i.e., the patent semantic information in the patent text), the technical competition intensity between the two target areas corresponding to the technical category is calculated; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined, i.e., the technical competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technical competition between the two target areas.

[0212] The present disclosure also provides an electronic device. The structural diagram of the electronic device is as follows: Figure 6 As shown, Figure 6 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure.

[0213] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.

[0214] Bus 4002 may include a path for transmitting information between the above components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0215] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.

[0216] The memory 4003 is used to store the computer program for executing the embodiments of the present disclosure, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiments.

[0217] Among them, electronic equipment includes but is not limited to: servers, etc.

[0218] The application of the embodiments of the present disclosure has at least the following beneficial effects:

[0219] A patent vector set of two target areas is obtained, wherein the patent vector set of each target area includes a subset of patent vectors corresponding to at least one technical category, the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; for each technical category, based on the similarity between the patent vectors of the technical category in the two target areas, the technical competition intensity corresponding to the technical category between the two target areas is determined; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined; in this way, based on the patent vector (feature representation of the patent text, i.e., the patent semantic information in the patent text), the technical competition intensity between the two target areas corresponding to the technical category is calculated; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined, i.e., the technical competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technical competition between the two target areas.

[0220] An embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.

[0221] The embodiments of the present disclosure further provide a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiments when executed by a processor.

[0222] It should be understood that, although the flowcharts of the embodiments of the present disclosure indicate the various operation steps by arrows, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be performed in other orders as required. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times. In scenarios where the execution times are different, the order of execution of these sub-steps or stages can be flexibly configured as required, and the embodiments of the present disclosure do not limit this.

[0223] The above description is only an optional implementation method for some implementation scenarios of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of the present disclosure, other similar implementation methods based on the technical ideas of the present disclosure also fall within the protection scope of the embodiments of the present disclosure.

Claims

1. A method for processing the technical competition intensity of a region, characterized by: include: Obtain patent vector sets for two target regions, where each patent vector set for the target region includes a patent vector subset corresponding to at least one technical category, and each patent vector subset corresponding to each technical category includes multiple patent vectors for that technical category. A patent vector for each technical category is a feature representation of a patent text for that technical category. For each technology category, determining the technology competition intensity corresponding to the technology category between the two target regions based on the similarity between the patent vectors of the technology category in the two target regions; determining the technological competition intensity between the two target regions based on the technological competition intensity of the two target regions corresponding to various technological categories; For each target region, the patent vector set of the target region is obtained by: Acquire a plurality of patent texts corresponding to each of the at least one technical category in the target region; For each patent text, input each patent text into the fine-tuned text embedding model for feature extraction to obtain a patent vector corresponding to each patent text; Based on the patent vectors of each patent text in each technology category, a patent vector set for the target area is constructed; The fine-tuned text embedding model is trained in the following way: Obtain multiple patent training texts; Input each patent training text into a pre-trained text embedding model to perform feature extraction to obtain a patent vector corresponding to each patent training text; Determine the first similarity between patent vectors corresponding to each pair of patent training texts; Obtaining the standard similarity between the two patent training texts obtained through expert scoring; Based on the first similarity and the corresponding standard similarity between the two patent training texts, the plurality of patent training texts are divided into positive samples and negative samples, and each positive sample is used as a positive sample set; Inputting each negative sample and the prompt word into the trained large language model to perform text enhancement processing to obtain patent text similar to each negative sample, wherein the prompt word is used to instruct the trained large language model to output patent text similar to each negative sample; Each negative sample and patent text similar to each negative sample are regarded as a negative sample set; The positive sample set and the negative sample set are used as fine-tuning data sets, and the pre-trained text embedding model is fine-tuned based on the fine-tuning data sets to obtain a fine-tuned text embedding model.

2. The method according to claim 1, characterized in that Based on the first similarity and the corresponding standard similarity between the two patent training texts, the plurality of patent training texts are divided into positive samples and negative samples, including: For each patent training text, if the difference between each first similarity corresponding to the patent training text and the corresponding standard similarity is greater than a preset threshold, the patent training text is determined as a negative sample; If the differences between each first similarity corresponding to the patent training text and the corresponding standard similarity are all less than or equal to a preset threshold, the patent training text is determined as a positive sample.

3. The method according to claim 1, characterized in that For each technology category, the patent vector subsets of the technology category in the two target regions are respectively used as a first vector set and a second vector set. Determining the technology competition intensity corresponding to the technology category between the two target regions includes: determining a second similarity between each patent vector in the first vector set and each patent vector in the second vector set; According to the sorting of the second similarities from largest to smallest, a preset number of second similarities ranked first among the second similarities are determined as a similarity set of the two target areas corresponding to the technology category; Based on the similarity set, the technical competition intensity of the two target areas corresponding to the technical category is determined.

4. The method according to claim 3, characterized in that The determining, based on the similarity set, the technical competition intensity of the two target regions corresponding to the technical category includes: A preset number of second similarities in the similarity set are fused, and a fusion result is determined as the technical competition intensity of the two target areas corresponding to the technical category.

5. The method according to claim 1, wherein Determining the technological competition intensity between the two target regions based on the technological competition intensity of the two target regions corresponding to various technological categories includes: Get the weight corresponding to each technology category; Based on the weights corresponding to various technology categories, the technology competition intensities of the two target areas corresponding to various technology categories are weightedly fused to obtain the technology competition intensities of the two target areas.

6. The method according to claim 1, characterized in that The obtaining of a plurality of patent texts corresponding to each of the at least one technical category in the target area includes: Acquire regional information of the target area, wherein the regional information includes a specific technical field of the target area; A plurality of patent texts of each technical category in the at least one technical category related to the specific technical field is obtained.

7. A device for processing regional technical competition intensity, characterized in that: include: A first processing module is configured to obtain patent vector sets for two target regions, wherein the patent vector set for each target region includes a subset of patent vectors corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of that technical category, wherein a patent vector for each technical category is a feature representation of a patent text of that technical category; A second processing module is configured to determine, for each technology category, the technology competition intensity corresponding to the technology category between the two target regions based on the similarity between the patent vectors of the technology category in the two target regions; a third processing module, configured to determine the technological competition intensity between the two target regions based on the technological competition intensity of the two target regions corresponding to various technological categories; For each target region, the patent vector set of the target region is obtained by the first processing module in the following manner: Acquire a plurality of patent texts corresponding to each of the at least one technical category in the target region; For each patent text, input each patent text into the fine-tuned text embedding model for feature extraction to obtain a patent vector corresponding to each patent text; Based on the patent vectors of each patent text in each technology category, a patent vector set for the target area is constructed; The fine-tuned text embedding model is trained by the first processing module in the following manner: Obtain multiple patent training texts; Input each patent training text into a pre-trained text embedding model to perform feature extraction to obtain a patent vector corresponding to each patent training text; Determine the first similarity between patent vectors corresponding to each pair of patent training texts; Obtaining the standard similarity between the two patent training texts obtained through expert scoring; Based on the first similarity and the corresponding standard similarity between the two patent training texts, the plurality of patent training texts are divided into positive samples and negative samples, and each positive sample is used as a positive sample set; Inputting each negative sample and the prompt word into the trained large language model to perform text enhancement processing to obtain patent text similar to each negative sample, wherein the prompt word is used to instruct the trained large language model to output patent text similar to each negative sample; Each negative sample and patent text similar to each negative sample are regarded as a negative sample set; The positive sample set and the negative sample set are used as fine-tuning data sets, and the pre-trained text embedding model is fine-tuned based on the fine-tuning data sets to obtain a fine-tuned text embedding model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Deep technology tracking method for high-tech companies

    CN110580261A

  • Competitive publication identification method based on publication content

    CN118193672A