Regional technical competition intensity processing method, device, equipment, medium and product

By acquiring and processing the patent vector sets of the two target areas, using the fine-tuned text embedding model to calculate the technical competition intensity, the problem of low accuracy in evaluation technology in the existing technology is solved, and more accurate technical competition evaluation is achieved.

CN120031683AActive Publication Date: 2025-05-23INST OF SCI & TECHN INFORMATION OF CHINA
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510164313.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-23
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The prior art lacks methods of utilizing patent semantic information when evaluating technological competition between two regions, resulting in low evaluation accuracy.

Method used

By obtaining the patent vector set of two target regions, the technical competition intensity is calculated based on the similarity of the patent vectors, the fine-tuned text embedding model is used to extract the feature representation of the patent text, and the technical competition intensity between the two regions is determined by weighting the competition intensity of different technical categories.

Benefits of technology

Improved accuracy in evaluating technological competition between two regions, and more accurately calculates technological competition intensity by utilizing patent semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031683A_ABST
    Figure CN120031683A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a regional technical competition intensity processing method and device, equipment, a medium and a product, and relates to the field of regional competition and the like, and application scenes include but are not limited to regional technical competition scenes. The method comprises the steps that patent vector sets of two target areas are obtained, the patent vector set of each target area comprises at least one patent vector sub-set corresponding to the technology category, and the patent vector sub-set corresponding to each technology category comprises multiple patent vectors of the technology category; one patent vector of each technical category is a feature representation of one patent text of the technical category; for each technology category, based on the similarity between the patent vectors of the technology category of the two target areas, determining the technology competition intensity corresponding to the technology category between the two target areas; and determining the technical competition intensity between the two target areas based on the technical competition intensity of the two target areas corresponding to the various technical categories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a method, device, equipment, medium and product for processing the technical competition intensity of a region. Background Art

[0002] In the prior art, the technological competition between two regions is evaluated based on patent texts. The feature representation method of patent texts is implemented through the traditional TF-IDF (Term Frequency–Inverse Document Frequency) or Word2Vec (word embedding model) based on keyword extraction, which does not make good use of patent semantic information. From the perspective of patent texts, there are few studies on the technological competition between two regions in the prior art. The technological competition between two regions is often more focused on the construction of indicators, which leads to lower accuracy in evaluating the technological competition between the two regions. Summary of the invention

[0003] In view of the shortcomings of existing methods, the present disclosure proposes a method, device, equipment, computer-readable storage medium and computer program product for processing the technical competition intensity of a region, which are used to solve the problem of how to improve the accuracy of evaluating the technical competition between two regions.

[0004] In a first aspect, the present disclosure provides a method for processing the technical competition intensity of a region, comprising: Obtain patent vector sets of two target regions, wherein the patent vector set of each target region includes a patent vector subset corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; For each technology category, based on the similarity between the patent vectors of the technology category in the two target regions, determine the technology competition intensity corresponding to the technology category between the two target regions; The technological competition intensity between the two target areas is determined based on the technological competition intensity of the two target areas corresponding to various technological categories.

[0005] In one embodiment, for each target region, the patent vector set of the target region is obtained by: Acquire a plurality of patent texts corresponding to each technical category of at least one technical category in the target area; For each patent text, input each patent text into the fine-tuned text embedding model for feature extraction to obtain the patent vector corresponding to each patent text; Based on the patent vectors of each patent text in each technical category, a patent vector set for the target area is constructed.

[0006] In one embodiment, the fine-tuned text embedding model is trained by: Obtain multiple patent training texts; Input each patent training text into the pre-trained text embedding model to extract features and obtain the patent vector corresponding to each patent training text; Determine the first similarity between patent vectors corresponding to each pair of patent training texts; Obtaining the standard similarity between two patent training texts obtained through expert scoring; Based on the first similarity and the corresponding standard similarity between the two patent training texts, the multiple patent training texts are divided into positive samples and negative samples, and each positive sample is used as a positive sample set; Input each negative sample and the prompt word into the trained large language model, perform text enhancement processing, and obtain patent text similar to each negative sample, wherein the prompt word is used to instruct the trained large language model to output patent text similar to each negative sample; Each negative sample and patent texts similar to each negative sample are taken as a negative sample set; The positive sample set and the negative sample set are used as fine-tuning datasets. Based on the fine-tuning datasets, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model.

[0007] In one embodiment, based on the first similarity and the corresponding standard similarity between the two patent training texts, a plurality of patent training texts are divided into positive samples and negative samples, including: For each patent training text, if there is a difference between each first similarity corresponding to the patent training text and the corresponding standard similarity that is greater than a preset threshold, the patent training text is determined as a negative sample; If the differences between each first similarity corresponding to the patent training text and the corresponding standard similarity are all less than or equal to a preset threshold, the patent training text is determined as a positive sample.

[0008] In one embodiment, for each technology category, the patent vector subsets of the technology category in the two target regions are respectively used as the first vector set and the second vector set, and the technology competition intensity corresponding to the technology category between the two target regions is determined, including: Determining a second similarity between each patent vector in the first vector set and each patent vector in the second vector set; According to the order of the second similarities from large to small, a preset number of second similarities ranked first among the second similarities are determined as a similarity set of the two target areas corresponding to the technology category; Based on the similarity set, the technical competition intensity of the two target areas corresponding to the technology category is determined.

[0009] In one embodiment, determining the technical competition intensity of two target areas corresponding to the technical category based on the similarity set includes: A preset number of second similarities in the similarity set are fused, and a fusion result is determined as the technical competition intensity of the two target areas corresponding to the technical category.

[0010] In one embodiment, determining the technical competition intensity between the two target areas based on the technical competition intensity of the two target areas corresponding to various technical categories includes: Get the weight corresponding to each technology category; Based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weightedly integrated to obtain the technology competition intensity of the two target areas.

[0011] In one embodiment, obtaining a plurality of patent texts corresponding to each technical category in the target area includes: Acquire regional information of the target area, where the regional information includes specific technical fields of the target area; Obtain multiple patent texts for each technology category in at least one technology category related to a specific technology field.

[0012] In a second aspect, the present disclosure provides a device for processing the technical competition intensity of a region, comprising: The first processing module is used to obtain patent vector sets of two target areas, each of which includes a patent vector subset corresponding to at least one technical category, and each patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; A second processing module is used to determine, for each technology category, the technology competition intensity corresponding to the technology category between the two target regions based on the similarity between the patent vectors of the technology category in the two target regions; The third processing module is used to determine the technical competition intensity between the two target areas based on the technical competition intensity of the two target areas corresponding to various technical categories.

[0013] In a third aspect, the present disclosure provides an electronic device, including: a processor, a memory, and a bus; A bus, used to connect the processor and memory; A memory, used for storing operation instructions; The processor is used to execute the method for processing the technical competition intensity of the area according to the first aspect of the present disclosure by calling an operation instruction.

[0014] In a fourth aspect, the present disclosure provides a computer-readable storage medium storing a computer program, wherein the computer program is used to execute the method for processing the technical competition intensity of a region according to the first aspect of the present disclosure.

[0015] In a fifth aspect, the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method for processing the technical competition intensity of a region in the first aspect of the present disclosure.

[0016] The technical solution provided by the embodiments of the present disclosure has at least the following beneficial effects: A patent vector set of two target areas is obtained, wherein the patent vector set of each target area includes a patent vector subset corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; for each technical category, based on the similarity between the patent vectors of the technical category of the two target areas, the technical competition intensity corresponding to the technical category between the two target areas is determined; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined; in this way, based on the patent vector (the feature representation of the patent text, i.e., the patent semantic information in the patent text), the technical competition intensity between the two target areas corresponding to the technical category is calculated, and based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined, that is, the technical competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technical competition between the two target areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for describing the embodiments of the present disclosure are briefly introduced below.

[0018] Figure 1 A schematic diagram of the architecture of a system for processing the technical competition intensity of a region provided by an embodiment of the present disclosure; Figure 2 A flowchart of a method for processing technical competition intensity of a region provided in an embodiment of the present disclosure; Figure 3 A schematic diagram of the technical competition intensity processing of a region provided by an embodiment of the present disclosure; Figure 4A flowchart of a method for processing technical competition intensity of a region provided in an embodiment of the present disclosure; Figure 5 A schematic diagram of the structure of a device for processing the technical competition intensity of a region provided in an embodiment of the present disclosure; Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] The embodiments of the present disclosure are described below in conjunction with the drawings in the present disclosure. It should be understood that the implementation methods described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions of the embodiments of the present disclosure.

[0020] It will be understood by those skilled in the art that, unless specifically stated, the singular forms "one", "an", "said" and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present disclosure refer to that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or it may refer to that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" indicates that it is implemented as "A", or implemented as "B", or implemented as "A and B".

[0021] It can be understood that in the specific implementation of the present disclosure, data related to the processing of regional technical competition intensity is involved. When the above embodiments of the present disclosure are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0022] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.

[0023] The disclosed embodiment is a method for processing the technical competition intensity of a region provided by a system for processing the technical competition intensity of a region. The method for processing the technical competition intensity of a region involves fields such as regional competition.

[0024] In order to better understand and illustrate the solutions of the embodiments of the present disclosure, some technical terms involved in the embodiments of the present disclosure are briefly explained below.

[0025] BGE model: The BGE model is a series of text embedding models that aims to convert text into low-dimensional dense vectors for efficient calculation and analysis.

[0026] Large language model: The large language model has demonstrated remarkable capabilities in text enhancement, which is mainly reflected in the understanding, generation and processing of multimodal content of text.

[0027] Vector database is a database system that specializes in storing and managing high-dimensional vector data. It is widely used in machine learning, image processing, natural language processing, and other fields. The current mainstream vector databases include: Faiss, Elasticsearch, and Milvus.

[0028] Faiss is an efficient similarity search library, mainly used for similarity search of dense vectors. Faiss provides a variety of index structures, including flat index, quantized index and hierarchical clustering index, which can perform efficient approximate nearest neighbor search on large-scale data sets. The core advantage of Faiss lies in its highly optimized implementation and support for GPU (Graphics Processing Unit) acceleration, which makes Faiss perform well in large-scale vector retrieval tasks. Faiss is suitable for scenarios with high precision requirements, especially in image retrieval and recommendation systems.

[0029] Elasticsearch is a full-text based open source search engine with powerful full-text retrieval and analysis capabilities. In recent years, Elasticsearch has introduced vector retrieval capabilities, enabling it to process high-dimensional vector data. By integrating a vector similarity calculation plug-in, Elasticsearch can perform vector-based approximate nearest neighbor search. Elasticsearch's advantage lies in its seamless integration with full-text retrieval capabilities, making it suitable for mixed query scenarios, such as applications that require both text and vector retrieval. Since its vector retrieval capabilities are based on plug-in implementations, performance and scalability may not be as good as specialized vector databases when processing ultra-large-scale vector data.

[0030] Milvus is an open source vector database that aims to provide a high-performance and highly scalable vector retrieval solution. Milvus is designed to handle large-scale, high-dimensional vector data, and is particularly suitable for application scenarios that require real-time response and efficient retrieval. As a specialized vector database, Milvus has shown significant advantages in processing high-dimensional vector data. When using Milvus, vector indexing is a key factor that requires special attention. This indexing mechanism relies on a specific mathematical model to construct a data structure that is more efficient in time and space. With the help of this type of index, multiple vectors with high similarity to the target vector can be efficiently retrieved. Most of the commonly used vector index types belong to the category of approximate nearest neighbor search. The core concept of this method is not to pursue absolutely accurate results, but to explore possible neighbor items and moderately sacrifice accuracy within an acceptable range, thereby significantly improving retrieval efficiency. The indexing method of Milvus vector database is shown in Table 1. In addition to the indexing method, the calculation method of comparing the distance between vectors is also an important aspect in the vector database. The distance calculation method directly affects the accuracy and efficiency of vector search. Currently, the methods for calculating the distance between vectors (vector distance calculation methods) that can be used include Euclidean distance, dot product (inner product), cosine similarity, etc. The vector distance calculation methods are shown in Table 2.

[0031] Table 1: Milvus vector database indexing method

[0032] Table 2: Vector distance calculation method

[0033] The solution provided by the embodiments of the present disclosure involves a technology for processing the technical competition intensity of a region. The technical solution of the present disclosure is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present disclosure will be described below in conjunction with the accompanying drawings.

[0034] In order to better understand the solution provided by the embodiment of the present disclosure, the solution is described below in conjunction with a specific application scenario.

[0035] In one embodiment, Figure 1 FIG. 1 shows a schematic diagram of the architecture of a system for processing the technical competition intensity of a region to which the embodiment of the present disclosure is applicable. It can be understood that the method for processing the technical competition intensity of a region provided by the embodiment of the present disclosure can be applied to, but not limited to, Figure 1 In the application scenario shown.

[0036] In this example, Figure 1As shown, the architecture of the regional technical competitiveness intensity processing system in this example may include but is not limited to a server 10, a terminal 20 and a database 30. The server 10, the terminal 20 and the database 30 may interact with each other through a network 40.

[0037] The server 10 obtains patent vector sets of two target areas, each of which includes a patent vector subset corresponding to at least one technical category, each of which includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; for each technical category, the server 10 determines the technical competition intensity between the two target areas corresponding to the technical category based on the similarity between the patent vectors of the technical category of the two target areas; the server 10 determines the technical competition intensity between the two target areas based on the technical competition intensity of the two target areas corresponding to various technical categories. The server 10 sends the technical competition intensity of the two target areas corresponding to various technical categories and the technical competition intensity between the two target areas to the terminal 20 for display, and the server 10 sends the technical competition intensity of the two target areas corresponding to various technical categories and the technical competition intensity between the two target areas to the database 30 for storage.

[0038] It is understandable that the above is only an example and is not limited to this embodiment.

[0039] Among them, terminals include but are not limited to smart phones (such as Android phones, iOS phones, etc.), mobile phone simulators, tablet computers, laptops, digital broadcast receivers, MIDs (Mobile Internet Devices), PDAs (Personal Digital Assistants), intelligent voice interaction devices, smart home appliances, car terminals, etc.

[0040] The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.

[0041] The above network may include but is not limited to: wired network, wireless network, wherein the wired network includes: local area network, metropolitan area network and wide area network, and the wireless network includes: Bluetooth, Wi-Fi and other networks that realize wireless communication. The specific can also be determined based on the actual application scenario requirements and is not limited here.

[0042] See also Figure 2 , Figure 2 A flow chart of a method for processing the technical competition intensity of a region provided by an embodiment of the present disclosure is shown, wherein the method can be executed by any electronic device, such as a server, etc.; as an optional implementation, the method can be executed by a server. For the convenience of description, in the description of some optional embodiments below, the method will be executed by a server as an example of the execution subject of the method. Figure 2 As shown, the method for processing the technical competition intensity of a region provided by the embodiment of the present disclosure includes the following steps: S201, obtain patent vector sets of two target areas, wherein the patent vector set of each target area includes a patent vector subset corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category.

[0043] Specifically, a target area is, for example, Beijing, and multiple target areas are, for example, Beijing, Shanghai, Wuhan, etc. For example, a patent vector set of a target area includes patent vector subsets corresponding to multiple technology categories, a technology category is, for example, brain-like chips, and multiple technology categories are, for example, GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), ASIC (Application-Specific Integrated Circuit), brain-like chips, NPU (Neural Processing Unit), etc. For example, a patent vector of a technology category is a feature representation of a patent text of the technology category, and the feature representation of the patent text corresponds to the patent semantic information in the patent text.

[0044] S202: For each technology category, based on the similarity between the patent vectors of the technology category in the two target regions, determine the technology competition intensity corresponding to the technology category between the two target regions.

[0045] Specifically, for example, the greater the intensity of technological competition between two target areas corresponding to a certain technology category, the closer the technological strengths between the two target areas for that technology category, and the two target areas become each other's main competitors.

[0046] For example, for a technology category, the two target areas are target area A and target area B, the patent vector set of target area A includes patent vector subset 1 corresponding to the technology category, the patent vector set of target area B includes patent vector subset 2 corresponding to the technology category, patent vector subset 1 includes multiple patent vectors, and patent vector subset 2 includes multiple patent vectors; determine the second similarity between a patent vector in vector set 1 and each patent vector in vector set 2, that is, obtain multiple second similarities corresponding to the patent vector in vector set 1, sort the multiple second similarities from large to small, and determine the top n second similarities among the multiple second similarities as the similarity set corresponding to the patent vector in vector set 1, where n is a positive integer; based on the similarity set corresponding to each patent vector in patent vector subset 1, determine the technical competition intensity between target area A and target area B corresponding to the technology category.

[0047] S203: Determine the technical competition intensity between the two target areas based on the technical competition intensity of the two target areas corresponding to various technical categories.

[0048] Specifically, for example, the greater the intensity of technological competition between two target areas, the closer the overall technological strength between the two target areas, and the two target areas are each other's main competitors.

[0049] For example, a specific technology field includes various technology categories. The greater the intensity of technological competition between two target areas, the closer the overall technological strength between the two target areas will be for this specific technology field. Among them, specific technology fields include artificial intelligence hardware platforms, and various technology categories include GPU, FPGA, ASIC, brain-like chips, NPU, etc.

[0050] For example, the weight corresponding to each technology category is obtained; based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weightedly integrated to obtain the technology competition intensity of the two target areas.

[0051] In an embodiment of the present disclosure, patent vector sets of two target areas are obtained, and the patent vector set of each target area includes a patent vector subset corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; for each technical category, based on the similarity between the patent vectors of the technical category of the two target areas, the technical competition intensity corresponding to the technical category between the two target areas is determined; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined; in this way, based on the patent vector (the feature representation of the patent text, i.e., the patent semantic information in the patent text), the technical competition intensity between the two target areas corresponding to the technical category is calculated, and based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined, that is, the technical competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technical competition between the two target areas.

[0052] In one embodiment, for each target region, the patent vector set of the target region is obtained by: Acquire a plurality of patent texts corresponding to each technical category of at least one technical category in the target area; For each patent text, input each patent text into the fine-tuned text embedding model for feature extraction to obtain the patent vector corresponding to each patent text; Based on the patent vectors of each patent text in each technical category, a patent vector set for the target area is constructed.

[0053] Specifically, for example, multiple patent texts corresponding to each of multiple technology categories in a certain target area are obtained; multiple technology categories such as GPU, FPGA, ASIC, brain-like chip, NPU, etc.

[0054] The fine-tuned text embedding model is, for example, a fine-tuned BGE model, and the BGE model is, for example, a BGE-large-zh-1.5 model. For example, for each patent text, each patent text is input into the fine-tuned BGE-large-zh-1.5 model for feature extraction to obtain a patent vector corresponding to each patent text, and the dimension of the patent vector corresponding to each patent text is 768 dimensions; the patent vector is stored in a vector database, such as Milvus, and the Milvus vector database index method is, for example, the graph-based index method HNSW, and the patent vector similarity calculation (vector distance calculation method) uses cosine similarity.

[0055] It should be noted that, based on the fine-tuning dataset, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model; a more accurate patent vector is output through the fine-tuned text embedding model, that is, the feature representation of the patent text is optimized.

[0056] In one embodiment, the fine-tuned text embedding model is trained by: Obtain multiple patent training texts; Input each patent training text into the pre-trained text embedding model to extract features and obtain the patent vector corresponding to each patent training text; Determine the first similarity between patent vectors corresponding to each pair of patent training texts; Obtaining the standard similarity between two patent training texts obtained through expert scoring; Based on the first similarity and the corresponding standard similarity between the two patent training texts, the multiple patent training texts are divided into positive samples and negative samples, and each positive sample is used as a positive sample set; Input each negative sample and the prompt word into the trained large language model, perform text enhancement processing, and obtain patent text similar to each negative sample, wherein the prompt word is used to instruct the trained large language model to output patent text similar to each negative sample; Each negative sample and patent texts similar to each negative sample are taken as a negative sample set; The positive sample set and the negative sample set are used as fine-tuning datasets. Based on the fine-tuning datasets, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model.

[0057] Specifically, for example, the feature representation optimization method of patent text based on expert feedback fine-tuning is as follows Figure 3As shown, (1) the model domain continued training and expert evaluation dataset construction include: based on a large-scale dataset, the initial pre-trained text embedding model is trained in the model domain to obtain a pre-trained text embedding model (domain-trained model); multiple patent training texts are obtained; each patent training text is input into the pre-trained text embedding model, feature extraction is performed, and the patent vector corresponding to each patent training text is obtained (the feature representation of the patent text); the first similarity between the patent vectors corresponding to the two patent training texts is determined, that is, the patent similarity calculation is performed; (2) the expert evaluation dataset construction includes: each first similarity and multiple patent training texts are stored in the database as similar patent data; the standard similarity between the two patent training texts is obtained through expert scoring (expert scoring and evaluation) (expert evaluation data); (3) the fine-tuning dataset construction includes: based on the first similarity between the two patent training texts and the corresponding standard similarity (expert evaluation data), the multiple patent training texts are divided into positive samples and negative samples, each positive sample is used as a positive sample set (positive sample dataset), and each negative sample is used as a negative sample dataset; each negative sample and the prompt word (Prompt) are input into the trained A large language model is used to perform text enhancement processing to obtain patent texts similar to each negative sample, wherein the prompt word is used to instruct the trained large language model to output patent texts similar to each negative sample, such as BLOOM, Baichuan, atom, etc.; (4) Model fine-tuning includes: taking each negative sample and patent texts similar to each negative sample as a negative sample set, wherein the patent texts similar to each negative sample are negative example data sets generated by the trained large language model, for example, negative example data sets are generated based on different trained large language models; the positive sample set and the negative sample set are combined into a negative sample set. This set is used as a fine-tuning dataset with balanced positive and negative examples (the ratio of positive and negative data is determined). In this way, fine-tuning datasets based on different pre-trained text embedding models can be obtained; pre-trained text embedding models are selected; based on the fine-tuning dataset and fine-tuning strategy, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model. Different pre-trained text embedding models correspond to different fine-tuning strategies, that is, fine-tuning strategies based on different pre-trained text embedding models; the performance of the fine-tuned text embedding model is evaluated based on the validation set, that is, the performance of the fine-tuned text embedding model is evaluated on the validation set.

[0058] It should be noted that, based on the fine-tuning dataset, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model; a more accurate patent vector is outputted by the fine-tuned text embedding model, that is, the feature representation of the patent text is optimized; the method provided in the embodiment of the present disclosure aims to solve the limitations of the traditional patent text similarity calculation method when considering the professional patent text context, that is, the patent text similarity calculation in a specific technical field (specific technical field such as artificial intelligence hardware platform) based on the pre-trained model (pre-trained text embedding model) obtained by corpus training has the limitation of fine-tuning data construction; (1) In the method provided in the embodiment of the present disclosure, the initial pre-trained model is continuously trained on a large scale to learn and obtain relevant field knowledge to obtain a pre-trained model; after obtaining the feature representation of the patent text through the pre-trained model, the similarity between patents is calculated based on the vector database, and experts score the similarity of similar patent pairs on this data and give reasons for the score, and a fine-tuning dataset is constructed through the feedback of experts on the similarity calculation results; since in the process of model fine-tuning, attention is paid to whether the similarity calculated by the pre-trained model can be The similarity (first similarity) of the patent is closer to the similarity (standard similarity) given by the expert, so the focus is on the changes in the fine-tuning of the patents with a large gap between the similarity (first similarity) calculated by the pre-trained model and the similarity (standard similarity) given by the expert; (2) through data enhancement, the data (patent text) with a large gap between the similarity given by the pre-trained model and the similarity given by the expert is used as a negative example (negative sample), and the data with a small gap is used as a positive example. The amount of positive and negative examples is unbalanced, so the negative examples are expanded. On the basis of maintaining the original semantic information, the expert's evaluation perspective is further integrated to obtain a richer and more accurate semantic representation and obtain a fine-tuning data set with balanced positive and negative examples; (3) fine-tuning one or more pre-trained models on the constructed fine-tuning data set with balanced positive and negative examples, specifically, using the positive and negative example data set that has been text-enhanced by the large language model to perform supervised training on the pre-trained model; usually a small learning rate and an appropriate number of training steps are used for fine-tuning to retain the semantic representation ability of the pre-trained model; the base selected in the method provided in the embodiment of the present disclosure is the BGE model.

[0059] In one embodiment, based on the first similarity and the corresponding standard similarity between the two patent training texts, a plurality of patent training texts are divided into positive samples and negative samples, including: For each patent training text, if there is a difference between each first similarity corresponding to the patent training text and the corresponding standard similarity that is greater than a preset threshold, the patent training text is determined as a negative sample; If the differences between each first similarity corresponding to the patent training text and the corresponding standard similarity are all less than or equal to a preset threshold, the patent training text is determined as a positive sample.

[0060] Specifically, for example, there are 10 patent training texts, and there are 9 first similarities between patent training text 1 in the 10 patent training texts and the other 9 patent training texts in the 10 patent training texts, each of the 9 first similarities corresponds to a standard similarity, and the difference between each of the 9 first similarities and the corresponding standard similarity is calculated to obtain 9 differences. If there is a difference greater than a preset threshold among the 9 differences, that is, there is a large gap between the first similarities and the corresponding standard similarities among the 9 first similarities, then patent training text 1 is determined as a negative sample; if there is no difference greater than the preset threshold among the 9 differences (the 9 differences are all less than or equal to the preset threshold), that is, the gap between each of the 9 first similarities and the corresponding standard similarities is small, then patent training text 1 is determined as a positive sample.

[0061] It should be noted that the patent training texts with a large gap between the first similarity and the standard similarity given by the expert are taken as negative samples, and the patent training texts with a small gap between the first similarity and the standard similarity given by the expert are taken as positive samples.

[0062] In one embodiment, for each technology category, the patent vector subsets of the technology category in the two target regions are respectively used as the first vector set and the second vector set, and the technology competition intensity corresponding to the technology category between the two target regions is determined, including: Determining a second similarity between each patent vector in the first vector set and each patent vector in the second vector set; According to the order of the second similarities from large to small, a preset number of second similarities ranked first among the second similarities are determined as a similarity set of the two target areas corresponding to the technology category; Based on the similarity set, the technical competition intensity of the two target areas corresponding to the technology category is determined.

[0063] Specifically, for example, the first vector set includes patent vector 1 and patent vector 2, and the second vector set includes patent vector 6, patent vector 7, patent vector 8, patent vector 9 and patent vector 10; determine the second similarity between each patent vector in the first vector set and each patent vector in the second vector set; the second similarities between patent vector 1 and patent vector 6, patent vector 7, patent vector 8, patent vector 9 and patent vector 10 are second similarity 1, second similarity 2, second similarity 3, second similarity 4 and second similarity 5, respectively, sort the second similarity 1, second similarity 2, second similarity 3, second similarity 4 and second similarity 5 from large to small, select the first two second similarities, for example, the first two second similarities 2 and the second Similarity 4, the preset number is 2; the second similarity 2 and the second similarity 4 are set in the similarity set corresponding to the technology category of the two target areas; the second similarities between patent vector 2 and patent vector 6, patent vector 7, patent vector 8, patent vector 9 and patent vector 10 are second similarity 6, second similarity 7, second similarity 8, second similarity 9 and second similarity 10 respectively, and the second similarity 6, second similarity 7, second similarity 8, second similarity 9 and second similarity 10 are sorted from large to small, and the first two second similarities are selected, and the first two second similarities are, for example, second similarity 6 and second similarity 9, the preset number is 2; the second similarity 6 and the second similarity 9 are set in the similarity set corresponding to the technology category of the two target areas.

[0064] For example, determine the intensity of technological competition between two target regions, including: (1) Data preprocessing.

[0065] Label the patent data in a certain field by year. For example, for each technology category t, collect patent data (patent text) in a certain field on technology category t, label each patent (patent text) by region (target region), and obtain a regional label, such as Beijing.

[0066] For example, patent information data processing divides patents into corresponding technology categories according to the "artificial intelligence classification name" field in the patent data and the artificial intelligence classification framework.

[0067] For example, in the data processing of patent owner information, since the patent owner information given for each patent is inconsistent, there are errors in some of the patent owner information data, especially in the fields of cities and districts and counties. Therefore, it is necessary to verify the patent owner information; in the information of cities, districts and counties to which the patent owner belongs, the area with the most appearances in all patents corresponding to the patent owner is selected as the final patent owner's area; then the data of the patent owners in the cities, districts and counties that need to be analyzed are extracted and merged to facilitate subsequent retrieval.

[0068] (2) Calculation of the intensity of technological competition between regions (target regions) by technology category (the intensity of technological competition between two target regions corresponding to technology category t), including steps 1 to 3: Step 1: Extract the patent sets of region A (target region A) and region B (target region B) in technology category t, denoted as and ; Step 2: For each patent in region A in technology category t, obtain the top n patents in region B that are most similar to it, and store the cosine similarity (the top n second similarities), where n is a positive integer.

[0069] Step 3: Repeat step 2 for all patents in region A in technology category t to obtain all patents and cosine similarities in region B related to region A; sum up each cosine similarity (each second similarity in the similarity set) to obtain the technical competition intensity between region A and region B. , that is, the technological competition intensity of the two target areas corresponding to technology category t ,calculate The formula (1) is as follows: Formula (1) in, and Represent the patent vector of patent p and the patent vector of patent q respectively; : The patent set of region A in technology category t; p: Each patent in region A; : The set of patents similar to patent p in region B; cos(Vec(p),Vec(q)): The cosine similarity between the patent vector of patent p and the patent vector of patent q, indicating the similarity between the patent vector of patent p and the patent vector of patent q in the vector space; : Find the top n patents in region B that are most similar to patent p.

[0070] (3) Regional (target region) technology competition intensity vector For any region, the technical competition intensity with a region B (assuming there are m regions in total) in all technical categories (assuming there are n technical categories in total) will form a technical competition intensity vector , As shown in formula (2): Formula (2) in, Represents the region index, Indicates the technology category index, represents the technical competition intensity between region i and region B in the jth technical category; there are n technical categories in total, The dimension is n.

[0071] (4) Calculation of inter-regional technological competition intensity that incorporates regional technological characteristics, that is, determining the technological competition intensity between two target regions.

[0072] Based on formula (1) and formula (2), determine the technical competition intensity vector of any region with a region B in n technology categories: ; In the technology competition intensity vector In the above formula, each component expresses the technical competition intensity between any region i and a region B in the jth technology category; the technical competition intensity between any region and a region B is calculated as , a simple calculation method is to vector The components of are summed up, as shown in formula (3): Formula (3) Among them, in the calculation of the technical competition intensity between any region and a certain region B, the technical competition intensity vector The amount are equally weighted.

[0073] In fact, the regional technology characteristics of any region in each technology category are not effectively expressed. For example, the distribution of patents in any region A in n technology categories is not uniform. Region A may be strong in some technology categories and have a large number of patents, but may have almost no patents in some technology categories. In order to express this difference, the regional technology characteristic factor is introduced. , The calculation method is shown in formula (4): Formula (4) Therefore, when the regional technology characteristic factors are introduced, the technological competition intensity between any region and a region B is The calculation method is changed to formula (5): Formula (5) in, represents the technical competition intensity between region i and region B in the jth technology category, Represents the regional technology characteristic factor of region i in the jth technology category.

[0074] It should be noted that the intensity of technological competition between the two target areas corresponding to the technology categories is calculated based on the patent vector, and the intensity of technological competition between the two target areas corresponding to various technology categories is determined based on the intensity of technological competition between the two target areas, that is, the technological competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technological competition between the two target areas.

[0075] In one embodiment, determining the technical competition intensity of two target areas corresponding to the technical category based on the similarity set includes: A preset number of second similarities in the similarity set are fused, and a fusion result is determined as the technical competition intensity of the two target areas corresponding to the technical category.

[0076] Specifically, for example, the preset number is 50, the 50 second similarities are summed to obtain a sum result, and the sum result is determined as the technical competition intensity of the two target areas corresponding to the technical category.

[0077] For example, as shown in formula (1), by summing up the cosine similarities (the second similarities in the similarity set), we can obtain the technical competition intensity of region A and region B (two target regions) corresponding to technology category t: .

[0078] In one embodiment, determining the technical competition intensity between the two target areas based on the technical competition intensity of the two target areas corresponding to various technical categories includes: Get the weight corresponding to each technology category; Based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weightedly integrated to obtain the technology competition intensity of the two target areas.

[0079] Specifically, the weight corresponding to each technology category is the regional technology characteristic factor shown in formula (4): For example, as shown in formula (5), based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weighted and fused to obtain the technology competition intensity of the two target areas (area i and area B): .

[0080] It should be noted that the intensity of technological competition between the two target areas corresponding to the technology categories is calculated based on the patent vector, and the intensity of technological competition between the two target areas corresponding to various technology categories is determined based on the intensity of technological competition between the two target areas, that is, the technological competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technological competition between the two target areas.

[0081] In one embodiment, obtaining a plurality of patent texts corresponding to each technical category in the target area includes: Acquire regional information of the target area, where the regional information includes specific technical fields of the target area; Obtain multiple patent texts for each technology category in at least one technology category related to a specific technology field.

[0082] Specifically, for example, a specific technology field includes various technology categories. The greater the intensity of technological competition between two target areas, the closer the overall technological strength between the two target areas for this specific technology field; among them, specific technology fields include artificial intelligence hardware platforms, and various technology categories include GPU, FPGA, ASIC, brain-like chips, NPU, etc.

[0083] The application of the embodiments of the present disclosure has at least the following beneficial effects: Based on the fine-tuning dataset, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model; a more accurate patent vector is output through the fine-tuned text embedding model, that is, the feature representation of the patent text is optimized; based on the patent vector (the feature representation of the patent text, that is, the patent semantic information in the patent text), the technical competition intensity between the two target areas corresponding to the technical categories is calculated, and based on the technical competition intensity between the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined, that is, the technical competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technical competition between the two target areas.

[0084] In order to better understand the method provided by the embodiment of the present disclosure, the solution of the embodiment of the present disclosure is further described below with reference to examples of specific application scenarios.

[0085] In one embodiment, for example, Table 3 shows that the “artificial intelligence hardware platform” includes several technology categories, such as GPU, FPGA, ASIC, brain-like chip, NPU, etc.; through the method provided by the embodiment of the present disclosure, the technology competition city and technology competition intensity of “Beijing” are obtained, and Table 3 is shown as follows: Table 3: Technology competition cities and technology competition intensity of “Beijing”

[0086] In a specific application scenario embodiment, such as a regional technology competition scenario, see Figure 4 , shows the processing flow of a method for processing the technical competition intensity of a region, such as Figure 4 As shown, the processing flow of the method for processing the technical competition intensity of a region provided by the embodiment of the present disclosure includes the following steps: S401, the server obtains a fine-tuning dataset.

[0087] Specifically, for example, Figure 3 As shown, based on the first similarity between each pair of patent training texts and the corresponding standard similarity (expert evaluation data), multiple patent training texts are divided into positive samples and negative samples, each positive sample is used as a positive sample set (positive example data set), and each negative sample is used as a negative example data set; each negative sample and a prompt word (Prompt) are input into the trained large language model, and text enhancement processing is performed to obtain a patent text similar to each negative sample, wherein the prompt word is used to instruct the trained large language model to output a patent text similar to each negative sample; each negative sample and a patent text similar to each negative sample are used as a negative sample set, and the positive sample set and the negative sample set are used as fine-tuning data sets.

[0088] S402: The server fine-tunes the pre-trained text embedding model based on the fine-tuning dataset to obtain a fine-tuned text embedding model.

[0089] Specifically, for example, a pre-trained text embedding model is fine-tuned using a smaller learning rate and an appropriate number of training steps to retain the semantic representation ability of the pre-trained text embedding model.

[0090] S403, the server obtains a plurality of patent texts corresponding to each of the plurality of technical categories in each of the two target areas.

[0091] Specifically, for example, regional information of the target area is obtained, where the regional information includes a specific technical field of the target area; and multiple patent texts of each technical category in multiple technical categories related to the specific technical field are obtained.

[0092] S404, the server performs feature extraction based on multiple patent texts corresponding to various technical categories in each target area through a fine-tuned text embedding model to obtain a patent vector set for the target area.

[0093] Specifically, for example, for each patent text, each patent text is input into a fine-tuned text embedding model for feature extraction to obtain a patent vector corresponding to each patent text; based on the patent vectors of each patent text in each technical category, a patent vector set for the target area is constructed; and the patent vector set for the target area is stored in a vector database.

[0094] S405, the server determines, for each technology category, the technology competition intensity corresponding to the technology category between the two target regions based on the similarity between the patent vectors of the technology category in the two target regions.

[0095] Specifically, for example, as shown in formula (1), the sum of each cosine similarity (each second similarity in the similarity set) is obtained to obtain the technical competition intensity of region A and region B (two target regions) corresponding to technology category t: .

[0096] S406: The server performs weighted fusion of the technical competition intensities of the two target areas corresponding to the various technical categories based on the weights corresponding to the various technical categories, to obtain the technical competition intensities of the two target areas.

[0097] Specifically, the weight corresponding to each technology category is the regional technology characteristic factor shown in formula (4): For example, as shown in formula (5), based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weighted and fused to obtain the technology competition intensity of the two target areas (area i and area B): .

[0098] The application of the embodiments of the present disclosure has at least the following beneficial effects: From the perspective of method and technology, based on patent data, we explore the precise identification of regional technology competitors; the basis for calculating regional technology competitors is the feature representation optimization method of patent texts based on expert feedback, and similar patents with any patent are obtained through patent text similarity, and then the technical competition intensity between regions is obtained based on the regional labels of similar patents; from the perspective of practical application of regional technology competitor identification, the provided patent recommendation and analysis methods have brought more accurate Chinese patent analysis tools to enterprises, scientific research institutions and science and technology management departments; in the domestic market, patent recommendations achieved through model optimization can help enterprises and scientific research institutions to more comprehensively grasp the technical layout and innovation direction of regional competitors and clarify their relative technical strength in the industry; at the national level, this method provides a global perspective for science and technology management departments, making it easier to identify major competitors and potential challengers in key technology fields, thereby supporting the formulation of science and technology policies and resource allocation; by accurately identifying technology competitors at the regional level, enterprises and countries can better allocate R&D resources, avoid repeated investment, improve innovation efficiency, and promote the overall improvement of domestic technological strength.

[0099] The embodiment of the present disclosure also provides a device for processing the technical competition intensity of a region. The structural schematic diagram of the device for processing the technical competition intensity of a region is as follows: Figure 5 As shown, the device 50 for processing the technical competition intensity of a region includes a first processing module 501 , a second processing module 502 and a third processing module 503 .

[0100] The first processing module 501 is used to obtain patent vector sets of two target areas, where the patent vector set of each target area includes a patent vector subset corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; The second processing module 502 is used to determine, for each technology category, the technology competition intensity corresponding to the technology category between the two target regions based on the similarity between the patent vectors of the technology category in the two target regions; The third processing module 503 is used to determine the technical competition intensity between the two target areas based on the technical competition intensity of the two target areas corresponding to various technical categories.

[0101] In one embodiment, for each target region, the patent vector set of the target region is obtained by the first processing module 501 in the following manner: Acquire a plurality of patent texts corresponding to each technical category of at least one technical category in the target area; For each patent text, input each patent text into the fine-tuned text embedding model for feature extraction to obtain the patent vector corresponding to each patent text; Based on the patent vectors of each patent text in each technical category, a patent vector set for the target area is constructed.

[0102] In one embodiment, the fine-tuned text embedding model is trained by the first processing module 501 in the following manner: Obtain multiple patent training texts; Input each patent training text into the pre-trained text embedding model to extract features and obtain the patent vector corresponding to each patent training text; Determine the first similarity between patent vectors corresponding to each pair of patent training texts; Obtaining the standard similarity between two patent training texts obtained through expert scoring; Based on the first similarity and the corresponding standard similarity between the two patent training texts, the multiple patent training texts are divided into positive samples and negative samples, and each positive sample is used as a positive sample set; Input each negative sample and the prompt word into the trained large language model, perform text enhancement processing, and obtain patent text similar to each negative sample, wherein the prompt word is used to instruct the trained large language model to output patent text similar to each negative sample; Each negative sample and patent texts similar to each negative sample are taken as a negative sample set; The positive sample set and the negative sample set are used as fine-tuning datasets. Based on the fine-tuning datasets, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model.

[0103] In one embodiment, the first processing module 501 is specifically configured to: For each patent training text, if there is a difference between each first similarity corresponding to the patent training text and the corresponding standard similarity that is greater than a preset threshold, the patent training text is determined as a negative sample; If the differences between each first similarity corresponding to the patent training text and the corresponding standard similarity are all less than or equal to a preset threshold, the patent training text is determined as a positive sample.

[0104] In one embodiment, for each technology category, the patent vector subsets of the technology category in the two target areas are respectively used as the first vector set and the second vector set, and the second processing module 502 is specifically used to: Determining a second similarity between each patent vector in the first vector set and each patent vector in the second vector set; According to the order of the second similarities from large to small, a preset number of second similarities ranked first among the second similarities are determined as a similarity set of the two target areas corresponding to the technology category; Based on the similarity set, the technical competition intensity of the two target areas corresponding to the technology category is determined.

[0105] In one embodiment, the second processing module 502 is specifically configured to: A preset number of second similarities in the similarity set are fused, and a fusion result is determined as the technical competition intensity of the two target areas corresponding to the technical category.

[0106] In one embodiment, the third processing module 503 is specifically configured to: Get the weight corresponding to each technology category; Based on the weights corresponding to various technology categories, the technology competition intensity of the two target areas corresponding to various technology categories is weightedly integrated to obtain the technology competition intensity of the two target areas.

[0107] In one embodiment, the first processing module 501 is specifically configured to: Acquire regional information of the target area, where the regional information includes specific technical fields of the target area; Obtain multiple patent texts for each technology category in at least one technology category related to a specific technology field.

[0108] The application of the embodiments of the present disclosure has at least the following beneficial effects: A patent vector set of two target areas is obtained, wherein the patent vector set of each target area includes a patent vector subset corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; for each technical category, based on the similarity between the patent vectors of the technical category of the two target areas, the technical competition intensity corresponding to the technical category between the two target areas is determined; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined; in this way, based on the patent vector (the feature representation of the patent text, i.e., the patent semantic information in the patent text), the technical competition intensity between the two target areas corresponding to the technical category is calculated, and based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined, that is, the technical competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technical competition between the two target areas.

[0109] The present disclosure also provides an electronic device. The structural diagram of the electronic device is as follows: Figure 6 As shown, Figure 6 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure.

[0110] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0111] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0112] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.

[0113] The memory 4003 is used to store the computer program for executing the embodiment of the present disclosure, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiment.

[0114] Among them, electronic equipment includes but is not limited to: servers, etc.

[0115] The application of the embodiments of the present disclosure has at least the following beneficial effects: A patent vector set of two target areas is obtained, wherein the patent vector set of each target area includes a patent vector subset corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; for each technical category, based on the similarity between the patent vectors of the technical category of the two target areas, the technical competition intensity corresponding to the technical category between the two target areas is determined; based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined; in this way, based on the patent vector (the feature representation of the patent text, i.e., the patent semantic information in the patent text), the technical competition intensity between the two target areas corresponding to the technical category is calculated, and based on the technical competition intensity of the two target areas corresponding to various technical categories, the technical competition intensity between the two target areas is determined, that is, the technical competition between the two target areas is evaluated, thereby improving the accuracy of evaluating the technical competition between the two target areas.

[0116] An embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.

[0117] The embodiments of the present disclosure also provide a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiments when executed by a processor.

[0118] It should be understood that, although the flowchart of the embodiment of the present disclosure indicates each operation step by arrows, the implementation order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiment of the present disclosure, the implementation steps in each flowchart can be executed in other orders according to demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times. In scenarios with different execution times, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present disclosure does not limit this.

[0119] The above is only an optional implementation method for some implementation scenarios of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the scheme of the present disclosure, other similar implementation methods based on the technical ideas of the present disclosure are also within the protection scope of the embodiments of the present disclosure.

Claims

1. A method for processing the technical competition intensity of a region, characterized in that: include: Obtain patent vector sets of two target regions, wherein the patent vector set of each target region includes a patent vector subset corresponding to at least one technical category, and the patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; For each technology category, based on the similarity between the patent vectors of the technology category in the two target regions, determine the technology competition intensity corresponding to the technology category between the two target regions; Based on the technology competition intensities of the two target areas corresponding to various technology categories, the technology competition intensity between the two target areas is determined.

2. The method according to claim 1, characterized in that: For each of the target regions, the patent vector set of the target region is obtained by: Acquire a plurality of patent texts corresponding to each of the at least one technical category in the target area; For each patent text, input each patent text into the fine-tuned text embedding model for feature extraction to obtain a patent vector corresponding to each patent text; Based on the patent vectors of each patent text in each technical category, a patent vector set for the target area is constructed.

3. The method according to claim 2, characterized in that The fine-tuned text embedding model is trained in the following way: Obtain multiple patent training texts; Input each patent training text into the pre-trained text embedding model to perform feature extraction to obtain a patent vector corresponding to each patent training text; Determine the first similarity between patent vectors corresponding to each pair of patent training texts; Obtaining the standard similarity between the two patent training texts obtained through expert scoring; Based on the first similarity and the corresponding standard similarity between the two patent training texts, the plurality of patent training texts are divided into positive samples and negative samples, and each positive sample is used as a positive sample set; Input each negative sample and the prompt word into the trained large language model, perform text enhancement processing, and obtain a patent text similar to each negative sample, wherein the prompt word is used to instruct the trained large language model to output a patent text similar to each negative sample; Each negative sample and patent texts similar to each negative sample are taken as a negative sample set; The positive sample set and the negative sample set are used as fine-tuning data sets, and based on the fine-tuning data sets, the pre-trained text embedding model is fine-tuned to obtain a fine-tuned text embedding model.

4. The method according to claim 3, characterized in that The method of dividing the plurality of patent training texts into positive samples and negative samples based on the first similarity between the two patent training texts and the corresponding standard similarity comprises: For each patent training text, if there is a difference between each first similarity corresponding to the patent training text and the corresponding standard similarity that is greater than a preset threshold, the patent training text is determined as a negative sample; If the differences between each first similarity corresponding to the patent training text and the corresponding standard similarity are all less than or equal to a preset threshold, the patent training text is determined as a positive sample.

5. The method according to claim 1, characterized in that For each technology category, the patent vector subsets of the technology category in the two target regions are respectively used as the first vector set and the second vector set, and the determining of the technology competition intensity corresponding to the technology category between the two target regions includes: Determining a second similarity between each patent vector in the first vector set and each patent vector in the second vector set; According to the order of the second similarities from large to small, a preset number of second similarities ranked first among the second similarities are determined as a similarity set of the two target areas corresponding to the technology category; Based on the similarity set, the technical competition intensity of the two target areas corresponding to the technical category is determined.

6. The method according to claim 5, characterized in that The determining, based on the similarity set, the technical competition intensity of the two target areas corresponding to the technical category comprises: A preset number of second similarities in the similarity set are fused, and a fusion result is determined as the technical competition intensity of the two target areas corresponding to the technical category.

7. The method according to claim 1, characterized in that The determining the technical competition intensity between the two target areas based on the technical competition intensity of the two target areas corresponding to various technical categories includes: Get the weight corresponding to each technology category; Based on the weights corresponding to various technology categories, the technology competition intensities of the two target areas corresponding to various technology categories are weightedly fused to obtain the technology competition intensities of the two target areas.

8. The method according to claim 2, characterized in that: The step of obtaining a plurality of patent texts corresponding to each of the at least one technical category in the target area includes: Acquire regional information of the target area, wherein the regional information includes a specific technical field of the target area; A plurality of patent texts of each technical category in the at least one technical category related to the specific technical field is obtained.

9. A device for processing the technical competition intensity of a region, characterized in that: include: The first processing module is used to obtain patent vector sets of two target areas, each of which includes a patent vector subset corresponding to at least one technical category, and each patent vector subset corresponding to each technical category includes multiple patent vectors of the technical category, and a patent vector of each technical category is a feature representation of a patent text of the technical category; A second processing module is used to determine, for each technology category, the technology competition intensity corresponding to the technology category between the two target regions based on the similarity between the patent vectors of the technology category in the two target regions; The third processing module is used to determine the technical competition intensity between the two target areas based on the technical competition intensity of the two target areas corresponding to various technical categories.

10. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Method and system for calculating competitiveness betweens objects

    CN101393550A

  • Technical competition and patent early warning analysis method based on knowledge discovery

    CN106897392A

  • Deep technology tracking method for high-tech companies

    CN110580261A

  • Method and device for determining core degree of patent technology, electronic equipment and storage medium

    CN114331766A

  • Method and device for mining competitors based on patent heterogeneous information network

    CN115641009A