Patent data management method and system based on large natural language model
Patent Information
- Application Number
- CN202311825450.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2043-12-26
AI Technical Summary
[0005]基于此,本申请实施例提供了一种基于大型自然语言模型的专利数据管理方法及其系统,以解决现有技术中专利侵权风险排查效率较低的问题
[0010]The beneficial effects compared with the prior art are as follows: The patent data management method based on a large natural language model provided in this application allows the terminal device to first obtain the first patent set information of the target enterprise and the second patent set information of at least one designated enterprise based on a preset patent database. Then, based on the first patent set information and the second patent set information, technology development tree information is generated. Based on the technology development tree information, infringement risk warning information is generated. Thus, by using the technology development tree information, it is possible to quickly check whether there are potential patent infringement risks in the publicly disclosed patents of other enterprises, and to find out the patents of other enterprises with potential patent infringement risks, thereby improving the efficiency of patent infringement risk screening and solving the problem of low efficiency in current patent infringement risk screening to a certain extent.
Smart Images

Figure CN117874026B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to a patent data management method and system based on a large-scale natural language model. Background Technology
[0002] Companies typically apply for patents for existing or future planned products or technologies. For example, they may create a 1:1 patent design covering all the technical features of an existing product, thus establishing a close relationship between the product and the patent. However, patent infringement disputes can pose significant economic risks to companies, such as litigation costs or hefty compensation claims. To mitigate these potential risks, companies usually conduct patent infringement risk assessments of their existing patents to identify any patents that conflict with their products or technologies and take appropriate preventative measures.
[0003] A common method for identifying patent infringement risks is for companies to first identify patents from other companies that are highly similar to their own, and then consider these as potential risks to the products corresponding to their own patents.
[0004] Currently, companies typically use manual methods to investigate patent infringement risks. When a company already owns a large number of patents, this process takes a considerable amount of time and is inefficient, requiring further improvement. Summary of the Invention
[0005] Based on this, embodiments of this application provide a patent data management method and system based on a large-scale natural language model to solve the problem of low efficiency in patent infringement risk screening in the prior art.
[0006] In a first aspect, embodiments of this application provide a patent data management method based on a large-scale natural language model, the method comprising:
[0007] Based on a preset patent database, information on a first patent set of a target company and information on a second patent set of at least one designated company are obtained. The patent database stores information on the first patent set and information on the second patent set. The first patent set information is used to describe the set of first published patents of the target company, and the second patent set information is used to describe the set of second published patents of the at least one designated company.
[0008] Based on the information from the first patent set and the information from the second patent set, generate technology development tree information;
[0009] Based on the aforementioned technology development tree information, infringement risk warning information is generated.
[0010] The beneficial effects compared with the prior art are as follows: The patent data management method based on a large natural language model provided in this application allows the terminal device to first obtain the first patent set information of the target enterprise and the second patent set information of at least one designated enterprise based on a preset patent database. Then, based on the first patent set information and the second patent set information, technology development tree information is generated. Based on the technology development tree information, infringement risk warning information is generated. Thus, by using the technology development tree information, it is possible to quickly check whether there are potential patent infringement risks in the publicly disclosed patents of other enterprises, and to find out the patents of other enterprises with potential patent infringement risks, thereby improving the efficiency of patent infringement risk screening and solving the problem of low efficiency in current patent infringement risk screening to a certain extent.
[0011] Secondly, embodiments of this application provide a patent data management system based on a large-scale natural language model, the system comprising:
[0012] First patent set information acquisition module: used to acquire first patent set information of target enterprise and second patent set information of at least one designated enterprise based on a preset patent database, wherein the patent database stores the first patent set information and the second patent set information, the first patent set information is used to describe the set of first published patents of the target enterprise, and the second patent set information is used to describe the set of second published patents of the at least one designated enterprise;
[0013] Technology development tree information generation module: used to generate technology development tree information based on the first patent set information and the second patent set information;
[0014] Infringement risk warning information generation module: used to generate infringement risk warning information based on the technology development tree information.
[0015] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described in the first aspect above.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in the first aspect above.
[0017] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0019] Figure 1 This is a schematic flowchart of a patent data management method provided in an embodiment of this application;
[0020] Figure 2 This is a flowchart illustrating step S200 in a patent data management method provided in an embodiment of this application;
[0021] Figure 3 This is a flowchart illustrating step S230 of a patent data management method provided in an embodiment of this application;
[0022] Figure 4 This is a schematic diagram of the target patent sequence information provided in an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of the backbone information provided in an embodiment of this application;
[0024] Figure 6 This is a flowchart illustrating step S260 in a patent data management method provided in an embodiment of this application;
[0025] Figure 7 This is a schematic diagram of the technology development tree information provided in one embodiment of this application;
[0026] Figure 8 This is a flowchart illustrating step S300 in a patent data management method provided in an embodiment of this application;
[0027] Figure 9 This is a block diagram of a patent data management system provided in one embodiment of this application;
[0028] Figure 10 This is a schematic diagram of a terminal device provided in an embodiment of this application. Detailed Implementation
[0029] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0030] In the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0031] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0032] To illustrate the technical solution described in this application, specific embodiments are provided below.
[0033] Please see Figure 1 , Figure 1 This is a flowchart illustrating the patent data management method based on a large-scale natural language model provided in this application embodiment. In this embodiment, the execution subject of the patent data management method is a terminal device. It is understood that the types of terminal devices include, but are not limited to, mobile phones, tablets, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. This application embodiment does not impose any restrictions on the specific type of terminal device.
[0034] Please see Figure 1 The patent data management method provided in this application includes, but is not limited to, the following steps:
[0035] In S100, based on a preset patent database, information on the first patent set of the target company and information on the second patent set of at least one designated company are obtained.
[0036] Without loss of generality, the patent database stores information on a first patent set and information on a second patent set; the first patent set information describes the set of first-published patents of the target company, and the first-published patents describe the patents that the target company has disclosed; the second patent set information describes the set of second-published patents of at least one designated company, and the designated companies describe other companies besides the target company. Both the target company and the designated companies can be preset manually, and the second-published patents describe the patents that the designated companies have disclosed.
[0037] Specifically, the terminal device can obtain the first patent set information of the target company and the second patent set information of at least one designated company based on a preset patent database.
[0038] In S200, technology development tree information is generated based on the information from the first patent set and the second patent set.
[0039] Specifically, the terminal device can generate technology development tree information based on the first patent set information and the second patent set information. The technology development tree information is used to describe the relationship and development trend between the technologies represented by different patents. It is similar to the shape of a tree, starting from a starting node and branching out different child nodes. Each child node is a patent and represents a technology.
[0040] In some possible implementations, the technology development tree information is divided into trunk information and multiple branch information. The trunk information is the core part of the technology development tree information, and the branch information is an extension of the trunk information.
[0041] For information on possible implementations that facilitate efficient and accurate identification of potential infringement risks, please refer to [link / reference]. Figure 2 Step S200 includes, but is not limited to, the following steps:
[0042] In S210, based on a preset technical term database and a preset word matching algorithm, the first identical word set information corresponding to each first published patent and the second identical word set information corresponding to each second patent set information are obtained.
[0043] For example, this patent data management method can be based on a large-scale natural language model. This model can access a technical term database, which can be pre-created by maintenance personnel. The database stores multiple technical term entries, which describe words associated with the target company's existing products or future planned products, such as "transformer," "GIS disconnect switch," and "silicon steel core." The word matching algorithm is used to find identical words; it can be a forward maximum matching algorithm, a backward maximum matching algorithm, or a bidirectional maximum matching algorithm.
[0044] Without loss of generality, the first set of identical words information is used to describe the set of first identical words information, which describes words in the first published patent that are identical to any one of the technical terms. The second set of identical words information is used to describe the set of second identical words information, which describes words in the second published patent that are identical to any one of the technical terms. For example, when the first published patent contains the term "GIS disconnect switch," and one of the technical terms pre-stored in the technical term database is "GIS disconnect switch," then the term "GIS disconnect switch" in the first published patent is the first set of identical words information.
[0045] Specifically, the terminal device can use a preset word matching algorithm to match technical terms in a preset technical term database with terms in the first disclosed patent, and simultaneously match technical terms in the technical term database with terms in the second disclosed patent, to obtain first identical word set information corresponding to each first disclosed patent and second identical word set information corresponding to each second patent set. The first identical word set information can be "lightning characteristics", "GIS disconnect switch", "electrical withstand level", "power grid lightning damage", and "lightning fault database".
[0046] In S220, the number of first identical words corresponding to each first disclosed patent is determined based on the first identical word set information corresponding to each first disclosed patent, and the number of second identical words corresponding to each second disclosed patent is determined based on the second identical word set information corresponding to each second patent set information.
[0047] Without loss of generality, the first identical word quantity information is used to describe the quantity of first identical word information in the first disclosed patent. For example, when the first identical word set information corresponding to a certain first disclosed patent is "lightning characteristics", "GIS disconnect switch", "electrical withstand level", "grid lightning damage", and "lightning fault database", the first identical word quantity information is 5. The second identical word quantity information is used to describe the quantity of second identical word information in the second disclosed patent.
[0048] Specifically, the terminal device can determine the number of first identical words corresponding to each first disclosed patent based on the first identical word set information corresponding to each first disclosed patent, and at the same time determine the number of second identical words corresponding to each second disclosed patent based on the second identical word set information corresponding to each second patent set information.
[0049] In S230, for each application year: target patent information is determined based on the first identical word count information and the second identical word count information.
[0050] Specifically, the terminal device can first determine the relevant application years based on the specific application years corresponding to each first-disclosure patent and each second-disclosure patent. For example, when the specific application years corresponding to several first-disclosure patents are "2020", "2022", and "2023", and the specific application years corresponding to several second-disclosure patents are "2019", "2021", and "2022", the relevant application years are "2019", "2020", "2021", "2022", and "2023". The terminal device can perform the following processing for each relevant application year: determine the target patent information based on the first and second identical word count information. Since the technical terms provided by the technical term database are strongly associated with existing products or products planned for future development, and the target patent information is used to describe the patents most associated with the technical term database, the target patent information can characterize the patents most associated with existing products or products planned for future development.
[0051] In some possible implementations, for the purpose of facilitating the identification of valid target patent information, please refer to [link / reference needed]. Figure 3 Step S230 includes, but is not limited to, the following steps:
[0052] In S231, the number of first identical words corresponding to each first disclosed patent is compared sequentially to determine the first candidate patent information.
[0053] Specifically, the first candidate patent information describes the first published patent with the highest number of identical first words in the first patent set information. The terminal device can sequentially compare the number of identical first words corresponding to each first published patent to determine the first candidate patent information. For example, when the number of identical first words corresponding to one first published patent is 15, another first published patent has 18, and yet another first published patent has 16, the terminal device can determine that the first published patent with 18 identical first words is the first candidate patent information.
[0054] In S232, the number of second identical words corresponding to each second disclosed patent is compared sequentially to determine the second candidate patent information.
[0055] Specifically, the second candidate patent information is used to describe the second published patent with the most identical second words in the second patent set information. The terminal device can sequentially compare the number of identical second words corresponding to each second published patent to determine the second candidate patent information. The specific determination process can be referred to in step S231 regarding similar content.
[0056] In S233, the number of third identical words in the first candidate patent information is compared with the number of fourth identical words in the second candidate patent information.
[0057] Specifically, the third identical word count information describes the number of words in the first candidate patent information that are identical to any one of the technical terms, and the fourth identical word count information describes the number of words in the second candidate patent information that are identical to any one of the technical terms. The terminal device can compare the third identical word count information of the first candidate patent information with the fourth identical word count information of the second candidate patent information.
[0058] In S234, if the number of third identical words is greater than the number of fourth identical words, then the first candidate patent information is determined as the target patent information.
[0059] Specifically, if the number of identical words in the third category is greater than the number of identical words in the fourth category, the terminal device can determine the first candidate patent information as the target patent information.
[0060] In S235, if the number of third identical words is less than the number of fourth identical words, then the second candidate patent information is determined as the target patent information.
[0061] Specifically, if the number of identical words in the third category is less than the number of identical words in the fourth category, the terminal device can determine the second alternative patent information as the target patent information.
[0062] In S236, if the number of third identical words is equal to the number of fourth identical words, then the first or second candidate patent information is determined as the target patent information.
[0063] Specifically, if the number of identical words in the third category is equal to the number of identical words in the fourth category, the terminal device can determine the first or second candidate patent information as the target patent information.
[0064] In S240, target patent sequence information is generated based on the target patent information.
[0065] Specifically, the target patent sequence information is used to describe multiple target patents arranged in chronological order according to their application years. For example, please refer to... Figure 4 The terminal device can generate target patent sequence information based on multiple target patent information. Figure 4 The terms “first target patent information,” “second target patent information,” and “third target patent information” all refer to target patent information.
[0066] In S250, the backbone information is generated based on the third identical word set information corresponding to each target patent information in the target patent sequence information.
[0067] Specifically, the terminal device can generate backbone information based on the third set of identical words corresponding to each target patent in the target patent sequence information. For example, please refer to [link to example]. Figure 5 When the third identical term set information corresponding to the first target patent information is "lightning characteristics, GIS disconnect switch, withstand voltage level, power grid lightning damage, lightning fault database, transmission line structure and power grid voltage level", the third identical term set information corresponding to the second target patent information is "power transformer, fiber optic grating generator, winding deformation meter, transformer winding temperature, winding internal coil temperature", and the third identical term set information corresponding to the third target patent information is "large-scale photovoltaic power station, photovoltaic cell, grid side voltage, circuit equivalent impedance, diode leakage resistance", the backbone information can be "lightning characteristics, GIS disconnect switch, withstand voltage level, power grid lightning damage, lightning fault database, transmission line structure and power grid voltage level", "power transformer, fiber optic grating generator, winding deformation meter, transformer winding temperature, winding internal coil temperature", and "large-scale photovoltaic power station, photovoltaic cell, grid side voltage, circuit equivalent impedance, diode leakage resistance".
[0068] In S260, branch information is generated based on the total remaining patent information.
[0069] Specifically, the total remaining patent information includes first remaining patent set information and second remaining patent set information. The first remaining patent set information describes the set of first remaining patents after removing the target patent information from the first patent set information, and the second remaining patent set information describes the set of second remaining patents after removing the target patent information from the second patent set information. The terminal device can generate branch information based on the total remaining patent information.
[0070] In some possible implementations, for efficient determination of valid branch information, please refer to [link / reference]. Figure 6 Step S260 includes, but is not limited to, the following steps:
[0071] In S261, the first remaining patent is bound to the target patent information of the same application year in the target patent sequence information, and the second remaining patent is bound to the target patent information of the same application year in the target patent sequence information.
[0072] Specifically, the branch information includes first branch information and second branch information. The terminal device can bind the first remaining patent with the target patent information of the same application year in the target patent sequence information, and bind the second remaining patent with the target patent information of the same application year in the target patent sequence information, thereby realizing the division and classification of different application years.
[0073] In S262, the number of fifth identical words corresponding to the first remaining patent is compared with the preset first quantity threshold information.
[0074] Specifically, the terminal device can compare the number of fifth identical words corresponding to the first remaining patent with the preset first quantity threshold information. The first quantity threshold information is used as a measure of the degree of association between the remaining patent and the technical term database, and as a benchmark for defining the degree of association between the patent and existing products or products planned to be developed in the future.
[0075] In S263, if the number of fifth identical words is greater than or equal to the first quantity threshold information, then the first branch information is generated based on the number of fifth identical words corresponding to the first remaining patent and the third identical word set information in the main information.
[0076] Specifically, the fifth identical word quantity information is used to describe the number of words in the first remaining patent that are identical to any one of the technical terms. If the fifth identical word quantity information is greater than or equal to the first quantity threshold information, the terminal device can generate the first branch information based on the fifth identical word quantity information corresponding to the first remaining patent and the third identical word set information in the main information.
[0077] For example, please refer to Figure 7 The terminal device generates first branch information based on the fifth identical word quantity information corresponding to the first remaining patent and the third identical word set information in the main information.
[0078] In S264, if the number of fifth identical words is less than the first quantity threshold, the number of fifth identical words corresponding to the first remaining patent is compared with the preset second quantity threshold.
[0079] Specifically, the second quantity threshold information is less than the first quantity threshold information; if the fifth identical word quantity information is less than the first quantity threshold information, the terminal device can compare the fifth identical word quantity information corresponding to the first remaining patent with the preset second quantity threshold information.
[0080] In S265, if the number of fifth identical words is greater than or equal to the second number threshold information, then the second branch information is generated based on the number of fifth identical words corresponding to the first remaining patent and the third identical word set information in the main information.
[0081] For example, please refer to Figure 7 If the number of fifth identical words is greater than or equal to the second number threshold information, the terminal device can generate second branch information based on the number of fifth identical words corresponding to the first remaining patent and the third identical word set information in the main information.
[0082] In S266, the number of sixth identical words corresponding to the second remaining patent is compared with the first number threshold information.
[0083] Specifically, the terminal device can compare the number of sixth identical words corresponding to the second remaining patent with the first number threshold information.
[0084] In S267, if the number of sixth identical words is greater than or equal to the first quantity threshold information, then the first branch information is generated based on the number of sixth identical words corresponding to the second remaining patent and the third identical word set information in the main information.
[0085] Specifically, if the number of sixth identical words is greater than or equal to the first quantity threshold information, the terminal device can generate the first branch information based on the number of sixth identical words corresponding to the second remaining patent and the third identical word set information in the main information.
[0086] In S268, if the number of sixth identical words is less than the first number threshold, the number of sixth identical words and the second number threshold corresponding to the second remaining patent are compared.
[0087] Specifically, if the number of sixth identical words is less than the first quantity threshold, the terminal device can compare the number of sixth identical words corresponding to the second remaining patent with the second quantity threshold.
[0088] In S269, if the number of sixth identical words is greater than or equal to the second number threshold information, then the second branch information is generated based on the number of sixth identical words corresponding to the second remaining patent and the third identical word set information in the main information.
[0089] Specifically, if the number of sixth identical words is greater than or equal to the second quantity threshold information, the terminal device can generate second branch information based on the number of sixth identical words corresponding to the second remaining patent and the third identical word set information in the main information.
[0090] In S300, infringement risk warning information is generated based on the technology development tree information.
[0091] Specifically, terminal devices can generate infringement risk warning information based on technology development tree information. This information describes potential infringement risks, enabling rapid matching of different patents related to the same technology based on the technology development tree information. This allows for quick identification of potential infringement risks and significantly improves the efficiency of patent infringement risk identification. It should be noted that the more patents other companies have around a particular technology, the higher the potential infringement risk will be.
[0092] In some possible implementations, to further improve the efficiency of patent infringement risk screening and to intuitively understand the severity of infringement risks, please refer to [link to relevant documentation]. Figure 8 Step S300 includes, but is not limited to, the following steps:
[0093] In S310, the target company's first branch quantity information, the target company's second branch quantity information, the designated company's first branch quantity information, and the designated company's second branch quantity information are obtained from the technology development tree information.
[0094] Without loss of generality, infringement risk warning information includes high-risk information and general-risk information. High-risk information describes serious infringement risks, while general-risk information describes ordinary infringement risks. The target company's first branch quantity information describes the number of the target company's first branches, i.e., the number of first branch information generated based on the first remaining patent. The target company's second branch quantity information describes the number of the target company's second branches, i.e., the number of second branch information generated based on the first remaining patent. The designated company's first branch quantity information describes the number of the designated company's first branches, i.e., the number of first branch information generated based on the second remaining patent. The designated company's second branch quantity information describes the number of the designated company's second branches, i.e., the number of second branch information generated based on the second remaining patent.
[0095] Specifically, the terminal device can obtain information on the number of first branches of the target company, the number of second branches of the target company, the number of first branches of the specified company, and the number of second branches of the specified company in the technology development tree information.
[0096] In S320, the first branch difference information is generated by subtracting the first branch quantity information of the specified enterprise from the first branch quantity information of the target enterprise.
[0097] Specifically, the terminal device can generate the first branch difference information by subtracting the first branch number information of the specified enterprise from the first branch number information of the target enterprise.
[0098] In S330, the first branch difference information is compared with the preset first branch number difference threshold information.
[0099] Specifically, the terminal device can compare the first branch difference information with the preset first branch number difference threshold information.
[0100] In S340, if the first branch difference information is greater than or equal to the first branch number difference threshold information, then high-risk information is generated.
[0101] Specifically, if the difference information of the first branch is greater than or equal to the threshold information of the difference in the number of first branches, then high-risk information is generated.
[0102] In S350, the difference between the number of second branches of the target enterprise and the number of second branches of the specified enterprise is used to generate the second branch difference information.
[0103] Specifically, the terminal device can generate second branch difference information by subtracting the second branch number information of the specified enterprise from the second branch number information of the target enterprise.
[0104] In S360, the difference information of the second branch is compared with the preset difference threshold information of the number of second branches.
[0105] Specifically, the terminal device can compare the second branch difference information with the preset second branch number difference threshold information. It should be noted that the specific value of the second branch number difference threshold information is not related to the first branch number difference threshold information. That is, the second branch number difference threshold information can be greater than, equal to or less than the first branch number difference threshold information.
[0106] In S370, if the second branch difference information is greater than or equal to the second branch number difference threshold information, then general risk information is generated.
[0107] Specifically, if the difference information of the second branch is greater than or equal to the difference threshold information of the number of second branches, the terminal device can generate general risk information.
[0108] The implementation principle of the patent data management method based on a large-scale natural language model in this application embodiment is as follows: The terminal device can first obtain the first patent set information of the target enterprise and the second patent set information of at least one designated enterprise based on a preset patent database. Then, based on the first patent set information and the second patent set information, technology development tree information is generated. Based on the technology development tree information, infringement risk warning information is generated. This enables the use of technology development tree information to quickly search for patents of other enterprises, find patents of other enterprises with potential patent infringement risks, quickly and accurately identify patent infringement risks, and improve the efficiency of patent infringement risk identification.
[0109] It should be noted that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0110] Embodiments of this application also provide a patent data management system based on a large-scale natural language model. For ease of explanation, only the parts relevant to this application are shown, such as... Figure 9 As shown, the system 90 includes:
[0111] First patent set information acquisition module 91: is used to acquire first patent set information of a target enterprise and second patent set information of at least one designated enterprise based on a preset patent database. The patent database stores first patent set information and second patent set information. The first patent set information is used to describe the set of first published patents of the target enterprise, and the second patent set information is used to describe the set of second published patents of at least one designated enterprise.
[0112] Technology development tree information generation module 92: used to generate technology development tree information based on the information of the first patent set and the second patent set;
[0113] Infringement Risk Warning Information Generation Module 93: Used to generate infringement risk warning information based on technology development tree information.
[0114] Optionally, the technology development tree information is divided into trunk information and multiple branch information, all of which are related to the trunk information; the aforementioned technology development tree information generation module 92 includes:
[0115] The first identical word set information acquisition submodule is used to acquire the first identical word set information corresponding to each first published patent and the second identical word set information corresponding to each second patent set information based on a preset technical word database and a preset word matching algorithm. The technical word database stores multiple technical word information. The first identical word set information is used to describe the set of first identical word information and the first identical word information is used to describe the words in the first published patent that are the same as any one of the technical word information. The second identical word set information is used to describe the set of second identical word information and the second identical word information is used to describe the words in the second published patent that are the same as any one of the technical word information.
[0116] The first identical word quantity information determination submodule is used to determine the first identical word quantity information corresponding to each first published patent based on the first identical word set information corresponding to each first published patent, and to determine the second identical word quantity information corresponding to each second published patent based on the second identical word set information corresponding to each second patent set information. The first identical word quantity information is used to describe the quantity of first identical word information in the first published patent, and the second identical word quantity information is used to describe the quantity of second identical word information in the second published patent.
[0117] Target Patent Information Determination Submodule: Used for each application year to determine target patent information based on the first and second identical word count information;
[0118] The target patent sequence information generation submodule is used to generate target patent sequence information based on target patent information. The target patent sequence information describes multiple target patent information arranged in chronological order according to the application year.
[0119] The main information generation submodule is used to generate main information based on the third identical word set information corresponding to each target patent information in the target patent sequence information;
[0120] Branch information generation submodule: used to generate branch information based on the total remaining patent information, wherein the total remaining patent information includes the first remaining patent set information and the second remaining patent set information. The first remaining patent set information is used to describe the set of first remaining patents after removing the target patent information from the first patent set information, and the second remaining patent set information is used to describe the set of second remaining patents after removing the target patent information from the second patent set information.
[0121] Optionally, the above-mentioned target patent information determination submodule includes:
[0122] First candidate patent information determination unit: used to sequentially compare the number of first identical words corresponding to each first published patent to determine the first candidate patent information, wherein the first candidate patent information is used to describe the first published patent with the most first identical words in the first patent set information;
[0123] Second candidate patent information determination unit: used to sequentially compare the number of second identical words corresponding to each second published patent to determine the second candidate patent information, wherein the second candidate patent information is used to describe the second published patent with the most second identical words in the second patent set information;
[0124] Third identical word quantity information comparison unit: used to compare the third identical word quantity information of the first candidate patent information with the fourth identical word quantity information of the second candidate patent information;
[0125] The first unit for determining target patent information is used to determine the first candidate patent information as the target patent information if the number of third identical words is greater than the number of fourth identical words.
[0126] The second unit for determining target patent information is used to determine the second candidate patent information as the target patent information if the number of third identical words is less than the number of fourth identical words.
[0127] The third unit for determining target patent information is used to determine the first or second candidate patent information as the target patent information if the third number of identical words is equal to the fourth number of identical words.
[0128] Optionally, the branch information includes first branch information and second branch information; the above branch information generation submodule includes:
[0129] Target patent information binding unit: used to bind the first remaining patent with the target patent information of the same application year in the target patent sequence information, and to bind the second remaining patent with the target patent information of the same application year in the target patent sequence information;
[0130] Fifth identical word quantity information comparison unit: used to compare the fifth identical word quantity information corresponding to the first remaining patent with the preset first quantity threshold information;
[0131] First branch information first generation unit: used to generate first branch information based on the fifth identical word quantity information corresponding to the first remaining patent and the third identical word set information in the main information if the fifth identical word quantity information is greater than or equal to the first quantity threshold information;
[0132] The second comparison unit for the fifth identical word quantity information is used to compare the fifth identical word quantity information corresponding to the first remaining patent with the preset second quantity threshold information if the fifth identical word quantity information is less than the first quantity threshold information, wherein the second quantity threshold information is less than the first quantity threshold information;
[0133] The first generation unit of the second branch is used to generate second branch information based on the number of fifth identical words corresponding to the first remaining patent and the third identical word set information in the main information if the number of fifth identical words is greater than or equal to the second quantity threshold information.
[0134] The first comparison unit for the sixth identical word quantity information is used to compare the sixth identical word quantity information and the first quantity threshold information corresponding to the second remaining patent.
[0135] The first branch information second generation unit is used to generate first branch information based on the sixth identical word quantity information corresponding to the second remaining patent and the third identical word set information in the main information if the sixth identical word quantity information is greater than or equal to the first quantity threshold information.
[0136] The second comparison unit for the sixth identical word quantity information is used to compare the sixth identical word quantity information and the second quantity threshold information corresponding to the second remaining patent if the sixth identical word quantity information is less than the first quantity threshold information.
[0137] The second branch second generation unit is used to generate second branch information based on the sixth identical word quantity information corresponding to the second remaining patent and the third identical word set information in the main information if the sixth identical word quantity information is greater than or equal to the second quantity threshold information.
[0138] Optionally, the infringement risk warning information includes high-risk information and general-risk information; the aforementioned infringement risk warning information generation module 93 includes:
[0139] The target company's first branch quantity information acquisition submodule is used to acquire the target company's first branch quantity information, target company's second branch quantity information, and the specified company's first branch quantity information and second branch quantity information for the technology development tree.
[0140] First Branch Difference Information Generation Submodule: Used to generate first branch difference information based on the difference between the first branch quantity information of the target enterprise and the first branch quantity information of the specified enterprise.
[0141] First branch difference information comparison submodule: used to compare the first branch difference information with the preset first branch number difference threshold information;
[0142] High-risk information generation submodule: Used to generate high-risk information if the first branch difference information is greater than or equal to the first branch number difference threshold information;
[0143] The second branch difference information generation submodule is used to generate second branch difference information by subtracting the second branch quantity information of the specified enterprise from the second branch quantity information of the target enterprise.
[0144] Second branch difference information comparison submodule: used to compare the second branch difference information with the preset second branch number difference threshold information;
[0145] The general risk information generation submodule is used to generate general risk information if the difference information of the second branch is greater than or equal to the difference threshold information of the number of second branches.
[0146] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0147] This application also provides a terminal device, such as... Figure 10 As shown, the terminal device 100 of this embodiment includes: a processor 101, a memory 102, and a computer program 103 stored in the memory 102 and executable on the processor 101. When the processor 101 executes the computer program 103, it implements the steps described in the above-described traffic processing method embodiment, for example... Figure 1 The steps S100 to S300 are shown; or, when the processor 101 executes the computer program 103, it implements the functions of each module in the above-described device, for example... Figure 9 The functions of modules 91 to 93 are shown.
[0148] The terminal device 100 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. The terminal device 100 includes, but is not limited to, a processor 101 and a memory 102. Those skilled in the art will understand that... Figure 10This is merely an example of terminal device 100 and does not constitute a limitation on terminal device 100. It may include more or fewer components than shown, or combine certain components, or different components. For example, terminal device 100 may also include input / output devices, network access devices, buses, etc.
[0149] The processor 101 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.; the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0150] The memory 102 can be an internal storage unit of the terminal device 100, such as the hard disk or memory of the terminal device 100. The memory 102 can also be an external storage device of the terminal device 100, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal device 100. Furthermore, the memory 102 can include both internal storage units and external storage devices of the terminal device 100. The memory 102 can also store computer program 103 and other programs and data required by the terminal device 100. The memory 102 can also be used to temporarily store data that has been output or will be output.
[0151] One embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0152] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the methods, principles and structures of this application should be covered within the scope of protection of this application.
Claims
1. A patent data management method based on a large-scale natural language model, characterized in that, The method includes: Based on a preset patent database, information on a first patent set of a target company and information on a second patent set of at least one designated company are obtained. The patent database stores information on the first patent set and information on the second patent set. The first patent set information is used to describe the set of first published patents of the target company, and the second patent set information is used to describe the set of second published patents of the at least one designated company. Based on the first patent set information and the second patent set information, a technology development tree information is generated. The technology development tree information is used to describe the relationship and development trend between the technologies represented by different patents. The technology development tree information is divided into main information and multiple branch information, and the multiple branch information are all related to the main information. Based on the aforementioned technology development tree information, generate infringement risk warning information; The step of generating technology development tree information based on the first patent set information and the second patent set information includes: Based on a preset technical term database and a preset word matching algorithm, first identical word set information corresponding to each first disclosed patent and second identical word set information corresponding to each second patent set information are obtained. The technical term database stores multiple technical term information. The first identical word set information is used to describe the set of first identical word information. The first identical word information is used to describe words in the first disclosed patent that are identical to any one of the technical term information. The second identical word set information is used to describe the set of second identical word information. The second identical word information is used to describe words in the second disclosed patent that are identical to any one of the technical term information. Based on the first identical word set information corresponding to each of the first disclosed patents, the number of first identical words corresponding to each of the first disclosed patents is determined, and based on the second identical word set information corresponding to each of the second patent sets, the number of second identical words corresponding to each of the second disclosed patents is determined, wherein the number of first identical words is used to describe the number of first identical words in the first disclosed patents, and the number of second identical words is used to describe the number of second identical words in the second disclosed patents. For each application year: Target patent information is determined based on the first identical word count information and the second identical word count information; Based on the target patent information, target patent sequence information is generated, wherein the target patent sequence information is used to describe multiple target patent information arranged in chronological order according to the application year; Generate backbone information based on the third identical word set information corresponding to each of the target patent information in the target patent sequence information; The branch information is generated based on the total remaining patent information, wherein the total remaining patent information includes a first remaining patent set information and a second remaining patent set information. The first remaining patent set information is used to describe the set of first remaining patents after removing the target patent information from the first patent set information, and the second remaining patent set information is used to describe the set of second remaining patents after removing the target patent information from the second patent set information.
2. The method according to claim 1, characterized in that, The step of determining the target patent information based on the first identical word count information and the second identical word count information includes: The number of identical words corresponding to each of the first disclosed patents is compared sequentially to determine the first candidate patent information, wherein the first candidate patent information is used to describe the first disclosed patent with the most identical words in the first patent set information. By sequentially comparing the number of identical words corresponding to each of the second disclosed patents, second candidate patent information is determined, wherein the second candidate patent information is used to describe the second disclosed patent with the most identical words in the second patent set information; Compare the number of third identical words in the first candidate patent information with the number of fourth identical words in the second candidate patent information; If the number of identical words in the third category is greater than the number of identical words in the fourth category, then the first candidate patent information is determined to be the target patent information. If the number of identical words in the third category is less than the number of identical words in the fourth category, then the second candidate patent information is determined to be the target patent information. If the number of identical words in the third category is equal to the number of identical words in the fourth category, then the first candidate patent information or the second candidate patent information is determined as the target patent information.
3. The method according to claim 1, characterized in that, The branch information includes first branch information and second branch information; generating the branch information based on the total remaining patent information includes: The first remaining patent is bound to the target patent information of the same application year in the target patent sequence information, and the second remaining patent is bound to the target patent information of the same application year in the target patent sequence information; Compare the number of fifth identical words corresponding to the first remaining patent with the preset first number threshold information; If the fifth identical word quantity information is greater than or equal to the first quantity threshold information, then the first branch information is generated based on the fifth identical word quantity information corresponding to the first remaining patent and the third identical word set information in the trunk information; If the number of fifth identical words is less than the first number threshold information, then compare the number of fifth identical words corresponding to the first remaining patent with the preset second number threshold information, wherein the second number threshold information is less than the first number threshold information; If the fifth identical word quantity information is greater than or equal to the second quantity threshold information, then the second branch information is generated based on the fifth identical word quantity information corresponding to the first remaining patent and the third identical word set information in the trunk information; Compare the number of sixth identical words corresponding to the second remaining patent with the first number threshold information; If the sixth identical word quantity information is greater than or equal to the first quantity threshold information, then the first branch information is generated based on the sixth identical word quantity information corresponding to the second remaining patent and the third identical word set information in the trunk information; If the number of sixth identical words is less than the first number threshold, then compare the number of sixth identical words corresponding to the second remaining patent with the second number threshold. If the sixth identical word quantity information is greater than or equal to the second quantity threshold information, then the second branch information is generated based on the sixth identical word quantity information corresponding to the second remaining patent and the third identical word set information in the trunk information.
4. The method according to claim 3, characterized in that, The infringement risk warning information includes high-risk information and general risk information; The step of generating infringement risk warning information based on the technology development tree information includes: The technology development tree information includes the number of first branches of the target enterprise, the number of second branches of the target enterprise, the number of first branches of the specified enterprise, and the number of second branches of the specified enterprise. The first branch difference information is generated by subtracting the first branch number information of the specified enterprise from the first branch number information of the target enterprise. Compare the first branch difference information with the preset first branch number difference threshold information; If the first branch difference information is greater than or equal to the first branch number difference threshold information, then high-risk information is generated; The second branch difference information is generated by subtracting the second branch number information of the specified enterprise from the second branch number information of the target enterprise. Compare the second branch difference information with the preset second branch number difference threshold information; If the second branch difference information is greater than or equal to the second branch number difference threshold information, then general risk information is generated.
5. A patent data management system based on a large-scale natural language model, characterized in that, The system includes: First patent set information acquisition module: used to acquire first patent set information of target enterprise and second patent set information of at least one designated enterprise based on a preset patent database, wherein the patent database stores the first patent set information and the second patent set information, the first patent set information is used to describe the set of first published patents of the target enterprise, and the second patent set information is used to describe the set of second published patents of the at least one designated enterprise; Technology development tree information generation module: used to generate technology development tree information based on the first patent set information and the second patent set information, wherein the technology development tree information is used to describe the relationship and development trend between the technologies represented by different patents, and the technology development tree information is divided into main trunk information and multiple branch information, and the multiple branch information are all related to the main trunk information; Infringement risk warning information generation module: used to generate infringement risk warning information based on the technology development tree information; The technology development tree information generation module includes: The first identical word set information acquisition submodule is used to acquire, based on a preset technical term database and a preset word matching algorithm, the first identical word set information corresponding to each of the first disclosed patents and the second identical word set information corresponding to each of the second patent sets. The technical term database stores multiple technical term information; the first identical word set information describes the set of first identical word information; the first identical word information describes words in the first disclosed patent that are identical to any one of the technical term information; the second identical word set information describes the set of second identical word information; the second identical word information describes words in the second disclosed patent that are identical to any one of the technical term information. The first identical word quantity information determination submodule is used to determine the first identical word quantity information corresponding to each of the first disclosed patents based on the first identical word set information corresponding to each of the first disclosed patents, and to determine the second identical word quantity information corresponding to each of the second disclosed patents based on the second identical word set information corresponding to each of the second patent set information. The first identical word quantity information is used to describe the quantity of the first identical word information in the first disclosed patent, and the second identical word quantity information is used to describe the quantity of the second identical word information in the second disclosed patent. Target patent information determination submodule: for each application year, it determines the target patent information based on the first identical word count information and the second identical word count information; Target patent sequence information generation submodule: used to generate target patent sequence information based on the target patent information, wherein the target patent sequence information is used to describe multiple target patent information arranged in chronological order according to the application year; The backbone information generation submodule is used to generate backbone information based on the third identical word set information corresponding to each of the target patent information in the target patent sequence information; Branch information generation submodule: used to generate the branch information based on the total remaining patent information, wherein the total remaining patent information includes a first remaining patent set information and a second remaining patent set information, the first remaining patent set information is used to describe the set of first remaining patents after removing the target patent information from the first patent set information, and the second remaining patent set information is used to describe the set of second remaining patents after removing the target patent information from the second patent set information.
6. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.