Patent classification method and system based on pre-training model bert

The patent documents are divided into multiple technical solution entity pairs by using the pre-trained BERT model, and the keyword score of each technical solution entity pair is calculated by the keyword score calculation model. The intelligent level of classification based on the common technical field is used to accurately classify the patent documents.

CN121233765APending Publication Date: 2025-12-30SHUHAO INFORMATION TECHNOLOGY (WUHAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410165300.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-05
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

There is no intelligent method in the existing technology that can accurately classify patent documents according to the technical field.

Method used

The pre-trained BERT model is used to divide patent documents into multiple technical solution entity pairs. The keyword score of each word is calculated by the keyword score calculation model. The similarity comparison is performed using technical classification words in IPC to determine the technical classification of the patent documents.

Benefits of technology

It achieves accurate classification of patent documents, improves the intelligence and automation of classification, and can accurately classify patent documents into the corresponding patent documents according to the technical field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121233765A_ABST
    Figure CN121233765A_ABST
Patent Text Reader

Abstract

The invention discloses a pre-training model bert-based patent classification method and system. The method comprises the steps of dividing patent documents into a plurality of technical scheme entity pairs through a pre-training model bert; and setting a keyword score calculation model, calculating the keyword score of each vocabulary in each technical scheme entity pair, performing similarity comparison on the vocabulary with the highest keyword score and the technical classification vocabulary in the IPC, and taking the technical classification vocabulary with the highest similarity in the IPC as the technical classification of the patent literature.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of patent classification, and more particularly relates to a patent classification method and system based on a pre-trained model bert. BACKGROUND

[0002] Information classification technology refers to the use of algorithms and technical means to effectively organize and classify large amounts of information in order to better understand and utilize these information. In the context of today's technological development, information classification technology has undergone a series of progress and innovation, some of the main trends and technologies include:

[0003] Machine learning and deep learning: Machine learning and deep learning technologies play a key role in information classification. Through training models, systems can learn to extract features from data and make decisions, making classification more accurate and automated. Deep learning models such as convolutional neural networks (CNN) and recurrent neural networks (RNN) have achieved remarkable results in the fields of text, image and speech.

[0004] Data mining and pattern recognition: Data mining techniques are used to discover hidden patterns and associations in large-scale data sets, helping to gain a deeper understanding and classification of information.

[0005] However, there is no technical solution in the prior art that can intelligently classify patent documents according to technical fields. SUMMARY

[0006] To solve the above technical problems, the present application provides a patent classification method based on a pre-trained model bert, comprising:

[0007] Divide the patent documents into a plurality of technical solution entity pairs by the pre-trained model bert;

[0008] Set a keyword score calculation model to calculate the keyword score of each word in each technical solution entity pair, compare the keyword score of the highest keyword score with the technical classification words in IPC, and select the technical classification words in IPC with the highest similarity as the technical classification of the patent document.

[0009] Further, the keyword score calculation model comprises:

[0010]

[0011] Wherein, S ik is the keyword score of the i-th word in the k-th technical solution entity pair, f ik is the frequency of the i-th word appearing in the k-th technical solution entity pair, n kH is the number of words in the kth technical solution entity pair, ik H is the information entropy of the ith word in the kth technical solution entity pair, max L is the maximum information entropy of the ith word in all technical solution entity pairs, ik L is the length of the ith word in the kth technical solution entity pair, avg β is the average length of all words in all technical solution entity pairs, A is the first adjustment factor, ik A is the value of the term frequency heterogeneity of the ith word in the kth technical solution entity pair, max γ is the maximum value of the term frequency heterogeneity of the ith word in all technical solution entity pairs, and γ is the second adjustment factor.

[0012] Further, the value A of the term frequency heterogeneity of the ith word in the kth technical solution entity pair ik comprises:

[0013]

[0014] P is the number of technical solution entity pairs in which the ith word appears, P ik C is the probability that the ith word appears in the kth technical solution entity pair, ik C is the context information value of the ith word in the kth technical solution entity pair, max C is the maximum context information value of the ith word in all technical solution entity pairs.

[0015] Further, the context information value C of the ith word in the kth technical solution entity pair is obtained by constructing a co-occurrence matrix ik and the maximum context information value C of the ith word in all technical solution entity pairs max .

[0016] Further, before comparing the highest-scoring word with the technical classification words in the IPC, the method further comprises: establishing a technical noun dictionary, filtering the highest-scoring word according to the technical noun dictionary, deleting the words in the non-technical noun dictionary, and comparing the filtered words with the technical classification words in the IPC.

[0017] The application also provides a patent classification system based on a pre-trained model bert, comprising:

[0018] The entity pair division module is used to divide the patent literature into multiple technical solution entity pairs by the pre-trained model bert;

[0019] The classification module is used to set up a keyword score calculation model, calculate the keyword score of each word in each entity pair of the technical solution, compare the word with the highest keyword score with the technical classification words in the IPC, and take the technical classification words in the IPC with the highest similarity as the technical classification of the patent document.

[0020] Furthermore, the keyword score calculation model includes:

[0021]

[0022] Among them, S ik To calculate the keyword score for the i-th word in the k-th technical solution entity pair, f ik Let n be the frequency of the i-th word in the k-th technical solution entity pair. k H represents the number of words in the entity pair of the k-th technical solution. ik Let H be the information entropy of the i-th word in the k-th technical solution entity pair. max To maximize the information entropy of the i-th word in all entity pairs of technical solutions, L ik Let L be the length of the i-th word in the k-th technical solution entity pair. avg The average length of all words in all entity pairs of all technical solutions, β is the first adjustment factor, and A ik Let A be the value of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair. max γ is the maximum value of the word frequency heterogeneity of the i-th word in all technical solution entity pairs, and γ is the second adjustment factor.

[0023] Furthermore, the value A of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair ik include:

[0024]

[0025] Where m is the number of technical solution entity pairs in which the i-th word appears in different technical solution entity pairs, and P ik Let C be the probability that the i-th word appears in the k-th technical solution entity pair. ik For the context information value of the i-th word in the k-th technical solution entity pair, C max This represents the maximum value of the contextual information for the i-th word in all entity pairs of technical solutions.

[0026] Furthermore, the contextual information value C of the i-th word in the k-th technical solution entity pair is obtained by constructing a co-occurrence matrix. ik And the maximum value of contextual information C for the i-th word in all technical solution entity pairs. max .

[0027] Furthermore, before comparing the word with the highest keyword score with the technical category words in IPC, the process includes: establishing a dictionary of technical terms, filtering the word with the highest keyword score based on the dictionary of technical terms, deleting words from the dictionary of non-technical terms, and comparing the filtered words with the technical category words in IPC.

[0028] Compared with the prior art, the above-described technical solutions conceived in this invention have the following beneficial effects:

[0029] This invention uses a pre-trained BERT model to divide patent documents into multiple technical solution entity pairs. A keyword score calculation model is set up to calculate the keyword score for each word in each technical solution entity pair. The word with the highest keyword score is compared with the technical category words in the IPC (Integrated Patent Classification), and the technical category word in the IPC with the highest similarity is taken as the technical category of the patent document. Through the above technical solution, this invention can accurately map patent documents to IPC classifications, thereby completing the classification of patent documents. Attached Figure Description

[0030] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention;

[0031] Figure 2 This is a system structure diagram of Embodiment 2 of the present invention. Detailed Implementation

[0032] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0033] The method provided by this invention can be implemented in a terminal environment that may include one or more of the following components: a processor, a storage medium, and a display screen. The storage medium stores at least one instruction, which is loaded and executed by the processor to implement the method described in the following embodiments.

[0034] A processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts of the terminal, and performs various functions and processes data by running or executing instructions, programs, code sets or instruction sets stored in the storage medium, and by calling data stored in the storage medium.

[0035] Storage media can include random access memory (RAM) or read-only memory (ROM). Storage media can be used to store instructions, programs, code, code sets, or instructions.

[0036] The display screen is used to show the user interface of each application.

[0037] In the formula of this invention, all subscripts are only used to distinguish parameters and have no actual meaning.

[0038] In addition, those skilled in the art will understand that the structure of the terminal described above does not constitute a limitation on the terminal. The terminal may include more or fewer components, or combine certain components, or have different component arrangements. For example, the terminal may also include radio frequency circuits, input units, sensors, audio circuits, power supplies, and other components, which will not be described in detail here.

[0039] Example 1

[0040] like Figure 1 As shown, this embodiment of the invention provides a patent classification method based on the pre-trained model BERT, including:

[0041] Step 101: The patent document is divided into multiple technical solution entity pairs by using the pre-trained model BERT (by using the authorized patent 202310594616.5 - a method, device, equipment and medium for extracting patent text entities, the patent document is divided into multiple technical solution entity pairs).

[0042] Step 102: Set up a keyword score calculation model, calculate the keyword score of each word in each entity pair of the technical solution, compare the similarity of the word with the technical classification words in the IPC, and take the technical classification words in the IPC with the highest similarity as the technical classification of the patent document.

[0043] Specifically, before comparing the word with the highest keyword score with the technical category words in IPC, the process includes: establishing a dictionary of technical terms, filtering the word with the highest keyword score based on the dictionary of technical terms, deleting words from the dictionary of non-technical terms, and comparing the filtered words with the technical category words in IPC.

[0044] Specifically, the keyword score calculation model includes:

[0045]

[0046] Among them, S ik To calculate the keyword score for the i-th word in the k-th technical solution entity pair, f ik Let n be the frequency of the i-th word in the k-th technical solution entity pair. k H represents the number of words in the entity pair of the k-th technical solution. ik Let H be the information entropy of the i-th word in the k-th technical solution entity pair.max To maximize the information entropy of the i-th word in all entity pairs of technical solutions, L ik Let L be the length of the i-th word in the k-th technical solution entity pair. avg The average length of all words in all entity pairs of all technical solutions, β is the first adjustment factor, and A ik Let A be the value of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair. max γ is the maximum value of the word frequency heterogeneity of the i-th word in all technical solution entity pairs, and γ is the second adjustment factor.

[0047] Specifically, the value A of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair. ik include:

[0048]

[0049] Where m is the number of technical solution entity pairs in which the i-th word appears in different technical solution entity pairs, and P ik Let C be the probability that the i-th word appears in the k-th technical solution entity pair. ik For the context information value of the i-th word in the k-th technical solution entity pair, C max This represents the maximum value of the contextual information for the i-th word in all entity pairs of technical solutions.

[0050] Specifically, the contextual information value C of the i-th word in the k-th technical solution entity pair is obtained by constructing a co-occurrence matrix. ik And the maximum value of contextual information C for the i-th word in all technical solution entity pairs. max .

[0051] Example 2

[0052] like Figure 2 As shown, this embodiment of the invention also proposes a patent classification system based on the pre-trained model BERT, comprising:

[0053] The entity pair segmentation module is used to segment patent documents into multiple technical solution entity pairs using the pre-trained BERT model.

[0054] The classification module is used to set up a keyword score calculation model, calculate the keyword score of each word in each entity pair of the technical solution, compare the word with the highest keyword score with the technical classification words in the IPC, and take the technical classification words in the IPC with the highest similarity as the technical classification of the patent document.

[0055] Specifically, before comparing the word with the highest keyword score with the technical category words in IPC, the process includes: establishing a dictionary of technical terms, filtering the word with the highest keyword score based on the dictionary of technical terms, deleting words from the dictionary of non-technical terms, and comparing the filtered words with the technical category words in IPC.

[0056] Specifically, the keyword score calculation model includes:

[0057]

[0058] Among them, S ik To calculate the keyword score for the i-th word in the k-th technical solution entity pair, f ik Let n be the frequency of the i-th word in the k-th technical solution entity pair. k H represents the number of words in the entity pair of the k-th technical solution. ik Let H be the information entropy of the i-th word in the k-th technical solution entity pair. max To maximize the information entropy of the i-th word in all entity pairs of technical solutions, L ik Let L be the length of the i-th word in the k-th technical solution entity pair. avg The average length of all words in all entity pairs of all technical solutions, β is the first adjustment factor, and A ik Let A be the value of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair. max γ is the maximum value of the word frequency heterogeneity of the i-th word in all technical solution entity pairs, and γ is the second adjustment factor.

[0059] Specifically, the value A of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair. ik include:

[0060]

[0061] Where m is the number of technical solution entity pairs in which the i-th word appears in different technical solution entity pairs, and P ik Let C be the probability that the i-th word appears in the k-th technical solution entity pair. ik For the context information value of the i-th word in the k-th technical solution entity pair, C max This represents the maximum value of the contextual information for the i-th word in all entity pairs of technical solutions.

[0062] Specifically, the contextual information value C of the i-th word in the k-th technical solution entity pair is obtained by constructing a co-occurrence matrix. ik And the maximum value of contextual information C for the i-th word in all technical solution entity pairs. max .

[0063] Example 3

[0064] This invention also proposes a storage medium storing multiple instructions for implementing the patent classification method based on the pre-trained BERT model.

[0065] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0066] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: Step 101, dividing the patent document into multiple technical solution entity pairs using a pre-trained model BERT;

[0067] Step 102: Set up a keyword score calculation model, calculate the keyword score of each word in each entity pair of the technical solution, compare the similarity of the word with the technical classification words in the IPC, and take the technical classification words in the IPC with the highest similarity as the technical classification of the patent document.

[0068] Specifically, before comparing the word with the highest keyword score with the technical category words in IPC, the process includes: establishing a dictionary of technical terms, filtering the word with the highest keyword score based on the dictionary of technical terms, deleting words from the dictionary of non-technical terms, and comparing the filtered words with the technical category words in IPC.

[0069] Specifically, the keyword score calculation model includes:

[0070]

[0071] Among them, S ik To calculate the keyword score for the i-th word in the k-th technical solution entity pair, f ik Let n be the frequency of the i-th word in the k-th technical solution entity pair. k H represents the number of words in the entity pair of the k-th technical solution. ik Let H be the information entropy of the i-th word in the k-th technical solution entity pair. max To maximize the information entropy of the i-th word in all entity pairs of technical solutions, L ik Let L be the length of the i-th word in the k-th technical solution entity pair. avg The average length of all words in all entity pairs of all technical solutions, β is the first adjustment factor, and A ik Let A be the value of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair.max γ is the maximum value of the word frequency heterogeneity of the i-th word in all technical solution entity pairs, and γ is the second adjustment factor.

[0072] Specifically, the value A of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair. ik include:

[0073]

[0074] Where m is the number of technical solution entity pairs in which the i-th word appears in different technical solution entity pairs, and P ik Let C be the probability that the i-th word appears in the k-th technical solution entity pair. ik For the context information value of the i-th word in the k-th technical solution entity pair, C max This represents the maximum value of the contextual information for the i-th word in all entity pairs of technical solutions.

[0075] Specifically, the contextual information value C of the i-th word in the k-th technical solution entity pair is obtained by constructing a co-occurrence matrix. ik And the maximum value of contextual information C for the i-th word in all technical solution entity pairs. max .

[0076] Example 4

[0077] This invention also proposes an electronic device, including a processor and a storage medium connected to the processor. The storage medium stores multiple instructions, which can be loaded and executed by the processor to enable the processor to execute a patent classification method based on a pre-trained model BERT.

[0078] Specifically, the electronic device in this embodiment can be a computer terminal, which may include one or more processors and a storage medium.

[0079] The storage medium can be used to store software programs and modules, such as the patent classification method based on the pre-trained BERT model in this embodiment of the invention. The corresponding program instructions / modules are executed by the processor through running the software programs and modules stored in the storage medium, thereby performing various functional applications and data processing, thus realizing the aforementioned patent classification method based on the pre-trained BERT model. The storage medium may include high-speed random access storage media, and may also include non-volatile storage media, such as one or more magnetic storage systems, flash memory, or other non-volatile solid-state storage media. In some instances, the storage medium may further include storage media remotely configured relative to the processor, which can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0080] The processor can call the information and application stored in the storage medium through the transmission system to execute the following steps: Step 101, dividing the patent document into multiple technical solution entity pairs through the pre-trained model BERT;

[0081] Step 102: Set up a keyword score calculation model, calculate the keyword score of each word in each entity pair of the technical solution, compare the similarity of the word with the technical classification words in the IPC, and take the technical classification words in the IPC with the highest similarity as the technical classification of the patent document.

[0082] Specifically, before comparing the word with the highest keyword score with the technical category words in IPC, the process includes: establishing a dictionary of technical terms, filtering the word with the highest keyword score based on the dictionary of technical terms, deleting words from the dictionary of non-technical terms, and comparing the filtered words with the technical category words in IPC.

[0083] Specifically, the keyword score calculation model includes:

[0084]

[0085] Among them, S ik To calculate the keyword score for the i-th word in the k-th technical solution entity pair, f ik Let n be the frequency of the i-th word in the k-th technical solution entity pair. k H represents the number of words in the entity pair of the k-th technical solution. ik Let H be the information entropy of the i-th word in the k-th technical solution entity pair. max To maximize the information entropy of the i-th word in all entity pairs of technical solutions, L ik Let L be the length of the i-th word in the k-th technical solution entity pair. avg The average length of all words in all entity pairs of all technical solutions, β is the first adjustment factor, and A ik Let A be the value of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair. max γ is the maximum value of the word frequency heterogeneity of the i-th word in all technical solution entity pairs, and γ is the second adjustment factor.

[0086] Specifically, the value A of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair. ik include:

[0087]

[0088] Where m is the number of technical solution entity pairs in which the i-th word appears in different technical solution entity pairs, and P ik Let C be the probability that the i-th word appears in the k-th technical solution entity pair. ik For the context information value of the i-th word in the k-th technical solution entity pair, C max This represents the maximum value of the contextual information for the i-th word in all entity pairs of technical solutions.

[0089] Specifically, the contextual information value C of the i-th word in the k-th technical solution entity pair is obtained by constructing a co-occurrence matrix. ik And the maximum value of contextual information C for the i-th word in all technical solution entity pairs. max .

[0090] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0091] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0092] In the several embodiments provided by this invention, it should be understood that the disclosed technical content can be implemented in other ways. The system embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between units or modules, and may be electrical or other forms.

[0093] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0094] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0095] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, read-only storage media (ROM), random access storage media (RAM), portable hard drives, magnetic disks, optical disks, and other media capable of storing program code.

[0096] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A patent classification method based on a pre-trained model bert, characterized in that, The application comprises the following steps: The patent document is divided into multiple technical scheme entity pairs by a pre-trained model bert; A keyword score calculation model is set to calculate the keyword score of each word in each technical scheme entity pair, and the word with the highest keyword score is compared with the technical classification words in IPC in terms of similarity, and the technical classification word in IPC with the highest similarity is taken as the technical classification of the patent document.

2. The patent classification method based on the pre-trained model bert of claim 1, wherein, The keyword score calculation model comprises the following steps: where S ik is the keyword score of the i-th word in the k-th technical solution entity pair, f ik is the frequency of the i-th word in the k-th technical solution entity pair, n k is the number of words in the k-th technical solution entity pair, H ik is the information entropy of the i-th word in the k-th technical solution entity pair, H max is the maximum information entropy of the i-th word in all technical solution entity pairs, L ik is the length of the i-th word in the k-th technical solution entity pair, L avg is the average length of all words in all technical solution entity pairs, β is the first adjustment factor, A ik is the value of the term frequency heterogeneity of the i-th word in the k-th technical solution entity pair, A max is the maximum value of the term frequency heterogeneity of the i-th word in all technical solution entity pairs, γ is the second adjustment factor.

3. The patent classification method based on pre-trained model bert of claim 2, wherein, The value A of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair ik Comprising: wherein m is the number of technical solution entity pairs in which the i-th vocabulary appears, P ik is the probability of the i-th vocabulary appearing in the k-th technical solution entity pair, C ik is the context information value of the i-th vocabulary in the k-th technical solution entity pair, C max is the maximum context information value of the i-th vocabulary in all technical solution entity pairs.

4. The patent classification method based on the pre-trained model bert of claim 3, wherein, The context information value C of the i-th word in the k-th technical solution entity pair is obtained by constructing a co-occurrence matrix ik and the maximum context information value C of the i-th word in all technical solution entity pairs max .

5. The patent classification method based on pre-trained model bert of claim 1, wherein, Before comparing the word with the highest keyword score with the technical classification words in IPC in terms of similarity, the following steps are further included: a technical noun dictionary is established, the word with the highest keyword score is filtered according to the technical noun dictionary, the words in the non-technical noun dictionary are deleted, and the filtered words are compared with the technical classification words in IPC in terms of similarity. 6.A patent classification system based on a pre-trained model bert, characterized in that, The application comprises the following steps: The entity pair division module is used for dividing the patent document into multiple technical scheme entity pairs by a pre-trained model bert; The classification module is used for setting a keyword score calculation model, calculating the keyword score of each word in each technical scheme entity pair, comparing the word with the highest keyword score with the technical classification words in IPC in terms of similarity, and taking the technical classification word in IPC with the highest similarity as the technical classification of the patent document.

7. The patent classification system based on pre-trained model bert of claim 6, wherein, The keyword score calculation model comprises the following steps: wherein S ik is the keyword score of the i-th word in the k-th technical solution entity pair, f ik is the frequency of the i-th word in the k-th technical solution entity pair, n k is the number of words in the k-th technical solution entity pair, H ik is the information entropy of the i-th word in the k-th technical solution entity pair, H max is the maximum information entropy of the i-th word in all technical solution entity pairs, L ik is the length of the i-th word in the k-th technical solution entity pair, L avg is the average length of all words in all technical solution entity pairs, β is a first adjustment factor, A ik is the value of the term frequency heterogeneity of the i-th word in the k-th technical solution entity pair, A max is the maximum value of the term frequency heterogeneity of the i-th word in all technical solution entity pairs, γ is a second adjustment factor.

8. The patent classification system based on pre-trained model bert of claim 7, wherein, The value A of the word frequency heterogeneity of the i-th word in the k-th technical solution entity pair ik Comprising: wherein m is the number of technical solution entity pairs in which the i-th vocabulary appears, P ik is the probability of the i-th vocabulary appearing in the k-th technical solution entity pair, C ik is the context information value of the i-th vocabulary in the k-th technical solution entity pair, C max is the maximum context information value of the i-th vocabulary in all technical solution entity pairs.

9. The patent classification system based on pre-trained model bert of claim 8, wherein, The context information value Cik of the i-th word in the k-th technical scheme entity pair and the maximum context information value Cmax of the i-th word in all technical scheme entity pairs are obtained by constructing a co-occurrence matrix.

10. The patent classification system based on pre-trained model bert of claim 6, wherein, Before comparing the word with the highest keyword score with the technical classification words in IPC in terms of similarity, the following steps are further included: a technical noun dictionary is established, the word with the highest keyword score is filtered according to the technical noun dictionary, the words in the non-technical noun dictionary are deleted, and the filtered words are compared with the technical classification words in IPC in terms of similarity.