Innovative analysis method and device for document, equipment, storage medium and product

Through the combination of document database and risk point identification model, the innovation of documents is automatically evaluated, which solves the long cycle and subjectivity problems caused by manual audits, and improves the audit efficiency and quality.

CN120508665APending Publication Date: 2025-08-19AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510635874.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

In the existing technology, the enterprise patent management department relies on manual evaluation for innovative document review of technical solutions, resulting in a long audit cycle and subjectiveness, which is difficult to meet the rapidly growing audit needs.

Method used

By obtaining the technical fields of documents to be analyzed, the document database is used to perform screening and automated analysis of similar documents, and combining pre-trained risk point identification models to evaluate the innovation and risk points of documents.

Benefits of technology

It realizes automated and innovative analysis of documents, reduces manual workload, improves audit efficiency and document quality, and improves the objectivity of innovative judgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508665A_ABST
    Figure CN120508665A_ABST
Patent Text Reader

Abstract

The invention discloses an innovative document analysis method and device, equipment, a storage medium and a product, and the method comprises the steps: obtaining a to-be-analyzed document, and determining a technical field corresponding to the to-be-analyzed document; wherein the to-be-analyzed document comprises the technical scheme; screening in a document database according to the technical field corresponding to the to-be-analyzed document to obtain similar documents of the to-be-analyzed document; and performing innovative analysis on the to-be-analyzed document according to the similar document. According to the innovative analysis method for the document, the document to be analyzed is subjected to automatic innovative analysis, so that related workers can be helped to judge the innovativeness of the document, the workload of workers is reduced, the document review efficiency is improved, and the document quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an innovative document analysis method, device, equipment, storage medium and product. Background Art

[0002] With the rapid development of science and technology, new technical solutions are constantly emerging. For enterprises, there is a great demand to convert technical solutions into patents. Generally speaking, the inventor of the technical solution will write the technical solution into a document and then submit the document to the enterprise's patent management department for review.

[0003] In the existing technology, when a company's patent management department conducts an audit, patent management personnel are required to manually evaluate the innovativeness and risk points of the document. The audit process mainly relies on the personal experience and professional knowledge of professionals.

[0004] However, because this traditional manual review method relies on the personal experience and expertise of patent managers, the review cycle is long and may be subject to certain subjectivity. With the increase in technical solutions, this manual review method has become difficult to meet actual needs. Summary of the Invention

[0005] The present invention provides a method, device, equipment, storage medium and product for innovative analysis of documents, so as to realize automated innovative analysis of documents containing technical solutions.

[0006] According to one aspect of the present invention, there is provided an innovative document analysis method, comprising:

[0007] Obtaining a document to be analyzed and determining the technical field corresponding to the document to be analyzed; wherein the document to be analyzed contains a technical solution;

[0008] Screening a document database according to the technical field corresponding to the document to be analyzed to obtain similar documents to the document to be analyzed;

[0009] Perform innovative analysis on the document to be analyzed based on the similar documents.

[0010] Furthermore, determining the technical field corresponding to the document to be analyzed includes:

[0011] Extracting keywords from the document to be analyzed to obtain at least one characteristic keyword;

[0012] The technical field is classified according to the characteristic keywords to obtain the technical field corresponding to the document to be analyzed.

[0013] Furthermore, screening is performed in a document database according to the technical field corresponding to the document to be analyzed to obtain similar documents to the document to be analyzed, including:

[0014] extracting at least one candidate document from the document database according to the technical field corresponding to the document to be analyzed;

[0015] The similarities between the document to be analyzed and each candidate document are compared respectively, and the candidate documents with similarities higher than a set threshold are taken as the similar documents.

[0016] Furthermore, the method for establishing the document database includes:

[0017] Collect original texts containing technical solutions;

[0018] After performing data cleaning and data standardization on each of the original texts, converting them into text vectors;

[0019] The text vectors are classified and stored according to the technical fields corresponding to the original texts.

[0020] Furthermore, performing innovative analysis on the document to be analyzed based on the similar documents includes:

[0021] Determining the similarity score of each of the similar documents with respect to the document to be analyzed;

[0022] In combination with the pre-trained risk point identification model, the risk points for the document to be analyzed in each of the similar documents are determined respectively;

[0023] The innovativeness analysis result of the document to be analyzed is determined based on the similarity scores and risk points corresponding to each of the similar documents.

[0024] Furthermore, respectively determining a similarity score of each of the similar documents with respect to the document to be analyzed includes:

[0025] Scoring the similarity of each of the similar documents to the document to be analyzed based on preset analysis dimensions; wherein the preset analysis dimensions include technical features, technical fields, and document repetition rate;

[0026] For each of the similar documents, a similarity score with the document to be analyzed is determined according to the scoring results corresponding to each of the preset analysis dimensions.

[0027] According to another aspect of the present invention, there is provided an innovative document analysis device, comprising:

[0028] A technical field determination module is used to obtain a document to be analyzed and determine the technical field corresponding to the document to be analyzed; wherein the document to be analyzed contains a technical solution;

[0029] A similar document screening module is used to screen the document database according to the technical field corresponding to the document to be analyzed, and obtain similar documents to the document to be analyzed;

[0030] The innovation analysis module is used to perform innovation analysis on the document to be analyzed based on the similar documents.

[0031] Optionally, the technical field determination module is further configured to:

[0032] Extracting keywords from the document to be analyzed to obtain at least one characteristic keyword;

[0033] The technical field is classified according to the characteristic keywords to obtain the technical field corresponding to the document to be analyzed.

[0034] Optionally, the similar document filtering module is also used to:

[0035] extracting at least one candidate document from the document database according to the technical field corresponding to the document to be analyzed;

[0036] The similarities between the document to be analyzed and each candidate document are compared respectively, and the candidate documents with similarities higher than a set threshold are taken as the similar documents.

[0037] Optionally, the device further includes a document database establishment module, which is used to:

[0038] Collect original texts containing technical solutions;

[0039] After performing data cleaning and data standardization on each of the original texts, converting them into text vectors;

[0040] The text vectors are classified and stored according to the technical fields corresponding to the original texts.

[0041] Optionally, the innovative analysis module is also used to:

[0042] Determining the similarity score of each of the similar documents with respect to the document to be analyzed;

[0043] In combination with the pre-trained risk point identification model, the risk points for the document to be analyzed in each of the similar documents are determined respectively;

[0044] The innovativeness analysis result of the document to be analyzed is determined based on the similarity scores and risk points corresponding to each of the similar documents.

[0045] Optionally, the innovative analysis module is also used to:

[0046] Scoring the similarity of each of the similar documents to the document to be analyzed based on preset analysis dimensions; wherein the preset analysis dimensions include technical features, technical fields, and document repetition rate;

[0047] For each of the similar documents, a similarity score with the document to be analyzed is determined according to the scoring results corresponding to each of the preset analysis dimensions.

[0048] According to another aspect of the present invention, an electronic device is provided, comprising:

[0049] at least one processor; and

[0050] a memory communicatively connected to the at least one processor; wherein,

[0051] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the innovative document analysis method described in any embodiment of the present invention.

[0052] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the innovative document analysis method described in any embodiment of the present invention when executed.

[0053] According to another aspect of the present invention, a computer program product is provided, which includes a computer program / instruction, which, when executed by a processor, implements the steps of the innovative document analysis method according to any embodiment of the present invention.

[0054] The method for analyzing the innovativeness of a document disclosed in the present invention first obtains a document to be analyzed and determines the technical field to which the document corresponds; the document to be analyzed contains a technical solution; then, based on the technical field corresponding to the document to be analyzed, a document database is screened to obtain similar documents to the document to be analyzed; and finally, the document to be analyzed is analyzed for innovativeness based on the similar documents. The method for analyzing the innovativeness of a document disclosed in the present invention, by performing automated innovativeness analysis on the document to be analyzed, can help relevant personnel determine the innovativeness of the document, reduce manual workload, improve document review efficiency, and enhance document quality.

[0055] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0057] Figure 1 This is a flow chart of an innovative document analysis method provided according to the first embodiment of the present invention;

[0058] Figure 2 This is a flow chart of an innovative document analysis method provided in accordance with the second embodiment of the present invention;

[0059] Figure 3 1 is a schematic structural diagram of an innovative document analysis device provided in accordance with a third embodiment of the present invention;

[0060] Figure 4 It is a structural diagram of an electronic device for implementing the innovative document analysis method of the fourth embodiment of the present invention. DETAILED DESCRIPTION

[0061] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0062] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0063] Example 1

[0064] Figure 1This is a flowchart of a method for analyzing the innovativeness of a document provided in the first embodiment of the present invention. This embodiment is applicable to analyzing the innovativeness of a document containing a technical solution. The method can be executed by a device for analyzing the innovativeness of a document. The device for analyzing the innovativeness of a document can be implemented in the form of hardware and / or software. The device for analyzing the innovativeness of a document can be configured in an electronic device. Figure 1 As shown, the method includes:

[0065] S110: Obtain a document to be analyzed, and determine the technical field corresponding to the document to be analyzed.

[0066] Among them, the document to be analyzed contains a technical solution.

[0067] In this embodiment, the document to be analyzed is the document currently undergoing innovation analysis and contains a technical solution. A technical solution is a targeted and systematic method, response, and related countermeasure proposed to address various technical issues. Based on the specific content of the technical solution, the technical field corresponding to the document to be analyzed can be determined.

[0068] Optionally, a method for determining the technical field corresponding to the document to be analyzed may be to extract keywords from the document to be analyzed, perform semantic analysis on the extracted keywords, and then match them with set technical field categories, and use the technical field corresponding to the keywords as the technical field corresponding to the document to be analyzed.

[0069] S120 , screening is performed in the document database according to the technical field corresponding to the document to be analyzed to obtain documents similar to the document to be analyzed.

[0070] The document database is a pre-established document database that contains relevant technical documents in various technical fields. Similar documents are documents that are screened out from the document database and have a certain degree of similarity with the content of the document to be analyzed.

[0071] In this embodiment, each document can be classified and stored in the document database according to the technical field to which each document corresponds. After determining the technical field of the document to be analyzed, documents in the corresponding technical field can be extracted from the document database, and the similarity between the extracted documents and the document to be analyzed can be further evaluated, and the documents with high similarity can be used as similar documents to the document to be analyzed.

[0072] S130. Perform innovative analysis on the document to be analyzed based on similar documents.

[0073] In this embodiment, after obtaining similar documents to the document to be analyzed, each similar document can be further compared in depth with the document to be analyzed. The degree of semantic similarity between each similar document and the document to be analyzed can be analyzed from multiple dimensions. The portions of each similar document that are semantically similar to the document to be analyzed can be identified, and the document to be analyzed can be scored for innovation. Subsequently, a risk analysis can be performed on the document to be analyzed based on the portions of each similar document that are semantically similar to the document to be analyzed, identifying possible risk characteristics in the document to be analyzed, such as infringement risk, invalidity risk, etc. Preferably, the innovation analysis results of the document to be analyzed can include information such as the innovation score for the document to be analyzed, the number and category of possible risks, and so on.

[0074] The method for analyzing the innovativeness of a document provided by an embodiment of the present invention first obtains a document to be analyzed and determines the technical field to which the document to be analyzed corresponds; wherein the document to be analyzed contains a technical solution; then, based on the technical field corresponding to the document to be analyzed, a document database is screened to obtain similar documents to the document to be analyzed; and finally, based on the similar documents, an innovativeness analysis is performed on the document to be analyzed. The method for analyzing the innovativeness of a document disclosed by the present invention, by performing automated innovativeness analysis on the document to be analyzed, can help relevant personnel determine the innovativeness of a document, reduce manual workload, improve document review efficiency, and enhance document quality.

[0075] Example 2

[0076] Figure 2 This is a flowchart of an innovative document analysis method provided in Example 2 of the present invention. This example is a refinement of the above example. Figure 2 As shown, the method includes:

[0077] S210: Obtain a document to be analyzed, extract keywords from the document to be analyzed, and obtain at least one characteristic keyword.

[0078] Among them, the document to be analyzed contains a technical solution.

[0079] In this embodiment, the document to be analyzed is a document currently undergoing innovation analysis and contains a technical solution. After obtaining the document to be analyzed, a keyword extraction method based on natural language processing technology, such as a statistical method, a semantic method, or a supervised learning method, can be used to extract characteristic keywords from the document to be analyzed.

[0080] S220 , classify technical fields according to characteristic keywords to obtain the technical field corresponding to the document to be analyzed, and extract at least one candidate document from the document database according to the technical field corresponding to the document to be analyzed.

[0081] The document database is a pre-established document database that contains relevant technical documents in various technical fields. The candidate documents are documents extracted from the document database that are in the same or similar technical fields as the document to be analyzed.

[0082] In this embodiment, the document database can classify and store each document according to its corresponding technical field. After determining the technical field of the document to be analyzed, documents in the same or similar technical fields can be extracted from the document database as candidate documents.

[0083] Among them, the method for establishing a document database can be: collecting original texts containing technical solutions; converting each original text into a text vector after data cleaning and data standardization; and classifying and storing the text vectors according to the technical fields corresponding to the original texts.

[0084] Preferably, the document database can be used to store technical documents in various technical fields, and to provide data support for the subsequent innovative analysis of the documents to be analyzed. When establishing a document database, the original text containing the technical solutions can be collected through various channels, for example, patent text data can be collected from patent databases and patent office websites, and technical documents such as journals, conferences and dissertations can be collected from HowNet and other websites. After collecting these original texts, they can be preprocessed, and the processing methods include removing special characters, word segmentation, removing stop words, sentence segmentation, segmentation, etc., and then the original text is cleaned and normalized, thereby providing a better data basis for subsequent text analysis tasks. Then, the text can be represented as a vector, for example, Word2vec (a production model for word vectors) can be used to represent the text as a vector.

[0085] Furthermore, the original documents need to be classified to determine the technical fields. For patents, this can be determined by their IPC (International Patent Classification) codes. For papers, a random forest algorithm can be used to classify keywords, and then the technical fields can be determined based on the journals or conferences in which they were published. Finally, the text vectors are classified and stored in a document database based on the technical fields corresponding to the original documents.

[0086] S230 : Compare the similarities between the document to be analyzed and each candidate document respectively, and select candidate documents with similarities higher than a set threshold as similar documents.

[0087] In this embodiment, after the candidate documents are extracted, the similarities between the document to be analyzed and each candidate document can be compared in sequence for judgment, and then the candidate documents with similarities higher than a set threshold are regarded as similar documents.

[0088] Preferably, when determining the similarity between the document to be analyzed and the candidate documents, a cosine similarity algorithm may be used to represent the document to be analyzed and the candidate documents as vectors, and then calculate the cosine similarity between the vectors.

[0089] S240: Determine the similarity score of each similar document with respect to the document to be analyzed.

[0090] In this embodiment, after obtaining similar documents to the document to be analyzed, each similar document can be further compared with the document to be analyzed in depth, and the semantic similarity between each similar document and the document to be analyzed can be analyzed from multiple dimensions and expressed in the form of a similarity score.

[0091] Optionally, the method for separately determining the similarity score of each similar document with respect to the document to be analyzed may be: scoring the similarity of each similar document with respect to the document to be analyzed according to preset analysis dimensions; wherein the preset analysis dimensions include technical features, technical fields, and document repetition rate; for each similar document, determining the similarity score with the document to be analyzed according to the scoring results corresponding to each preset analysis dimension.

[0092] Specifically, for each similar document, when calculating the similarity score between the similar document and the document to be analyzed, the similar document and the document to be analyzed can be respectively input into the BERT model for training to obtain the BERT word vector. Among them, BERT (Bidirectional Encoder Representation from Transformers) is the full name of the bidirectional encoder representation. It is a pre-training model. Its main feature is to use the Encoder part of the Transformer to achieve two-way language representation learning. It can capture language features through two pre-training tasks, namely Masked Language Model (MLM) and Next Sentence Prediction (NSP). After obtaining the BERT word vectors of the similar document and the document to be analyzed, they can be input into the LSTM network, and the semantic representation is learned through the LSTM network. Content recognition is performed on multiple dimensions including technical features, technical fields and document repetition rates, and semantically similar parts are identified. LSTM (Long Short-Term Memory) is a deep learning model suitable for sequential data, particularly for processing long sequences. It has the ability to capture long-term dependencies within sequences. In text similarity calculation tasks, it can be used to learn semantic information between texts and map it into a low-dimensional space, thereby achieving similarity calculation between texts. The semantic recognition results of the LSTM network can then be used to score the similarity of each dimension. Finally, the similarity scores between similar documents and the document being analyzed are determined based on the scores corresponding to each dimension.

[0093] S250: Determine the risk points for the document to be analyzed in each similar document by combining the pre-trained risk point identification model.

[0094] In this embodiment, for each similar document, it can be input into a pre-trained risk point identification model together with the document to be analyzed, and the risk point identification model can be used to perform risk analysis on the document to be analyzed to identify possible risk points in the document to be analyzed, such as infringement risk, invalidity risk, etc.

[0095] Furthermore, based on the risk point analysis results, the documents to be analyzed can be classified into risk-free texts, suspected infringing texts, suspected invalid texts, etc., and risk warning opinions can be given based on the number, category, severity, etc. of risk points.

[0096] S260: Determine the innovative analysis result of the document to be analyzed based on the similarity scores and risk points corresponding to each similar document.

[0097] The innovation analysis results may include an innovation score and risk point prompt information for the document to be analyzed.

[0098] In this embodiment, the document to be analyzed can be scored for its innovation based on the similarity scores of each similar document to the document to be analyzed. A higher score indicates a higher innovation of the document to be analyzed. For example, if the document to be analyzed corresponds to two similar documents, and both documents have a similarity score of 80, then the innovation score of the document to be analyzed is 20.

[0099] Furthermore, the risk point prompt information of the document to be analyzed may include information such as the number, category, severity, etc. of the risk points of the document to be analyzed. Preferably, after the innovation analysis result is generated based on the innovation score of the document to be analyzed and the risk point prompt information, it can be output in a visual form.

[0100] The document innovation analysis method provided by the embodiment of the present invention can help relevant staff determine the innovation of documents by performing automated innovation analysis on the documents to be analyzed, reduce manual workload, improve document review efficiency, and enhance document quality.

[0101] Example 3

[0102] Figure 3 This is a schematic diagram of the structure of an innovative document analysis device provided in Example 3 of the present invention, such as Figure 3 As shown, the device includes: a technical field determination module 310, a similar document screening module 320 and an innovation analysis module 330.

[0103] The technical field determination module 310 is used to obtain a document to be analyzed and determine the technical field corresponding to the document to be analyzed.

[0104] Among them, the document to be analyzed contains a technical solution.

[0105] The similar document screening module 320 is used to screen the document database according to the technical field corresponding to the document to be analyzed, and obtain similar documents to the document to be analyzed.

[0106] The innovation analysis module 330 is used to perform innovation analysis on the document to be analyzed based on similar documents.

[0107] Optionally, the technical field determination module 310 is further configured to:

[0108] Keywords are extracted from the document to be analyzed to obtain at least one characteristic keyword; technical fields are classified according to the characteristic keywords to obtain the technical field corresponding to the document to be analyzed.

[0109] Optionally, the similar document screening module 320 is further configured to:

[0110] At least one candidate document is extracted from the document database according to the technical field corresponding to the document to be analyzed; the similarity between the document to be analyzed and each candidate document is compared respectively, and the candidate document with a similarity higher than a set threshold is regarded as a similar document.

[0111] Optionally, the device further includes a document database establishment module, which is used to:

[0112] Collect the original texts containing technical solutions; convert the original texts into text vectors after data cleaning and data standardization; and classify and store the text vectors according to the technical fields corresponding to the original texts.

[0113] Optionally, the innovation analysis module 330 is further configured to:

[0114] Determine the similarity score of each similar document with respect to the document to be analyzed; combine with the pre-trained risk point identification model to determine the risk points in each similar document with respect to the document to be analyzed; and determine the innovative analysis results of the document to be analyzed based on the similarity scores and risk points corresponding to each similar document.

[0115] Optionally, the innovation analysis module 330 is further configured to:

[0116] Based on preset analysis dimensions, each similar document is scored for its similarity to the document to be analyzed; wherein the preset analysis dimensions include technical features, technical fields, and document repetition rate; for each similar document, the similarity score with the document to be analyzed is determined based on the scoring results corresponding to each preset analysis dimension.

[0117] The innovative document analysis device provided by the embodiment of the present invention can execute the innovative document analysis method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0118] Example 4

[0119] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0120] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0121] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0122] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the innovative document analysis method.

[0123] In some embodiments, the innovative document analysis method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the innovative document analysis described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the innovative document analysis method in any other appropriate manner (e.g., by means of firmware).

[0124] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0125] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0126] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0127] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0128] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0129] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

Claims

1. An innovative document analysis method, characterized in that: include: Obtaining a document to be analyzed and determining the technical field corresponding to the document to be analyzed; wherein the document to be analyzed contains a technical solution; Screening a document database according to the technical field corresponding to the document to be analyzed to obtain similar documents to the document to be analyzed; Perform innovative analysis on the document to be analyzed based on the similar documents.

2. The method according to claim 1, characterized in that Determine the technical field corresponding to the document to be analyzed, including: Extracting keywords from the document to be analyzed to obtain at least one characteristic keyword; The technical field is classified according to the characteristic keywords to obtain the technical field corresponding to the document to be analyzed.

3. The method according to claim 1, characterized in that Screening the document database based on the technical field corresponding to the document to be analyzed to obtain similar documents to the document to be analyzed includes: extracting at least one candidate document from the document database according to the technical field corresponding to the document to be analyzed; The similarities between the document to be analyzed and each candidate document are compared respectively, and the candidate documents with similarities higher than a set threshold are taken as the similar documents.

4. The method according to claim 3, characterized in that The method for establishing the document database includes: Collect original texts containing technical solutions; After performing data cleaning and data standardization on each of the original texts, converting them into text vectors; The text vectors are classified and stored according to the technical fields corresponding to the original texts.

5. The method according to claim 1, wherein Performing innovative analysis on the document to be analyzed based on the similar documents, including: Determining the similarity score of each of the similar documents with respect to the document to be analyzed; In combination with the pre-trained risk point identification model, the risk points for the document to be analyzed in each of the similar documents are determined respectively; The innovativeness analysis result of the document to be analyzed is determined based on the similarity scores and risk points corresponding to each of the similar documents.

6. The method according to claim 5, characterized in that Determining the similarity score of each of the similar documents with respect to the document to be analyzed includes: Scoring the similarity of each of the similar documents to the document to be analyzed based on preset analysis dimensions; wherein the preset analysis dimensions include technical features, technical fields, and document repetition rate; For each of the similar documents, a similarity score with the document to be analyzed is determined according to the scoring results corresponding to each of the preset analysis dimensions.

7. An innovative document analysis device, characterized in that: include: A technical field determination module is used to obtain a document to be analyzed and determine the technical field corresponding to the document to be analyzed; wherein the document to be analyzed contains a technical solution; A similar document screening module is used to screen the document database according to the technical field corresponding to the document to be analyzed, and obtain similar documents to the document to be analyzed; The innovation analysis module is used to perform innovation analysis on the document to be analyzed based on the similar documents.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the innovative document analysis method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a processor to implement the innovative document analysis method according to any one of claims 1 to 6 when the instructions are executed.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the innovative document analysis method according to any one of claims 1 to 6 are implemented.