Batch document processing method, device and computer equipment
By calculating the word difference degree and differentiated keywords in multiple batches of documents, the problem of inaccurate keyword extraction caused by the failure to consider document differences in the existing technology is solved, and more accurate document summary extraction is achieved.
Patent Information
- Application Number
- CN202110965078.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-20
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2041-08-20
AI Technical Summary
Existing technologies fail to consider the differences between multiple documents or batches when extracting keywords, resulting in inaccurate keyword extraction and affecting the accuracy of document processing results.
By acquiring a set of words, calculating the inverse document frequency and term frequency of words in different batches of documents, determining the degree of difference of words in different batches of documents, and determining differentiated keywords based on the degree of difference, these keywords are used to extract the summary of the current document.
It improves the accuracy of keyword extraction, enabling the abstract to better reflect the core content of the current document.
Smart Images

Figure CN113627177B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of document processing, and more specifically relates to a method, apparatus and computer equipment for multi-batch document processing. Background Technology
[0002] With the rapid development of the internet, all kinds of documents are being generated. Currently, the approach involves processing a single document to extract its keywords, or processing a batch of documents to extract their keywords. These extracted keywords are then used to process the current document.
[0003] However, when realizing the inventive concept of this invention, the inventors discovered at least the following technical problems in the related technology: when extracting keywords, the related technology does not consider the differences between multiple documents or multiple batches of documents, so that the extracted keywords cannot characterize the differences in multiple documents or multiple batches of documents, and thus the results obtained when processing the current document are not accurate enough.
[0004] Therefore, it is necessary to provide a multi-batch document processing method to solve the above problems. Summary of the Invention
[0005] (a) Technical problems to be solved
[0006] The present invention aims to solve the problems in related technologies that do not consider the differences between multiple documents or batches of documents when extracting keywords, so that the extracted keywords cannot represent the differences in multiple documents or batches of documents, and thus the results obtained when processing the current document are not accurate enough.
[0007] (II) Technical Solution
[0008] To address the aforementioned technical problems, one aspect of the present invention proposes a multi-batch document processing method, comprising: obtaining a word set based on the multi-batch documents; determining the degree of difference of each word in the word set across different batches of documents based on the inverse document frequency and word frequency of each word in the word set across different batches of documents; determining differentiated keywords in different batches of documents based on the degree of difference of each word in the word set across different batches of documents; and extracting a summary of the current document based on the differentiated keywords in different batches of documents to obtain a summary of the current document.
[0009] According to a preferred embodiment of the present invention, obtaining the word set based on the multiple batches of documents includes: performing word segmentation on each batch of documents in the multiple batches of documents to obtain words for each batch of documents; removing stop words from the words in each batch of documents according to a stop word list to obtain keywords in each batch of documents; and obtaining the word set based on the keywords in each batch of documents.
[0010] According to a preferred embodiment of the present invention, before determining the degree of difference of each word in the word set in different batches of documents, the method further includes: calculating the inverse document frequency of each word in the word set in different batches of documents; and calculating the word frequency of each word in the word set in different batches of documents.
[0011] According to a preferred embodiment of the present invention, determining the differentiated keywords in different batches of documents based on the degree of difference of each word in the word set in different batches of documents includes: sorting the degree of difference of each word in the word set in different batches of documents; and obtaining the differentiated keywords in different batches of documents based on the sorting results.
[0012] According to a preferred embodiment of the present invention, extracting a summary of the current document based on differentiated keywords in different batches of documents, and obtaining the summary of the current document includes: segmenting the current document into sentences to obtain multiple sentences; determining the differentiated keywords contained in each sentence based on the differentiated keywords in different batches of documents; calculating the weight sum of the differentiated keywords in each sentence based on the weight of the differentiated keywords in each sentence, wherein the weight of the differentiated keyword is the degree of difference of the differentiated keyword in different batches of documents; and determining the summary of the current document based on the weight sum of the differentiated keywords in each sentence.
[0013] According to a preferred embodiment of the present invention, the method further includes: scoring the current document based on differentiated keywords in different batches of documents to obtain a score for the current document; and determining the language quality level of the current document based on the score of the current document.
[0014] According to a preferred embodiment of the present invention, the method further includes: determining the differentiated keywords contained in the current document based on the differentiated keywords in different batches of documents; calculating the weight sum of the differentiated keywords in the current document based on the weight of each differentiated keyword, wherein the weight of the differentiated keyword is the degree of difference of the differentiated keyword in different batches of documents; and determining the category of the current document based on the weight sum of the differentiated keywords in the current document.
[0015] A second aspect of the present invention provides a multi-batch document processing apparatus, comprising: a word set acquisition module, configured to acquire a word set based on the multi-batch documents; a difference degree determination module, configured to determine the difference degree of each word in the word set across different batches of documents based on the inverse document frequency and word frequency of each word in the word set across different batches of documents; a differentiated keyword determination module, configured to determine differentiated keywords in different batches of documents based on the difference degree of each word in the word set across different batches of documents; and a summary extraction module, configured to extract a summary of the current document based on the differentiated keywords in different batches of documents.
[0016] A third aspect of the present invention provides a computer device including a processor and a memory, the memory being used to store a computer-executable program, wherein when the computer program is executed by the processor, the processor executes a multi-batch document processing method as described in any of the preceding claims.
[0017] A fourth aspect of the present invention provides a computer program product storing a computer-executable program, wherein when the computer-executable program is executed, it implements a multi-batch document processing method as described in any of the preceding claims.
[0018] (III) Beneficial Effects
[0019] Compared with related technologies, this invention obtains a word set from multiple batches of documents. Based on the inverse document frequency and word frequency of each word in the word set across different batches of documents, it determines the degree of difference of each word in the word set across different batches of documents. Then, based on the degree of difference of each word in the word set across different batches of documents, it identifies differentiated keywords in different batches of documents. Finally, based on these differentiated keywords, it extracts a summary from the current document. This invention considers the differences between multiple batches of documents when extracting keywords and utilizes the extracted differentiated keywords to extract the summary of the current document. The summary extracted in this way is more accurate and better reflects the core content of the current document. Attached Figure Description
[0020] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of embodiments of the present invention can be applied is shown;
[0021] Figure 2 This is a flowchart illustrating an example of a multi-batch document processing method according to an embodiment of the present invention;
[0022] Figure 3 This is a flowchart of another example of the multi-batch document processing method according to an embodiment of the present invention;
[0023] Figure 4 This is a flowchart of another example of the multi-batch document processing method according to an embodiment of the present invention;
[0024] Figure 5 This is a flowchart of another example of the multi-batch document processing method according to an embodiment of the present invention;
[0025] Figure 6 This is a flowchart of another example of the multi-batch document processing method according to an embodiment of the present invention;
[0026] Figure 7 This is a flowchart of another example of the multi-batch document processing method according to an embodiment of the present invention;
[0027] Figure 8 This is a schematic diagram of an example of a multi-batch document processing apparatus according to an embodiment of the present invention;
[0028] Figure 9 This is a schematic diagram of yet another example of a multi-batch document processing apparatus according to an embodiment of the present invention;
[0029] Figure 10 This is a schematic diagram of yet another example of a multi-batch document processing apparatus according to an embodiment of the present invention;
[0030] Figure 11 This is a schematic diagram of yet another example of a multi-batch document processing apparatus according to an embodiment of the present invention;
[0031] Figure 12 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0032] Figure 13 This is a schematic diagram of a computer program product according to an embodiment of the present invention. Detailed Implementation
[0033] In the description of specific embodiments, detailed descriptions of structures, performance, effects, or other features are provided to enable those skilled in the art to fully understand the embodiments. However, this does not preclude those skilled in the art from implementing the present invention with technical solutions that do not contain the aforementioned structures, performance, effects, or other features under specific circumstances.
[0034] The flowcharts in the accompanying drawings are merely illustrative examples and do not imply that the solution of this invention must include all the content, operations, and steps shown in the flowcharts, nor do they imply that the execution must be performed in the order shown in the diagrams. For example, some operations / steps in the flowcharts can be decomposed, some operations / steps can be combined or partially combined, etc. Without departing from the inventive spirit of this invention, the execution order shown in the flowcharts can be changed according to the actual situation.
[0035] The box in the attached diagram Figure 1 Generally, these refer to functional entities, and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processing unit devices and / or microcontroller devices.
[0036] The same reference numerals in the accompanying drawings denote the same or similar elements, components, or parts, and therefore, repeated descriptions of the same or similar elements, components, or parts may be omitted below. It should also be understood that although terms such as first, second, third, etc., indicating numbers may be used herein to describe various devices, elements, components, or parts, these devices, elements, components, or parts should not be limited by these terms. That is, these terms are only used to distinguish one from another. For example, a first device may also be referred to as a second device, without departing from the essential technical solution of the invention. Furthermore, the terms "and / or" and "and / or" refer to all combinations including any one or more of the listed items.
[0037] This invention proposes an image processing method that can more accurately locate images, such as clock images, within problem questions. Based on this, the method segments the image (e.g., a clock image) from the problem image and identifies it using a recognition model, enabling more efficient recognition processing. For example, a clock image recognition model and an abacus image recognition model can be trained to recognize the clock or abacus image respectively, obtaining the corresponding readings and acquiring the corresponding answer information for further automatic grading. This improves the recognition accuracy of various problem types related to clocks and abacuses, achieves more intelligent automatic grading, and also enhances robustness.
[0038] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0039] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of embodiments of the present invention can be applied is shown.
[0040] like Figure 1 As shown, system architecture 100 may include one or more of user terminals 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between user terminals 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0041] It should be understood that Figure 1The number of user terminals, networks, and servers shown is merely illustrative. Depending on implementation needs, there can be any number of user terminals, networks, and servers. For example, server 105 could be a server cluster composed of multiple servers.
[0042] Users can use user terminals 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. User terminals 101, 102, and 103 can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers, etc.
[0043] Server 105 can be a server providing various services. For example, server 105 can acquire multiple batches of documents from user terminal 103 (or user terminal 101 or 102) in real time, and based on these batches, obtain a word set. It then determines the degree of difference between each word in the word set across different batches of documents based on the inverse document frequency and word frequency of each word in the word set. Furthermore, based on this degree of difference, it identifies differentiated keywords in different batches of documents. Finally, it extracts a summary of the current document based on these differentiated keywords. This invention considers the differences between multiple batches of documents when extracting keywords and utilizes the extracted differentiated keywords to extract the summary of the current document. The summary extracted in this way is more accurate and better reflects the core content of the current document.
[0044] In some embodiments, the batch document processing method provided in this invention is generally executed by server 105, and correspondingly, the batch document processing device is generally located in server 105. In other embodiments, certain terminals may have similar functions to the server to execute this method. Therefore, the batch document processing method provided in this invention is not limited to execution on the server side.
[0045] Figure 2 This is a flowchart illustrating an example of a multi-batch document processing method according to an embodiment of the present invention.
[0046] like Figure 2 As shown, the multi-batch document processing method includes steps S210 to S240.
[0047] In step S210, a word set is obtained based on the multiple batches of documents.
[0048] In step S220, the degree of difference of each word in the word set in different batches of documents is determined based on the inverse document frequency of each word in the word set in different batches of documents and the word frequency of each word in the word set in different batches of documents.
[0049] In step S230, differentiated keywords in different batches of documents are determined based on the degree of difference of each word in the word set in different batches of documents.
[0050] In step S240, the current document is extracted based on the differentiated keywords in different batches of documents to obtain the summary of the current document.
[0051] This method can obtain a word set from multiple batches of documents. Based on the inverse document frequency and word frequency of each word in different batches of documents, it determines the degree of difference of each word in the word set across different batches of documents. Furthermore, based on the degree of difference of each word in the word set across different batches of documents, it identifies differentiated keywords in different batches of documents. This approach considers the differences between multiple batches of documents when extracting keywords, making the extracted keywords more representative. Then, based on the differentiated keywords in different batches of documents, it extracts a summary of the current document, resulting in a more accurate summary that better reflects the core content of the current document.
[0052] In some embodiments of the present invention, the aforementioned batches of documents can be documents of different types. For example, batches of documents may include sports-related documents, entertainment-related documents, military-related documents, etc., but are not limited thereto. As another example, in the field of customer service, batches of documents may include completed order documents and uncompleted order documents.
[0053] In some embodiments of the present invention, the aforementioned multiple batches of documents may also be conversation documents for the same issue at different times. For example, conversation documents formed by multiple teachers answering parents or students online at different times for the same issue. Another example is conversation documents formed by multiple customer service representatives of a merchant answering users online at different times for the same issue.
[0054] In some embodiments of the present invention, before determining the degree of difference of each word in the above-mentioned word set in different batches of documents, the method further includes: calculating the inverse document frequency of each word in the above-mentioned word set in different batches of documents; and calculating the word frequency of each word in the above-mentioned word set in different batches of documents.
[0055] In some embodiments of the present invention, the inverse document frequency of each word in the above word set in different batches of documents is calculated by formula (1), as shown below:
[0056]
[0057] Where IDF represents inverse document frequency, w represents a word in the above word set, and b represents the batch identifier in multiple batches of documents. This represents the inverse document frequency of word w in the word set within the batch of documents b. In this embodiment, This represents the inverse document frequency of word w in the set of words within the batch document a.
[0058] In some embodiments of the present invention, the word frequency of each word in the above word set in different batches of documents is calculated by formula (2), as shown below:
[0059]
[0060] Where TF represents word frequency, w represents a word in the above word set, and b represents the batch identifier in multiple batches of documents. This represents the word frequency of word w in the word set within the batch document b. In this embodiment, This indicates the word frequency of word w in the set of words within the batch document a.
[0061] In some embodiments of the present invention, the degree of difference of each word in the word set across different batches of documents is determined based on the inverse document frequency and the word frequency of each word in the word set across different batches of documents. For example, the degree of difference of word w in batch a relative to word w in batch b is calculated based on the inverse document frequency and the word frequency of word w in the word set across different batches of documents. As another example, the degree of difference of word w in batch b relative to word w in batch a is calculated based on the inverse document frequency and the word frequency of word w in the word set across different batches of documents.
[0062] In some embodiments of the present invention, the difference between words w in batch a and words w in batch b is calculated based on the inverse document frequency and word frequency of words w in different batches of documents. For example, the difference between words w in batch a and words w in batch b is calculated using formula (3), which is shown below:
[0063]
[0064] Where DP represents the degree of difference. This represents the degree of difference between word w in batch a documents and word w in batch b documents. This indicates the word frequency of word w in batch document a. This indicates the word frequency of word w in batch document b. This indicates the inverse document frequency of word w in batch document a. This indicates the inverse document frequency of word w in batch b documents.
[0065] In some embodiments of the present invention, the difference between words w in batch b and words w in batch a is calculated based on the inverse document frequency and word frequency of words w in different batches of documents. The difference between words w in batch b and words w in batch a is calculated using formula (4), which is shown below:
[0066]
[0067] Where DP represents the degree of difference. This indicates the degree of difference between word w in batch b documents and word w in batch a documents. This indicates the word frequency of word w in batch document a. This indicates the word frequency of word w in batch document b. This indicates the inverse document frequency of word w in batch document a. This indicates the inverse document frequency of word w in batch b documents.
[0068] Figure 3 This is a flowchart of another example of the multi-batch document processing method of the present invention.
[0069] like Figure 3 As shown, step S210 can specifically include steps S310 to S330.
[0070] In step S310, each batch of documents in the multiple batches of documents is segmented into words to obtain the words of each batch of documents.
[0071] In step S320, stop words are removed from the words in each batch of documents according to the stop word list to obtain the keywords in each batch of documents.
[0072] In step S330, the word set is obtained based on the keywords in each batch of documents.
[0073] This method can perform word segmentation on each batch of documents in multiple batches to obtain the words in each batch of documents. Then, it removes stop words from the words in each batch of documents according to the stop word list to obtain the keywords in each batch of documents. This can effectively avoid the need to process stop words in each batch of documents in the subsequent process and improve the efficiency of subsequent processing.
[0074] In some embodiments of the present invention, word segmentation is performed on each batch of documents in multiple batches to obtain the words for each batch of documents. For example, a word segmentation tool is used to perform word segmentation on each batch of documents in multiple batches to obtain the words for each batch of documents. Another example is to perform word segmentation on each batch of documents in multiple batches based on a self-built thesaurus to obtain the words for each batch of documents. In this embodiment, the self-built thesaurus can be set according to the actual application scenario.
[0075] In some embodiments of the present invention, stop words are removed from the words in each batch of documents according to a stop word list to obtain the keywords in each batch of documents. For example, by iterating through the words in each batch of documents according to the stop word list, the stop words to be deleted can be quickly filtered out from the words in each batch of documents. In this embodiment, the stop word list may include words unrelated to actual business. The stop word list can be set according to actual business requirements.
[0076] Figure 4 This is a flowchart of another example of the multi-batch document processing method of the present invention.
[0077] like Figure 4 As shown, step S230 may specifically include steps S410 to S420.
[0078] In step S410, the differences in the word set between different batches of documents are sorted.
[0079] In step S420, based on the sorting results, the differentiated keywords in different batches of documents are obtained.
[0080] This method can sort the differences of each word in the word set across different batches of documents, and obtain differentiated keywords from different batches of documents based on the sorting results. This allows keywords with greater differences across different batches of documents to be used as differentiated keywords. The differentiated keywords obtained in this way have greater distinguishability and are helpful for subsequent processing of the current document using these differentiated keywords.
[0081] In some embodiments of the present invention, the degree of difference of each word in the word set across different batches of documents is sorted. For example, the different batches of documents are batch a and batch b, and the word set contains five words, namely word u, word v, word w, word x, and word y. The five degrees of difference are calculated using the above formula (3). Arrange in descending order Sort the data, and the sorting result is: Based on the ranking results, the top K words can be selected as differentiating keywords in different batches of documents. For example, words w, x, and v can be used as differentiating keywords in different batches of documents. In this embodiment, the top K can be set according to the actual situation.
[0082] Figure 5 This is a flowchart of another example of the multi-batch document processing method of the present invention.
[0083] like Figure 5 As shown, step S240 may specifically include steps S510 to S540.
[0084] In step S510, the current document is segmented into sentences to obtain multiple sentences of the current document.
[0085] In step S520, the differentiated keywords contained in each sentence are determined based on the differentiated keywords in different batches of documents.
[0086] In step S530, the weights of the differentiated keywords in each sentence are calculated based on the weights of the differentiated keywords in each sentence.
[0087] In step S540, the summary of the current document is determined based on the weight of the differentiated keywords in each sentence.
[0088] This method can segment the current document into sentences, obtaining multiple sentences. Then, based on the differentiated keywords in different batches of documents, it can quickly and accurately determine the differentiated keywords contained in each sentence. Based on the weight of the differentiated keywords in each sentence, it calculates the weight sum of the differentiated keywords in each sentence, so as to determine the summary of the current document. The summary content extracted in this way can better reflect the core content of the current document.
[0089] In some embodiments of the present invention, the weight of the differentiated keywords for each sentence can be the difference degree DP of the keywords in different batches of documents calculated by formula (3) or formula (4) above.
[0090] In some embodiments of the present invention, the current document is segmented into sentences to obtain multiple sentences. For example, a sentence segmentation tool is used to segment the current document into sentences to obtain multiple sentences.
[0091] In some embodiments of the present invention, differentiated keywords contained in each sentence are determined based on differentiated keywords in different batches of documents. For example, based on the differentiated keywords in different batches of documents, the words in each sentence are traversed, and words in each sentence that are the same as the aforementioned differentiated keywords are taken as the differentiated keywords of that sentence.
[0092] In some embodiments of the present invention, the weights of the differentiated keywords in each sentence are calculated based on the weights of the differentiated keywords in each sentence. For example, the weights of the differentiated keywords in each sentence are summed to obtain the weights of the differentiated keywords in each sentence.
[0093] In some embodiments of the present invention, the summary of the current document is determined based on the sum of the weights of the differentiated keywords in each sentence. For example, the sums of the weights of the differentiated keywords in each sentence are sorted in descending order, and the top-ranked sentence is taken as the summary of the current document.
[0094] Figure 6 This is a flowchart of another example of the multi-batch document processing method of the present invention.
[0095] like Figure 6 As shown, the method further includes steps S610 to S620.
[0096] In step S610, the current document is scored based on the differentiated keywords in different batches of documents to obtain the score of the current document.
[0097] In step S620, the quality level of the current document is determined based on the current document's rating.
[0098] This method can accurately obtain the score of the current document based on the differentiated keywords in different batches of documents, and then determine the quality level of the current document based on the score. The quality level determined in this way is more accurate and better matches the actual expression effect of the current document.
[0099] In some embodiments of the present invention, the current document is scored based on differentiated keywords from different batches of documents to obtain a score for the current document. For example, the current document is segmented into words to obtain a word set. The word set of the current document is traversed, and words that are the same as the differentiated keywords from the different batches of documents are selected from the word set. The current document is then scored based on the degree of difference of these words. For example, the degree of difference of these words is summed, and the score of the current document is determined based on the sum of the degree of difference of these words.
[0100] In some embodiments of the present invention, the quality level of the current document is determined based on its rating. For example, if the current document is a conversation document in which a teacher answers questions for parents or students online, and the current document's rating is greater than or equal to a preset threshold, then the current document's quality level is determined to be good. Conversely, if the current document's rating is less than the preset threshold, then the current document's quality level is determined to be bad. In this embodiment, the quality level can be set according to the business type to which the current document belongs.
[0101] Figure 7 This is a flowchart of another example of the multi-batch document processing method of the present invention.
[0102] like Figure 7 As shown, the method further includes steps S710 to S730.
[0103] In step S710, the differentiated keywords contained in the current document are determined based on the differentiated keywords in different batches of documents.
[0104] In step S720, the weights of the differentiated keywords in the current document are calculated based on the weights of each differentiated keyword in the current document.
[0105] In step S730, the category of the current document is determined based on the weight sum of the differentiated keywords of the current document.
[0106] This method can determine the differentiated keywords contained in the current document based on the differentiated keywords in different batches of documents, and calculate the weight sum of the differentiated keywords in the current document based on the weight of each differentiated keyword in the current document. Based on the weight sum of the differentiated keywords in the current document, the category of the current document can be determined. In this way, the category of the current document can be determined quickly and accurately.
[0107] In some embodiments of the present invention, differentiated keywords contained in the current document are determined based on differentiated keywords in different batches of documents. For example, based on the differentiated keywords in different batches of documents, the words in the current document are traversed, and words that are the same as the aforementioned differentiated keywords are taken as differentiated keywords of the current document.
[0108] In some embodiments of the present invention, the weight of each differentiated keyword in the current document can be the difference degree DP of the keyword in different batches of documents calculated by formula (3) or formula (4) above.
[0109] In some embodiments of the present invention, the weights of the differentiated keywords in the current document are calculated based on the weight of each differentiated keyword in the current document. For example, the weights of each differentiated keyword in the current document are summed to obtain the weights of the differentiated keywords in the current document.
[0110] In some embodiments of the present invention, the category of the current document is determined based on the weighted sum of the differentiated keywords of the current document. For example, if the current document is a document in the field of customer service, and the weighted sum of the differentiated keywords of the current document is positive, the current document is determined to be a document of the "completed order" type. Conversely, if the weighted sum of the differentiated keywords of the current document is negative, the current document is determined to be a document of the "non-completed order" type. In this embodiment, the category of the current document can be set according to actual circumstances.
[0111] Figure 8 This is a schematic diagram of an example of a multi-batch document processing apparatus according to an embodiment of the present invention.
[0112] like Figure 8 As shown, the multi-batch document processing device 800 includes a word set acquisition module 801, a difference determination module 802, a differentiated keyword determination module 803, and a summary extraction module 804.
[0113] Specifically, the word set acquisition module 801 is used to acquire a word set based on the multiple batches of documents.
[0114] The difference determination module 802 is used to determine the difference degree of each word in the word set in different batches of documents based on the inverse document frequency of each word in the word set in different batches of documents and the word frequency of each word in the word set in different batches of documents.
[0115] The differentiated keyword determination module 803 is used to determine differentiated keywords in different batches of documents based on the degree of difference of each word in the word set in different batches of documents.
[0116] The abstract extraction module 804 is used to extract a summary of the current document based on the differentiated keywords in different batches of documents, and obtain a summary of the current document.
[0117] This multi-batch document processing device 800 can acquire a word set based on multiple batches of documents. It determines the degree of difference between each word in the word set across different batches of documents based on the inverse document frequency and word frequency of each word in the word set across different batches of documents. Furthermore, based on the degree of difference, it identifies differentiated keywords in different batches of documents. This approach considers the differences between multiple batches of documents when extracting keywords, making the extracted keywords more representative. Then, based on the differentiated keywords in different batches of documents, it extracts a summary of the current document, resulting in a more accurate summary that better reflects the core content of the current document.
[0118] According to an embodiment of the present invention, the multi-batch document processing device 800 can be used to implement Figure 2 The embodiments describe a method for processing multiple batches of documents.
[0119] In some embodiments of the present invention, the word set acquisition module 801 is further configured to: perform word segmentation on each batch of documents in the multiple batches of documents to obtain words for each batch of documents; remove stop words from the words in each batch of documents according to the stop word list to obtain keywords in each batch of documents; and acquire the word set according to the keywords in each batch of documents.
[0120] In some embodiments of the present invention, the differentiated keyword determination module 803 is further configured to: sort the degree of difference of each word in the word set in different batches of documents; and obtain differentiated keywords in different batches of documents based on the sorting results.
[0121] In some embodiments of the present invention, the above-mentioned summary extraction module 804 is further configured to: perform sentence segmentation on the current document to obtain multiple sentences of the current document; determine the differentiated keywords contained in each sentence based on the differentiated keywords in different batches of documents; calculate the weight sum of the differentiated keywords of each sentence based on the weight of the differentiated keywords of each sentence, wherein the weight of the differentiated keywords is the degree of difference of the differentiated keywords in different batches of documents; and determine the summary of the current document based on the weight sum of the differentiated keywords of each sentence.
[0122] Figure 9 This is a schematic diagram of an example of a multi-batch document processing apparatus according to an embodiment of the present invention.
[0123] like Figure 9 As shown, the multi-batch document processing device 800 also includes an inverse document frequency calculation module 805 and a word frequency calculation module 806.
[0124] Specifically, the inverse document frequency calculation module 805 is used to calculate the inverse document frequency of each word in the word set in different batches of documents.
[0125] The word frequency calculation module 806 is used to calculate the word frequency of each word in the word set in different batches of documents.
[0126] Figure 10 This is a schematic diagram of an example of a multi-batch document processing apparatus according to an embodiment of the present invention.
[0127] like Figure 10 As shown, the multi-batch document processing device 800 also includes a scoring acquisition module 807 and a quality grade determination module 808.
[0128] Specifically, the scoring acquisition module 807 is used to score the current document based on the differentiated keywords in different batches of documents, and to obtain the score of the current document.
[0129] The quality level determination module 808 is used to determine the language quality level of the current document based on the current document's rating.
[0130] The multi-batch document processing device 800 can accurately obtain the score of the current document based on the differentiated keywords in different batches of documents, and then determine the quality level of the current document based on the score. The quality level of the current document determined in this way is more accurate, and the quality level is more in line with the actual expression effect of the current document.
[0131] According to an embodiment of the present invention, the multi-batch document processing device 800 can be used to implement Figure 6 The embodiments describe a method for processing multiple batches of documents.
[0132] Figure 11 This is a schematic diagram of an example of a multi-batch document processing apparatus according to an embodiment of the present invention.
[0133] like Figure 11 As shown, the multi-batch document processing device 800 also includes a differentiated keyword determination module 809, a differentiated keyword weight and calculation module 810, and a document category determination module 811.
[0134] Specifically, the differentiated keyword determination module 809 is used to determine the differentiated keywords contained in the current document based on the differentiated keywords in different batches of documents.
[0135] The differential keyword weight calculation module 810 is used to calculate the weight sum of the differential keywords in the current document based on the weight of each differential keyword in the current document, wherein the weight of the differential keyword is the degree of difference of the differential keyword in different batches of documents.
[0136] The document category determination module 811 is used to determine the category of the current document based on the weight sum of the differentiated keywords of the current document.
[0137] The multi-batch document processing device 800 can determine the differentiated keywords contained in the current document based on the differentiated keywords in different batches of documents, and calculate the weight sum of the differentiated keywords in the current document based on the weight of each differentiated keyword in the current document. Based on the weight sum of the differentiated keywords in the current document, the category of the current document can be determined quickly and accurately.
[0138] According to an embodiment of the present invention, the multi-batch document processing device 800 can be used to implement Figure 7 The embodiments describe a method for processing multiple batches of documents.
[0139] Since each module of the multi-batch document processing apparatus 800 in the example embodiment of the present invention can be used to implement the above 2~ Figure 7 The steps of the example embodiment of the multi-batch document processing method described herein are as follows. Therefore, for details not disclosed in the device embodiments of the present invention, please refer to the embodiments of the multi-batch document processing method of the present invention described above.
[0140] It is understandable that the word set acquisition module 801, the difference determination module 802, the differentiated keyword determination module 803, the summary extraction module 804, the inverse document frequency calculation module 805, the word frequency calculation module 806, the score acquisition module 807, the quality level determination module 808, the differentiated keyword determination module 809, the differentiated keyword weight and calculation module 810, and the document category determination module 811 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of the present invention, at least one of the following modules can be implemented, at least partially, as hardware circuits: a word set acquisition module 801, a difference degree determination module 802, a differentiated keyword determination module 803, a summary extraction module 804, an inverse document frequency calculation module 805, a word frequency calculation module 806, a score acquisition module 807, a quality level determination module 808, a differentiated keyword determination module 809, a differentiated keyword weight and calculation module 810, and a document category determination module 811. These modules can be implemented, for example, as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or as hardware or firmware in any other reasonable manner of integrating or packaging circuits, or as a suitable combination of software, hardware, and firmware implementations. Alternatively, at least one of the following modules can be implemented, at least partially, as a computer program module: word set acquisition module 801, difference degree determination module 802, differentiated keyword determination module 803, abstract extraction module 804, inverse document frequency calculation module 805, word frequency calculation module 806, score acquisition module 807, quality level determination module 808, differentiated keyword determination module 809, differentiated keyword weight and calculation module 810, and document category determination module 811. When the program is run by a computer, it can perform the functions of the corresponding module.
[0141] The following describes embodiments of the computer device of the present invention, which can be considered as specific implementations of the methods and apparatus embodiments of the present invention described above. Details described in the computer device embodiments of the present invention should be considered as supplements to the methods or apparatus embodiments described above; details not disclosed in the computer device embodiments of the present invention can be implemented with reference to the methods or apparatus embodiments described above.
[0142] Figure 12 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. The computer device includes a processor and a memory. The memory is used to store a computer-executable program. When the computer program is executed by the processor, the processor performs the method described in any one of the embodiments, including but not limited to... Figure 2 The method.
[0143] like Figure 12 As shown, the computer device is represented in the form of a general-purpose computing device. There can be one or more processors working collaboratively. This invention also does not exclude distributed processing, meaning that processors can be distributed across different physical devices. The computer device of this invention is not limited to a single entity, but can also be the sum of multiple physical devices.
[0144] The memory stores a computer-executable program, typically machine-readable code. The computer-readable program can be executed by the processor to enable the computer device to perform the method of the present invention, or at least some steps of the method.
[0145] The memory includes volatile memory, such as random access memory (RAM) and / or cache memory, and may also be non-volatile memory, such as read-only memory (ROM).
[0146] Optionally, in this embodiment, the computer device further includes an I / O interface for exchanging data with external devices. The I / O interface can represent one or more of several bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0147] It should be understood that Figure 12 The computer device shown is merely an example of the present invention, and the computer device of the present invention may also include elements or components not shown in the above examples. For example, some computer devices also include display units such as screens, and some computer devices also include human-computer interaction elements such as buttons and keyboards. Any computer device capable of executing a computer-readable program in memory to implement the method of the present invention or at least some steps of the method can be considered as a computer device covered by the present invention.
[0148] Figure 13 This is a schematic diagram of a computer program product according to an embodiment of the present invention. Figure 13 As shown, the computer program product stores a computer-executable program, which, when executed, implements the method described above. The computer program product may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer program product may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer program product may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0149] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0150] From the above description of the embodiments, those skilled in the art will readily understand that the present invention can be implemented by hardware capable of executing specific computer programs, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. included in the system. The present invention can also be implemented by computer software executing the methods of the present invention, for example, by control software executed by a microprocessor, electronic control unit, client, server, etc. However, it should be noted that the computer software executing the methods of the present invention is not limited to execution in one or a specific set of hardware entities; it can also be implemented in a distributed manner by unspecified hardware. For computer software, the software product can be stored in a computer-readable storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) or distributed across a network, as long as it enables computer devices to execute the methods according to the present invention.
[0151] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or computer equipment, and various general-purpose devices can also implement the present invention. The above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for processing multiple batches of documents, characterized in that, include: Obtain a set of words based on multiple batches of documents; Based on the inverse document frequency and word frequency of each word in different batches of documents, determine the relative degree of difference of each word in the word set in different batches of documents; Based on the relative differences of each word in the word set across different batches of documents, identify the keywords with the largest relative differences across different batches of documents to obtain differentiated keywords across different batches of documents; Based on the differentiated keywords in different batches of documents, the current document is extracted to obtain a summary, which includes: segmenting the current document into sentences to obtain multiple sentences; traversing the words of each sentence based on the differentiated keywords in different batches of documents to determine the differentiated keywords contained in each sentence; and determining the summary of the current document based on the calculated weight of the differentiated keywords in each sentence.
2. The multi-batch document processing method according to claim 1, characterized in that, Based on multiple batches of documents, the word set obtained includes: Each batch of documents in multiple batches is segmented into words to obtain the words for each batch of documents; By removing stop words from the list of stop words in each batch of documents, the keywords in each batch of documents can be obtained. Obtain a set of words based on the keywords in each batch of documents.
3. The multi-batch document processing method according to claim 1, characterized in that, Before determining the degree of difference of each word in the word set across different batches of documents, the following steps are also included: Calculate the inverse document frequency of each word in the word set across different batches of documents; Calculate the word frequency of each word in the word set across different batches of documents.
4. The multi-batch document processing method according to claim 1, characterized in that, Based on the relative differences of each word in the word set across different batches of documents, keywords with high relative differences across different batches of documents are identified to obtain differentiated keywords across different batches of documents, including: Sort the words in the word set according to their degree of difference across different batches of documents; Based on the sorting results, extract the differentiated keywords from different batches of documents.
5. The multi-batch document processing method according to claim 1, characterized in that, The summary of the current document is determined based on the weights of the differentiated keywords in each sentence, including: Based on the weight of the differentiated keywords in each sentence, calculate the sum of the weights of the differentiated keywords in each sentence. The weight of a differentiated keyword represents the degree of difference of that differentiated keyword in different batches of documents. Sort the sum of the weights of the differentiated keywords in each sentence, and determine the sentence ranked first as the summary of the current document.
6. The multi-batch document processing method according to claim 1, characterized in that, Also includes: Based on the differentiated keywords in different batches of documents, the current document is scored to obtain its score. Determine the quality level of the current document based on its rating.
7. The multi-batch document processing method according to claim 1, characterized in that, Also includes: Based on the differentiated keywords in different batches of documents, determine the differentiated keywords contained in the current document; Based on the weight of each differentiating keyword in the current document, calculate the sum of the weights of the differentiating keywords in the current document. The weight of a differentiating keyword is the relative degree of difference of that differentiating keyword in different batches of documents. The category of the current document is determined based on the weight of its differentiated keywords.
8. A multi-batch document processing device, characterized in that, include: The word set acquisition module is used to obtain word sets from multiple batches of documents; The difference determination module is used to determine the relative difference of each word in the word set in different batches of documents based on the inverse document frequency and word frequency of each word in the word set in different batches of documents. The differentiated keyword identification module is used to identify keywords with high relative differences in different batches of documents based on the relative differences of each word in the word set in different batches of documents, so as to obtain differentiated keywords in different batches of documents; The abstract extraction module is used to extract a summary of the current document based on the differentiated keywords in different batches of documents. The summary of the current document includes: segmenting the current document into sentences to obtain multiple sentences; traversing the words of each sentence based on the differentiated keywords in different batches of documents to determine the differentiated keywords contained in each sentence; and determining the summary of the current document based on the weight of the differentiated keywords in each sentence.
9. A computer device comprising a processor and a memory, the memory being used to store a computer-executable program, characterized in that: When the computer executable program is executed by the processor, the processor performs the multi-batch document processing method as described in any one of claims 1-7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the multi-batch document processing method according to any one of claims 1-7.
Citation Information
Patent Citations
Information data processing method and apparatus
CN107368489A
Voice quality inspection method, device and system
CN111314566A