File content labeling method and related device

By identifying and segmenting files, searching and filtering keywords, automatic labeling of file content is achieved, the problem of lack of automatic labeling function in the existing technology is solved, and the user's reading experience and efficiency are improved.

CN120086363AInactive Publication Date: 2025-06-03BEIJING SHANGYIN MICRO CORE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213173.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art lacks automatic labeling function when processing and reading a large number of files, resulting in low user reading efficiency.

Method used

By text recognition of the target file, sentence segmentation of the text is based on the first segmentation object, target keywords are found and related statements are filtered, and finally, statements containing target matching words are further segmented and marked based on the second segmentation object.

Benefits of technology

It realizes automatic labeling of file content, improves user reading experience and efficiency, and enhances the positioning and understanding of key content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086363A_ABST
    Figure CN120086363A_ABST
Patent Text Reader

Abstract

The invention discloses a file content labeling method and a related device, and relates to the technical field of computers, and the method comprises the following steps: carrying out text recognition on a target file to obtain a recognized text; based on the first segmentation object, performing statement segmentation on the recognized text to obtain a first segmented text; searching for a target keyword from the first segmented text, and based on a keyword search range corresponding to the target keyword, screening from the first segmented text to obtain a statement containing the target keyword; selecting a statement containing a target matching word from the statements containing the target keyword, and performing statement segmentation on the statement containing the target matching word based on a second segmentation object to obtain a second segmentation text; and marking the second segmented text. According to the method, the file content can be automatically labeled, so that the purpose of improving the reading experience of a user is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method for annotating file content and related devices. Background Art

[0002] Currently, when users process and read a large number of files (such as PDF (Portable Document Format), WORD files), they often need to quickly locate and understand the key content. However, most of the existing processing tools only provide basic text recognition and reading functions, lacking an automatic annotation function for key content, resulting in low reading efficiency for users. Therefore, it is necessary to annotate file content to improve the user reading experience. Summary of the Invention

[0003] In view of the above problems, this application provides a method for annotating file content and related devices, which can automatically annotate file content to achieve the purpose of improving the user reading experience. The specific solutions are as follows:

[0004] The first aspect of this application provides a method for annotating file content, including:

[0005] Performing text recognition on a target file to obtain the recognized text;

[0006] Based on a first segmentation object, segmenting the recognized text into sentences to obtain a first segmented text;

[0007] Searching for a target keyword from the first segmented text, and based on the keyword search range corresponding to the target keyword, screening and obtaining sentences containing the target keyword from the first segmented text;

[0008] Selecting sentences containing a target matching word from the sentences containing the target keyword, and based on a second segmentation object, segmenting the sentences containing the target matching word into sentences to obtain a second segmented text;

[0009] Annotating the second segmented text.

[0010] In a possible implementation, the step of based on a first segmentation object, segmenting the recognized text into sentences to obtain a first segmented text includes:

[0011] Obtaining a first segmentation object; where the first segmentation object includes one or more first sentence segmentation symbols;

[0012] Selecting the first sentence segmentation symbol arranged at the top, and based on the selected symbol, segmenting the recognized text into sentences to obtain a first segmented sentence;

[0013] In the case where not all of the first sentence segmentation symbols are selected, select the next first sentence segmentation symbol, and then return the sentence segmentation of the recognized text based on the selected symbol to obtain the first segmented sentences;

[0014] In the case where all of the first sentence segmentation symbols are selected, obtain the first segmented text.

[0015] In a possible implementation, screening for sentences containing the target keyword from the first segmented text based on the keyword search range corresponding to the target keyword includes:

[0016] If the keyword search range is the content of the target page of the target document, select the sentences in the first segmented text whose page numbers are the page numbers of the target page as the sentences containing the target keyword;

[0017] If the keyword search range is the content of the target segment related to the target keyword, screen out the target segment from the first segmented text as the sentences containing the target keyword.

[0018] In a possible implementation, sentence segmentation of the sentences containing the target matching word based on the second segmentation object to obtain a second segmented text includes:

[0019] Obtain a second segmentation object; wherein, the second segmentation object includes a truncated phrase, and the truncated phrase includes two second sentence segmentation symbols;

[0020] Intercept the content between the two second sentence segmentation symbols from the sentences containing the target matching word to obtain the second segmented text.

[0021] In a possible implementation, after searching for the target keyword from the first segmented text, it further includes:

[0022] Determine whether the found target keyword contains non-keywords;

[0023] After deleting the target keywords containing non-keywords, obtain the processed keywords;

[0024] Label the processed keywords.

[0025] In a possible implementation, after screening for sentences containing the target keyword from the first segmented text and before selecting sentences containing the target matching word from the sentences containing the target keyword, the method further includes:

[0026] Remove the target symbols from the statement containing the target keyword to obtain a first processed statement;

[0027] Replace the content in the first processed statement that conforms to the regular expression to obtain a second processed statement.

[0028] The second aspect of the present application provides a file content annotation system, including:

[0029] An identification module for performing text identification on a target file to obtain the identified text;

[0030] A first segmentation module for segmenting the identified text based on a first segmentation object to obtain a first segmented text;

[0031] A screening module for searching for a target keyword from the first segmented text and screening out the statements containing the target keyword from the first segmented text based on the keyword search range corresponding to the target keyword;

[0032] A second segmentation module for selecting the statements containing a target matching word from the statements containing the target keyword, and segmenting the statements containing the target matching word based on a second segmentation object to obtain a second segmented text;

[0033] A marking module for marking the second segmented text.

[0034] The third aspect of the present application provides a computer program product, including computer-readable instructions, which when running on an electronic device, enable the electronic device to implement the file content annotation method in the first aspect or any implementation manner of the first aspect.

[0035] The fourth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:

[0036] The memory is used to store a computer program;

[0037] The processor is used to execute the computer program so that the electronic device can implement the file content annotation method in the first aspect or any implementation manner of the first aspect.

[0038] The fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the file content annotation method in the first aspect or any implementation manner of the first aspect.

[0039] With the above technical solution, the document content annotation method and related device provided by this application segment sentences of the text through the first segmentation object, and screen out the sentences containing the target keyword, so as to locate the text to be annotated and narrow the search range of the text to be annotated. Then, based on the second segmentation object, the sentences containing the target matching word are further segmented to obtain the second segmented text, realizing the automatic annotation of the second segmented text and enhancing the user reading experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.

[0041] Figure 1 It is a flowchart of a document content annotation method provided by this application;

[0042] Figure 2 It is a structural diagram of a document content annotation system provided by this application;

[0043] Figure 3 It is a schematic structural diagram of an electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The following describes the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only for explaining the specific embodiments of the present application, and are not intended to limit the present application.

[0045] The following describes the embodiments of the present application with reference to the accompanying drawings. Those skilled in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0046] The terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0047] An embodiment of the present application provides a method for annotating file content. The method for annotating file content in the embodiment of the present application will be introduced in detail below with reference to the accompanying drawings.

[0048] Refer to Figure 1 , Figure 1 which is a flowchart of a method for annotating file content provided by an embodiment of the present application. As Figure 1 shown, a method for annotating file content provided by an embodiment of the present application may include steps 101 to 106, and these steps will be described in detail below.

[0049] Step 101: Perform text recognition on the target file to obtain the recognized text.

[0050] The target file may be a bidding document, a contract document, a guarantee letter document, etc., and the file format of the target file may be PDF, WORD, a picture, etc. When performing text recognition on the target file, the OCR (Optical Character Recognition) method may be used to obtain the recognized text as editable and processable text.

[0051] Optionally, in order to improve the accuracy and efficiency of subsequent extraction, after obtaining the recognized text, preprocessing may be performed on the recognized text, which may specifically include, but is not limited to, denoising processing, segmentation processing, formatting processing, etc.

[0052] Step 102: Based on the first segmentation object, perform sentence segmentation on the recognized text to obtain the first segmented text.

[0053] When extracting target information from the recognized text, the recognized text may be intercepted before and after to form multiple truncated contents. Specifically, the recognized text may be split into multiple truncated contents by performing sentence segmentation on the recognized text through the first segmentation object, and the first segmented text obtained includes multiple truncated contents.

[0054] In practical applications, different first segmentation objects may be determined according to different types of target files, and different first segmentation objects may also be used for sentence segmentation of the same type of target file. Different first segmentation objects perform sentence segmentation on the recognized text to obtain different first segmented texts.

[0055] In a possible implementation, performing sentence segmentation on the recognized text based on the first segmentation object to obtain the first segmented text includes:

[0056] Obtain the first segmentation object; wherein, the first segmentation object includes one or more first sentence segmentation symbols;

[0057] Select the first sentence segmentation symbol ranked first, and perform sentence segmentation on the recognized text based on the selected symbol to obtain the first segmented sentence;

[0058] If not all the first sentence segmentation symbols are selected, select the next first sentence segmentation symbol, and then return to perform sentence segmentation on the recognized text based on the selected symbol to obtain the first segmented sentence;

[0059] If all the first sentence segmentation symbols are selected, obtain the first segmented text.

[0060] The first segmentation object can be one first sentence segmentation symbol or multiple first sentence segmentation symbols. The first sentence segmentation symbol can include but is not limited to line break, colon, comma, period, exclamation mark, question mark, etc.

[0061] When the first segmentation object is one first sentence segmentation symbol, such as a line break, the recognized text can be segmented by the line break to obtain multiple sentences with one paragraph as one first segmented sentence, that is, the first segmented text is obtained.

[0062] When the first segmentation object is multiple first sentence segmentation symbols, the multiple first sentence segmentation symbols are arranged in order. For example, two sequentially arranged first sentence segmentation symbols are a line break and a period. First, select the line break symbol to perform sentence segmentation on the recognized text to obtain multiple sentences with one paragraph as one first segmented sentence, and then select the period to perform sentence segmentation on the first segmented sentence again to obtain multiple sentences with the content between the line break and the comma as one first segmented sentence, or the content between two commas as one first segmented sentence, or the content between the comma and the line break as one first segmented sentence, that is, the first segmented text is obtained.

[0063] Step 103: Search for the target keyword from the first segmented text, and based on the keyword search range corresponding to the target keyword, filter and obtain the sentences containing the target keyword from the first segmented text.

[0064] The target keyword(s) can be one or more. When there are multiple target keywords, they can be searched for simultaneously or one by one. The target keyword(s) can be determined according to the content of the target document. If the target document is a guarantee letter document, the target keyword(s) can be the guarantee letter number, issue date, beneficiary, etc. The keyword search range can be the upper and lower paragraphs, the first page, the last page, the full text, etc. The keyword search ranges corresponding to different target keywords can be the same or different. If the target keyword is the guarantee letter number, the keyword search range can be the first page. If the target keyword is the issue date, the keyword search range can be the last page. If the target keyword is the beneficiary, the keyword search range can be the upper and lower paragraphs or the full text. Among them, if the keyword search range is the upper and lower paragraphs, the specific search range is the sentences containing the target keyword among the multiple sentences obtained based on the first segmentation object through step 102, as well as the previous one or more sentences and the next one or more sentences of the sentence containing the target keyword. In practical applications, if the first segmentation object is a line break, for the target keyword of beneficiary, the keyword search range is the paragraph containing the beneficiary and the previous and next paragraphs of this paragraph.

[0065] In one possible implementation, after searching for the target keyword(s) from the first segmented text, it further includes:

[0066] Determine whether the found target keyword(s) contain non-keywords;

[0067] After deleting the target keyword(s) containing non-keywords, obtain the processed keyword(s);

[0068] Annotate the processed keyword(s).

[0069] Non-keyword refers to a word that contains the target keyword but is not a non-keyword. For example, if the target keyword is the guarantor, the non-keyword can be the guaranteed. If the target keyword is the surety, and in the first segmented text, in addition to the surety, there is a counter-surety, and the counter-surety is not the target keyword to be searched, then it can be determined whether the found target keyword(s) contain non-keywords, that is, whether there is a counter-surety. After deleting the target keyword(s) containing non-keywords, obtain the keyword(s) only containing the surety, and this keyword(s) can be annotated. When annotating, one or more methods such as color, highlighting, underlining, bolding, etc. can be used to prominently mark the processed keyword(s).

[0070] In one possible implementation, based on the keyword search range corresponding to the target keyword(s), filter out the sentences containing the target keyword(s) from the first segmented text, including:

[0071] If the keyword search scope is the content of the target page of the target document, then select the sentences with the page number of the target page in the first split text as the sentences containing the target keyword;

[0072] If the keyword search scope is the content of the target segment related to the target keyword, then screen out the target segment from the first split text as the sentences containing the target keyword.

[0073] The keyword search scope is the target page of the target document, and the target page can be the first page, the last page, multiple consecutive pages or non-consecutive pages, the full text, etc. The first split text has a page number identifier, and the sentences of the first split text with the page number identifier being the page number of the target page can be selected as the sentences containing the target keyword. If the target page is the first page, the sentences of the target keyword are all the sentences on the first page.

[0074] The keyword search scope is the content of the target segment related to the target keyword. The target segment related to the target keyword can be the paragraph where the target keyword is located. In addition to the paragraph where the target keyword is located, the target segment can also include the adjacent paragraphs of the paragraph where the target keyword is located, such as the paragraph above and the paragraph below the paragraph where the target keyword is located. The selected target segment can be used as the sentences containing the target keyword, that is, in practical applications, the paragraph where the target keyword is located and the paragraph above and the paragraph below the paragraph where the target keyword is located can be selected. Among them, the division of paragraphs is based on the different sentences obtained by splitting the recognized text in step 102 as one paragraph.

[0075] For example, if the target document is a guarantee letter document, the target keyword is the guarantee letter number or number, and the keyword search scope is the first page, and the first split object is a line break, then the first page of the guarantee letter document will be truncated into specific sentences by line break, and the sentences with the guarantee letter number or number will be the sentences containing the target keyword.

[0076] In a possible implementation, after screening and obtaining the sentences containing the target keyword from the first split text, the method further includes:

[0077] Remove the target symbols in the sentences containing the target keyword to obtain the first processed sentence;

[0078] Replace the content in the first processed sentence that conforms to the regular expression to obtain the second processed sentence.

[0079] To prevent inaccurate subsequent statement splitting due to the presence of target symbols in the statements of target keywords, the target symbols in the statements containing target keywords can be removed. The target symbols can be line breaks, parentheses, etc. For example, if the bidder is a company and there is a line break due to the long company name, the line break character can be removed at this time to ensure the integrity of the company name. Another example is that the bidder is a company (XX address), and the content within the parentheses is not the company name, then the parentheses can be removed, and the content within the target symbol can also be removed, that is, (XX address) is removed entirely to obtain the first processed statement.

[0080] If the first processed statement contains an amount, and the amount requires specific numbers, the content that conforms to the regular expression can be "ten thousand". If the content of the first processed statement is "the insured amount is 60,000", then "ten thousand" can be replaced with 4 zeros, and the processed statement is "the insured amount is 60000". Of course, the content that conforms to the regular expression can also be set according to different target files, different target keywords, and different target matching words. This application does not make specific restrictions on this.

[0081] Step 104: Select the statements containing the target matching words from the statements containing the target keywords, and perform statement splitting on the statements containing the target matching words based on the second splitting object to obtain the second split text.

[0082] After obtaining the statements containing the target keywords, judge the target matching words in the statements containing the target keywords, and select the statements containing the target matching words.

[0083] In a possible implementation, performing statement splitting on the statements containing the target matching words based on the second splitting object to obtain the second split text includes:

[0084] Obtain the second splitting object; where the second splitting object includes truncated phrases, and the truncated phrases include two second statement splitting symbols;

[0085] Intercept the content between the two second statement splitting symbols from the statements containing the target matching words to obtain the second split text.

[0086] The second object to be segmented includes truncated phrases, and each truncated phrase includes two second sentence segmentation symbols, which may be the same or different. The second sentence segmentation symbol may include, but is not limited to, colon, line break, comma, period, parentheses, quotation marks, dash, exclamation mark, question mark, semicolon, etc. The second object to be segmented may be partially or completely different from the first object to be segmented. The number of truncated phrases may be one or more. When there is one truncated phrase, the sentence containing the target matching word is segmented by the two second sentence segmentation symbols. When there are multiple truncated phrases, sentence segmentation can be performed in the order of arrangement of the truncated phrases. First, use one truncated phrase for sentence segmentation, and then use the next truncated phrase to perform sentence segmentation again on the segmented sentence until all truncated phrases complete sentence segmentation, thereby obtaining the second segmented text.

[0087] For example, the sentence containing the target matching word is "Beneficiary: XXX Co., Ltd.", the second object to be segmented includes one truncated phrase, and this truncated phrase includes a colon and a line break. The target matching word is "company", so the content between the colon and the line break is intercepted as "XXX Co., Ltd.", and this "XXX Co., Ltd." is the second segmented text.

[0088] If the content intercepted between the two second sentence segmentation symbols from the sentence containing the target matching word is multiple, the content with a value and the shortest can be selected as the second segmented text.

[0089] Step 105: Annotate the second segmented text.

[0090] Annotate the second segmented text. When annotating, one or more methods such as color, highlighting, underlining, bolding, etc. can be used to highlight the second segmented text. If the processed keyword is also annotated while annotating the second segmented text, different annotation methods or different colors of the same annotation method can be used for distinction. Through classified highlighting, the reading efficiency and understanding depth of users can be improved.

[0091] After annotating the second segmented text, the target file before annotation, the target file after annotation can be viewed. The file after annotation can also be proofread and modified. The target file before annotation and the target file after annotation can be downloaded, and the results can be extracted and exported in formats such as Excel, CSV, etc., which is convenient for users' subsequent processing and analysis.

[0092] It should be noted that for the file content annotation method provided in this application, target keywords, target matching words, the first object to be segmented, the second object to be segmented, etc. can also be added to meet the continuously changing annotation requirements of users. Security measures such as permission management and data encryption can also be provided to ensure the security and privacy of user data.

[0093] A file content annotation method provided by an embodiment of the present application has the following advantages:

[0094] 1. High flexibility: Users can flexibly configure target keywords, target matching words, first segmentation objects, and second segmentation objects to meet different content annotation requirements.

[0095] 2. High accuracy: Ensure the accuracy of the extracted content through OCR technology.

[0096] 3. Strong adaptability: It can handle documents with various situations such as non-keyword content, target symbols, and complex formats.

[0097] 4. Good user experience: The annotated file content can be presented on the user interface, and the rich result display method reduces the learning cost and usage difficulty of users.

[0098] 5. Strong scalability and security: Support the addition of new configurations and security measures to ensure the sustainable development of the platform and the security of user data.

[0099] The above introduces a file content annotation method provided by an embodiment of the present application. Next, a system for executing the above file content annotation method will be introduced.

[0100] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a file content annotation system provided by an embodiment of the present application. As Figure 2 shown, the file content annotation system includes:

[0101] An identification module 201, configured to perform text recognition on a target file to obtain the recognized text.

[0102] A first segmentation module 202, configured to perform sentence segmentation on the recognized text based on a first segmentation object to obtain first segmented text.

[0103] A screening module 203, configured to search for target keywords from the first segmented text, and screen out sentences containing the target keywords from the first segmented text based on a keyword search range corresponding to the target keywords.

[0104] A second segmentation module 204, configured to select sentences containing target matching words from the sentences containing the target keywords, and perform sentence segmentation on the sentences containing the target matching words based on a second segmentation object to obtain second segmented text.

[0105] A marking module 205, configured to mark the second segmented text.

[0106] In a possible implementation, the first segmentation module 202 is specifically configured to:

[0107] Obtain a first segmentation object; wherein, the first segmentation object includes one or more first statement segmentation symbols;

[0108] Select the first statement segmentation symbol arranged at the first position, and perform statement segmentation on the recognized text based on the selected symbol to obtain a first segmented statement;

[0109] In the case where not all the first statement segmentation symbols are selected, select the next first statement segmentation symbol, and then return to perform statement segmentation on the recognized text based on the selected symbol to obtain a first segmented statement;

[0110] In the case where all the first statement segmentation symbols are selected, obtain a first segmented text.

[0111] In a possible implementation, the screening module 203 is specifically configured to:

[0112] If the keyword search range is the content of the target page of the target file, then select the statements whose statement page numbers are the page numbers of the target page from the first segmented text as the statements containing the target keyword;

[0113] If the keyword search range is the content of the target segment related to the target keyword, then screen out the target segment from the first segmented text as the statements containing the target keyword.

[0114] In a possible implementation, the second segmentation module 204 is specifically configured to:

[0115] Obtain a second segmentation object; wherein, the second segmentation object includes truncated phrases, and the truncated phrases include two second statement segmentation symbols;

[0116] Intercept the content between the two second statement segmentation symbols from the statements containing the target matching word to obtain a second segmented text.

[0117] In a possible implementation, the file content annotation system further includes:

[0118] A non-keyword processing module, configured to determine whether the found target keyword contains a non-keyword; after deleting the target keyword containing the non-keyword, obtain a processed keyword; and perform annotation on the processed keyword.

[0119] In a possible implementation, the file content annotation system further includes:

[0120] A deletion module, configured to remove the target symbol from the statement containing the target keyword to obtain a first processed statement;

[0121] A replacement module, configured to replace the content that conforms to the regular expression in the first processed statement to obtain a second processed statement.

[0122] In an embodiment of the present application, an electronic device is further provided. Refer to Figure 3 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 3 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiment of the present application.

[0123] As Figure 3 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0124] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a memory card, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.

[0125] The electronic device can implement the above-mentioned file content annotation method.

[0126] In an embodiment of the present application, a computer program product is further provided, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement any one of the file content annotation methods provided in the embodiment of the present application.

[0127] In an embodiment of the present application, a computer-readable storage medium is further provided. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can be enabled to implement any one of the file content annotation methods provided in the embodiment of the present application.

[0128] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software program implementation is a better implementation method in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0130] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0131] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. A method for marking file content, characterized in that: include: Perform text recognition on the target file to obtain the recognized text; Based on the first segmentation object, performing sentence segmentation on the recognized text to obtain a first segmented text; Searching for a target keyword from the first segmented text, and based on a keyword search range corresponding to the target keyword, filtering out sentences containing the target keyword from the first segmented text; Selecting a sentence containing a target matching word from the sentences containing the target keyword, and performing sentence segmentation on the sentence containing the target matching word based on the second segmentation object to obtain a second segmented text; The second segmented text is marked.

2. The file content marking method according to claim 1, characterized in that: The step of performing sentence segmentation on the recognized text based on the first segmentation object to obtain the first segmented text includes: Acquire a first segmentation object; wherein the first segmentation object includes one or more first sentence segmentation symbols; Selecting a first sentence segmentation symbol that is arranged first, and performing sentence segmentation on the recognized text based on the selected symbol to obtain a first segmented sentence; In the case that not all the first sentence segmentation symbols are selected, selecting the next first sentence segmentation symbol, and then returning to the method of performing sentence segmentation on the recognized text based on the selected symbol to obtain the first segmented sentence; When all the first sentence segmentation symbols are selected, the first segmented text is obtained.

3. The file content marking method according to claim 1, characterized in that: The step of filtering out sentences containing the target keyword from the first segmented text based on the keyword search range corresponding to the target keyword includes: If the keyword search range is the content of the target page of the target file, a sentence whose page number is the page number of the target page is selected from the first segmented text as a sentence containing the target keyword; If the keyword search range is the content of the target segment related to the target keyword, the target segment is screened out from the first segmented text as a sentence containing the target keyword.

4. The file content marking method according to claim 1, characterized in that: The step of performing sentence segmentation on the sentence containing the target matching word based on the second segmentation object to obtain a second segmented text includes: Acquire a second segmentation object; wherein the second segmentation object includes a truncated phrase, and the truncated phrase includes two second sentence segmentation symbols; The content between two second sentence segmentation symbols is intercepted from the sentence containing the target matching word to obtain the second segmented text.

5. The file content marking method according to any one of claims 1 to 4, characterized in that: After searching the target keyword from the first segmented text, the method further includes: Determine whether the searched target keywords contain non-keywords; After deleting the target keywords containing non-keywords, the processed keywords are obtained; The processed keywords are marked.

6. The file content marking method according to any one of claims 1 to 4, characterized in that: After selecting the sentences containing the target keyword from the first segmented text and before selecting the sentences containing the target matching words from the sentences containing the target keyword, the method further includes: Removing the target symbol from the sentence containing the target keyword to obtain a first processing sentence; The content that conforms to the regular expression in the first processing statement is replaced to obtain a second processing statement.

7. A file content annotation system, characterized in that: include: A recognition module is used to perform text recognition on the target file to obtain the recognized text; A first segmentation module, configured to perform sentence segmentation on the recognized text based on a first segmentation object to obtain a first segmented text; A screening module, configured to search for a target keyword from the first segmented text, and based on a keyword search range corresponding to the target keyword, screen the first segmented text to obtain a sentence containing the target keyword; A second segmentation module is used to select a sentence containing a target matching word from the sentence containing the target keyword, and perform sentence segmentation on the sentence containing the target matching word based on a second segmentation object to obtain a second segmented text; A marking module is used to mark the second segmented text.

8. A computer program product, characterized in that It includes computer-readable instructions, and when the computer-readable instructions are executed on an electronic device, the electronic device implements the file content marking method as described in any one of claims 1 to 6.

9. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the file content marking method as described in any one of claims 1 to 6.

10. A computer storage medium, characterized in that: The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the file content marking method as described in any one of claims 1 to 6.