Document processing method and apparatus

CN115495556BActive Publication Date: 2026-09-04GUANGZHOU KINGSOFT MOBILE TECH +3
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211195167.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2026-09-04
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

[0003]然而,目前一般是基于人工实现对单个或多个文档的分析处理,分析效率较低

Benefits of technology

[0073]The technical solution provided by this invention extracts a set of scene keywords that match a target scenario from at least one document, along with the scene text corresponding to each scene keyword in the set. It then determines at least one analysis dimension for the target scenario. Based on the scene text corresponding to each scene keyword in the set, it performs dimensional analysis on the scene keyword set from at least one analysis dimension to obtain the dimensional analysis results. Finally, it determines the data to be displayed based on the dimensional analysis results. This achieves intelligent extraction of scene-related text from one or more documents and analyzes it to obtain corresponding analysis results, thereby improving the efficiency of document analysis and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115495556B_ABST
    Figure CN115495556B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to a document processing method and device, comprising: extracting a scene keyword set conforming to a target scene and scene text corresponding to each scene keyword in the scene keyword set from at least one document; determining at least one analysis dimension for the target scene; performing dimension analysis on the scene keyword set from the at least one analysis dimension based on the scene text corresponding to each scene keyword in the scene keyword set, to obtain a dimension analysis result corresponding to the scene keyword set; and determining display data based on the dimension analysis result corresponding to the scene keyword set. Thus, the scene-related text in a single or multiple documents is intelligently extracted and analyzed to obtain corresponding analysis results, thereby improving the efficiency of analyzing and processing the documents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a document processing method and apparatus. Background Technology

[0002] In the daily office work of companies or individuals, a large number of documents (such as contracts, archives, etc.) are generated. In some application scenarios, users can perform analysis and processing based on single or multiple documents, such as analyzing the transaction situation within a certain time period based on multiple contracts within that time period.

[0003] However, currently, the analysis and processing of one or more documents is generally done manually, which results in low analysis efficiency. Summary of the Invention

[0004] In view of this, in order to improve the efficiency of document analysis and processing, embodiments of the present invention provide a document processing method and apparatus.

[0005] In a first aspect, embodiments of the present invention provide a document processing method, including:

[0006] Extract a set of scene keywords that match the target scenario from at least one document, and the scene text corresponding to each scene keyword in the set of scene keywords;

[0007] For the target scenario, at least one analysis dimension shall be determined;

[0008] Based on the scene text corresponding to each scene keyword in the scene keyword set, dimensional analysis is performed on the scene keyword set from at least one analysis dimension to obtain the dimensional analysis result corresponding to the scene keyword set.

[0009] The data to be displayed is determined based on the dimensional analysis results corresponding to the set of scenario keywords.

[0010] In one possible implementation, the scene text corresponding to each scene keyword is extracted from at least one document in the following manner:

[0011] For each of the aforementioned scenario keywords, the following processing is performed:

[0012] Extract the contextual text content of the scene keywords from at least one of the documents;

[0013] Based on the scene keywords and the contextual text content of the scene keywords, construct the scene text corresponding to the scene keywords.

[0014] In one possible implementation, extracting the contextual text content of the scene keywords from at least one of the documents includes:

[0015] The clause containing the scene keyword is determined as the context text content of the scene keyword; and / or,

[0016] The paragraph containing the scene keyword is defined as the context text content of the scene keyword.

[0017] In one possible implementation, constructing the scene text corresponding to the scene keyword based on the scene keyword and the contextual text content of the scene keyword includes:

[0018] When the scenario keyword is a number, extract the unit of measurement corresponding to the scenario keyword from at least one of the documents;

[0019] The scene keywords and their corresponding units of measurement are concatenated to obtain the concatenated text;

[0020] The scene text corresponding to the scene keyword is constructed based on the concatenated text and the contextual text content of the scene keyword.

[0021] In one possible implementation, the analysis dimensions include at least two; the step of performing dimensional analysis on the scene keyword set from the at least one analysis dimension based on the scene text corresponding to each scene keyword in the scene keyword set, to obtain the dimensional analysis result corresponding to the scene keyword set, includes:

[0022] Based on at least two of the aforementioned analytical dimensions, the scene text corresponding to each scene keyword in the scene keyword set is classified to obtain at least two scene text classes;

[0023] The at least two scene text classes are determined as the dimensional analysis results corresponding to the scene keyword set;

[0024] Different scenario text classes correspond to different analysis dimensions.

[0025] In one possible implementation, determining the display data based on the dimensional analysis results corresponding to the set of scene keywords includes:

[0026] Define the target analysis dimensions;

[0027] Determine the target scene text class corresponding to the target analysis dimension from the at least two scene text classes;

[0028] The display data is determined based on the text class of the target scene.

[0029] In one possible implementation, determining the display data based on the target scene text class includes:

[0030] For each of the target scene text classes, the following processing is performed:

[0031] The multiple scene texts in the target scene text class are sorted according to a preset sorting method to obtain the display data corresponding to the target scene text class.

[0032] In one possible implementation, the analysis dimensions include at least two; the step of performing dimensional analysis on the scene keyword set from the at least one analysis dimension based on the scene text corresponding to each scene keyword in the scene keyword set, to obtain the dimensional analysis result corresponding to the scene keyword set, includes:

[0033] For each of the scene keywords, the following processing is performed: extract the dimension analysis results corresponding to each of the analysis dimensions from the scene text corresponding to the scene keyword to obtain the multi-dimensional analysis results corresponding to the scene keyword;

[0034] The multi-dimensional analysis results corresponding to each of the scenario keywords are determined as the dimensional analysis results corresponding to the scenario keyword set.

[0035] In one possible implementation, determining the display data based on the dimensional analysis results corresponding to the set of scene keywords includes:

[0036] The dimensional analysis results corresponding to the set of scene keywords are analyzed according to the preset analysis strategy to obtain the target analysis results;

[0037] The displayed data is determined based on the target analysis results and the dimensional analysis results.

[0038] In a second aspect, embodiments of the present invention provide a document processing apparatus, comprising:

[0039] The extraction module is used to extract a set of scene keywords that match the target scenario and the scene text corresponding to each scene keyword in the set of scene keywords from at least one document;

[0040] The first determining module is used to determine at least one analysis dimension for the target scenario;

[0041] The analysis module is used to perform dimensional analysis on the scene keyword set from at least one analysis dimension based on the scene text corresponding to each scene keyword in the scene keyword set, and obtain the dimensional analysis result corresponding to the scene keyword set.

[0042] The second determining module is used to determine the display data based on the dimensional analysis results corresponding to the set of scene keywords.

[0043] In one possible implementation, the extraction module is specifically used for:

[0044] For each of the aforementioned scenario keywords, the following processing is performed:

[0045] Extract the contextual text content of the scene keywords from at least one of the documents;

[0046] Based on the scene keywords and the contextual text content of the scene keywords, construct the scene text corresponding to the scene keywords.

[0047] In one possible implementation, the extraction module is further configured to:

[0048] The clause containing the scene keyword is determined as the context text content of the scene keyword; and / or,

[0049] The paragraph containing the scene keyword is defined as the context text content of the scene keyword.

[0050] In one possible implementation, the extraction module is further configured to:

[0051] When the scenario keyword is a number, extract the unit of measurement corresponding to the scenario keyword from at least one of the documents;

[0052] The scene keywords and their corresponding units of measurement are concatenated to obtain the concatenated text;

[0053] The scene text corresponding to the scene keyword is constructed based on the concatenated text and the contextual text content of the scene keyword.

[0054] In one possible implementation, the analysis module is specifically used for:

[0055] Based on at least two of the aforementioned analytical dimensions, the scene text corresponding to each scene keyword in the scene keyword set is classified to obtain at least two scene text classes;

[0056] The at least two scene text classes are determined as the dimensional analysis results corresponding to the scene keyword set;

[0057] Different scenario text classes correspond to different analysis dimensions.

[0058] In one possible implementation, the second determining module is specifically used for:

[0059] Define the target analysis dimensions;

[0060] Determine the target scene text class corresponding to the target analysis dimension from the at least two scene text classes;

[0061] The display data is determined based on the text class of the target scene.

[0062] In one possible implementation, the second determining module is further configured to:

[0063] For each of the target scene text classes, the following processing is performed:

[0064] The multiple scene texts in the target scene text class are sorted according to a preset sorting method to obtain the display data corresponding to the target scene text class.

[0065] In one possible implementation, the analysis module is further configured to:

[0066] For each of the scene keywords, the following processing is performed: extract the dimension analysis results corresponding to each of the analysis dimensions from the scene text corresponding to the scene keyword to obtain the multi-dimensional analysis results corresponding to the scene keyword;

[0067] The multi-dimensional analysis results corresponding to each of the scenario keywords are determined as the dimensional analysis results corresponding to the scenario keyword set.

[0068] In one possible implementation, the second determining module is further configured to:

[0069] The dimensional analysis results corresponding to the set of scene keywords are analyzed according to the preset analysis strategy to obtain the target analysis results;

[0070] The displayed data is determined based on the target analysis results and the dimensional analysis results.

[0071] Thirdly, embodiments of the present invention provide an electronic device, including: a processor and a memory, wherein the processor is configured to execute a document processing program stored in the memory to implement the document processing method described in any one of the first aspects.

[0072] Fourthly, embodiments of the present invention provide a storage medium storing one or more programs, which can be executed by one or more processors to implement the document processing method described in any one aspect.

[0073] The technical solution provided by this invention extracts a set of scene keywords that match a target scenario from at least one document, along with the scene text corresponding to each scene keyword in the set. It then determines at least one analysis dimension for the target scenario. Based on the scene text corresponding to each scene keyword in the set, it performs dimensional analysis on the scene keyword set from at least one analysis dimension to obtain the dimensional analysis results. Finally, it determines the data to be displayed based on the dimensional analysis results. This achieves intelligent extraction of scene-related text from one or more documents and analyzes it to obtain corresponding analysis results, thereby improving the efficiency of document analysis and processing. Attached Figure Description

[0074] Figure 1 A flowchart illustrating an embodiment of a document processing method provided by this invention;

[0075] Figure 2 A flowchart illustrating an embodiment of another document processing method provided by the present invention;

[0076] Figure 3 This is an example of scene text in a digital scene provided in an embodiment of the present invention;

[0077] Figure 4 A flowchart illustrating another embodiment of a document processing method provided by the present invention;

[0078] Figure 5 This is an example of a multi-dimensional analysis result provided in an embodiment of the present invention;

[0079] Figure 6 A flowchart illustrating another embodiment of a document processing method provided by the present invention;

[0080] Figure 7 A block diagram illustrating an embodiment of a document processing apparatus provided by the present invention;

[0081] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0083] To facilitate understanding of the embodiments of the present invention, the following provides an exemplary description of the application scenarios involved in the document processing method provided by the embodiments of the present invention.

[0084] In one exemplary application scenario, a user needs to extract text from numerous documents containing similar scenarios and perform analysis based on this text. For example, they might need to extract data such as the amount and due date of all contracts from 100 contracts to make a budget estimate. According to existing technology, this requires the user to manually open each document, extract this data from each document, and then perform the subsequent analysis. This undoubtedly requires a significant investment of time and manpower, especially when the number of documents is too large, such as 100,000 or 1 million documents. In such cases, the user would incur enormous time and manpower costs, and might even struggle to achieve the final analysis goal.

[0085] Therefore, embodiments of the present invention provide a document processing method. Applying this method, it is possible to intelligently extract scene-related text from single or multiple documents and perform intelligent analysis on it.

[0086] The document processing method provided by the present invention will be explained and described below with reference to the accompanying drawings and specific embodiments. The embodiments do not constitute a limitation on the embodiments of the present invention.

[0087] See Figure 1 This is a flowchart illustrating an embodiment of a document processing method provided by the present invention. Figure 1 As shown, the process may include the following steps:

[0088] Step 101: Extract a set of scene keywords that match the target scenario from at least one document, and the scene text corresponding to each scene keyword in the set of scene keywords.

[0089] The above documents can be in formats such as .doc, .docx, .xlsx, .xls, .ppt, .txt, etc.

[0090] The aforementioned target scenario refers to the document analysis scenario in which the user is located when the document processing method provided in this embodiment of the invention is used. For example, if the user's document analysis scenario involves analyzing numbers in a document, then the aforementioned target scenario is a number scenario. Of course, in practical applications, other document analysis scenarios may exist besides number scenarios, such as name scenarios (analyzing names in a document), place name scenarios (analyzing place names in a document), and organizational structure scenarios (analyzing organizational structures in a document), etc., without specific limitations here.

[0091] In one embodiment, before executing the document processing method provided by the present invention, the user may set a target scenario, and then execute the document processing method provided by the present invention for the target scenario.

[0092] The aforementioned scenario keywords refer to the text content in the document that matches the target scenario. For example, if the target scenario is a number scenario, then the numbers in the document are scenario keywords; similarly, if the target scenario is a name scenario, then the names in the document are scenario keywords.

[0093] In one embodiment, the specific implementation of extracting scene keywords that match the target scene from at least one document may include: inputting at least one document into a pre-trained scene keyword extraction model to obtain scene keywords in the document that match the target scene.

[0094] Optionally, the aforementioned scene keyword extraction model can be used to extract scene keywords that match a specific scene. In other words, a separate scene keyword extraction model can be trained for each different scene to extract scene keywords that match that scene. Therefore, inputting the document into a pre-trained scene keyword extraction model means inputting the document into the scene keyword extraction model corresponding to the target scene.

[0095] Optionally, the above-mentioned scene keyword extraction model can be used to extract scene keywords that fit different scenes. In other words, a unified scene keyword extraction model can be trained for different scenes to extract scene keywords that fit different scenes.

[0096] Scene text corresponding to scene keywords refers to text related to those keywords. For example, text containing scene keywords, or annotation text corresponding to scene keywords. How the scene text corresponding to scene keywords is extracted will be explained below. Figure 2 The illustrated embodiments are provided as examples and will not be described in detail here.

[0097] Step 102: Determine at least one analysis dimension for the target scenario.

[0098] The aforementioned analytical dimensions refer to the dimensions used for analysis based on the target scenario. For example, the analytical dimensions corresponding to a "digital scenario" may include, but are not limited to: time dimension, date dimension, percentage dimension, monetary dimension, unit dimension, and count dimension; similarly, the analytical dimensions corresponding to a name scenario may include, but are not limited to: gender dimension and occupation dimension.

[0099] In one embodiment, different scenarios and analysis dimensions can be pre-configured, wherein each scenario can correspond to at least one analysis dimension. A specific implementation of determining at least one analysis dimension for a target scenario may include: determining at least one analysis dimension corresponding to the target scenario according to the pre-set scenario-analysis dimension correspondence.

[0100] In another embodiment, a machine learning model that can automatically match analysis dimensions for a scene can be pre-trained. Based on this, the specific implementation of determining at least one analysis dimension for a target scene may include: inputting scene information of the target scene into the pre-trained machine learning model, and having the machine learning model output the analysis dimension that matches the target scene.

[0101] The embodiments of the present invention do not impose specific limitations on the model structure of the above-described machine learning model.

[0102] Step 103: Based on the scene text corresponding to each scene keyword in the scene keyword set, perform dimensional analysis on the scene keyword set from at least one analysis dimension to obtain the dimensional analysis results corresponding to the scene keyword set.

[0103] As can be seen from the description of step 103, the document processing method provided by the embodiments of the present invention can analyze multiple scene keywords from at least one dimension to obtain the dimension analysis results corresponding to the scene keyword set.

[0104] Step 104: Determine the display data based on the dimensional analysis results corresponding to the set of scenario keywords.

[0105] The data displayed above can be all or part of the dimensional analysis results of the scenario keyword set obtained in step 103, or data obtained after further analysis and processing for display. By showing the above data to users, they can easily understand the analysis of the scenario keywords.

[0106] In one embodiment, the specific implementation of further analyzing and processing the dimensional analysis results of the scene keyword set to obtain the display data may include: analyzing the dimensional analysis results corresponding to the scene keyword set according to a preset analysis strategy to obtain the target analysis results, and determining the display data based on the target analysis results and the dimensional analysis results. The aforementioned analysis strategies may include, but are not limited to, cross-analysis, summation, and averaging.

[0107] In this embodiment, the dimensional analysis results corresponding to the set of scene keywords are further analyzed according to a preset analysis strategy to obtain the corresponding target analysis results. Then, the target analysis results and the dimensional analysis results corresponding to the set of scene keywords are integrated to obtain the display data, which can help users to deeply analyze scene keywords and improve user experience.

[0108] As for how to perform dimensional analysis on the set of scene keywords from at least one analytical dimension to obtain the dimensional analysis results of the set of scene keywords, and how to determine the display data based on the dimensional analysis results, these will be explained in the following text through different embodiments, and will not be detailed here.

[0109] The technical solution provided by this invention extracts a set of scene keywords that match a target scenario from at least one document, along with the scene text corresponding to each scene keyword in the set. It then determines at least one analysis dimension for the target scenario. Based on the scene text corresponding to each scene keyword in the set, it performs dimensional analysis on the set of scene keywords from at least one analysis dimension to obtain the dimensional analysis results. Finally, it determines the data to be displayed based on the dimensional analysis results. This achieves intelligent extraction of scene-related text from one or more documents and analyzes it to obtain corresponding analysis results, thereby improving the efficiency of document analysis and processing.

[0110] See Figure 2 This is a flowchart illustrating an embodiment of another document processing method provided by the present invention. Figure 2 The process shown above Figure 1 Based on the illustrated process, describe how to extract the scene text corresponding to each scene keyword from at least one document. For example... Figure 2 As shown, the process may include the following steps:

[0111] Step 201: For each scene keyword, extract the contextual text content of the scene keyword from at least one document.

[0112] Optionally, the contextual text of the scene keyword refers to the clause containing the scene keyword, or the paragraph containing the scene keyword, or the contextual text of the scene keyword includes both the clause and the paragraph containing the scene keyword.

[0113] Based on this, a specific implementation of extracting the contextual text content of scene keywords from at least one document may include: determining the clause containing the scene keyword as the contextual text content of the scene keyword; or, determining the paragraph containing the scene keyword as the contextual text content of the scene keyword; or, determining both the clause containing the scene keyword and the paragraph containing the scene keyword as the contextual text content of the scene keyword.

[0114] Step 202: Based on the scene keywords and the contextual text content of the scene keywords, construct the scene text corresponding to the scene keywords.

[0115] In one embodiment, scene keywords and their contextual text content can be used as the scene text corresponding to the scene keywords.

[0116] For example, when the scene keyword is a name (or place name), the context text content corresponding to the name (or place name) is directly used as the corresponding scene text.

[0117] Furthermore, in digital scenarios, since numbers usually have corresponding units of measurement, and these units of measurement can represent the meaning of the numbers, in digital scenarios, that is, when the scenario keyword is a number, the unit of measurement corresponding to the scenario keyword of the number type can be extracted from at least one document. Then, the scenario keyword and the unit of measurement corresponding to the scenario keyword are concatenated to obtain the concatenated text. This concatenated text and the context text content of the scenario keyword are used as the scenario text corresponding to the scenario keyword.

[0118] For example, if the scenario keyword is "20," the corresponding unit of measurement is "days," and the corresponding context text is "installation time 20 days," then the concatenated text would be "20 days," and the corresponding scenario text would be: "20 days" + "installation time 20 days." See also... Figure 3 The image shown is an example of scene text in a digital context.

[0119] pass Figure 2 The process shown extracts the contextual text content of each scenario keyword from at least one document. Based on the scenario keyword and its contextual text content, the scenario text corresponding to the scenario keyword is constructed. This enables the intelligent extraction of the scenario text corresponding to the scenario keyword from at least one document, assisting in subsequent intelligent analysis of the scenario keyword.

[0120] See Figure 4 This is a flowchart illustrating an embodiment of another document processing method provided by the present invention. Figure 4 The process shown above Figure 1 Based on the illustrated process, this paper focuses on describing a method for performing multi-dimensional analysis of the scene keyword set from at least two analytical dimensions, based on the scene text corresponding to each scene keyword in the scene keyword set, to obtain the dimensional analysis results corresponding to the scene keyword set. For example... Figure 4 As shown, the process may include the following steps:

[0121] Step 401: Extract from at least one document a set of scene keywords that match the target scenario and the scene text corresponding to each scene keyword in the set of scene keywords.

[0122] Step 402: Determine at least two analysis dimensions for the target scenario.

[0123] Figure 4 In the illustrated embodiment, a scene can correspond to at least two analysis dimensions. This setup allows for multi-dimensional analysis of scene keywords from multiple different dimensions, yielding multi-dimensional analysis results for the scene keywords.

[0124] For further descriptions of steps 401 to 402, please refer to the relevant descriptions in the above embodiments, which will not be repeated here.

[0125] Step 403: Classify the scene text corresponding to each scene keyword in the scene keyword set based on at least two analysis dimensions to obtain at least two scene text classes, and determine the at least two scene text classes as the dimension analysis results corresponding to the scene keyword set; wherein, different scene text classes correspond to different analysis dimensions.

[0126] In one embodiment, the multi-dimensional analysis of multiple scene keywords based on at least two analysis dimensions and scene texts corresponding to multiple scene keywords to obtain the multi-dimensional analysis results of multiple scene keywords specifically includes: for each analysis dimension, searching for scene texts corresponding to that analysis dimension in all scene texts, and classifying the scene texts into the corresponding scene text classes, thus obtaining the scene text classes corresponding to each analysis dimension, and then determining all scene text classes as the dimension analysis results (multi-dimensional analysis results) corresponding to the set of scene keywords.

[0127] For example, see Figure 5 In order to Figure 3 This is an example of a multi-dimensional analysis result obtained by performing multi-dimensional analysis on the example scene text and scene keywords.

[0128] like Figure 5 As shown, the scenarios are divided according to the time dimension, date dimension, percentage dimension, and money dimension, resulting in scenario text classes for the time dimension, date dimension, percentage dimension, and money dimension, respectively.

[0129] Step 404: Determine the target analysis dimensions.

[0130] In one embodiment, the aforementioned target analysis dimension refers to the analysis dimensions desired by the user, and the number of these dimensions can be one or more. Specifically, the user can flexibly set the target analysis dimensions according to actual needs to meet user requirements.

[0131] In another embodiment, if the user does not set a target analysis dimension, all analysis dimensions corresponding to the target scenario can be determined as the target analysis dimensions. This allows for a comprehensive display of the analysis results corresponding to all analysis dimensions, making it convenient for the user to view.

[0132] In another embodiment, if the user does not set a target analysis dimension this time, the target analysis dimension set by the user in the past can be determined as the target analysis dimension of the current application. Alternatively, the target analysis dimension set by the user in the past can be determined as the target analysis dimension of the current application if it appears the most times or meets certain conditions.

[0133] The aforementioned conditions could be that the number of times is set to be greater than a certain threshold, or that the number of times is set to be ranked relatively high (e.g., ranked in the top N).

[0134] Step 405: Determine the target scene text class corresponding to the target analysis dimension from at least two scene text classes.

[0135] In this embodiment of the invention, the scene text class corresponding to the target analysis dimension is determined as the target scene text class. It can be understood that when there is only one target analysis dimension, there is also only one target scene text class; conversely, when there are multiple target analysis dimensions, there are also multiple target scene text classes.

[0136] For example, if the target analysis dimensions include the percentage dimension and the money dimension in the example above, then the target scene text class includes the scene text class corresponding to the percentage dimension and the scene text class corresponding to the money dimension.

[0137] Step 406: Determine the display data based on the text class of the target scene.

[0138] In one embodiment, the target scene text class can be directly determined as the display data.

[0139] In another embodiment, for each target scene text class, multiple scene texts within the target scene text class can be sorted according to a preset sorting method, and the sorted target scene text class is determined as the display data. For example, for scene text classes under the time dimension, the scene texts can be sorted according to chronological order; similarly, for scene text classes under the date dimension, the scene texts can be sorted according to the order of dates from front to back. In this way, multiple scene texts in the display data can be arranged according to a certain pattern, thereby facilitating user viewing.

[0140] In another embodiment, the target scene text can be analyzed according to a preset analysis strategy, and the analysis results can be used as the displayed data. Optionally, the analysis strategy can be summary analysis, cluster analysis, analogy analysis, etc. This not only facilitates users to perform in-depth analysis of documents.

[0141] Figure 4The process described involves classifying scene text corresponding to multiple scene keywords based on at least two analytical dimensions, resulting in at least two scene text classes. From these at least two scene text classes, the target scene text class corresponding to the target analytical dimension is determined, and the data to be displayed is determined based on the target scene text class. This effectively divides scene text corresponding to different analytical dimensions, making it easier for users to view the scene text corresponding to each analytical dimension.

[0142] See Figure 6 This is a flowchart illustrating another embodiment of a document processing method provided by the present invention. Figure 6 The process shown above Figure 1 Based on the illustrated process, this paper focuses on describing an implementation method that uses the scene text corresponding to each scene keyword in the scene keyword set to perform dimensional analysis on the scene keyword set from at least two analytical dimensions, thereby obtaining the dimensional analysis results corresponding to the scene keyword set. For example... Figure 6 As shown, the process may include the following steps:

[0143] Step 601: Extract from at least one document a set of scene keywords that match the target scenario and the scene text corresponding to each scene keyword in the set of scene keywords.

[0144] Step 602: Determine at least two analysis dimensions for the target scenario.

[0145] The descriptions of steps 601 to 602 above can be found in the relevant descriptions in the above embodiments, and will not be repeated here.

[0146] Step 603: For each scene keyword, extract the dimensional analysis results corresponding to each analysis dimension from the scene text corresponding to the scene keyword to obtain the multi-dimensional analysis results corresponding to the scene keyword, and determine the multi-dimensional analysis results corresponding to each scene keyword as the dimensional analysis results corresponding to the scene keyword set.

[0147] The dimensional analysis result for each of the aforementioned analysis dimensions refers to the content extracted from the scene text corresponding to the scene keywords based on that analysis dimension. In one embodiment, a content extraction model can be trained for each analysis dimension, and the content extraction model can be used to extract the dimensional analysis result corresponding to the respective analysis dimension from the scene text. Specifically, for each content extraction model corresponding to an analysis dimension, the scene text corresponding to each scene keyword is used as the input to the content extraction model, and the content extraction model outputs the corresponding dimensional analysis result.

[0148] The aforementioned content extraction model can be implemented based on natural language processing and semantic recognition technologies. As one possible implementation, the content extraction model corresponding to each analysis dimension can be trained through the following steps: For each analysis dimension, pre-collected scene text is used as training data, and manually labeled dimension analysis results corresponding to that analysis dimension are used as labels for the training data. Then, the initial model is trained using the aforementioned training data until the currently trained model converges. At this point, training stops, and the currently trained model is determined as the successfully trained content extraction model.

[0149] For each scenario keyword, the multi-dimensional analysis results are obtained by summarizing the dimensional analysis results of each analysis dimension corresponding to that scenario keyword. Furthermore, the multi-dimensional analysis results corresponding to all scenario keywords in the scenario keyword set are determined as the dimensional analysis results corresponding to the scenario keyword set.

[0150] Table 1 shows an example of how, when the scene keywords are numbers, the dimensional analysis results corresponding to each analysis dimension are extracted to obtain multi-dimensional analysis results, and then the dimensional analysis results corresponding to the scene keyword set are obtained:

[0151] Table 1

[0152]

[0153] In Table 1 above, the "Numbers" column corresponds to scenario keywords, while "Unit," "Sub-scenario," "Count," "Development Expectations," "Original Sentence," and "Use" represent six analysis dimensions. Specifically, the "Unit" column contains the dimensional analysis results for the "Unit" dimension; the "Sub-scenario" column contains the dimensional analysis results for the "Sub-scenario" dimension; the "Count" column contains the dimensional analysis results for the "Count" dimension; the "Development Expectations" column contains the dimensional analysis results for the "Development Expectations" dimension; the "Original Sentence" column contains the dimensional analysis results for the "Original Sentence" dimension; and the "Use" column contains the dimensional analysis results for the "Use" dimension.

[0154] As shown in Table 1, the set of scene keywords includes "3", "5", "10", "20", "18", and "180". For each scene keyword in this set, the dimensional analysis results of all dimensions in its corresponding row constitute the multi-dimensional analysis results for that scene keyword. Table 1 as a whole represents the dimensional analysis results for this set of scene keywords.

[0155] As can be seen from Table 1 above, the method provided by the embodiments of the present invention can perform multi-dimensional analysis of text. Compared with the prior art, which generally only analyzes based on a single dimension (for example, analyzing the number "20" in the text "I walked 20 kilometers today", according to the existing analysis method, only the analysis result "20 kilometers" of the "distance" dimension can be obtained), the multi-dimensional analysis results can help users understand the document content in a more comprehensive and in-depth way.

[0156] Step 604: Analyze the dimensional analysis results corresponding to the set of scene keywords according to the preset analysis strategy to obtain the target analysis results, and determine the display data based on the target analysis results and dimensional analysis results.

[0157] The above-mentioned analytical strategies may include, but are not limited to, cross-analysis, summation, and averaging.

[0158] After obtaining the dimensional analysis results shown in Table 1 above, assuming the analysis strategy is based on a cross-analysis of the two dimensions "Scenario" and "Development Expectations" in Table 1, we can obtain target analysis results such as "the more the better" and "the fewer the distance, the better." Furthermore, these target analysis results and dimensional analysis results are presented together as display data. That is, the displayed data includes not only the dimensional analysis results but also the target analysis results obtained from further analysis based on the dimensional analysis results (such as "the more the better" and "the fewer the distance, the better"). In this way, while displaying the dimensional analysis results to the user, we can also simultaneously display the analysis content obtained from further analysis of the dimensional analysis results, which can assist users in deeply analyzing the document content.

[0159] Figure 6 The process described involves extracting dimensional analysis results for each analytical dimension from the corresponding scenario text for each scenario keyword, resulting in multi-dimensional analysis results for each scenario keyword. These multi-dimensional analysis results are then defined as the dimensional analysis results for the scenario keyword set. Following a preset analysis strategy, these dimensional analysis results are analyzed to obtain the target analysis result. This achieves multi-dimensional analysis results based on at least two dimensions, enabling users to gain a comprehensive and in-depth understanding of document content.

[0160] See Figure 7 This is a block diagram of an embodiment of a document processing device provided by an embodiment of the present invention.

[0161] like Figure 7 As shown, the device may include:

[0162] Extraction module 701 is used to extract a set of scene keywords that match the target scene and the scene text corresponding to each scene keyword in the set of scene keywords from at least one document;

[0163] The first determining module 702 is used to determine at least one analysis dimension for the target scenario;

[0164] Analysis module 703 is used to perform dimensional analysis on the scene keyword set from at least one analysis dimension based on the scene text corresponding to each scene keyword in the scene keyword set, and obtain the dimensional analysis result corresponding to the scene keyword set.

[0165] The second determining module 704 is used to determine the display data based on the dimensional analysis results corresponding to the set of scene keywords.

[0166] In one possible implementation, the extraction module is specifically used for:

[0167] For each of the aforementioned scenario keywords, the following processing is performed:

[0168] Extract the contextual text content of the scene keywords from at least one of the documents;

[0169] Based on the scene keywords and the contextual text content of the scene keywords, construct the scene text corresponding to the scene keywords.

[0170] In one possible implementation, the extraction module is further configured to:

[0171] The clause containing the scene keyword is determined as the context text content of the scene keyword; and / or,

[0172] The paragraph containing the scene keyword is defined as the context text content of the scene keyword.

[0173] In one possible implementation, the extraction module is further configured to:

[0174] When the scenario keyword is a number, extract the unit of measurement corresponding to the scenario keyword from at least one of the documents;

[0175] The scene keywords and their corresponding units of measurement are concatenated to obtain the concatenated text;

[0176] The scene text corresponding to the scene keyword is constructed based on the concatenated text and the contextual text content of the scene keyword.

[0177] In one possible implementation, the analysis module is specifically used for:

[0178] Based on at least two of the aforementioned analytical dimensions, the scene text corresponding to each scene keyword in the scene keyword set is classified to obtain at least two scene text classes;

[0179] The at least two scene text classes are determined as the dimensional analysis results corresponding to the scene keyword set;

[0180] Different scenario text classes correspond to different analysis dimensions.

[0181] In one possible implementation, the second determining module is specifically used for:

[0182] Define the target analysis dimensions;

[0183] Determine the target scene text class corresponding to the target analysis dimension from the at least two scene text classes;

[0184] The display data is determined based on the text class of the target scene.

[0185] In one possible implementation, the second determining module is further configured to:

[0186] For each of the target scene text classes, the following processing is performed:

[0187] The multiple scene texts in the target scene text class are sorted according to a preset sorting method to obtain the display data corresponding to the target scene text class.

[0188] In one possible implementation, the analysis module is further configured to:

[0189] For each of the scene keywords, the following processing is performed: extract the dimension analysis results corresponding to each of the analysis dimensions from the scene text corresponding to the scene keyword to obtain the multi-dimensional analysis results corresponding to the scene keyword;

[0190] The multi-dimensional analysis results corresponding to each of the scenario keywords are determined as the dimensional analysis results corresponding to the scenario keyword set.

[0191] In one possible implementation, the second determining module is further configured to:

[0192] The dimensional analysis results corresponding to the set of scene keywords are analyzed according to the preset analysis strategy to obtain the target analysis results;

[0193] The displayed data is determined based on the target analysis results and the dimensional analysis results.

[0194] The technical solution provided by this invention extracts a set of scene keywords that match a target scenario from at least one document, along with the scene text corresponding to each scene keyword in the set. It then determines at least one analysis dimension for the target scenario. Based on the scene text corresponding to each scene keyword in the set, it performs dimensional analysis on the scene keyword set from at least one analysis dimension to obtain the dimensional analysis results. Finally, it determines the data to be displayed based on the dimensional analysis results. This achieves intelligent extraction of scene-related text from one or more documents and analyzes it to obtain corresponding analysis results, thereby improving the efficiency of document analysis and processing.

[0195] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Figure 8 The illustrated electronic device 800 includes at least one processor 801, a memory 802, at least one network interface 804, and other user interfaces 803. The various components in the electronic device 800 are coupled together via a bus system 805. It is understood that the bus system 805 is used to implement communication between these components. In addition to a data bus, the bus system 805 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 8 The general labeled all buses as Bus System 805.

[0196] The user interface 803 may include a display, keyboard or clicking device (e.g., mouse, trackball), touchpad or touch screen, etc.

[0197] It is understood that the memory 802 in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 802 described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0198] In some implementations, memory 802 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 8021 and application programs 8022.

[0199] The operating system 8021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 8022 includes various applications, such as a media player and a browser, used to implement various application functions. The program implementing the method of this embodiment can be included in the application program 8022.

[0200] In this embodiment of the invention, by calling the program or instructions stored in memory 802, specifically the program or instructions stored in application program 8022, processor 801 executes the method steps provided in each method embodiment, including, for example:

[0201] Extract a set of scene keywords that match the target scenario from at least one document, and the scene text corresponding to each scene keyword in the set of scene keywords;

[0202] For the target scenario, at least one analysis dimension shall be determined;

[0203] Based on the scene text corresponding to each scene keyword in the scene keyword set, dimensional analysis is performed on the scene keyword set from at least one analysis dimension to obtain the dimensional analysis result corresponding to the scene keyword set.

[0204] The data to be displayed is determined based on the dimensional analysis results corresponding to the set of scenario keywords.

[0205] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 801. Processor 801 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 801 or by instructions in the form of software. The processor 801 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 802. Processor 801 reads the information in memory 802 and, in conjunction with its hardware, completes the steps of the above method.

[0206] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.

[0207] For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0208] The electronic device provided in this embodiment may be as follows: Figure 8 The electronic device shown can perform the following: Figures 1-2 as well as Figure 4 , Figure 6 All steps of the Chinese document processing method, thereby achieving Figures 1-2 as well as Figure 4 , Figure 6 For details on the technical effects of the Chinese document processing method, please refer to [link / reference]. Figures 1-2 as well as Figure 4 , Figure 6 The relevant descriptions are presented concisely and will not be elaborated upon here.

[0209] This invention also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; the memory may also include combinations of the above types of memory.

[0210] When one or more programs in the storage medium can be executed by one or more processors to implement the document processing method described above that is executed on the electronic device side.

[0211] The processor is used to execute a document processing program stored in the memory to implement the following steps of a document processing method executed on the electronic device side:

[0212] Extract a set of scene keywords that match the target scenario from at least one document, and the scene text corresponding to each scene keyword in the set of scene keywords;

[0213] For the target scenario, at least one analysis dimension shall be determined;

[0214] Based on the scene text corresponding to each scene keyword in the scene keyword set, dimensional analysis is performed on the scene keyword set from at least one analysis dimension to obtain the dimensional analysis result corresponding to the scene keyword set.

[0215] The data to be displayed is determined based on the dimensional analysis results corresponding to the set of scenario keywords.

[0216] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0217] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0218] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A document processing method, characterized in that, include: Extract a set of scene keywords that match the target scenario from at least one document, and the scene text corresponding to each scene keyword in the set of scene keywords; For the target scenario, at least two analytical dimensions are determined, which are used to describe the scenario keywords from different perspectives; Based on the scene text corresponding to each scene keyword in the scene keyword set, multi-dimensional analysis is performed on the scene keywords in the scene keyword set from the at least two analysis dimensions to obtain the dimensional analysis results corresponding to the scene keyword set and respectively associated with the at least two analysis dimensions; The data to be displayed is determined based on the dimensional analysis results corresponding to the set of scenario keywords.

2. The method according to claim 1, characterized in that, Extract the scene text corresponding to each scene keyword from at least one document using the following method: For each of the aforementioned scenario keywords, the following processing is performed: Extract the contextual text content of the scene keywords from at least one of the documents; Based on the scene keywords and the contextual text content of the scene keywords, construct the scene text corresponding to the scene keywords.

3. The method according to claim 2, characterized in that, The extraction of contextual text content of the scene keywords from at least one of the documents includes: The clause containing the scene keyword is determined as the context text content of the scene keyword; and / or, The paragraph containing the scene keyword is defined as the context text content of the scene keyword.

4. The method according to claim 2, characterized in that, The step of constructing the scene text corresponding to the scene keyword based on the scene keyword and the context text content of the scene keyword includes: When the scenario keyword is a number, extract the unit of measurement corresponding to the scenario keyword from at least one of the documents; The scene keywords and their corresponding units of measurement are concatenated to obtain the concatenated text; The scene text corresponding to the scene keyword is constructed based on the concatenated text and the contextual text content of the scene keyword.

5. The method according to claim 1, characterized in that, The step involves performing multi-dimensional analysis on the scene keywords in the scene keyword set based on the scene text corresponding to each scene keyword in the scene keyword set, from at least two analytical dimensions, to obtain dimensional analysis results corresponding to the scene keyword set and respectively associated with the at least two analytical dimensions, including: Based on at least two of the aforementioned analytical dimensions, the scene text corresponding to each scene keyword in the scene keyword set is classified to obtain at least two scene text classes; The at least two scene text classes are determined as the dimensional analysis results corresponding to the scene keyword set; Different scenario text classes correspond to different analysis dimensions.

6. The method according to claim 5, characterized in that, The determination of display data based on the dimensional analysis results corresponding to the set of scenario keywords includes: Define the target analysis dimensions; Determine the target scene text class corresponding to the target analysis dimension from the at least two scene text classes; The display data is determined based on the text class of the target scene.

7. The method according to claim 6, characterized in that, The step of determining the display data based on the target scene text class includes: For each of the target scene text classes, the following processing is performed: The multiple scene texts in the target scene text class are sorted according to a preset sorting method to obtain the display data corresponding to the target scene text class.

8. The method according to claim 1, characterized in that, The step involves performing multi-dimensional analysis on the scene keywords in the scene keyword set based on the scene text corresponding to each scene keyword in the scene keyword set, from at least two analytical dimensions, to obtain dimensional analysis results corresponding to the scene keyword set and respectively associated with the at least two analytical dimensions, including: For each of the scene keywords, the following processing is performed: extract the dimension analysis results corresponding to each of the analysis dimensions from the scene text corresponding to the scene keyword to obtain the multi-dimensional analysis results corresponding to the scene keyword; The multi-dimensional analysis results corresponding to each of the scenario keywords are determined as the dimensional analysis results corresponding to the scenario keyword set.

9. The method according to claim 1 or 8, characterized in that, The determination of display data based on the dimensional analysis results corresponding to the set of scenario keywords includes: The dimensional analysis results corresponding to the set of scene keywords are analyzed according to the preset analysis strategy to obtain the target analysis results; The displayed data is determined based on the target analysis results and the dimensional analysis results.

10. A document processing apparatus, characterized in that, include: The extraction module is used to extract a set of scene keywords that match the target scenario and the scene text corresponding to each scene keyword in the set of scene keywords from at least one document; The first determining module is used to determine at least two analysis dimensions corresponding to the target scenario, wherein the at least two analysis dimensions are used to describe the scenario keywords from different perspectives; The analysis module is used to perform multi-dimensional analysis on the scene keywords in the scene keyword set from at least two analysis dimensions based on the scene text corresponding to each scene keyword in the scene keyword set, and to obtain the dimensional analysis results corresponding to the scene keyword set and respectively associated with the at least two analysis dimensions; The second determining module is used to determine the display data based on the dimensional analysis results corresponding to the set of scene keywords.

Citation Information

Patent Citations

  • Demand document processing method and device, computer equipment and storage medium

    CN112580363A