Document Management Method, Apparatus, Device, and Medium
By identifying and establishing analysis indicator labels in the document, the problem of unused analysis indicators in the document is solved, and efficient document management and data asset optimization are achieved.
Patent Information
- Application Number
- CN202210516691.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-05-11
AI Technical Summary
In the prior art, the value of analytical indicators in documents and materials has not been fully explored and utilized, making it difficult to effectively manage and utilize data assets.
By leveraging the first artificial intelligence model based on natural language processing and machine learning, analytical metrics in the document are identified and metric labels are established based on these metrics, including attribute information such as criticality and type.
It realizes multi-dimensional retrieval and analysis of documents, improves the availability and analysis efficiency of data assets, and optimizes the efficiency of document management.
Smart Images

Figure CN114819697B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and more particularly to a document management method, apparatus, device, medium, and program product. Background Art
[0002] Currently, people have gradually realized that data assets are becoming increasingly important for technological development, product research and development, production decision-making, etc. Among them, in some analysis and research work, documents are usually used as carriers for storing and transmitting data assets. For example, various research reports, academic articles, information, etc.
[0003] In the process of implementing the concept of the present disclosure, the inventors found that: when forming document materials, for the problems in the analyzed theme or field, some analysis indicators are usually used for qualitative or quantitative analysis. With the help of these analysis indicators, the current state, change trend, or evolution direction of the problems concerned in the theme or field can be judged. It can be seen that these analysis indicators have very important value for decision-making. However, in the past, when forming document-based data assets, only the documents themselves were often stored, and in some platforms, abstracts or keywords were also stored, but the analysis indicators were not refined and utilized as an important part of the documents, which made it difficult to explore and utilize the value of the analysis indicators in the document materials. Summary of the Invention
[0004] In view of the above problems, the present disclosure provides a document management method, apparatus, device, medium, and program product that can manage document-based data assets from the perspective of analysis indicators.
[0005] In the first aspect of the embodiments of the present disclosure, a document management method is provided. The method includes: obtaining a first document; identifying a first analysis indicator that appears in the statements of the first document; and establishing an indicator label for the first document based on the first analysis indicator.
[0006] According to an embodiment of the present disclosure, the identifying a first analysis indicator that appears in the statements of the first document includes: using a first artificial intelligence model to identify the first analysis indicator in the first document, where the first artificial intelligence model is obtained based on natural language processing and machine learning technologies.
[0007] According to an embodiment of the present disclosure, the identifying of the first analysis index in the first document by using the first artificial intelligence model includes: performing word segmentation on the sentences in the first document; using the first artificial intelligence model to identify the relationship between each word in the word-segmented first document and the first analysis index; and based on the relationship between each word and the first analysis index identified by the first artificial intelligence model, combining and outputting one word or a continuous plurality of words related to the first analysis index to obtain the first analysis index.
[0008] According to an embodiment of the present disclosure, the relationship between each word identified by the first artificial intelligence model and the first analysis index includes: being related to the first analysis index or being unrelated to the first analysis index. Among them, being related to the first analysis index includes at least one of the following: being at the beginning of the first analysis index, being in the middle of the first analysis index, or being at the end of the first analysis index.
[0009] According to an embodiment of the present disclosure, the first artificial intelligence model is trained in the following manner: obtaining at least one second document; using the sentences in the second document as training data and performing word segmentation on the training data; based on the relationship between each word in the word-segmented training data and the first analysis index, annotating each word in the training data; and using the annotated training data to train the first artificial intelligence model.
[0010] According to an embodiment of the present disclosure, the first artificial intelligence model adopts a conditional random field model.
[0011] According to an embodiment of the present disclosure, before establishing the index label of the first document, the method further includes: when multiple first analysis indexes are identified, calculating the similarity between every two of the first analysis indexes based on the semantic analysis of the first analysis indexes; and combining every two of the first analysis indexes with a similarity greater than the similarity threshold; and / or counting the number of occurrences of each identified first analysis index in the first document, and removing the first analysis indexes whose number of occurrences meets the removal condition.
[0012] According to an embodiment of the present disclosure, the method further includes: identifying the attribute information of the first analysis index, where the attribute information includes at least one of the following: key importance or index type in the first document; where the key importance is used to indicate whether the first analysis index is a key index in the first document. Then, establishing the index label of the first document based on the first analysis index further includes: constructing the content of the index label based on the first analysis index and the attribute information of the first analysis index.
[0013] According to an embodiment of the present disclosure, the attribute information for identifying the first analysis metric includes: obtaining numerical values of M evaluation factors for evaluating the criticality of the first analysis metric in the first document, where M is an integer greater than or equal to 2; obtaining a first feature vector of the first analysis metric based on the numerical values of the M evaluation factors; and using the first feature vector as an input to an index evaluation regression model and determining the criticality of the first analysis metric in the first document based on the output of the index evaluation regression model.
[0014] According to an embodiment of the present disclosure, the M evaluation factors include at least one of the following: the position where the first analysis metric appears in the first document; the analysis space of the first analysis metric in the first document; or the number of times the first analysis metric appears in the first document.
[0015] According to an embodiment of the present disclosure, obtaining the numerical values of the M evaluation factors for evaluating the criticality of the first analysis metric in the first document includes obtaining a numerical value for characterizing the position where the first analysis metric appears in the first document. Specifically, it includes: retrieving the first occurrence positions of N first analysis metrics identified from the first document in the first document, where N is an integer greater than or equal to 2; numbering the N first analysis metrics based on the order of the first occurrence positions; and determining a numerical value for characterizing the position where each first analysis metric appears in the first document based on the number of each first analysis metric.
[0016] According to an embodiment of the present disclosure, obtaining the numerical values of the M evaluation factors for evaluating the criticality of the first analysis metric in the first document includes: obtaining a numerical value for characterizing the analysis space of the first analysis metric in the first document. Specifically, it includes: obtaining the title level of the title to which the first analysis metric belongs in the first document to obtain a target title level; where the title level is determined according to a title hierarchy; and obtaining a numerical value for characterizing the analysis space of the first analysis metric in the first document based on the target title level.
[0017] According to an embodiment of the present disclosure, obtaining the title level of the title to which the first analysis metric belongs in the first document includes: when the first analysis metric appears in the title of the first document, obtaining the title level of the title where the first analysis metric is located; or when the first analysis metric does not appear in the title of the first document, determining the title to which the paragraph where the first analysis metric is located belongs and obtaining the title level of that title.
[0018] According to an embodiment of the present disclosure, obtaining the value representing the analysis space of the first analysis metric in the first document based on the target title level includes: converting the highest title level in the first document into a first value based on a preset conversion relationship between the title level and the value, where the highest title level is the level of the title located at the topmost layer in the title hierarchy; converting the target title level into a second value based on the conversion correspondence between the title level and the value; and using the first value as a parameter of a preset normalization model and the second value as a variable of the normalization model to calculate the value representing the analysis space of the first analysis metric in the first document.
[0019] According to an embodiment of the present disclosure, the method further includes: setting a conversion relationship between the title level and the value, where the higher the position of the title level in the title hierarchy, the larger the converted value.
[0020] According to an embodiment of the present disclosure, identifying the attribute information of the first analysis metric includes: using a second artificial intelligence model to identify the metric type of the first analysis metric, where the second artificial intelligence model is a multi-classification model obtained based on machine learning technology.
[0021] According to an embodiment of the present disclosure, the second artificial intelligence model is trained in the following manner: obtaining at least one second analysis metric; performing word segmentation on the second analysis metric and converting it into a word vector to obtain a second feature vector of the second analysis metric; annotating the metric type of the second analysis metric; and using the second feature vector as the input of the second artificial intelligence model and the annotated metric type of the second analysis metric as the output reference of the second artificial intelligence model to train the second artificial intelligence model.
[0022] According to an embodiment of the present disclosure, using the second feature vector as the input of the second artificial intelligence model and the annotated metric type of the second analysis metric as the output reference of the second artificial intelligence model to train the second artificial intelligence model further includes: performing manual review on the output of the second artificial intelligence model; and training the second artificial intelligence model based on the difference between the output result after manual review and the annotated metric type of the second analysis metric.
[0023] According to an embodiment of the present disclosure, the second artificial intelligence model adopts a BERT model.
[0024] According to an embodiment of the present disclosure, the metric type includes the type obtained by dividing the first analysis metric based on the analysis object of the metric.
[0025] According to an embodiment of the present disclosure, the index type includes at least one of the following: an index for the product itself, an index for the customer, or an index for the partner.
[0026] In a second aspect of the embodiments of the present disclosure, a document management device is provided. The device includes a first acquisition module, a first recognition module, and an index label establishment module. The first acquisition module is configured to acquire a first document. The first recognition module is configured to recognize a first analysis index that appears in the statements of the first document. The index label establishment module is configured to establish an index label for the first document based on the first analysis index.
[0027] According to an embodiment of the present disclosure, the first recognition module is configured to: use a first artificial intelligence model to recognize the first analysis index in the first document, where the first artificial intelligence model is obtained based on natural language processing and machine learning technologies.
[0028] According to an embodiment of the present disclosure, the device further includes a second recognition module. The second recognition module is configured to recognize attribute information of the first analysis index, where the attribute information includes at least one of the following: criticality or index type in the first document; where the criticality is used to indicate whether the first analysis index is a key index in the first document. The index label establishment module is further configured to construct the content of the index label based on the first analysis index and the attribute information of the first analysis index.
[0029] According to an embodiment of the present disclosure, the second recognition module includes a key index recognition module. The key index recognition module includes an evaluation factor acquisition sub-module, a feature vector combination sub-module, and an index evaluation regression model. The evaluation factor acquisition sub-module is configured to acquire numerical values of M evaluation factors for evaluating the criticality of the first analysis index in the first document, where M is an integer greater than or equal to 2. The feature vector combination sub-module is configured to obtain a first feature vector of the first analysis index based on the numerical values of the M evaluation factors. The index evaluation regression model is configured to use the first feature vector as an input and predict the criticality of the first analysis index in the first document based on the first feature vector.
[0030] According to an embodiment of the present disclosure, the second recognition module includes an index type recognition module. The index type recognition module includes a second artificial intelligence model. The second artificial intelligence model is configured to recognize the index type of the first analysis index, where the second artificial intelligence model is a multi-classification model obtained based on machine learning technologies.
[0031] In a third aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned document management method.
[0032] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is further provided, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned document management method.
[0033] In a fifth aspect of the embodiments of the present disclosure, a computer program product is further provided, including a computer program, and when the computer program is executed by a processor, the above-mentioned document management method is implemented. Description of the Drawings
[0034] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above-mentioned content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0035] Figure 1 Schematically shown is an application system architecture diagram of a document management method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;
[0036] Figure 2 Schematically shown is a schematic diagram of establishing index tags in the document management method according to an embodiment of the present disclosure;
[0037] Figure 3 Schematically shown is a flowchart of a document management method according to an embodiment of the present disclosure;
[0038] Figure 4 Schematically shown is a flowchart of a document management method according to another embodiment of the present disclosure;
[0039] Figure 5 Schematically shown is a flowchart of identifying analysis indicators in a document using a first artificial intelligence model according to an embodiment of the present disclosure;
[0040] Figure 6 Schematically shown is a method flowchart for training a first artificial intelligence model according to an embodiment of the present disclosure;
[0041] Figure 7 Schematically shown is a flowchart of a document management method according to another embodiment of the present disclosure;
[0042] Figure 8 Schematically shown is a flowchart of a document management method according to still another embodiment of the present disclosure;
[0043] Figure 9Schematically shows a flowchart of identifying key metrics among analysis metrics in a document management method according to an embodiment of the present disclosure;
[0044] Figure 10 Schematically shows a flowchart of using an index evaluation regression model to identify whether it is a key metric according to an embodiment of the present disclosure;
[0045] Figure 11 Schematically shows a flowchart of obtaining a value representing the occurrence position of an analysis metric during the process of identifying key metrics among analysis metrics according to an embodiment of the present disclosure;
[0046] Figure 12 Schematically shows a flowchart of obtaining a value representing the analysis space of an analysis metric during the process of identifying key metrics among analysis metrics according to an embodiment of the present disclosure;
[0047] Figure 13 Schematically shows a flowchart of obtaining a value representing the analysis space of an analysis metric according to another embodiment of the present disclosure;
[0048] Figure 14 Schematically shows a method flowchart of using a second artificial intelligence model to identify the metric type of analysis metrics in a document management method according to another embodiment of the present disclosure;
[0049] Figure 15 Schematically shows a method flowchart of training a second artificial intelligence model according to an embodiment of the present disclosure;
[0050] Figure 16 Schematically shows a method flowchart of training a second artificial intelligence model according to another embodiment of the present disclosure;
[0051] Figure 17 Schematically shows a flowchart of a document management method according to still another embodiment of the disclosure;
[0052] Figure 18 Schematically shows a structural block diagram of a document management apparatus according to an embodiment of the present disclosure;
[0053] Figure 19 Schematically shows a structural block diagram of a first identification module in a document management apparatus according to an embodiment of the present disclosure;
[0054] Figure 20 Schematically shows a structural block diagram of a key metric identification module in a document management apparatus according to an embodiment of the present disclosure;
[0055] Figure 21 Schematically shows a structural block diagram of a metric type identification module in a document management apparatus according to an embodiment of the present disclosure; and
[0056] Figure 22 A block diagram schematically showing an electronic device suitable for implementing a document management method according to an embodiment of the present disclosure is shown. Detailed implementation manners
[0057] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0058] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0059] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0060] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0061] In this document, it should be understood that any number of elements in the specification and the drawings is for illustration rather than limitation, and any naming (for example, first, second) is only for distinction and does not have any limiting meaning.
[0062] Embodiments of the present disclosure provide a document management method, apparatus, device, medium, and program product. The method includes first obtaining a first document, then identifying a first analysis index that appears in the statements of the first document, and then establishing an index label for the first document based on the first analysis index.
[0063] The first document can be a document in any field and on any topic. In this text, the "first document" is used to indicate the document to be analyzed. Herein, a document is mainly a text material written in natural language, which may also include data information such as figures and tables. For example, product operation analysis reports, product R & D reports, scientific research reports, academic papers, news reports, information, etc.
[0064] The first analysis indicator is an analysis indicator used to quantitatively or qualitatively analyze problems in the topic or field described by the first document, and is an instrumental term for analyzing problems in this topic or field. In this text, the "first analysis indicator" is used to indicate the analysis indicators appearing in the document to be analyzed. A document may have one analysis indicator or multiple analysis indicators, and the present disclosure does not limit this.
[0065] Taking the product operation analysis report as an example, the important value of the analysis indicator is briefly described. A product operation analysis report is a report prepared by an enterprise to analyze the operation conditions of a product, such as its market performance, competition situation of peer products, and customer usage effect. In the process of analyzing product operation, it is inevitable to pre-define various product operation indicators (i.e., the above-mentioned "first analysis indicators"). Product operation indicators are analysis indicators used to evaluate the product operation situation. Taking Internet products as an example, common product operation indicators include the number of registered users, the number of daily active users, the number of monthly active users, or the channel conversion rate, etc. By analyzing the magnitude, change trend, influencing factors, etc. of product operation indicators, the operation situation of the product can be comprehensively understood. Different product operation indicators often imply product characteristics in different dimensions. In addition, product operation indicators often reflect the product operation analysis ideas and frameworks, as well as the future product operation strategies of the enterprise. Thus, it can be seen that for the product operation of an enterprise, product operation indicators are a very important data asset.
[0066] The embodiments of the present disclosure can automatically identify product operation indicators from the product operation analysis report and establish indicator tags for the product operation analysis report based on this. In this way, taking the product operation indicators as a part of the product operation analysis report, the product operation indicators can be used as data assets of the enterprise. In this way, the enterprise can retrieve, count, and analyze the product operation analysis report based on the product operation indicator dimension, improving the usability of the enterprise's data analysis assets and optimizing the efficiency of product operation analysis.
[0067] It can be understood that the above-exemplified product operation analysis report and the product operation indicators therein are only exemplary. When managing documents of different topics or fields, there are corresponding analysis indicators according to the topic or field described by the document, and the embodiments of the present disclosure can establish indicator tags adapted to the managed documents. Thus, a way to collect, compare, and analyze documents from the dimension of analysis indicators is provided, improving the analysis and use efficiency of the documents.
[0068] It should be noted that the document management method, apparatus, device, medium and program product determined in the embodiments of the present disclosure can be used in the financial field (such as, Internet product analysis, internal document management, product development management, etc.), and can also be used in any field other than the financial field (for example, technology, education, medicine, military industry, logistics, etc.). The present disclosure does not limit the application field.
[0069] Figure 1 Schematically shows an application system architecture diagram of the document management method, apparatus, device, medium and program product according to the embodiments of the present disclosure. It should be noted that Figure 1 The shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0070] As Figure 1 shown, the system architecture 100 according to this embodiment may include at least one terminal device (three are shown in the figure, terminal devices 101, 102, 103), a server 104, and at least one database system (three are shown in the figure, database systems 105, 106, 107).
[0071] The terminal devices 101, 102, 103 can be connected to the server 104 through a network. Users can use the terminal devices 101, 102, 103 to interact with the server 104 through the network to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0072] The server 104 can be communicatively connected to the database systems 105, 106, 107 through a network. Among them, a large amount of document materials can be stored in the database systems 105, 106, 107.
[0073] The document management method of the embodiments of the present disclosure can generally be executed by the server 104. Among them, the server 104 can determine the first document to be analyzed by interacting with the terminal devices 101, 102, 103, obtain the first document from the database systems 105, 106, 107, and then process the first document according to the methods provided in the various embodiments of the present disclosure to establish the index tags of the first document. Correspondingly, the document management apparatus, device, medium and program product provided by the embodiments of the present disclosure can generally be set in the server 104.
[0074] Of course, the document management method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 104 and capable of communicating with terminal devices 101, 102, 103, database systems 105, 106, 107, and / or server 104. Correspondingly, the document management device, equipment, medium, and program product provided by the embodiments of the present disclosure can also be disposed in a server or a server cluster different from server 104 and capable of communicating with terminal devices 101, 102, 103, database systems 105, 106, 107, and / or server 104.
[0075] It should be understood that Figure 1 the numbers of terminal devices, networks, servers, and database systems in
[0076] Figure 2 FIG. schematically shows a schematic diagram of establishing an index label in the document management method according to an embodiment of the present disclosure.
[0077] As Figure 2 shown, according to an embodiment of the present disclosure, an index label 202 can be established by identifying and processing analysis indexes in the first document 201.
[0078] The index label 202 may include the names of the analysis indexes that appear in the first document 201 (such as, the first analysis index 1, the first analysis index 2, the first analysis index 3, etc.).
[0079] According to other embodiments of the present disclosure, the index label 202 may further include the attribute information of each analysis index that appears in the first document 201 (such as, the index type, and / or whether it is a key index or a non-key index).
[0080] The index type can be a type divided according to any one or more dimensions; for example, it can be a type obtained by dividing the first analysis index based on the analysis object of the index (such as, product, customer, or partner); or it can be a type obtained by dividing the first analysis index based on the use of the index (for analyzing input, output, efficiency, or quality); or it can be a type divided based on the qualitative or quantitative characteristics of the index. For example, an index for qualitative analysis or an index for quantitative analysis. It can be understood that for the division of the index type, there can be multiple division methods according to the research problem focus in different fields or topics, and no further examples are given here.
[0081] The key index is the core of the analysis in the content described in the first document 201, and often the content of the first document 201 can be developed based on the key index.
[0082] Taking the first document 201 as an example of a product operation analysis report, the value of key indicators is illustrated. In production practice, when analyzing the operation of a certain product, usually the key indicators of the product are first determined, and then the quality of the product operation is judged by analyzing the quality of the key indicators. Therefore, it can be understood that in each product operation analysis report, the analysis often revolves around one or more key indicators.
[0083] The key indicators in the product operation analysis report are one or more of the most critical indicators during the operation of the product, which are used to guide the development direction of the product. For example, the key indicator of e-banking business is, for example, the number of effective accounts. Through the number of effective accounts, the ability of e-banking business to attract customers can be evaluated, and the improvement direction of e-banking business can also be provided accordingly. Another example is that the key indicators of the online payment acceptance product can include the number of payment acceptance merchants, based on which the acceptance degree of the online payment acceptance product among merchants can be judged.
[0084] It can be seen that according to the embodiments of the present disclosure, when the indicator label 202 further includes information such as indicator type and / or key indicator / non-key indicator, it can help users make more effective use of the value of the indicators, and the research focus of a document can be initially determined based on the indicator type and / or key indicator in the document, improving the efficiency of the user's research and analysis work.
[0085] According to some other embodiments of the present disclosure, the indicator label 202 may further include other information related to the first analysis indicator 1, the first analysis indicator 2, the first analysis indicator 3, etc., such as the number of occurrences of each analysis indicator, the order of the first occurrence of these indicators, etc.
[0086] According to the embodiments of the present disclosure, when the indicator label 202 is established for each of a large number of documents in the manner shown in the first document 201, multi-dimensional retrieval of these documents can be supported. For example, all documents with the same indicator can be retrieved, or all documents with the same key indicator can be retrieved, or documents with the same indicator type and analyzed as key indicators can be obtained, etc., facilitating the horizontal and vertical comparison of document materials from the aspect of analysis indicators, conveniently mining and utilizing the value of analysis indicators, reducing the research time of analysis, and improving the efficiency of analysis.
[0087] The following will be based on Figure 1 the described system architecture, combined with Figure 2 the schematic illustration of the indicator label, and through Figures 3 to 17 a detailed description of the document management method of the embodiments of the present disclosure will be given.
[0088] Figure 3 A flowchart of a document management method according to an embodiment of the present disclosure is schematically shown.
[0089] As Figure 3 shown, the document management method 300 according to this embodiment may include operation S310 to operation S330.
[0090] In operation S310, the first document 201 is obtained. The user may specify one or more documents to be analyzed through the terminal devices 101, 102, or 103, or may also obtain the documents to be analyzed in batches at regular intervals through program settings of the server 104.
[0091] In operation S320, the first analysis indicators appearing in the statements of the first document 201 are identified. Among them, the first analysis indicators identified in operation S320 may be one or multiple.
[0092] In one embodiment, by collecting and organizing a large number of documents in the theme or field described by the first document 201, the analysis indicators used in this theme or field can be extracted, and then these indicators are used to perform intelligent matching on the first document 201 in operation S320.
[0093] In another embodiment, the analysis indicators in the first document 201 can be identified through natural language processing and machine learning techniques. For example, machine learning training can be performed on the first artificial intelligence model constructed based on natural language processing, and after training, the first artificial intelligence model is used to identify the first analysis indicators in the first document 201. Among them, the training data used when training the first artificial intelligence model can come from the same theme or field as the first document 201. For example, corresponding first artificial intelligence models can be trained for different themes or fields. Then, in operation S320, according to the theme or field described by the first document 201, the corresponding first artificial intelligence model is selected to identify the first analysis indicators.
[0094] In operation S330, based on the first analysis indicators, the indicator label 202 of the first document 201 is established. For example, the identified first analysis indicators can be used as the content of the indicator label 202 of the first document 201. In some embodiments, the content of the indicator label 202 may include, in addition to the identified first analysis indicators, statistical information such as the occurrence times, occurrence frequencies, occurrence orders, and occurrence positions of each first analysis indicator. In other embodiments, the content of the indicator label 202 may also include information obtained by further analyzing and processing the identified analysis indicators (such as attribute information such as indicator types and whether they are key indicators, as Figure 2 shown).
[0095] It can be seen that, according to the embodiments of the present disclosure, establishing index tags for the first document 201 based on the first analysis index can facilitate document retrieval, statistics, and comparative analysis from the dimension of the analysis index, improve the usability of the document, and optimize the efficiency of analyzing the document. For example, when a user wants to refer to existing documents, by retrieving and analyzing the analysis index, they can understand the existing research results and progress in the research topic or field, effectively improving the usage efficiency of the document and the efficiency of utilizing data assets.
[0096] Figure 4 The flowchart of the document management method according to another embodiment of the present disclosure is schematically shown.
[0097] As Figure 4 shown, the document management method 400 according to this embodiment may include operation S310, operations S421 to S423, and operation S330. Among them, operation S310 and operation S330 are the same as those in method 300, and operations S421 to S423 are a specific embodiment of operation S320 in method 300.
[0098] First, in operation S310, the first document 201 is obtained.
[0099] Then, in operation S421, word segmentation is performed on the statements in the first document 201. For example, when the first document 201 is an English document, word segmentation can be roughly performed through steps such as splitting words according to spaces, excluding stop words, and extracting stems. Another example is that when the first document 201 is a Chinese document, word segmentation can be performed using any one or more Chinese word segmentation tools, and the word segmentation algorithm can roughly include a word segmentation method based on string matching, a word segmentation method based on understanding, or a word segmentation method based on statistics, etc. Correspondingly, for documents written in other languages, word segmentation is performed using the corresponding word segmentation tools or algorithms for that language.
[0100] Next, in operation S422, the relationship between each word in the word-segmented first document 201 and the first analysis index is identified using the first artificial intelligence model.
[0101] The relationship between each word and the first analysis index can specifically be related to the first analysis index or unrelated to the first analysis index.
[0102] According to an embodiment of the present disclosure, a word being related to the first analysis index can be further divided into at least one of the following: being at the beginning of the first analysis index, being in the middle of the first analysis index, or being at the end of the first analysis index.
[0103] When training the first artificial intelligence model, the relationship between the words in the training data (e.g., the second document) and the first analysis index can be labeled, so that the first artificial intelligence model can learn the relationship between each word and the first analysis index when it appears in a sentence according to the context and language environment of the second document.
[0104] Next, in operation S423, based on the relationship between each word identified by the first artificial intelligence model and the first analysis index, a word or a continuous plurality of words related to the first analysis index are combined and output to obtain the first analysis index.
[0105] Thereafter, in operation S330, based on the first analysis index, the index label 202 of the first document 201 is established.
[0106] Figure 5 A flowchart showing the identification of analysis indexes in a document using the first artificial intelligence model according to an embodiment of the present disclosure is schematically shown.
[0107] As Figure 5 shown, in combination with Figure 4 , in process 500, the first document 201 can be segmented and the order of the segmentation is retained, and then input into the first artificial intelligence model 501. In one embodiment, the first artificial intelligence model 501 can adopt a conditional random field model (CRF).
[0108] The first artificial intelligence model 501 can output the relationship between the words in the first document 201 and the first analysis index in the form of a part-of-speech tag 502, where the relationship between each word and the first analysis index can be marked in the form of a tag in the part-of-speech tag 502.
[0109] For example, in one embodiment, the relationship between each identified word and the first analysis index can be marked in the following form of tags (B, M, E, and A); where B represents a word at the beginning of the first analysis index; M represents a word in the middle of the first analysis index; E represents a word at the end of the first analysis index; A represents a word unrelated to the first analysis index.
[0110] Next, based on the part-of-speech tag 502, the first analysis index can be output in the same way as in operation S423. For example, for a word labeled as B, find the nearest E behind it. After finding it, output all the words from B to E as the first analysis index; if no E is found behind the word labeled as B, it is considered that the word labeled as B alone constitutes an analysis index, and thus B is directly output as the first analysis index. Thus, the set 503 of the first analysis indexes can be obtained.
[0111] Here is a specific example of e-banking business to illustrate. For example, the content of the document is: "The number of newly launched partners is XX, and it has completed XX.X% of the annual target; affected by the new regulations on Internet deposits, the number of newly opened e-accounts and the amount of newly deposited funds have not achieved the sequential progress."
[0112] After segmenting the above document and inputting it into the first artificial intelligence model 501, the first artificial intelligence model 501 can output the following content:
[0113] The number of newly launched partners (A) is XX (A), and it has completed (A) XX.X% (A) of the annual (A) target (A); affected by the new regulations on Internet (A) deposits (A), the number of newly opened e-(B) accounts (M) and the amount of newly deposited e-(B) funds (M) have not (A) achieved (A) the sequential (A) progress (A).
[0114] Based on the above recognition results, the analysis indicators that can be obtained are: the number of newly opened e-accounts, the amount of deposited funds.
[0115] Figure 6 The method flow chart for training the first artificial intelligence model according to an embodiment of the present disclosure is schematically shown.
[0116] As Figure 6 shown, in combination with Figure 5 , according to the embodiments of the present disclosure, the method 600 for training the first artificial intelligence model 501 may include operation S610 to operation S640.
[0117] First, in operation S610, obtain at least one second document. The second document may be a document that is consistent with the theme or the field to which the first document 201 pertains. Here, the "second document" is used to indicate the document used when training the first artificial intelligence model.
[0118] Then, in operation S620, use the sentences in the second document as training data and segment the training data.
[0119] Next, in operation S630, based on the relationship between each word in the segmented training data and the first analysis indicator, annotate each word in the training data.
[0120] Taking the product operation analysis report as an example to describe the annotation process. For example, a large number of historical product operation reports can be selected as the basic data set, and each sentence in the reports of the basic data set is used as a training data. Then, the existing product operation index thesaurus can be utilized, and each training data is segmented by a segmentation tool (such as jieba segmentation). Next, the segmented corpus is annotated. For example, each word is annotated according to the above-mentioned tags B, M, E, and A. If a word is related to the product operation index, the corresponding tag (one of B, M, or E) is annotated for the word; if a word is not related to the product operation index, A is annotated for the word.
[0121] After that, in operation S640, the first artificial intelligence model 501 is trained using the annotated training data.
[0122] During the training process, the output result of each round of the first artificial intelligence model 501 is compared with the annotation of the training data, and the parameters of the first artificial intelligence model 501 are optimized in reverse according to the comparison differences. Repeated training is carried out until the output result of the first artificial intelligence model 501 converges and the accuracy meets the requirements. After that, the trained first artificial intelligence model 501 can be used to automatically extract the analysis indicators in the document.
[0123] According to an embodiment of the present disclosure, during the training process of the first artificial intelligence model 501 in operation S640, a manual review step can be added, that is, the output of the first artificial intelligence model training process is reviewed and adjusted through manual review, and then the result after review and adjustment is compared with the annotation of the training data to optimize the model in reverse, thereby improving the training accuracy of the first artificial intelligence model 501, especially in the case where it is difficult to ensure that the training data volume is large enough or the sample distribution is diverse enough.
[0124] Figure 7 The flowchart of the document management method according to another embodiment of the present disclosure is schematically shown.
[0125] As Figure 7 shown, in addition to operation S310, operation S320, and operation S330, the document management method 700 according to this embodiment may further include operations S721 to S722, where operations S721 to S722 may be executed before operation S330.
[0126] First, in operation S310, the first document 201 is obtained.
[0127] Then, in operation S320, the first analysis indicators appearing in the sentences of the first document 201 are identified.
[0128] Among them, operations S310 and S320 can refer to the relevant descriptions in method 300 or 400, which will not be elaborated here.
[0129] Next, in operation S721, count the number of occurrences of each identified first analysis index in the first document 201, and exclude the first analysis indexes whose number of occurrences meets the exclusion condition.
[0130] The exclusion condition can be, for example, that the number of occurrences is lower than the threshold, or when sorted by the number of occurrences, it is at the end; or when compared with the analysis index in front of it, the decrease in the number of occurrences is greater than a predetermined value, etc. In this way, the possible errors in the recognition result can be reduced.
[0131] Then, in operation S722, when multiple first analysis indexes are identified, calculate the similarity between every two first analysis indexes based on the semantic analysis of the first analysis indexes. For example, segment the first analysis indexes and convert them into word vectors, and calculate the similarity between the word vectors corresponding to every two first analysis indexes (for example, the angle, cosine similarity, etc.). Then, merge every two first analysis indexes whose similarity is greater than the similarity threshold, that is, regard two first analysis indexes whose similarity is greater than the similarity threshold as the same analysis index.
[0132] During the writing process of the first document 201, the compilation personnel may define some analysis indexes according to their own experience or fuzzy memory, which may lead to problems such as non-standardization, repetition, or non-uniformity of some analysis indexes. According to the embodiments of the present disclosure, similar analysis indexes are merged, so as to reduce the chaos caused by the non-standard use of analysis indexes by the compilation personnel when writing the document, and improve the standardization and uniformity of the analysis indexes. Moreover, it also helps to make the retrieval and utilization of the analysis index dimension simple, unified, and standardized.
[0133] It should be noted that Figure 7 the order of operations S721 and S722 is only exemplary, and the present disclosure does not limit the order of the two. In addition, in some embodiments, method 700 may also include only one of operations S721 and S722.
[0134] Thereafter, in operation S330, based on the first analysis index, establish the index label 202 of the first document 201. Among them, operation S330 can refer to the relevant description in method 300, which will not be elaborated here.
[0135] Figure 8 Schematically shows a flowchart of a document management method according to another embodiment of the present disclosure.
[0136] As Figure 8As shown, in addition to operation S310 and operation S320, the document management method 800 according to this embodiment may further include operation S820 and operation S830. Among them, operation S820 and operation S830 are executed after operation S320.
[0137] First, in operation S310, the first document 201 is obtained.
[0138] Then, in operation S320, the first analysis indicators appearing in the statements of the first document 201 are identified.
[0139] Among them, operation S310 and operation S320 may refer to the relevant descriptions in method 300 or 400, which will not be elaborated here.
[0140] Next, in operation S820, the attribute information of the first analysis indicator is identified. The attribute information may include at least one of the following: criticality or indicator type in the first document 201. Among them, the criticality is used to indicate whether the first analysis indicator is a key indicator in the first document 201.
[0141] Then, in operation S830, based on the first analysis indicator and the attribute information of the first analysis indicator, the content of the indicator label 202 is constructed. For example, the first analysis indicator and the attribute information of the first analysis indicator may be used as the content of the indicator label 202. For example, refer to Figure 2 for illustration. Of course, in addition to the first analysis indicator and the attribute information of the first analysis indicator, the content of the indicator label 202 may further include other information, which is not limited in this disclosure.
[0142] Figure 9 Schematically shows a flowchart of identifying key indicators among analysis indicators in a document management method according to an embodiment of the present disclosure.
[0143] As Figure 9 shown, operation S820 may include operations S901 to S903. Among them, through operations S901 to S903, it can be identified whether the first analysis indicator is a key indicator of the first document 201.
[0144] Specifically, in operation S901, the numerical values of M evaluation factors for evaluating the criticality of the first analysis indicator in the first document 201 are obtained, where M is an integer greater than or equal to 2.
[0145] The writing requirements, writing methods, or formats for each field, or each subject, or each platform may result in some differences in the appearance of key indicators and non-key indicators in the document. For example, in a product operation analysis report, the differences between key indicators and non-key indicators can be any one or more of the following: 1) The analysis of key indicators often appears more prominently in the report, such as in a more prominent position like the title, so that readers can quickly understand the most core operation analysis conclusions; 2) Key indicators often have more detailed analysis, so the analysis space is longer than that of non-key indicators; 3) The number or frequency of key indicators appearing in the report is often higher than that of non-key indicators. Of course, only several differences are listed here as examples. For documents in different subjects, different fields, or different platforms, the difference points between key indicators and non-key indicators can be different.
[0146] According to the embodiments of the present disclosure, corresponding parameters can be refined as evaluation factors based on the differences in the appearance characteristics of key indicators and non-key indicators in the document, that is, the above-mentioned M evaluation factors.
[0147] Then in operation S902, based on the values of the M evaluation factors, the first feature vector of the first analysis indicator is obtained.
[0148] Next, in operation S903, the first feature vector is used as the input of the index evaluation regression model, and the key nature of the first analysis indicator in the first document 201 is determined based on the output of the index evaluation regression model, that is, it is determined whether the first analysis indicator is a key indicator.
[0149] As described above, the key indicator is the core of the analysis in the content described in the first document 201, and often the content of the first document 201 can be developed based on the key indicator. In the embodiments of the present disclosure, automatically extracting the key indicators of the first document 201 can greatly improve the usage efficiency of the first document. For example, when the first document 201 is a product operation R & D report, after distinguishing the key indicators, valuable information can be provided for the R & D and improvement of the product. Correspondingly, when the first document 201 is a document in other subjects or fields, automatically distinguishing the key indicators therein is of great help for grasping the research context and development direction in that subject or field.
[0150] Figure 10 A flowchart showing the identification of whether it is a key indicator using the index evaluation regression model according to an embodiment of the present disclosure is schematically shown.
[0151] As Figure 10 shown, in combination with Figure 9, in this process 1000, based on the differential analysis of key indicators and non-key indicators in documents such as product analysis reports, the M evaluation factors determined may include: the occurrence position 4 of the first analysis indicator in the first document 201; the analysis space A2 of the first analysis indicator in the first document 201, and / or the occurrence times A3 of the first analysis indicator in the first document 201.
[0152] The occurrence times A3 of the first analysis indicator in the first document 201 can be obtained by counting the occurrence times of the corresponding analysis indicator in the whole analysis report.
[0153] Regarding the value of the occurrence position A1 of the first analysis indicator in the first document 201, according to an embodiment of the present disclosure, reference can be made to the following Figure 11 method 1100 shown to obtain.
[0154] Regarding the value of the analysis space A2 of the first analysis indicator in the first document 201, according to an embodiment of the present disclosure, reference can be made to the following Figure 12 method 1200 shown to obtain.
[0155] In the process 1000, for a first analysis indicator, the corresponding A1, A2, and A3 can be combined to form a feature vector of the first analysis indicator. Then, the feature vector is input into the index evaluation regression model 1001, and according to the value (e.g., 0 or 1) output by the index evaluation regression model 1001, it is determined whether the first analysis indicator is a key indicator or a non-key indicator. Among them, the index evaluation regression model 1001 can be an SVM algorithm model, a logistic regression model, a GBDT model, etc.
[0156] Taking the case where the index evaluation regression model 1001 adopts the logistic regression algorithm as an example, the process 1000 is described as follows. In an embodiment, the formula of the logistic regression algorithm adopted by the index evaluation regression model 1001 can be listed as follows
[0157] z = α0 + α1A1 + α2A2 + α3A3 (1)
[0158]
[0159]
[0160] Among them, in formula (1), A1, A2, and A3 respectively represent the occurrence position, analysis length, and number of occurrences of the analysis index. α0, α1, α2, and α3 respectively represent the weight parameters of the model, which need to be obtained through model training calculation. Formula (2) is the index evaluation regression model 1001. Through the training of this model, the output result of the finally obtained model is shown in formula (3). Among them, Y represents whether it is a key index. Among them, Y = 1 means it is a key index, and Y = 0 means it is not a key operation index. After the model training is completed, for the document to be predicted and analyzed and the analysis index therein, the above calculation formula of Y can be substituted, and based on the obtained Y value, it can be judged whether a certain analysis index is a key index.
[0161] Figure 11 Schematically shows a flowchart of obtaining a value representing the occurrence position of an analysis index in the process of identifying a key index among analysis indexes according to an embodiment of the present disclosure.
[0162] As Figure 11 shown, in combination with Figure 10 , the method 1100 for obtaining a value representing the occurrence position of the first analysis index in the first document 201 according to this embodiment may include operation S1101 to operation S1103.
[0163] First, in operation S1101, retrieve the first occurrence positions of N first analysis indexes identified from the first document 201, where N is an integer greater than or equal to 2.
[0164] Next, in operation S1102, number the N first analysis indexes based on the order of the first occurrence positions.
[0165] Then, in operation S1103, based on the number of each first analysis index, determine the value representing the occurrence position of each first analysis index in the first document 201.
[0166] For example, Figure 2 among the three indexes identified in the first document 201: the first analysis index 1, the first analysis index 2, and the first analysis index 3, assume that the order of the first occurrences is the first analysis index 2, the first analysis index 3, and the first analysis index 1. Thus, the first analysis index 1, the first analysis index 2, and the first analysis index 3 can be numbered 3, 1, and 2 respectively according to the first occurrence position. Accordingly, the A1 corresponding to the first analysis index 1, the first analysis index 2, and the first analysis index 3 can be assigned 3, 1, and 2 respectively.
[0167] Figure 12 Schematically shows a flowchart of obtaining a value representing the analysis length of an analysis index in the process of identifying a key index among analysis indexes according to an embodiment of the present disclosure.
[0168] As shown Figure 12 in combination with Figure 10 , the method 1200 for obtaining a value representing the analysis space of the first analysis index in the first document 201 according to this embodiment may include operation S1201 to operation S1202.
[0169] First, in operation S1201, obtain the title level of the title to which the first analysis index belongs in the first document 201 to obtain the target title level. In this embodiment, the "target title level" is only for easy distinction in subsequent descriptions and has no defined meaning.
[0170] According to an embodiment of the present disclosure, the title level is determined according to the title hierarchy. For example, in a Word document, there are usually first-level headings, second-level headings, third-level headings, etc. Among them, in a Word document, the larger the number of the title level, the lower the hierarchy of the title. It can be understood that the assignment methods of the number of title levels in different document editors may be different. However, reflected in the title hierarchy with the document name as the root node, it can be unified that the lower the title hierarchy (that is, the farther away from the root node), the smaller the covered space.
[0171] In an embodiment of the present disclosure, when obtaining the title level of a word document, the python-docx library can be used, or the POI interface (Point of Interface) can be used to obtain it. Among them, the POI interface is an open-source API written in Java (Application Programming Interface) for supporting Java programs to read and write word documents.
[0172] In some embodiments, in operation S1201, according to whether the first analysis index appears in the title of the first document 201, it can be divided into two cases to obtain the title level of the title to which the first analysis index belongs in the first document 201.
[0173] When the first analysis index appears in the title of the first document 201, obtain the title level of the title where the first analysis index is located as the target title level.
[0174] When the first analysis index does not appear in the title of the first document 201, determine the title to which the paragraph where the first analysis index is located belongs, and obtain the title level of this title. Generally, the title will be before the paragraph. You can find the nearest title before the paragraph where the first analysis index is located, and then obtain the title level of this title as the target title level.
[0175] When the first analysis indicator appears in multiple headings or in paragraphs under multiple headings, the heading level of the highest position in the heading hierarchy among the multiple headings can be used as the target heading level, or the heading level of the heading to which the paragraph with the most occurrences of the first analysis indicator belongs can also be used as the target heading level.
[0176] Then, in operation S1202, based on the target heading level, a value is obtained that characterizes the analysis space of the first analysis indicator in the first document 201.
[0177] For example, the conversion relationship between the heading level and the value can be preset in advance, and then the target heading level is converted into the corresponding value according to this conversion relationship. Among them, in the conversion relationship between the heading level and the value, it can be set that the higher the position of the heading level in the heading hierarchy, the larger the converted value, which means the larger the corresponding analysis space.
[0178] According to some other embodiments of the present disclosure, considering that there are differences in the number of heading levels of different documents, for example, some documents only have 2-level headings, and some documents have 5-level headings. For unified analysis, the target heading level can be normalized in operation S1202. Specifically, reference can be made to Figure 13 for the introduction.
[0179] Figure 13 Schematically shows a flowchart for obtaining a value characterizing the analysis space of an analysis indicator according to another embodiment of the present disclosure.
[0180] As Figure 13 shown, according to the embodiment of the present disclosure, operation S1202 may include operations S1212 to S1232.
[0181] In operation S1212, based on the preset conversion relationship between the heading level and the value, the highest heading level in the first document 201 is converted to obtain a first value (for example, denoted as B max ), where the highest heading level is the level of the heading located at the top layer in the heading hierarchy.
[0182] In operation S1222, based on the conversion correspondence between the heading level and the value, the target heading level is converted to obtain a second value (for example, denoted as B).
[0183] In operation S1232, using the first value as the parameter of the preset normalization model and the second value as the variable of the normalization model, a value is calculated that characterizes the analysis space of the first analysis indicator in the first document 201 (for example, denoted as B′).
[0184] In one embodiment, the formula for normalization processing can be as shown in the following formula (4).
[0185]
[0186] B′ = 0.5 (if B max = 1)
[0187] Figure 14 Schematically shows a method flowchart for identifying the type of analysis metrics using a second artificial intelligence model in a document management method according to another embodiment of the present disclosure.
[0188] As Figure 14 shown, in combination with Figure 8 , when identifying the type of the first analysis metric (operation S820) in process 1400, the second artificial intelligence model 1401 can be used to identify that the type of the first analysis metric is one of metric type 1, metric type 2, or metric type 3 (only as an example), where the second artificial intelligence model 1401 is a multi-classification model obtained based on machine learning technology.
[0189] In one embodiment, the second artificial intelligence model 1401 can adopt a Bidirectional Encoder Representations from Transformers (BERT). Among them, the BERT model is a type of keyword extraction model. According to the embodiments of the present disclosure, a classification layer for multi-classification can be added to the BERT model. Input the preprocessed word vectors of various types of text data into the BERT model, extract the semantic information of the text through the Chinese encoder of the BERT model, and obtain the sequence encoding of the input text; on this basis, perform parameter layer feature extraction, pooling layer processing, and random masking processing respectively. After training the model to a convergent state, for subsequent newly identified operation analysis metrics, the BERT model can be used to analyze the type to which the metric belongs.
[0190] According to an embodiment of the present disclosure, the metric type can be a type divided based on the analysis object. For example, in the process of product operation analysis, the operation situation of the product is often comprehensively analyzed from perspectives such as the product itself, customers, and partners, so as to provide a decision-making basis for work such as product marketing promotion, function optimization, and partner screening. Therefore, different types of metrics can be established for different analysis objects or analysis perspectives to better characterize the characteristics of the product.
[0191] In one embodiment, the indicator types include at least one of the following: indicators for the product itself, indicators for customers, or indicators for partners. For example, for financial products, the indicator types in the product operation analysis report may include, but are not limited to, the product itself, customers (including individual customers, corporate customers, etc.), and partners. Among them, each analysis indicator in the product operation analysis report can belong to one of the above three categories. Thus, in process 1400, the second artificial intelligence model 1401 can be used to identify whether an analysis indicator is an indicator for the product itself, an indicator for customers, or an indicator for partners.
[0192] Figure 15 Schematically shows a flowchart of a method for training a second artificial intelligence model according to an embodiment of the present disclosure.
[0193] As Figure 15 shown, in combination with Figure 14 , the method 1500 for training the second artificial intelligence model 1401 according to this embodiment may include operation S1501 to operation S1504.
[0194] In operation S1501, obtain at least one second analysis indicator. Herein, the "second analysis indicator" refers to the analysis indicator used when training the second artificial intelligence model 1401, and can be, for example, an analysis indicator collected, sorted out, or extracted from the topics described in the first document 201 or a large number of documents in the relevant field.
[0195] In operation S1502, segment the second analysis indicator and convert it into a word vector to obtain a second feature vector of the second analysis indicator. For example, word segmentation processing can be performed through a general word segmentation library, and the second analysis indicator can be converted into a word vector.
[0196] In operation S1503, label the indicator type of the second analysis indicator. For example, when the indicator types are divided into three types: indicators for the product itself, indicators for customers, and indicators for partners, corresponding label information can be labeled for the second analysis indicator according to the indicator type to which the second analysis indicator belongs.
[0197] In operation S1504, use the second feature vector as the input of the second artificial intelligence model 1401, and use the labeled indicator type of the second analysis indicator as the output reference of the second artificial intelligence model 1401 to train the second artificial intelligence model 1401. When the second artificial intelligence model 1401 converges (for example, when after a certain number of training rounds, the output results do not change or the change ratio is less than a preset ratio in several consecutive training rounds), and the accuracy rate meets the requirements (for example, the accuracy rate > 90%), it can be considered that the training of the second artificial intelligence model 1401 is completed.
[0198] On the premise of a large amount of training data, a good recognition accuracy can be achieved (for example, when the number of training samples > 5000, the accuracy can be > 90%). And as the training data volume and sample diversity increase, the accuracy of the second artificial intelligence model 1401 can be further improved.
[0199] Figure 16 Schematically shows a flowchart of training a second artificial intelligence model according to another embodiment of the present disclosure.
[0200] As Figure 16 shown, the method 1600 for training the second artificial intelligence model according to this embodiment may include operation S1501 to operation S1503, and operation S1604 to operation S1606.
[0201] In operation S1501, obtain at least one second analysis metric.
[0202] In operation S1502, segment the second analysis metric and convert it into a word vector to obtain a second feature vector of the second analysis metric.
[0203] In operation S1503, label the metric type of the second analysis metric.
[0204] Among them, operations S1501 to S1503 can refer to the introduction in method 1500, which will not be elaborated here.
[0205] Then in operation S1604, use the second feature vector as the input of the second artificial intelligence model 1401 to obtain the output of the second artificial intelligence model 1401.
[0206] Next in operation S1605, manually review the output of the second artificial intelligence model 1401.
[0207] After that in operation S1606, based on the difference between the output result after manual review and the metric type labeled for the second analysis metric, train the second artificial intelligence model 1401.
[0208] According to the embodiment of the present disclosure, in order to ensure the accuracy of the second artificial intelligence model 1401 and facilitate the retrieval and reuse of documents, the result automatically recognized by the second artificial intelligence model 1401 during the training process can be further optimized by setting up manual review. On the one hand, this can improve the accuracy of the final presentation result; on the other hand, the second artificial intelligence model 1401 can be trained based on the result adjusted by manual review to further improve the prediction accuracy of the second artificial intelligence model 1401.
[0209] Figure 17 Schematically shows a flowchart of a document management method according to still another embodiment of the disclosure.
[0210] As Figure 17 shown, the document management method 1700 according to this embodiment may include operations S1 to S4. Among them, the method 1700 is used to manage product operation analysis reports.
[0211] First, in operation S1, product operation indicators are identified from the product operation analysis report. In this embodiment, the product operation analysis report is a specific embodiment of the "first document" described above. Correspondingly, the product operation indicators are a specific embodiment of the "first analysis indicators". Therefore, the specific implementation of operation S1 can refer to Figures 3 to 7 the relevant introduction in [reference] regarding the identification of the first analysis indicators from the first document.
[0212] Then, in operation S2, for each identified product operation indicator, its indicator type is determined. The specific implementation can refer to the relevant introduction in the foregoing Figures 14 to 16 [reference] regarding the identification of indicator types.
[0213] Next, in operation S3, key indicators are determined from the obtained product operation indicators. The specific implementation can refer to the relevant introduction in the foregoing [reference] regarding the key nature of the identification of the first analysis indicators in the first document 201. Figures 9 to 13 [reference] regarding the key nature of the identification of the first analysis indicators in the first document 201.
[0214] Finally, in operation S4, indicator labels are established for the product operation analysis report for index retrieval of different reports. The indicator labels may include, for example, the names of the product operation indicators in the product operation analysis report, the names of the key indicators, and / or the indicator types of the key indicators, etc.
[0215] For example, for a certain product operation analysis report, through the analysis and processing of operations S1 to S4, the content of the finally obtained indicator labels may include: (1) The product operation indicators are: the number of effective accounts, the number of partners going online, the amount of capital precipitation, and the transaction volume of the customer group; (2) The indicator types of the above product operation indicators are: product, partner, product, customer; (3) The key indicator is: the number of effective accounts.
[0216] It can be understood that an enterprise can store the indicator labels of many product operation analysis reports and support multi-dimensional retrieval. For example, retrieving all relevant reports with the same key indicator; obtaining all key indicators of the same indicator type, etc. Therefore, for product operation analysts, they can conveniently retrieve the analysis indicators of existing reports, thereby reducing the research time of operation analysis and improving the efficiency of operation analysis.
[0217] It can be seen that, according to the embodiments of the present disclosure, the method 1700 realizes the effect of automatically identifying product operation indicators and their indicator types from operation analysis reports through artificial intelligence technology, and establishes indicator tags, and stores the product operation indicators and the indicator tags together as part of the data analysis assets. Further, after establishing the indicator tags, it can support the retrieval of indicators, greatly improving the convenience of product operation analysts in finding the indicators of previous analysis reports, finding similar indicators, etc., thereby improving the usability of the data analysis assets and optimizing the efficiency of product operation analysis.
[0218] Based on the document management methods of the above various embodiments, the embodiments of the present disclosure also provide a document management device. The following will be combined with Figures 18 to 21 to describe the document management device of the embodiments of the present disclosure in detail.
[0219] Figure 18 The structural block diagram of a document management device 1800 according to an embodiment of the present disclosure is schematically shown.
[0220] As Figure 18 shown, according to the embodiments of the present disclosure, the device 1800 may include a first acquisition module 1810, a first identification module 1820, and an indicator tag establishment module 1830. According to another embodiment of the present disclosure, the device 1800 may further include a second identification module 1840. The device 1800 may be used to implement the method described with reference to Figures 2 to 17 the foregoing.
[0221] The first acquisition module 1810 is used to acquire a first document. In one embodiment, the first acquisition module 1810 may perform the operation S310 described above.
[0222] The first identification module 1820 is used to identify the first analysis indicators that appear in the statements of the first document. In one embodiment, the first identification module 1820 may perform the operation S320 described above.
[0223] The indicator tag establishment module 1830 is used to establish indicator tags for the first document based on the first analysis indicators. In one embodiment, the indicator tag establishment module 1830 may perform the operation S330 described above.
[0224] The second recognition module 1840 is used to recognize the attribute information of the first analysis metric, where the attribute information includes at least one of the following: criticality or metric type in the first document; where the criticality is used to indicate whether the first analysis metric is a key metric in the first document. Accordingly, the metric label establishment module 1830 is further used to construct the content of the metric label based on the first analysis metric and the attribute information of the first analysis metric. In one embodiment, the second recognition module 1840 may perform operation S820 described above. Correspondingly, the metric label establishment module 1830 may perform operation S830 described above.
[0225] According to an embodiment of the present disclosure, the second recognition module 1840 may include at least one of a key metric recognition module 1841 or a metric type recognition module 1842.
[0226] The key metric recognition module 1841 is used to recognize the criticality of the first analysis metric in the first document, that is, to recognize whether the first analysis metric is a key metric.
[0227] The metric type recognition module 1842 is used to recognize the metric type of the first analysis metric.
[0228] Figure 19 A structural block diagram of the first recognition module 1820 in the document management device according to an embodiment of the present disclosure is schematically shown.
[0229] As Figure 19 shown, according to this embodiment, the first recognition module 1820 may include a word segmentation sub-module 1921, a first recognition sub-module 1922, a metric output sub-module 1923, and a first training sub-module 1924. Among them, the first recognition module 1820 may be used to recognize the first analysis metric in the first document by using a first artificial intelligence model, where the first artificial intelligence model is obtained based on natural language processing and machine learning technologies.
[0230] Specifically, the word segmentation sub-module 1921 is used to perform word segmentation processing on the statements in the first document. In one implementation, the word segmentation sub-module 1921 may perform operation S421 described above.
[0231] The first recognition sub-module 1922 includes a first artificial intelligence model and is used to recognize the relationship between each word in the word-segmented first document and the first analysis metric by using the first artificial intelligence model. In one embodiment, the first recognition sub-module 1922 may perform operation S422 described above.
[0232] The index output sub-module 1923 is configured to output a word or a combination of consecutive words related to the first analysis index based on the relationship between each word identified by the first artificial intelligence model and the first analysis index, so as to obtain the first analysis index. In one embodiment, the index output sub-module 1923 may perform the operation S423 described above.
[0233] The first training sub-module 1924 can be used to train the first artificial intelligence model. The specific training process may include: obtaining at least one second document; performing word segmentation on the statements in the second document as training data; annotating each word in the training data based on the relationship between each word in the word-segmented training data and the first analysis index; and training the first artificial intelligence model using the annotated training data. In one embodiment, the first training sub-module 1924 may perform the operations S610 to S640 described above.
[0234] For a detailed description of the functions of the various modules in the first recognition module 1820 in this embodiment, reference may be made to the relevant introduction above Figure 4 and Figure 6 are not elaborated here.
[0235] Figure 20 Schematically shows a structural block diagram of the key index recognition module 1841 in a document management device according to an embodiment of the present disclosure.
[0236] As Figure 20 shown, according to this embodiment, the key index recognition module 1841 may include an evaluation factor acquisition sub-module 2001, a feature vector combination sub-module 2002, an index evaluation regression model 2003, and a second training sub-module 2004.
[0237] The evaluation factor acquisition sub-module 2001 is configured to obtain the numerical values of M evaluation factors for evaluating the criticality of the first analysis index in the first document, where M is an integer greater than or equal to 2. In one embodiment, the M evaluation factors include at least one of the following: the position where the first analysis index appears in the first document; the analysis space of the first analysis index in the first document; or the number of times the first analysis index appears in the first document. In one embodiment, the evaluation factor acquisition sub-module 2001 may perform the operation S901 described above.
[0238] The feature vector combination sub-module 2002 is configured to obtain a first feature vector of the first analysis index based on the numerical values of the M evaluation factors. In one embodiment, the feature vector combination sub-module 2002 may perform the operation S902 described above.
[0239] The index evaluation regression model 2003 is used to take the first feature vector as input and predict the criticality of the first analysis index in the first document. In one embodiment, the index evaluation regression model 2003 may perform the operation S903 described above.
[0240] The second training sub-module 2004 is used to train the index evaluation regression model 2003.
[0241] For a detailed introduction to the functions of the various modules in the key index identification module 1841 in this embodiment, reference may be made to the above description of Figure 9 and will not be elaborated here.
[0242] Figure 21 Schematically shows a structural block diagram of the index type identification module 1842 in a document management device according to an embodiment of the present disclosure.
[0243] As Figure 21 shown, the index type identification module 1842 according to this embodiment may include a second artificial intelligence model 2101 and a third training sub-module 2102.
[0244] The second artificial intelligence model 2101 is used to identify the index type of the first analysis index, where the second artificial intelligence model is a multi-classification model obtained based on machine learning technology.
[0245] The third training sub-module 2102 is used to train the second artificial intelligence model 2101. The specific training process includes: obtaining at least one second analysis index; segmenting the second analysis index and converting it into a word vector to obtain a second feature vector of the second analysis index; labeling the index type of the second analysis index; and using the second feature vector as the input of the second artificial intelligence model and the labeled index type of the second analysis index as the output reference of the second artificial intelligence model to train the second artificial intelligence model 2101. In one embodiment, the third training sub-module 2102 may perform operations S1501 to S1504.
[0246] According to another embodiment of the present disclosure, during the training process of the second artificial intelligence model 2101 by the third training sub-module 2102, manual review may be set, where during the training process, the output of the second artificial intelligence model 2101 is manually reviewed; and based on the difference between the output result after manual review and the labeled index type of the second analysis index, the second artificial intelligence model is trained. In one embodiment, the third training sub-module 2102 may also perform operations S1604 to S1606.
[0247] For a detailed description of the functions of the various modules in the index type identification module 1842 in this embodiment, reference may be made to the above description of Figure 15 andFigure 16 For the introduction, it will not be elaborated here.
[0248] According to embodiments of the present disclosure, any plurality of modules among the first acquisition module 1810, the first recognition module 1820 and / or its sub-modules, the index label establishment module 1830, the key index recognition module 1841 and / or its sub-modules, and the index type recognition module 1842 and / or its sub-modules can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to embodiments of the present disclosure, at least one of the first acquisition module 1810, the first recognition module 1820 and / or its sub-modules, the index label establishment module 1830, the key index recognition module 1841 and / or its sub-modules, and the index type recognition module 1842 and / or its sub-modules can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as hardware or firmware for integrating or packaging circuits, or implemented in any one of the three implementation manners of software, hardware, and firmware or in an appropriate combination of any several of them. Alternatively, at least one of the first acquisition module 1810, the first recognition module 1820 and / or its sub-modules, the index label establishment module 1830, the key index recognition module 1841 and / or its sub-modules, and the index type recognition module 1842 and / or its sub-modules can be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions can be executed.
[0249] Figure 22 A block diagram of an electronic device suitable for implementing the document management method according to an embodiment of the present disclosure is schematically shown.
[0250] As Figure 22 shown, the electronic device 2200 according to an embodiment of the present disclosure includes a processor 2201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 2202 or a program loaded from a storage section 2208 into a random access memory (RAM) 2203. The processor 2201 can include, for example, a general microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 2201 can also include on-board memory for caching purposes. The processor 2201 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0251] In the RAM 2203, various programs and data required for the operation of the electronic device 2200 are stored. The processor 2201, the ROM 2202, and the RAM 2203 are connected to each other via a bus 2204. The processor 2201 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 2202 and / or the RAM 2203. It should be noted that the programs may also be stored in one or more memories other than the ROM 2202 and the RAM 2203. The processor 2201 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0252] According to an embodiment of the present disclosure, the electronic device 2200 may further include an input / output (I / O) interface 2205, and the input / output (I / O) interface 2205 is also connected to the bus 2204. The electronic device 2200 may further include one or more of the following components connected to the I / O interface 2205: an input portion 2206 including a keyboard, a mouse, etc.; an output portion 2207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 2208 including a hard disk, etc.; and a communication portion 2209 including a network interface card such as a LAN card, a modem, etc. The communication portion 2209 performs communication processing via a network such as the Internet. A drive 2210 is also connected to the I / O interface 2205 as needed. A removable medium 2211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 2210 as needed so that a computer program read from it can be installed into the storage portion 2208 as needed.
[0253] The present disclosure also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiments; or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0254] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 2202 and / or RAM 2203 and / or one or more memories other than ROM 2202 and RAM 2203.
[0255] An embodiment of the present disclosure further includes a computer program product, which includes a computer program, and the computer program contains program codes for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program codes are used to enable the computer system to implement the method provided by the embodiment of the present disclosure.
[0256] When the computer program is executed by the processor 2201, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0257] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program can also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 2209, and / or be installed from the removable medium 2211. The program codes included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0258] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 2209, and / or be installed from the removable medium 2211. When the computer program is executed by the processor 2201, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0259] In accordance with embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0260] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0261] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0262] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A document management method, comprising: Obtaining a first document; Using a first artificial intelligence model to identify a first analysis metric that appears in the statements of the first document; wherein, the first artificial intelligence model is obtained based on natural language processing and machine learning technologies; Identifying attribute information of the first analysis metric, the attribute information including a metric type and the criticality of the first analysis metric in the first document; wherein, the metric type of the first analysis metric is a type divided according to any one or more of the following dimensions: a type obtained by dividing the first analysis metric based on the analysis object of the metric, a type obtained by dividing the first analysis metric based on the use of the metric, or a type divided based on the qualitative or quantitative characteristics of the metric; the criticality is used to indicate whether the first analysis metric is a key metric in the first document; and Based on the first analysis metric and the attribute information of the first analysis metric, establishing an index label for the first document, wherein the index label includes the name, magnitude, metric type, and criticality of the first analysis metric; Wherein, the identifying the attribute information of the first analysis metric further includes: Obtaining numerical values of M evaluation factors for evaluating the criticality of the first analysis metric in the first document, where M is an integer greater than or equal to 2; Based on the numerical values of the M evaluation factors, obtaining a first feature vector of the first analysis metric; and Using the first feature vector as an input to an index evaluation regression model, and determining the criticality of the first analysis metric in the first document based on the output of the index evaluation regression model.
2. The method according to claim 1, wherein The using the first artificial intelligence model to identify the first analysis metric in the first document includes: Performing word segmentation processing on the statements in the first document; Using the first artificial intelligence model to identify the relationship between each word in the word-segmented first document and the first analysis metric; and Based on the relationship between each word identified by the first artificial intelligence model and the first analysis metric, combining one word or a continuous plurality of words related to the first analysis metric and outputting them to obtain the first analysis metric.
3. The method according to claim 2, wherein The relationship between each word identified by the first artificial intelligence model and the first analysis metric includes: Related to the first analysis metric or unrelated to the first analysis metric; Wherein, Related to the first analysis metric includes at least one of the following: located at the beginning of the first analysis metric, located in the middle of the first analysis metric, or located at the end of the first analysis metric.
4. The method according to any one of claims 1 to 3, wherein The first artificial intelligence model is trained through the following method: Obtaining at least one second document; Using the statements in the second document as training data and performing word segmentation on the training data; Based on the relationship between each word in the word-segmented training data and the first analysis metric, annotating each word in the training data; And Using the annotated training data to train the first artificial intelligence model.
5. The method according to claim 4, wherein The first artificial intelligence model adopts a conditional random field model.
6. The method according to claim 1, wherein, Before establishing the index label for the first document, the method further includes: When multiple of the first analysis indicators are identified, based on the semantic analysis of the first analysis indicators, calculate the similarity between every two of the first analysis indicators; and merge every two of the first analysis indicators whose similarity is greater than the similarity threshold; and / or Count the number of occurrences of each of the identified first analysis indicators in the first document, and eliminate the first analysis indicators whose number of occurrences meets the elimination condition.
7. The method according to claim 1, wherein, The M evaluation factors include at least one of the following: The position where the first analysis indicator appears in the first document; The analysis space of the first analysis indicator in the first document; or The number of occurrences of the first analysis indicator in the first document.
8. The method according to claim 7, wherein, The obtaining the values of the M evaluation factors for evaluating the criticality of the first analysis indicator in the first document includes: Obtaining the value for characterizing the position where the first analysis indicator appears in the first document, specifically including: Retrieve the first appearance positions of the N first analysis indicators identified from the first document in the first document, where N is an integer greater than or equal to 2; Number the N first analysis indicators based on the order of the first appearance positions; and Based on the number of each first analysis indicator, determine the value for characterizing the position where each first analysis indicator appears in the first document.
9. The method according to claim 7, wherein, The obtaining the values of the M evaluation factors for evaluating the criticality of the first analysis indicator in the first document includes: Obtaining the value for characterizing the analysis space of the first analysis indicator in the first document, specifically including: Obtain the title level of the title to which the first analysis indicator belongs in the first document to obtain the target title level; where the title level is determined according to the title hierarchy structure; and Based on the target title level, obtain the value for characterizing the analysis space of the first analysis indicator in the first document.
10. The method according to claim 9, wherein, The obtaining the title level of the title to which the first analysis indicator belongs in the first document includes: When the first analysis indicator appears in the title of the first document, obtain the title level of the title where the first analysis indicator is located; or When the first analysis indicator does not appear in the title of the first document, determine the title to which the paragraph where the first analysis indicator is located belongs, and obtain the title level of this title.
11. The method according to claim 9, wherein, The obtaining the value for characterizing the analysis space of the first analysis indicator in the first document based on the target title level includes: Based on the preset conversion relationship between the title level and the value, convert the highest title level in the first document to obtain a first value; the highest title level is the level of the title located at the top layer in the title hierarchy structure; Based on the conversion correspondence between the title level and the value, convert the target title level to obtain a second value; and Using the first value as a parameter of a preset normalization model and the second value as a variable of the normalization model, a value for characterizing the analyzed length of the first analysis metric in the first document is calculated.
12. The method according to claim 11, wherein, The method further includes: Setting a conversion relationship between the title level and the value, where the higher the position of the title level in the title hierarchy structure, the larger the converted value.
13. The method according to claim 1, wherein, The identifying the attribute information of the first analysis metric includes: Using a second artificial intelligence model to identify the metric type of the first analysis metric, where the second artificial intelligence model is a multi-classification model obtained based on machine learning technology.
14. The method according to claim 13, wherein The second artificial intelligence model is trained in the following manner: Obtaining at least one second analysis metric; Performing word segmentation on the second analysis metric and converting it into a word vector to obtain a second feature vector of the second analysis metric; Annotating the metric type of the second analysis metric; and Using the second feature vector as the input of the second artificial intelligence model and using the metric type annotated for the second analysis metric as the output reference of the second artificial intelligence model to train the second artificial intelligence model.
15. The method according to claim 14, wherein, The using the second feature vector as the input of the second artificial intelligence model and using the metric type annotated for the second analysis metric as the output reference of the second artificial intelligence model to train the second artificial intelligence model further includes: Performing manual review on the output of the second artificial intelligence model; and Training the second artificial intelligence model based on the difference between the output result after manual review and the metric type annotated for the second analysis metric.
16. The method according to any one of claims 13 to 15, wherein, The second artificial intelligence model adopts a BERT model.
17. The method according to any one of claims 13 to 15, wherein The types obtained by dividing the first analysis metric based on the analysis object of the metric include at least one of the following: Metrics for the product itself, metrics for customers, or metrics for partners.
18. A document management device, including: A first acquisition module, configured to acquire a first document; A first identification module, configured to use a first artificial intelligence model to identify a first analysis metric that appears in the statements of the first document, where the first artificial intelligence model is obtained based on natural language processing and machine learning technologies; A second identification module, configured to identify the attribute information of the first analysis metric, where the attribute information includes the metric type and the criticality of the first analysis metric in the first document; where the metric type of the first analysis metric is a type divided according to any one or more of the following dimensions: the type obtained by dividing the first analysis metric based on the analysis object of the metric, the type obtained by dividing the first analysis metric based on the use of the metric, or the type divided based on the qualitative or quantitative characteristics of the metric; the criticality is used to indicate whether the first analysis metric is a key metric in the first document; and An index label establishment module, configured to establish an index label for the first document based on the first analysis index and the attribute information of the first analysis index, where the index label includes the name, size, index type, and the criticality of the first analysis index; Wherein, the second identification module is further configured to: obtain numerical values of M evaluation factors for evaluating the criticality of the first analysis index in the first document, where M is an integer greater than or equal to 2; obtain a first feature vector of the first analysis index based on the numerical values of the M evaluation factors; and use the first feature vector as an input to an index evaluation regression model, and determine the criticality of the first analysis index in the first document based on the output of the index evaluation regression model.
19. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 17.
20. A computer-readable storage medium, having executable instructions stored thereon, which when executed by a processor cause the processor to execute the method according to any one of claims 1 to 17.
21. A computer program product, comprising a computer program, which when executed by a processor implements the method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Abstraction text extraction method and device, storage medium and electronic equipment
CN112052308A
Unstructured data document processing method and related equipment
CN113642569A
Document numerical index extraction method and device based on machine learning
CN114398853A