Data analysis report multiplexing method, device, equipment, medium and program product

CN115688762BActive Publication Date: 2026-09-25INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210639246.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-07
Publication Date
2026-09-25
Estimated Expiration
2042-06-07

AI Technical Summary

Benefits of technology

[0021](1)本公开通过构建结构化的段落实体,并将数据分析报告的各个段落所分析的数据类型进行分类,最终得到结构化的知识图谱,便于后续对数据分析报告的零散化复用。本公开的知识图谱构建方法通用,可以将不同类型的数据分析报告进行结构化整合,降低了数据分析报告的复用门槛,提高了复用效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115688762B_ABST
    Figure CN115688762B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data analysis report reuse method, which can be applied to the fields of artificial intelligence and financial technology. The method comprises: extracting nouns in a data analysis report in units of paragraphs; determining entity objects according to the occurrence frequency of the nouns and the positions of the nouns in the paragraphs; obtaining label attributes and content attributes corresponding to the entity objects, wherein the label attributes are used to represent the form of analyzing the entity objects, and the content attributes are used to represent the content of analyzing the entity objects; constructing paragraph entities of each paragraph according to the label attributes and the content attributes; classifying each paragraph according to the hierarchical structure among the paragraphs to obtain analysis topics; constructing a knowledge graph according to the paragraph entities and the analysis topics; and reusing the data analysis report according to the knowledge graph. The present disclosure also provides a data analysis report reuse device, equipment, storage medium and program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of artificial intelligence and finance, and specifically to a data analysis report reuse method, apparatus, device, medium, and program product. Background Technology

[0002] With the accumulation of data assets and the widespread application of big data technology, the demand for data analysis is growing rapidly across various industries. By generating scientific, effective, and accurate data reports, it is possible to understand the current state of affairs and gain insights into future trends. However, currently, producing scientifically sound, comprehensive, and accurate data analysis reports still requires a certain level of knowledge from the writers, making it difficult to meet the demand for large-scale and frequent data analysis reports. Furthermore, training specialized data analysis personnel is costly for enterprises and other organizations, often focusing on theoretical aspects with low returns.

[0003] Currently, relevant data analysis technologies suffer from low efficiency and accuracy, necessitating the development of specialized data analysis frameworks and processes for specific business needs. On one hand, the content logic chains are relatively rigid and lack automation, resulting in a lack of flexibility. On the other hand, they fail to integrate multi-dimensional analytical perspectives or utilize various tools to assist data analysis. Summary of the Invention

[0004] In view of the above problems, this disclosure provides a data analysis report reuse method, apparatus, device, medium and program product to at least partially solve the above technical problems.

[0005] According to the first aspect of this disclosure, a method for reusing data analysis reports is provided, comprising: extracting nouns from the data analysis report on a paragraph-by-paragraph basis; determining entity objects based on the frequency of occurrence of the nouns and their position within the paragraphs; obtaining the tag attributes and content attributes corresponding to the entity objects, wherein the tag attributes represent the form of the entity objects and the content attributes represent the content of the entity objects; constructing paragraph entities for each paragraph based on the tag attributes and content attributes; classifying each paragraph according to the hierarchical structure between paragraphs to obtain analysis topics; constructing a knowledge graph based on the paragraph entities and analysis topics; and reusing the data analysis report based on the knowledge graph.

[0006] According to embodiments of this disclosure, extracting nouns from a data analysis report, on a paragraph-by-paragraph basis, includes: constructing word vectors; performing part-of-speech tagging on the word vectors; and extracting the subject and object to obtain the nouns.

[0007] According to embodiments of this disclosure, determining entity objects based on the frequency of noun occurrence and the position of noun in a paragraph includes: obtaining a tag extraction model; inputting nouns into the tag extraction model to obtain nodes of the tag extraction model; setting a sliding window based on the paragraph length; calculating the weight of the node based on the number of times the noun co-occurs in the sliding window and the position of the noun in the paragraph; and determining entity objects based on the weight of the node.

[0008] According to embodiments of this disclosure, calculating the weight of the edge between nodes based on the number of times the noun co-occurs in the sliding window and the position of the noun in the paragraph includes: assigning the weight of the edge between nodes to the co-occurring nodes when the nouns co-occur in the sliding window; and assigning weight to the nodes when the noun is located in the first two sentences or the last two sentences of the paragraph.

[0009] According to embodiments of this disclosure, obtaining the tag attributes and content attributes corresponding to an entity object includes: obtaining chart tag attributes, data tag attributes, method tag attributes, and code tag attributes respectively; wherein, the chart tag attribute is the description content, data dimension, and visualization method of the chart corresponding to the paragraph; the data tag attribute is the data source tag associated with the paragraph; the method tag attribute is the analysis method and / or analysis model used by the paragraph; and the code tag attribute is the code block corresponding to the paragraph.

[0010] According to embodiments of this disclosure, obtaining chart label attributes includes: establishing an index relationship between paragraph text and chart; using a convolutional neural network model to classify and identify images to obtain image category labels; and constructing chart label attributes based on the index relationship and image category labels.

[0011] According to embodiments of this disclosure, obtaining data tag attributes includes: obtaining a pre-built data lineage dictionary; matching data source tags from the data lineage dictionary based on the text of the paragraph to obtain data tag attributes.

[0012] According to embodiments of this disclosure, obtaining method tag attributes includes: using natural language processing methods to obtain method tag attributes from the text of a paragraph.

[0013] According to embodiments of this disclosure, obtaining code tag attributes includes: obtaining code comment content; matching the text of a paragraph with the code comment content to obtain code tag attributes.

[0014] According to embodiments of this disclosure, classifying paragraphs based on their hierarchical structure to obtain an analysis topic includes: calculating a distance coefficient based on the number of intervals between paragraphs to obtain contextual relationships; extracting a table of contents or outline to obtain hierarchical relationships between paragraphs; and classifying paragraphs based on contextual and hierarchical relationships to obtain an analysis topic.

[0015] According to embodiments of this disclosure, reusing a data analysis report based on a knowledge graph includes: storing the knowledge graph in a graph database; constructing a search engine based on the graph database; using the search engine to search for the title of the data analysis report to obtain paragraph entities and analysis topics corresponding to the title; and directly referencing the searched paragraph entities and analysis topics or generating a report template.

[0016] The second aspect of this disclosure provides a data analysis report reuse device, comprising: a noun acquisition module for extracting nouns from the data analysis report on a paragraph-by-paragraph basis; an object determination module for determining entity objects based on the frequency of noun occurrence and the position of nouns in paragraphs; an attribute acquisition module for acquiring the tag attributes and content attributes corresponding to the entity objects, wherein the tag attributes represent the form of the analyzed entity objects and the content attributes represent the content of the analyzed entity objects; an entity construction module for constructing paragraph entities for each paragraph based on the tag attributes and content attributes; a topic classification module for classifying each paragraph according to the hierarchical structure between paragraphs to obtain the analysis topics; a graph construction module for constructing a knowledge graph based on the paragraph entities and the analysis topics; and a report reuse module for reusing the data analysis report based on the knowledge graph.

[0017] A third aspect of this disclosure provides an electronic device, including: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the data analysis report reuse method of any of the above embodiments.

[0018] A fourth aspect of this disclosure also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the data analysis report reuse method of any of the above embodiments.

[0019] The fifth aspect of this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the data analysis report reuse method of any of the above embodiments.

[0020] Compared with the prior art, the methods, apparatus, electronic devices, storage media, and program products provided in this disclosure have at least the following beneficial effects:

[0021] (1) This disclosure constructs structured paragraph entities and classifies the data types analyzed in each paragraph of the data analysis report to obtain a structured knowledge graph, which facilitates the subsequent reuse of fragmented data analysis reports. The knowledge graph construction method of this disclosure is universal and can structurally integrate different types of data analysis reports, reducing the reuse threshold of data analysis reports and improving reuse efficiency.

[0022] (2) This disclosure combines the writing characteristics of data analysis reports and incorporates the positional information of nouns into the calculation of the edge weights between nodes, thereby improving the accuracy of tag extraction.

[0023] (3) This disclosure constructs structured display attributes of entity objects from aspects such as chart label attributes, data label attributes, method label attributes and code label attributes. It can flexibly extract one or more attribute contents and extract analysis content from multiple dimensions from the data source, while taking into account its visualization display and code environment creation, thus improving the efficiency of data analysis report reuse or writing. Attached Figure Description

[0024] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0025] Figure 1 The illustration schematically depicts application scenarios of data analysis report reuse methods, apparatus, devices, media, and program products according to embodiments of this disclosure;

[0026] Figure 2 A flowchart illustrating a data analysis report reuse method according to an embodiment of the present disclosure is shown schematically.

[0027] Figure 3 A flowchart illustrating a method for extracting nouns according to an embodiment of the present disclosure is shown schematically.

[0028] Figure 4 A flowchart illustrating a method for determining entity objects according to embodiments of the present disclosure is shown schematically.

[0029] Figure 5 A flowchart illustrating a method for obtaining chart label attributes according to an embodiment of the present disclosure is shown schematically.

[0030] Figure 6 A flowchart illustrating a method for obtaining data tag attributes according to an embodiment of the present disclosure is shown schematically.

[0031] Figure 7 A flowchart illustrating a method for obtaining method tag attributes according to an embodiment of the present disclosure is shown schematically.

[0032] Figure 8 A flowchart illustrating a method for obtaining code tag attributes according to an embodiment of the present disclosure is shown schematically.

[0033] Figure 9 A flowchart illustrating a method for obtaining analysis topics according to embodiments of the present disclosure is shown schematically.

[0034] Figure 10 A knowledge graph according to an embodiment of the present disclosure is illustrated schematically;

[0035] Figure 11 A flowchart illustrating a method for reusing data analysis reports according to an embodiment of this disclosure is shown schematically.

[0036] Figure 12 A schematic diagram illustrating the structure of a data analysis report multiplexing apparatus according to embodiments of the present disclosure is shown; and

[0037] Figure 13 A block diagram schematically illustrates an electronic device suitable for implementing a data analysis report reuse method according to an embodiment of the present disclosure. Detailed Implementation

[0038] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0039] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0040] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0041] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).

[0042] This disclosure provides a data analysis report reuse method, apparatus, device, medium, and program product, which can be used in the financial field or other fields. It should be noted that the data analysis report reuse method, apparatus, device, medium, and program product of this disclosure can be used in the financial field, as well as in any field other than the financial field; the application field of the data analysis report reuse method, apparatus, device, medium, and program product of this disclosure is not limited.

[0043] In the technical solution disclosed herein, the acquisition, storage, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.

[0044] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0045] Figure 1 The illustration schematically depicts application scenarios of data analysis report reuse methods, apparatus, devices, media, and program products according to embodiments of the present disclosure.

[0046] like Figure 1 As shown, application scenario 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as a medium for providing a communication link between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0047] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0048] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0049] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0050] It should be noted that the data analysis report multiplexing method provided in this embodiment can generally be executed by server 105. Correspondingly, the data analysis report multiplexing device provided in this embodiment can generally be located in server 105. The data analysis report multiplexing method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the data analysis report multiplexing device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0051] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0052] The following will be based on Figure 1 The described scene, through Figures 2 to 11 The data analysis report reuse method of the disclosed embodiments is described in detail.

[0053] Figure 2 A flowchart illustrating a data analysis report reuse method according to an embodiment of this disclosure is shown schematically.

[0054] like Figure 2 As shown, embodiments of this disclosure provide a data analysis report reuse method, which includes, for example:

[0055] S210: Extract nouns from the data analysis report, paragraph by paragraph.

[0056] S220, determine the entity object based on the frequency of noun occurrence and the position of the noun in the paragraph.

[0057] S230, obtain the tag attribute and content attribute corresponding to the entity object. The tag attribute is used to represent the form of the entity object being analyzed, and the content attribute is used to represent the content of the entity object being analyzed.

[0058] S240, construct paragraph entities for each paragraph based on tag attributes and content attributes.

[0059] S250: Classify paragraphs according to their hierarchical structure to obtain the analysis topic.

[0060] S260, construct a knowledge graph based on paragraph entities and analysis topics.

[0061] S270 reuses data analysis reports based on knowledge graphs.

[0062] According to embodiments of this disclosure, the analysis theme and key points of the data analysis report are first identified by extracting key terms from each paragraph. For example, in the analysis of "forecast of new customer signings in the next quarter," the number of customers is the key term, which is also the entity that needs to be analyzed. Then, the paragraphs are used to find charts, data sources, analysis methods, code blocks, etc., associated with the key term, representing how the key term is analyzed, as well as the specific content-related attributes corresponding to these tag attributes, such as the content itself, content source, and content quantity. Finally, a structured paragraph entity is constructed by combining the tag attributes and content attributes for easy storage and retrieval. There are many types of data analysis reports, such as those analyzing user data, operational data, product data, industry data, etc. The same analysis report may also analyze data on one or more of these themes in different paragraphs. Therefore, it is necessary to categorize paragraphs analyzing different themes by paragraph to clearly present the analysis report structurally, and specifically display it in the form of a knowledge graph for subsequent retrieval.

[0063] Figure 3 A flowchart illustrating a method for extracting nouns according to an embodiment of the present disclosure is shown schematically.

[0064] According to embodiments of this disclosure, such as Figure 3 As shown, for example, key words in each paragraph can be extracted by operating S211 to S213.

[0065] S211, Construct word vectors.

[0066] According to embodiments of this disclosure, data analysis reports typically analyze one topic or type of data per paragraph, and thus, for example, text tag extraction preprocessing is performed on the data analysis report on a paragraph-by-paragraph basis.

[0067] Specifically, stop words are removed from the text content of the paragraph and word segmentation is performed. Then, word embedding technology (such as word2vec) is used to construct word vectors. Based on this, according to business practice, all word vectors are labeled with part-of-speech tags, and only the nouns representing the subject and object of the text are retained as candidate words. These candidate words are then used as input to a label extraction model (such as textRank).

[0068] S212, perform part-of-speech tagging on word vectors.

[0069] According to embodiments of this disclosure, after segmenting and tagging the parts of speech of each word in a paragraph, there are words with multiple parts of speech, such as nouns, verbs, adjectives, and adverbs, and they also play different roles in the original sentence, such as serving as subjects, predicates, objects, and adverbials. By focusing on subject-object nouns (i.e., nouns corresponding to subjects and objects), the objects of data analysis can be found more quickly, facilitating subsequent classification and retrieval.

[0070] S213, extract the subject and object to obtain the noun.

[0071] According to embodiments of this disclosure, the nouns obtained here as subjects and objects are the key nouns that need to be found.

[0072] Figure 4 A flowchart illustrating a method for determining entity objects according to an embodiment of the present disclosure is shown schematically.

[0073] According to embodiments of this disclosure, such as Figure 4 As shown, for example, operations S221 to S225 are used to determine the entity objects that need to be focused on expansion.

[0074] S221, Obtain the label extraction model.

[0075] According to embodiments of this disclosure, after obtaining the nouns, due to their large number, not all nouns are suitable for representing data analysis objects. Therefore, it is necessary to calculate the weight of each noun with respect to the data analysis topic and extract the nouns with high weights as the entity objects to be expanded. This disclosure, for example, uses textRank as a tag extraction model, treating the preprocessed nouns as semantic units and considering them as nodes in the tag extraction graph model.

[0076] S222, input the nouns into the label extraction model to obtain the nodes of the label extraction model.

[0077] S223, Set the sliding window according to the paragraph length.

[0078] According to embodiments of this disclosure, nouns are acquired on a paragraph-by-paragraph basis. When analyzing the weight of nouns within a paragraph, the sliding window length is also set based on the paragraph length. For example, the sliding window length is set to the length of a paragraph. If semantic units co-occur within a sliding window, these semantic units are considered to have a strong semantic relationship, and the weights of edges between nodes are assigned based on the number of co-occurrences.

[0079] S224, calculate the node weight based on the number of times the noun appears in the sliding window and the position of the noun in the paragraph.

[0080] According to embodiments of this disclosure, regarding weight calculation, for example, when nouns co-occur in a sliding window, weights are assigned to the co-occurring nodes. Also, when a noun is located in the first two or last two sentences of a paragraph, weights are assigned to the nodes. To improve the accuracy of tag extraction, this disclosure incorporates the positional information of semantic units into the calculation of the weights of the edges between nodes. For example, if a semantic unit exists in the first two sentences of a paragraph or the last two sentences of a paragraph, the positional weight is assigned a value of 1; otherwise, it is assigned a value of 0. After obtaining the edge weights between nodes, the weight of the semantic unit can be calculated using the PageRank algorithm.

[0081] S225, determine the entity object based on the node weight.

[0082] According to embodiments of this disclosure, based on business needs, for example, the semantic unit with the highest weight can be selected as the entity object for extraction. Obtaining the high-weight entity object that needs to be expanded greatly reduces the workload and also reduces the complexity of the system.

[0083] According to embodiments of this disclosure, after obtaining an entity object, there can be many dimensions for analyzing the entity object, such as analysis from a chart perspective, analysis using different data sources, analysis using different analysis methods, and analysis from the perspective of code comment content. Therefore, obtaining the tag attributes and content attributes corresponding to the entity object includes, for example, obtaining chart tag attributes, data tag attributes, method tag attributes, and code tag attributes. Specifically, chart tag attributes include, for example, the description content, data dimensions, and visualization method of the chart corresponding to the paragraph. Data tag attributes include, for example, the data source tag associated with the paragraph. Method tag attributes include, for example, the analysis method and / or analysis model used by the paragraph. Code tag attributes include, for example, the code block corresponding to the paragraph.

[0084] Figure 5 A flowchart illustrating a method for obtaining chart label attributes according to an embodiment of the present disclosure is shown.

[0085] According to embodiments of this disclosure, such as Figure 5 As shown, for example, the chart label attributes corresponding to the entity object can be obtained by operating S231 to S2333.

[0086] S231, Establish the index relationship between paragraph text and charts.

[0087] According to embodiments of this disclosure, firstly, based on the chart index keywords in the text, such as: "as shown in the table below", "as shown in the figure above", "as shown in the figure above", etc. Figure 5 As shown, establish an index relationship between paragraph text and charts.

[0088] S232 uses a convolutional neural network model to classify and recognize images, obtaining image category labels.

[0089] According to embodiments of this disclosure, for visual content such as charts associated with paragraphs, image classification and recognition can be performed using a convolutional neural network (CNN) model to obtain image category labels, such as histograms, bar charts, time series graphs, scatter plots, maps, and other category labels.

[0090] S233, construct chart label attributes based on index relationships and image category labels.

[0091] According to embodiments of this disclosure, by establishing an index relationship, the content of the chart description can be linked to the paragraph content. By combining the chart title with word segmentation, the chart description content (data category, data source, time dimension, etc.) can be obtained, and tag attributes such as the description content, data dimension, and visualization method of the paragraph chart can be constructed.

[0092] Figure 6 A flowchart illustrating a method for obtaining data tag attributes according to an embodiment of the present disclosure is shown.

[0093] According to embodiments of this disclosure, such as Figure 6 As shown, for example, the data tag attributes corresponding to the entity object are obtained by operating S234 to S235.

[0094] S234, obtain the pre-built data lineage dictionary.

[0095] According to embodiments of this disclosure, the data lineage dictionary fully reflects the relationships between the analyzed data, providing a clear understanding of the data's origin and development. The data lineage dictionary can be established simultaneously during the data accumulation process.

[0096] S235: Based on the text of the paragraph, match the data source tags from the data lineage dictionary to obtain the data tag attributes.

[0097] According to embodiments of this disclosure, paragraph-related data is matched using a pre-constructed data lineage dictionary. Based on information such as analysis content, data charts, and data definitions extracted from the paragraph content, the closest data source tags in the data warehouse can be automatically matched from the data lineage dictionary. By obtaining data tag attributes, the utilization of big data-based data analysis reports can be achieved, improving the depth and breadth of data analysis, thereby enhancing the effectiveness and reliability of data analysis.

[0098] Figure 7 A flowchart illustrating a method for obtaining method tag attributes according to an embodiment of the present disclosure is shown.

[0099] According to embodiments of this disclosure, such as Figure 7 As shown, for example, the method tag attribute corresponding to the entity object can be obtained by operating S236.

[0100] S236 uses natural language processing methods to obtain method label attributes from the text of a paragraph.

[0101] According to embodiments of this disclosure, natural language processing methods such as keyword recognition and semantic recognition can be used to extract analytical method features from paragraphs. For example, the phrase "compared with the same period last year" can indicate that a paragraph uses a comparative analysis method, and the combination of keywords such as "p-value less than 0.05" and "independence hypothesis" can indicate that the paragraph uses an independence test. Analytical method features may include, for example, data analysis methods and data analysis models.

[0102] Figure 8 A flowchart illustrating a method for obtaining code tag attributes according to an embodiment of the present disclosure is shown.

[0103] According to embodiments of this disclosure, such as Figure 8 As shown, for example, the code tag attribute corresponding to the entity object can be obtained by operating S237 to S238.

[0104] S237, retrieve code comment content.

[0105] S238 matches the paragraph text with the code comment content to obtain the code tag attribute.

[0106] According to embodiments of this disclosure, the extracted tag information, such as paragraph association data, methods, and models, is matched with the comments in the code accompanying the report to achieve mutual binding between paragraphs and code modules. The next step is to extract the language format and dependent environment of the code module based on its code format and header file. After obtaining the code tag attributes, the content related to the data to be analyzed can be directly called at the code level.

[0107] After extracting the attributes of the paragraph-related element tags through steps S231 to S238, the paragraph entities of each paragraph can be constructed based on the tag attributes and content attributes.

[0108] Specifically, the analysis content, methods, data, models, code, and charts obtained from the paragraphs above are used as tags for paragraph entities. The specific content extracted from each tag is used as the entity attribute (i.e., content attribute) corresponding to that tag. This allows for the construction of a unique class for each paragraph entity, facilitating the storage of data processing oriented towards entity nodes and relationships in the graph database. A typical data structure for a paragraph entity is shown below:

[0109]

[0110]

[0111] Figure 9A flowchart illustrating a method for obtaining analysis topics according to an embodiment of this disclosure is shown schematically.

[0112] According to embodiments of this disclosure, such as Figure 9 As shown, for example, the analysis topic can be obtained by operating S251 to S253.

[0113] S251, calculate the distance coefficient based on the number of intervals in each paragraph to obtain the context relationship.

[0114] According to embodiments of this disclosure, the paragraph hierarchy of the analysis report is used to extract mapping relationships between paragraphs from the context and article hierarchy, including, for example, contextual relationships and hierarchical relationships. For contextual relationships, a distance coefficient is generated, for example, using the number of paragraphs between them. The distance between paragraphs represents the standardized weighted distance coefficient between two paragraphs, distributed between [0, 1], with smaller numbers indicating a stronger correlation between the two paragraphs.

[0115] S252, extract the table of contents or outline to obtain the hierarchical relationship of each paragraph.

[0116] According to embodiments of this disclosure, for hierarchical relationships, methods such as extracting directories or report outlines are used to mine paragraph hierarchical or parallel relationships to obtain relationship categories. Relationship categories represent the association between paragraph B and paragraph A, such as parallel, inclusive, or subordinate.

[0117] S253, classify each paragraph according to context and hierarchy to obtain the analysis topic.

[0118] Figure 10 A knowledge graph according to an embodiment of the present disclosure is illustrated schematically.

[0119] According to embodiments of this disclosure, entity relationship attributes are formed by combining hierarchical and contextual relationships. The extracted paragraph relationships, for example, include the paragraph IDs, relationship categories, and distances of the two described paragraphs. Wherein: paragraph A represents the unique ID of paragraph A, and paragraph B represents the unique ID of paragraph B. A complete paragraph relationship representation is shown in the following example:

[0120] {Paragraph A, Paragraph B, parallel, 0.5}

[0121] {Paragraph A, Paragraph C, Contains, 0}

[0122] {Paragraph B, Paragraph C, Contains, 0}.

[0123] After categorizing the paragraphs, and combining them with their main themes, paragraphs on different themes can be grouped and summarized to form categories such as... Figure 10The knowledge graph shown is illustrated. Paragraph topics can be any one or more of the following: business analysis, user analysis, product analysis, and industry analysis. The paragraph entity models constructed through the above process are automatically matched according to the dimensions of report entities, paragraph entities, relationships, and attributes to complete the construction of the knowledge graph, which is then stored in a graph database. A search engine built based on this graph database enables the simple reuse of data analysis reports and analysis process assets.

[0124] Figure 11 A flowchart illustrating a method for reusing data analysis reports according to an embodiment of this disclosure is shown.

[0125] According to embodiments of this disclosure, such as Figure 11 As shown, for example, data analysis reports can be reused through operations S271 to S274.

[0126] S271, store the knowledge graph in a graph database.

[0127] S272, Build a search engine based on a graph database.

[0128] S273, use a search engine to search for the title of the data analysis report to obtain the paragraph entity and analysis topic corresponding to the title.

[0129] According to embodiments of this disclosure, after constructing the knowledge graph, it can be stored in a graph database and retrieved by topic when needed. Topic retrieval mainly refers to using data analysis process assets (i.e., structured paragraph entities constructed based on data analysis reports) as search results for basic browsing and viewing. Search results may be displayed as a single data entry, such as the report title. Clicking on a result displays the tag attributes, content attributes, and paragraph topic categories of all paragraph entities in the report. The display methods include, for example, visual node display and structured text display. By default, all entity nodes related to the search topic and their corresponding tag attributes are displayed. If the user customizes the tags during the search, the display can be personalized. For example, selecting only the "Analysis Content" tag and clicking "Search" will only display all "Analysis Content" sections from the relevant analysis reports.

[0130] S274 allows for direct referencing of searched paragraph entities and analysis topics, or the generation of report templates.

[0131] According to embodiments of this disclosure, in addition to browsing and viewing search results, users can also quickly reference or generate templates. Quick referencing refers to the original, unprocessed direct reference of data analysis process assets, suitable for writing comparative analysis reports such as periodic trend analysis. For paragraph entities in the topic search results, they can be directly placed in the document page in the right column of the search results page by dragging and dropping, and displayed as structured text. Users can directly reuse them or edit the text content according to their personalized needs. Based on this, users can quickly reuse existing data analysis process assets in comparative analysis and other applications.

[0132] Template generation refers to the visual display of data analysis process assets based on a specific theme, which can automatically generate a draft data analysis report based on the current search results. After clicking on the report title to display a single report based on the search results for a specific theme, the client selects tag attributes (where "Analysis Time Range" is mandatory, while "Data Source," "Analysis Methods and Models," etc., are optional). Clicking the "Generate Template" button generates a complete report template according to the tag attributes and relationship attributes of the current report paragraph entities. Charts, data sources, models, and analysis methods are automatically replaced in the data analysis process assets according to the user's selected analysis time range and other settings (if not specified, the original assets are reused). The generated draft report is displayed in the right column of the search results page. Users can continue searching, selecting other paragraph entities that meet their analysis needs and dragging them to the document editing bar on the right to fine-tune the existing analysis framework. The rapid reuse of analysis process assets provides data analysts with reference and decision-making support for quickly outputting data analysis reports.

[0133] In summary, this disclosure provides a data analysis report reuse method. Based on knowledge graph technology, it offers a one-stop packaged service from dimension sorting, data source matching, model application, and chart output. This enables the extraction and asset accumulation of multi-dimensional analysis processes in data analysis reports, greatly facilitating junior and intermediate data analysts to quickly output scientific, standardized, multi-dimensional, and chart-rich data analysis reports, thereby improving the reuse efficiency of data analysis reports.

[0134] Based on the above-described data analysis report reuse method, this disclosure also provides a data analysis report reuse device. The following will be combined with... Figure 12 The device is described in detail.

[0135] Figure 12 A schematic block diagram of a data analysis report multiplexing apparatus according to an embodiment of the present disclosure is shown.

[0136] like Figure 12As shown, the data analysis report reuse device 1200 of this embodiment includes, for example, a noun acquisition module 1210, an object determination module 1220, an attribute acquisition module 1230, an entity construction module 1240, a topic classification module 1250, a graph construction module 1260, and a report reuse module 1270.

[0137] The noun acquisition module 1210 is used to extract nouns from the data analysis report on a paragraph-by-paragraph basis. In one embodiment, the noun acquisition module 1210 can be used to perform the operation S210 described above, which will not be repeated here.

[0138] The object determination module 1220 is used to determine entity objects based on the frequency of occurrence of nouns and the position of nouns in a paragraph. In one embodiment, the object determination module 1220 can be used to perform the operation S220 described above, which will not be repeated here.

[0139] The attribute acquisition module 1230 is used to acquire the tag attributes and content attributes corresponding to the entity object. The tag attributes represent the form of the entity object being analyzed, and the content attributes represent the content of the entity object being analyzed. In one embodiment, the attribute acquisition module 1230 can be used to perform the operation S230 described above, which will not be repeated here.

[0140] The entity construction module 1240 is used to construct paragraph entities for each paragraph based on tag attributes and content attributes. In one embodiment, the entity construction module 1240 can be used to perform the operation S240 described above, which will not be repeated here.

[0141] The topic classification module 1250 is used to classify paragraphs according to the hierarchical structure between paragraphs to obtain the analysis topics. In one embodiment, the topic classification module 1250 can be used to perform the operation S250 described above, which will not be repeated here.

[0142] The graph construction module 1260 is used to construct a knowledge graph based on paragraph entities and analysis topics. In one embodiment, the graph construction module 1260 can be used to perform the operation S260 described above, which will not be repeated here.

[0143] The report reuse module 1270 is used to reuse data analysis reports based on a knowledge graph. In one embodiment, the report reuse module 1270 can be used to perform the operation S270 described above, which will not be repeated here.

[0144] According to embodiments of this disclosure, any and multiple modules among the noun acquisition module 1210, object determination module 1220, attribute acquisition module 1230, entity construction module 1240, topic classification module 1250, graph construction module 1260, and report reuse module 1270 can be merged into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the noun acquisition module 1210, object determination module 1220, attribute acquisition module 1230, entity construction module 1240, topic classification module 1250, map construction module 1260, and report multiplexing module 1270 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the noun acquisition module 1210, object determination module 1220, attribute acquisition module 1230, entity construction module 1240, topic classification module 1250, map construction module 1260, and report multiplexing module 1270 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.

[0145] Figure 13 A block diagram schematically illustrates an electronic device suitable for implementing a data analysis report reuse method according to an embodiment of the present disclosure.

[0146] like Figure 13 As shown, an electronic device 1300 according to an embodiment of the present disclosure includes a processor 1301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage portion 1308 into a random access memory (RAM) 1303. The processor 1301 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1301 may also include onboard memory for caching purposes. The processor 1301 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0147] RAM 1303 stores various programs and data required for the operation of electronic device 1300. Processor 1301, ROM 1302, and RAM 1303 are interconnected via bus 1304. Processor 1301 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 1302 and / or RAM 1303. It should be noted that the programs may also be stored in one or more memories other than ROM 1302 and RAM 1303. Processor 1301 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0148] According to embodiments of this disclosure, the electronic device 1300 may further include an input / output (I / O) interface 1305, which is also connected to a bus 1304. The electronic device 1300 may also include one or more of the following components connected to the I / O interface 1305: an input section 1306 including a keyboard, mouse, etc.; an output section 1307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN card, modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the I / O interface 1305 as needed. A removable medium 1311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1310 as needed so that computer programs read from it can be installed into the storage section 1308 as needed.

[0149] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0150] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 1302 and / or RAM 1303 and / or one or more memories other than ROM 1302 and RAM 1303 described above.

[0151] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the item recommendation method provided in the embodiments of this disclosure.

[0152] When the computer program is executed by the processor 1301, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0153] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 1309, and / or installed from the removable medium 1311. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0154] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1309, and / or installed from the removable medium 1311. When the computer program is executed by the processor 1301, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0155] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0157] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0158] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. A method for reusing data analysis reports, characterized in that, include: Extract nouns from the data analysis report, paragraph by paragraph; The entity object is determined based on the frequency of occurrence of the noun and the position of the noun in the paragraph; Obtain the tag attribute and content attribute corresponding to the entity object. The tag attribute is used to represent the form of the entity object being analyzed, and the content attribute is used to represent the content of the entity object being analyzed. Construct paragraph entities for each paragraph based on the tag attributes and the content attributes; The paragraphs are classified according to their hierarchical structure to obtain the analysis topic; Construct a knowledge graph based on the paragraph entities and the analysis topic; The data analysis report can be reused based on the knowledge graph. The classification of paragraphs based on their hierarchical structure yields analysis topics including: The distance coefficient is calculated based on the number of intervals between each paragraph to obtain the contextual relationship; Extract the table of contents or outline to obtain the hierarchical relationship of each paragraph; The paragraphs are categorized according to the contextual and hierarchical relationships to obtain the analysis topic.

2. The data analysis report reuse method according to claim 1, characterized in that, The nouns extracted from the data analysis report, grouped by paragraph, include: Construct word vectors; Perform part-of-speech tagging on the word vectors; Extract the subject and object to obtain the noun.

3. The data analysis report reuse method according to claim 1, characterized in that, The process of determining entity objects based on the frequency of occurrence of the noun and the position of the noun in the paragraph includes: Obtain the label extraction model; Input the nouns into the tag extraction model to obtain the nodes of the tag extraction model; Set the sliding window based on paragraph length; The weight of the node is calculated based on the number of times the noun appears in the sliding window and the position of the noun in the paragraph; The entity object is determined based on the weight of the node.

4. The data analysis report reuse method according to claim 3, characterized in that, The calculation of the weight of the edge between nodes based on the co-occurrence frequency of the noun in the sliding window and the position of the noun in the paragraph includes: When the nouns co-occur in the sliding window, assign edge weights between the co-occurring nodes; If the noun is located in the first two sentences or the last two sentences of the paragraph, assign a weight to the node.

5. The data analysis report reuse method according to claim 1, characterized in that, The step of obtaining the tag attributes and content attributes corresponding to the entity object includes: Retrieve the chart label attribute, data label attribute, method label attribute, and code label attribute respectively; The chart label attributes include the description content, data dimensions, and visualization method of the chart corresponding to the paragraph. The data tag attribute is a data source tag associated with the paragraph; The method tag attribute refers to the analysis method and / or analysis model used in the paragraph; The code tag attribute is the code block corresponding to the paragraph.

6. The data analysis report reuse method according to claim 5, characterized in that, The process of obtaining chart label attributes includes: Establish an index relationship between the text of the paragraph and the chart; Image classification and recognition are performed using a convolutional neural network model to obtain image category labels; The chart label attributes are constructed based on the index relationship and the image category labels.

7. The data analysis report reuse method according to claim 5, characterized in that, The data tag attributes to be acquired include: Obtain a pre-built data lineage dictionary; Based on the text of the paragraph, the data source tag is matched from the data lineage dictionary to obtain the data tag attribute.

8. The data analysis report reuse method according to claim 5, characterized in that, The acquisition method tag attributes include: The method tag attribute is obtained from the text of the paragraph using natural language processing methods.

9. The data analysis report reuse method according to claim 5, characterized in that, The acquisition of code tag attributes includes: Get the code comment content; The text of the paragraph is matched with the code comment content to obtain the code tag attribute.

10. The data analysis report reuse method according to claim 1, characterized in that, The reuse of the data analysis report based on the knowledge graph includes: The knowledge graph is stored in a graph database; A search engine is constructed based on the graph database; The title of the data analysis report is searched using the search engine to obtain the paragraph entity and the analysis topic corresponding to the title; The searched paragraph entities and analysis topics can be directly referenced or used to generate report templates.

11. A data analysis report reuse device, characterized in that, include: The noun extraction module is used to extract nouns from data analysis reports, organized by paragraph. An object determination module is used to determine entity objects based on the frequency of occurrence of the noun and the position of the noun in the paragraph; The attribute acquisition module is used to acquire the tag attribute and content attribute corresponding to the entity object. The tag attribute is used to represent the form of the entity object being analyzed, and the content attribute is used to represent the content of the entity object being analyzed. An entity construction module is used to construct paragraph entities for each paragraph based on the tag attributes and the content attributes; The topic classification module is used to classify each paragraph according to the hierarchical structure between them to obtain the analysis topic; The knowledge graph construction module is used to construct a knowledge graph based on the paragraph entities and the analysis topic. as well as A report reuse module is used to reuse the data analysis report based on the knowledge graph; The classification of paragraphs based on their hierarchical structure yields analysis topics including: The distance coefficient is calculated based on the number of intervals between each paragraph to obtain the contextual relationship; Extract the table of contents or outline to obtain the hierarchical relationship of each paragraph; The paragraphs are categorized according to the contextual and hierarchical relationships to obtain the analysis topic.

12. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors execute the data analysis report reuse method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, cause the processor to perform the data analysis report reuse method according to any one of claims 1 to 10.

14. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the data analysis report reuse method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Key word extraction method based on graph

    CN106202042A

  • Method and device for recognizing picture by combining text and picture

    CN107862239A

  • Semantic recognition method and device combined with knowledge graph entity information and related equipment

    CN112818690A

  • Article processing method and device, electronic equipment and medium

    CN114239588A