Intelligent Generation Method for Data Report Based on Multidimensional Feature Mining and Related Devices
Through the intelligent data report generation method based on multi-dimensional feature mining, the problems of low efficiency and poor timeliness of traditional report generation are solved, and efficient and accurate report generation is achieved to adapt to the rapidly changing decision-making environment.
Patent Information
- Application Number
- CN202411333517.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-09-24
AI Technical Summary
Traditional methods are inefficient and time-consuming when generating analysis reports of political and legal clues, making it difficult to adapt to the rapid changes in data characteristics.
The intelligent generation method of data report based on multi-dimensional feature mining is adopted. By obtaining the clue data to be processed, determining entities and attributes, establishing the association relationship between entities, building a prompt word input large language model, extracting key data, performing statistical analysis, dynamically adjusting the analysis template, and generating an analysis report.
Improve the efficiency and accuracy of analysis reports to ensure that the content of the report is updated as the data is updated and meets the needs of a rapidly changing decision-making environment.
Smart Images

Figure CN119203973B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an intelligent data report generation method based on multi-dimensional feature mining and related devices. Background Art
[0002] With the advent of the information age, the quantity and complexity of data have been increasing day by day. Especially in the fields of social governance and law, in the face of a large number of events and related clues, traditional manual processing and analysis methods can no longer meet the requirements of rapid response and processing. Especially when generating analysis reports for specific political and legal clues, traditional methods are often inefficient, error-prone, and difficult to adapt to the rapid changes in data characteristics.
[0003] In addition, manually compiling reports is not only time-consuming and laborious, but also difficult to ensure the timeliness and depth of the report content. With the explosion of data volume, extracting key information from this data and generating insightful reports has become a major challenge. These reports usually need to reflect the latest data insights and trends to support decision-making, while manual methods are often powerless in dealing with a large amount of complex data. Summary of the Invention
[0004] In view of this, the purpose of this application is to propose an intelligent data report generation method based on multi-dimensional feature mining and related devices to solve the problems of low efficiency and poor timeliness in generating analysis reports.
[0005] Based on the above purpose, the first aspect of this application provides an intelligent data report generation method based on multi-dimensional feature mining, including:
[0006] Obtain the clue data to be processed;
[0007] Determine the entities and attributes to be extracted from the clue data according to the pre-extracted core analysis elements, and establish the association relationships between the entities; wherein, the core analysis elements are determined according to the historical report data;
[0008] Construct a first prompt word according to the entities, the attributes, the association relationships and the clue data, input the first prompt word into the large language model, and output the key data corresponding to the clue data through the large language model; wherein, the first prompt word is used to instruct the large language model to extract key data from the clue data according to the entities, the attributes and the association relationships;
[0009] Use a statistical algorithm to perform statistical analysis on the key data to obtain a statistical analysis result;
[0010] Update the pre-constructed first analysis template according to the statistical analysis result to obtain a second analysis template; wherein, the first analysis template is determined according to historical report data;
[0011] Populate the second analysis template based on the key data to generate the analysis report.
[0012] The second aspect of the present application provides a data report intelligent generation device based on multi-dimensional feature mining, including:
[0013] An acquisition module configured to acquire clue data to be processed;
[0014] A determination module configured to determine entities and attributes to be extracted from the clue data according to pre-extracted core analysis elements, and establish an association relationship between the entities; wherein, the core analysis elements are determined according to historical report data;
[0015] An extraction module configured to construct a first prompt word according to the entities, the attributes, the association relationship and the clue data, input the first prompt word into a large language model, and output key data corresponding to the clue data through the large language model; wherein, the first prompt word is used to instruct the large language model to extract key data from the clue data according to the entities, the attributes and the association relationship;
[0016] An analysis module configured to perform statistical analysis on the key data using a statistical algorithm to obtain a statistical analysis result;
[0017] An update module configured to update the pre-constructed first analysis template according to the statistical analysis result to obtain a second analysis template; wherein, the first analysis template is determined according to historical report data;
[0018] A generation module configured to populate the second analysis template based on the key data to generate the analysis report.
[0019] The third aspect of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor, and the processor implements the method as described in the first aspect when executing the computer program.
[0020] The present application also provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method as described in the first aspect.
[0021] As can be seen from the above, the intelligent data report generation method and related devices based on multi-dimensional feature mining provided by this application. The method includes obtaining clue data to be processed; determining entities and attributes to be extracted from the clue data according to pre-extracted core analysis elements, and establishing association relationships between the entities; wherein, the core analysis elements are determined according to historical report data, which is convenient for accurately determining the specific content to be extracted from the clue data. Construct a first prompt word according to the entity, the attribute, the association relationship and the clue data, input the first prompt word into a large language model, and output key data corresponding to the clue data through the large language model; wherein, the first prompt word is used to instruct the large language model to extract key data from the clue data according to the entity, the attribute and the association relationship. Utilize the understanding ability of the large language model, and instruct the large language model to extract key data through the first prompt word, so as to provide a data basis for subsequent generation of an analysis report. Use a statistical algorithm to perform statistical analysis on the key data, mine new trends and new content in the key data, and obtain a statistical analysis result. Update a pre-constructed first analysis template according to the statistical analysis result to obtain a second analysis template; wherein, the first analysis template is determined according to historical report data; realizing dynamic adjustment of the analysis template to adapt to the real-time change trend of the clue data. Fill the second analysis template based on the key data to generate the analysis report. Automatically generating the analysis report helps to improve the generation rate and accuracy of the analysis report, ensures that the report content is updated as the data is updated, and meets the needs of a rapidly changing decision-making environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in this application or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a schematic flowchart of the intelligent data report generation method based on multi-dimensional feature mining according to an embodiment of this application;
[0024] Figure 2 It is a schematic flowchart of the method for determining core analysis elements according to an embodiment of this application;
[0025] Figure 3 It is a schematic structural diagram of the intelligent data report generation device based on multi-dimensional feature mining according to an embodiment of this application;
[0026] Figure 4 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed implementation manners
[0027] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0028] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the field to which the present application belongs. The "first", "second" and similar terms used in the embodiments of the present application do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "include" or "comprise" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connect" or "couple" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0029] As described in the background art, there are problems of low efficiency and poor timeliness when generating an analysis report of political and legal clues. The present application proposes an intelligent data report generation method based on multi-dimensional feature mining according to the above problems, which can effectively extract key information from a large amount of data, conduct in-depth statistical analysis, and dynamically adjust the report content according to the analysis results. The method of the present application can not only improve the efficiency and accuracy of generating the analysis report, but also ensure that the report content is updated as the data is updated, so as to truly meet the needs of a rapidly changing decision-making environment.
[0030] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0031] Figure 1 There is shown an intelligent data report generation method based on multi-dimensional feature mining, with reference to Figure 1 , which includes the following steps:
[0032] Step 102: Obtain the clue data to be processed.
[0033] Specifically, clue data belongs to unstructured data, which is data described in natural language. Clue data can include various types of data. Taking the data in the political and legal fields as an example, clue data can include housing-related data related to real estate, disaster-related data related to disasters, traffic accident data, and so on. These clue data are usually stored in databases, information systems or information platforms, which may include internal databases, public records, and other relevant information systems. When obtaining clue data, it is necessary to ensure that you have the data access right and obtain the clue data on the premise of complying with data protection regulations.
[0034] Step 104: Determine the entities and attributes to be extracted from the clue data according to the pre-extracted core analysis elements, and establish the association relationships between the entities; wherein, the core analysis elements are determined according to the historical report data.
[0035] Specifically, the core analysis elements are determined according to the historical report data. The core analysis elements reflect the core content included in the historical report data and are also the content that needs to be reflected in the analysis report corresponding to the clue data. Determine the entities and attributes involved in the business process corresponding to the clue data according to the core analysis elements. The core analysis elements may specifically include clue data types, time, location, participant information, and the number of participants, etc. Exemplarily, in the scenario of analyzing housing-related events, the entities may include "real estate developers", "event participants", "time", and "location", etc. Among them, the attributes of "event participants" include "name", "gender", and "contact information", etc.
[0036] Define the relationships between entities according to the business logic corresponding to the clue data. The relationships between entities are crucial for data analysis and report generation, which can help identify the internal relevance of clue data and better complete data mining and data analysis. Exemplarily, in the scenario of analyzing housing-related events, the possible relationships include "events and participants", "events and facility properties", etc.
[0037] Identified entities and attributes can be stored in a database through a pre-built data model. When the historical report data changes, the entities and attributes stored in the database will change accordingly. Data model construction is the process of designing and defining how to organize, store, and manage data in a database or data platform. Data model construction aims to ensure that the data structure can support business logic and analysis requirements while maintaining data scalability, security, and maintainability. The data model also sets data integrity rules, such as primary keys, foreign keys, and other business rules, to ensure data accuracy and consistency. A primary key is a field or combination of fields in a database table that is uniquely used to identify a record. Each table can have only one primary key, and the value of the primary key is unique and cannot be a null value. A foreign key is a field in one table whose value references the primary key in another table, and users establish and maintain the relationship between the two tables to ensure data consistency and integrity. In the database, the ID of each lead data is unique. The definition of the relationship between entities not only helps to understand how data is related but also supports database operations such as join queries.
[0038] At the same time, it is also necessary to determine the security and access mechanisms of the database, determine the access rights of each user and the access operations that can be performed, and protect the sensitive data in the database from being accessed by unauthorized users. In addition, data security measures, such as encrypting sensitive fields, can be implemented to comply with data protection regulations and maintain user trust.
[0039] Step 106: Construct a first prompt word based on the entity, the attribute, the association relationship, and the lead data, and input the first prompt word into the large language model. Output the key data corresponding to the lead data through the large language model; wherein, the first prompt word is used to instruct the large language model to extract key data from the lead data according to the entity, the attribute, and the association relationship.
[0040] Specifically, after determining the entity, attribute, and association relationship, the key data in the lead data can be extracted through the large language model. The key data is the data used to generate the analysis report. In this embodiment, the large language model is used to process the data described in natural language (lead data) and extract the key data. Specifically, a first prompt word is constructed according to the entity, attribute, and association relationship, so that the large language model can understand the specific content to be extracted and extract the key data in the lead data according to the first prompt word. For example, key information such as event type, event occurrence time, and location is extracted.
[0041] After the extraction of critical data, it is also necessary to verify the critical data through sample inspection or automated testing to verify whether the extracted critical data accurately reflects the core analysis elements. Further, data verification includes randomly sampling and inspecting the critical data, as well as comparing and verifying it with the labeled data. The data verification results are used to feedback and optimize the large language model to ensure the accuracy and efficiency of critical data extraction.
[0042] The extracted critical data can be converted into structured data and stored to form a dataset convenient for analysis. Each piece of critical data exists in a formatted manner, such as in JSON format, which contains detailed information such as event type, occurrence date, involved property, participants, etc. Exemplarily, the data structure of a piece of critical data regarding a property-related event is as follows:
[0043] {
[0044] "Event ID":"E12345",
[0045] "Event type":"Delivery delay",
[0046] "Occurrence date":"2024-08-10",
[0047] "Involved property":{
[0048] "Property ID":"P9876",
[0049] "Address":"No. 123, City Center Road",
[0050] "Property type":"Commercial"
[0051] },
[0052] "Participants":
[0053] {"Participant ID":"C123",
[0054] "Name":"Zhang San",
[0055] "Role":"Home buyer",
[0056] "Contact information":"123456789"},
[0057] {"Participant ID":"C456",
[0058] "Name":"Li Si",
[0059] "Role":"Developer",
[0060] "Contact information":"987654321"}
[0062] }.
[0063] Step 108: Use a statistical algorithm to perform statistical analysis on the key data to obtain a statistical analysis result.
[0064] In this step, first, necessary preprocessing of the above-structured key data is required to adapt to the statistical analysis requirements, including data type conversion and missing value handling, etc. Then, based on the preprocessed key data, statistical analysis is performed, such as calculating the number of occurrences of each type of event in different unit times, the number of participants, the average time to resolve the event, etc., to obtain a statistical analysis result. These statistical analysis results reveal important patterns in the key data and trends that change over time and social dynamics, ensuring the depth and breadth of data analysis.
[0065] Step 110: Update the pre-constructed first analysis template according to the statistical analysis result to obtain a second analysis template; wherein, the first analysis template is determined based on historical report data.
[0066] Specifically, the first analysis template is determined based on historical report data. The first analysis template includes fixed report content, chapters, etc. corresponding to the historical report data. Different types of historical report data correspond to different first analysis templates. For example, housing-related events and disaster-related events both have corresponding first analysis templates. If the statistical analysis result contains content not included in the first analysis template, the first analysis template needs to be updated according to the statistical analysis result to obtain a second analysis template, so that the second analysis template can display the trends and insights of the data over time.
[0067] Step 112: Fill the second analysis template based on the key data to generate the analysis report.
[0068] Specifically, determine the corresponding content for filling the second analysis template from the key data. The filling process includes, but is not limited to, converting the format of the key data to meet the report requirements, generating visual elements, etc. Scripts or report generation tools can be used to automatically fill the key data into the reserved positions of the template. The second analysis template includes multiple template elements, and the template elements are used to indicate the content that needs to be filled into the template. Each template element has a corresponding reserved position. For each template element in the second analysis template, fill the key data corresponding to the template element into the reserved position corresponding to the template element in a predetermined form. The predetermined form includes visual charts or analysis texts, etc. For visual charts, chart automatic generation tools can be used to automatically generate them according to the key data. For analysis texts, large language models can be used to automatically generate descriptive and analytical texts, including background information, explaining trends, and possible reasons and suggestions. Finally, a complete analysis report is generated.
[0069] Through the method of this embodiment, the analysis report can not only reflect the current data status, but also show the trends and insights of data changes over time. The dynamic adjustment of the analysis template ensures that the analysis report always maintains relevance and practicality, supporting enterprises to make informed decisions in the ever-changing market environment. In this embodiment, the method of extracting key information according to entities, relationships, and association relationships expands the depth of data analysis, the statistical analysis of key data expands the breadth of data analysis, and the real-time update and personalized display of the content of the analysis report greatly improve the strategic value and operational efficiency of the analysis report.
[0070] Based on the above steps 102 to 112, the intelligent data report generation method provided by this application based on multi-dimensional feature mining includes obtaining clue data to be processed; determining entities and attributes to be extracted from the clue data according to pre-extracted core analysis elements, and establishing association relationships between the entities; wherein, the core analysis elements are determined according to historical report data, which is convenient for accurately determining the specific content to be extracted from the clue data. Construct a first prompt word according to the entity, the attribute, the association relationship, and the clue data, input the first prompt word into the large language model, and output the key data corresponding to the clue data through the large language model; wherein, the first prompt word is used to instruct the large language model to extract key data from the clue data according to the entity, the attribute, and the association relationship. Utilize the understanding ability of the large language model, and instruct the large language model to extract key data through the first prompt word, so as to provide a data basis for subsequent generation of the analysis report. Use statistical algorithms to perform statistical analysis on the key data, mine new trends and new content in the key data, and obtain statistical analysis results. Update the pre-constructed first analysis template according to the statistical analysis results to obtain a second analysis template; wherein, the first analysis template is determined according to historical report data; realizing the dynamic adjustment of the analysis template to adapt to the real-time change trend of the clue data. Fill the second analysis template based on the key data to generate the analysis report. Automatically generating the analysis report helps to improve the generation rate and accuracy of the analysis report, ensure that the report content is updated as the data is updated, and meet the needs of the rapidly changing decision-making environment.
[0071] The innovation of this application lies in its combination of data structured extraction, statistical analysis, dynamic template adjustment, and data-driven report generation, providing an efficient, accurate, and dynamic report generation solution for the political and legal fields. By using advanced large language models to extract key data from unstructured data and convert it into a structured format, the efficiency of data processing and the accuracy of analysis are significantly improved. In addition, through in-depth statistical analysis of the extracted key data, not only can new patterns with statistical significance be identified, but the report template can also be dynamically adjusted based on these insights. This ensures that as time goes by and data characteristics change, the generated reports can continuously reflect the latest trends and insights, thus supporting more effective decision-making. Through intelligent data processing and report generation, the operational efficiency and response ability of political and legal organs are enhanced.
[0072] In some embodiments, the method for determining the core analysis elements refers to Figure 2 , and includes the following steps:
[0073] Step 202: Obtain historical report data and perform data cleaning on the historical report data;
[0074] Step 204: Determine an initial analysis template based on the historical report data that has undergone data cleaning; wherein, multiple template elements are included in the initial analysis template;
[0075] Step 206: Extract all the template elements in the initial analysis template, summarize and generalize all the template elements, and determine the core analysis elements.
[0076] Specifically, historical report data are analysis reports for a historical period. Data cleaning of historical report data is to ensure the accuracy and consistency of the data, providing a high-quality data foundation for subsequent data mining and analysis. In a specific embodiment, step 202 further includes: ① Removing unnecessary characters and interfering characters from the historical report data: By automatically identifying and removing unnecessary characters and interfering characters in the text, such as redundant spaces and special symbols, etc., to remove characters irrelevant to the text content. ② Standardizing the punctuation marks in the historical report data: According to the language context of the text, converting Chinese punctuation marks to English punctuation marks, or converting English punctuation marks to Chinese punctuation marks, to ensure the consistency of the text. At the same time, unnecessary repeated punctuation marks in the text can also be removed through regular expressions, deleting redundant punctuation marks. ③ Using natural language processing technology to correct spelling mistakes in the historical report data: In the process of data cleaning, correcting spelling mistakes is a crucial step to ensure the text quality and data accuracy. Using natural language processing technology to automatically detect spelling mistakes in the text and assist in spelling correction through context analysis to ensure the use of appropriate words in the context. For example, correcting the misspelled "defense" to "real estate-related". Text error correction not only improves the readability of the text but also enhances the accuracy of data analysis, laying a solid foundation for subsequent automated processing and report generation.
[0077] After cleaning the historical report data, an initial analysis template is determined according to the cleaned historical report data. Step 204 specifically includes:
[0078] Step 2042: Divide the historical report data that has undergone data cleaning into paragraphs to obtain multiple text paragraphs.
[0079] Specifically, each historical report is automatically divided into paragraphs. When segmenting, it can be achieved by identifying specific format markers, such as line breaks, paragraph numbers, or headings, etc. Exemplarily, if there are 100 historical report data, each report is segmented, and finally 810 text paragraphs are obtained.
[0080] Step 2044: Use a clustering algorithm to cluster multiple text paragraphs to obtain multiple clusters.
[0081] Specifically, a text embedding model or other vectorization techniques are used to convert text data into feature vectors in numerical form for machine learning processing. If the feature dimension of the text paragraph data is very high, methods such as Principal Components Analysis (PCA) can be applied to reduce the dimension to simplify subsequent clustering analysis. Select a suitable clustering algorithm according to the nature of the data, such as the K-means clustering algorithm, hierarchical clustering, or Density-Based Spatial Clustering of Applications with Noise (DBSCAN). Apply the clustering algorithm to the feature vectors of the text paragraphs, group the text paragraphs analyzing the same problem into the same cluster, and each cluster represents the analysis of a type of problem or content pattern. Evaluate the clustering effect according to metrics such as the silhouette coefficient, and adjust the parameters of the clustering algorithm (such as the number of clusters) to optimize the algorithm.
[0082] Step 2046: For the text paragraphs in each cluster, use a large language model to extract the core analysis information corresponding to the text paragraphs in the cluster, including:
[0083] For each cluster, randomly select multiple text paragraphs from the cluster, generate a third prompt word based on the multiple text paragraphs, input the third prompt word into the large language model, and output the core analysis information through the large language model; wherein, the third prompt word is used to instruct the large language model to summarize and analyze the multiple text paragraphs to obtain the core analysis information.
[0084] Specifically, each cluster contains several text paragraphs. Randomly select multiple text paragraphs from the several text paragraphs, such as three text paragraphs. By constructing a clear prompt (i.e., the third prompt word), the large language model automatically learns and extracts the common core analysis information from the multiple text paragraphs. Exemplarily, the third prompt word can be:
[0085] "
[0086] Please analyze the following text paragraphs from the same cluster and summarize their common core analysis points:
[0087] [Insert text paragraph 1]
[0088] [Insert text paragraph 2]
[0089] [Insert text paragraph 3]
[0090] "
[0091] Send the above third prompt to the large language model for processing and receive the content returned by the large language model. For example, if the large language model outputs: "Analysis of the proportion of clue types", check whether the large language model accurately captures the core theme and concerns of the text. If the check passes, the core analysis information is "Analysis of the proportion of clue types". If the check fails, the design of the third prompt can be adjusted according to the content returned by the large language model to optimize the output result of the large language model.
[0092] Step 2048: Construct a second prompt based on the core analysis information and the historical report data, input the second prompt into the large language model, and output the initial analysis template corresponding to the historical report data through the large language model; wherein, the second prompt is used to instruct the large language model to replace the variable information in the historical report data with corresponding template elements.
[0093] Specifically, the initial analysis template defines the basic framework and content requirements of the subsequent analysis report. The large language model is used to automatically identify and replace specific data and specific identifiers in the text paragraph, and convert them into general template elements (or called placeholders) to create a flexible and reusable initial analysis template. Specifically, the large language model is instructed to replace the variable information in the historical report data with corresponding template elements by constructing a second prompt.
[0094] Exemplarily, the second prompt can be:
[0095] "
[0096] 'Since the beginning of this year, 2000 intelligence clues have been collected and disposed of. Among them, 1200 are housing-related clues.'
[0097] Please replace any specific data and specific identifiers therein with appropriate placeholders to form a reusable template.
[0098] "
[0099] Among them, 'Since the beginning of this year, 2000 intelligence clues have been collected and disposed of. Among them, 1200 are housing-related clues.' is part of the text paragraph in the historical report data. At the same time, the core analysis information can also be added to the second prompt to ensure that the large language model can accurately understand the core idea of the text paragraph in the historical report data, which helps the large language model to identify information such as the event type in the text paragraph.
[0100] Send the constructed second prompt to the large language model for processing. The large language model automatically replaces the specific variable information identified in the text paragraph with template elements such as [Total number of events throughout the year], [Number of events related to the housing category throughout the year], etc. Continuing the previous example, the initial analysis template output by the large language model is: "Since the beginning of this year, [Total number of events throughout the year] intelligence clue information has been collected and processed. Among them, the housing-related clues are [Number of events related to the housing category throughout the year]." When generating the subsequent analysis report, the values of "Total number of events throughout the year" and "Housing-related events" obtained through extraction can be directly filled into the corresponding positions of the template elements.
[0101] After determining the initial analysis template, it is necessary to extract the core analysis elements from the initial analysis template. The core analysis elements are the key basis for clue data extraction and analysis report generation. Exemplarily, for example, in the template of the housing-related event analysis report, although multiple data-related template elements such as "Total number of events", "Number of housing-related events", and "Proportion of various events" are filled in the initial analysis template, after analysis, it is found that "Clue type" is the common basis for these template elements, that is, the core analysis element. Correctly classifying each actual clue data into its "Clue type" can automatically calculate the information of template elements such as "Number of housing-related events" and "Event proportion" through statistical analysis. Determining these core analysis elements is conducive to formulating clear guidance for subsequent data extraction and processing work, ensuring the accuracy of the analysis report content and the effectiveness of in-depth analysis.
[0102] Furthermore, collect all the template elements in the initial analysis template, such as [Total number of events throughout the year], [Number of events related to the housing category throughout the year], [Proportion of housing-related events], etc. Analyze the data types and contents represented by each template element, including the classification of data such as numerical values, percentages, names, etc.
[0103] Exemplarily, the specific initial analysis template for safety events is:
[0104] "A total of [Total number of events throughout the year] safety events have been processed this year. Among them, there are [Number of fire events] fire events, accounting for [Proportion of fire events]%; [Number of traffic accidents] traffic accidents, accounting for [Proportion of traffic accidents]%. Among these events, the most common is [Most common event type], with the number of related events being [Number of events related to the most common event]."
[0105] Extract all the template elements from the above initial analysis template, including: [Total number of events throughout the year] [Number of fire events] [Proportion of fire events] [Number of traffic accidents] [Proportion of traffic accidents] [Most common event type] [Number of events related to the most common event].
[0106] When summarizing and extracting core analysis elements, analyze the relationships between various template elements and the theme of the analysis report. First, observe whether these template elements change around a specific theme or classification, such as "housing-related leads". Then, identify the dependencies between the template elements. For example, if multiple template elements such as [Number of housing-related category events for the whole year] and [Percentage of housing-related category events] both depend on the "housing-related" lead category, this indicates that the "lead category" is a key element. Another example is that in the aforementioned initial analysis template, fire events and traffic accidents are specifically mentioned, indicating that "type" is a key element, and each type has corresponding event numbers and percentages, showing the dependencies between the template elements. If multiple important template elements are associated with a specific category (such as the lead category), then this category is very likely to be the core analysis element that needs special attention when constructing the template. Then, verify the summarized core analysis element to ensure that it conforms to the actual data content and report requirements. The verification process can be confirmed with the data manager or business analyst.
[0107] In the aforementioned example, since all important template elements are directly related to the event type, the core analysis element is determined to be "event type" because the classification and analysis of the data are based on different event types (such as fires and traffic accidents).
[0108] After determining the core analysis element, it is also necessary to further update the initial analysis template to make the determined analysis template more flexible and practical to adapt to different types of data analysis requirements. In some embodiments, the method for determining the first analysis template includes: determining new template elements according to the core analysis element; adding the new template elements to the initial analysis template to obtain the first analysis template.
[0109] Specifically, record the identified core analysis element and its application in the initial analysis template in detail as a reference for future template updates and maintenance. Then update the template according to the core analysis element to ensure that all relevant template elements can accurately reflect the changes in the core analysis element. Continuing the aforementioned example, the core analysis element is "event type", and the new template element is [Event Type]. Add the new template element to the initial analysis template to obtain the first analysis template:
[0110] "A total of [Total number of events for the whole year] safety events were processed this year. Among them, [Event Type 1] events occurred [Number of Event Type 1] times, accounting for [Percentage of Event Type 1]%; [Event Type 2] events occurred [Number of Event Type 2] times, accounting for [Percentage of Event Type 2]%. Among these events, the most common is [Most common event type], involving [Number of events involving the most common event type] times."
[0111] Through the above-mentioned update process of the analysis template, not only the core analysis elements in the analysis template are identified, but also the design of the analysis template can be optimized and standardized according to the core analysis elements, making the analysis template more flexible and practical, and adaptable to different types of data analysis requirements.
[0112] In some embodiments, the first analysis template includes multiple template elements; updating the pre-constructed first analysis template according to the statistical analysis result to obtain a second analysis template includes:
[0113] Determining newly added template elements according to the statistical analysis result, and adding the newly added template elements to the first analysis template to obtain the second analysis template.
[0114] Specifically, in the foregoing embodiments, unstructured clue data has been processed through technical methods such as large language models, key data has been extracted, and structured storage has been performed. Since the clue data changes in real time, there may be a problem of mismatch between the currently obtained clue data and each template element in the determined first analysis template. Therefore, it is also necessary to update the first analysis template according to the statistical analysis result of the key data, that is, to dynamically adjust the first analysis template, so as to ensure that the subsequent generated analysis report can accurately reflect the current data state, capture and display the deficiencies that change with time and social dynamics. The statistical analysis result includes new patterns or trends with statistical significance. Exemplarily, the new trend includes an increase in the number of delayed housing deliveries within a specific season. A special chapter for analyzing this phenomenon can be added to the analysis template as a newly added template element to deeply explore the reasons and put forward suggestions. Adding the newly added template element to the first analysis template, and after updating the first analysis template, the second analysis template is obtained. At the same time, the charts, data visualization elements, and text descriptions in the analysis template can also be updated adaptively to ensure that the second analysis template can reflect the latest data and analysis results. Then, filling each template element in the second analysis template with key data to generate an analysis report, so that the analysis report can continuously reflect the latest trends and insights, thereby supporting more effective decision-making.
[0115] It should be noted that the method of the embodiment of the present application can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiment of the present application, and these multiple devices will interact with each other to complete the described method.
[0116] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0117] Based on the same inventive concept, corresponding to any of the above-described method embodiments, the present application further provides an intelligent analysis report generation device based on multi-dimensional feature mining.
[0118] Refer to Figure 3 , the intelligent analysis report generation device based on multi-dimensional feature mining includes:
[0119] An acquisition module 302, configured to acquire clue data to be processed;
[0120] A determination module 304, configured to determine entities and attributes to be extracted from the clue data according to pre-extracted core analysis elements, and establish association relationships between the entities; wherein, the core analysis elements are determined according to historical report data;
[0121] An extraction module 306, configured to construct a first prompt word according to the entities, the attributes, the association relationships, and the clue data, input the first prompt word into a large language model, and output key data corresponding to the clue data through the large language model; wherein, the first prompt word is used to instruct the large language model to extract key data from the clue data according to the entities, the attributes, and the association relationships;
[0122] An analysis module 308, configured to perform statistical analysis on the key data using a statistical algorithm to obtain a statistical analysis result;
[0123] An update module 310, configured to update a pre-constructed first analysis template according to the statistical analysis result to obtain a second analysis template; wherein, the first analysis template is determined according to historical report data;
[0124] A generation module 312, configured to fill the second analysis template based on the key data to generate the analysis report.
[0125] In some embodiments, the first analysis template includes multiple template elements; the update module 310 is further configured to determine newly added template elements according to the statistical analysis result, and add the newly added template elements to the first analysis template to obtain the second analysis template.
[0126] In some embodiments, the determination module 304 is further configured to obtain historical report data and perform data cleaning on the historical report data;
[0127] Determine an initial analysis template according to the historical report data that has undergone data cleaning; wherein, a plurality of template elements are included in the initial analysis template;
[0128] Extract all the template elements in the initial analysis template, summarize and generalize all the template elements, and determine the core analysis elements.
[0129] In some embodiments, the determination module 304 is further configured to divide the historical report data that has undergone data cleaning into paragraphs to obtain a plurality of text paragraphs;
[0130] Use a clustering algorithm to cluster the plurality of text paragraphs to obtain a plurality of clusters;
[0131] For the text paragraphs in each cluster, use a large language model to extract the core analysis information corresponding to the text paragraphs in the cluster;
[0132] Construct a second prompt word according to the core analysis information and the historical report data, input the second prompt word into the large language model, and output the initial analysis template corresponding to the historical report data through the large language model; wherein, the second prompt word is used to instruct the large language model to replace the variable information in the historical report data with the corresponding template elements.
[0133] In some embodiments, the determination module 304 is further configured to remove unnecessary characters and interfering characters from the historical report data; standardize the punctuation marks in the historical report data; use natural language processing technology to correct the spelling mistakes in the historical report data.
[0134] In some embodiments, the determination module 304 is further configured to, for each cluster, randomly select a plurality of text paragraphs from the cluster, generate a third prompt word according to the plurality of text paragraphs, input the third prompt word into the large language model, and output the core analysis information through the large language model; wherein, the third prompt word is used to instruct the large language model to summarize and analyze the plurality of text paragraphs to obtain the core analysis information.
[0135] In some embodiments, the determination module 304 is further configured to determine newly added template elements according to the core analysis elements; add the newly added template elements to the initial analysis template to obtain the first analysis template.
[0136] For the convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0137] The device in the above embodiment is used to implement the corresponding intelligent generation method of data reports based on multi-dimensional feature mining in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0138] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the intelligent generation method of data reports based on multi-dimensional feature mining described in any of the above embodiments.
[0139] Figure 4 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0140] The processor 1010 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0141] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0142] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input devices may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output devices may include a display, a speaker, a vibrator, an indicator light, etc.
[0143] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. The communication module can achieve communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0144] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0145] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0146] The electronic device in the above embodiment is used to implement the corresponding intelligent data report generation method based on multi-dimensional feature mining in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0147] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the intelligent data report generation method based on multi-dimensional feature mining as described in any of the foregoing embodiments.
[0148] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0149] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the intelligent generation method of data reports based on multi-dimensional feature mining described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0150] Based on the same concept, corresponding to the method of any of the above embodiments, the present application also provides a computer program product, including computer program instructions. When the computer program instructions run on a computer, the computer is caused to execute the method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0151] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0152] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0153] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window. The prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information by the electronic device.
[0154] It should be understood that the above notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0155] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application is limited to these examples; under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above. For the sake of brevity, they are not provided in detail.
[0156] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections of integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in the form of a block diagram in order not to make the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation manner of these block diagram devices are highly dependent on the platform on which the embodiments of the present application will be implemented (that is, these details should be completely within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0157] Although the present application has been described in conjunction with specific embodiments of the present application, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) can be used with the embodiments discussed.
[0158] The embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the present application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present application.
Claims
1. A method for intelligently generating data reports based on multidimensional feature mining, characterized in that: include: Get the clue data to be processed; Determine entities and attributes to be extracted in the clue data according to the pre-extracted core analysis elements, and establish association relationships between entities; wherein the core analysis elements are determined according to the historical report data; Construct a first prompt word according to the entity, the attribute, the association relationship and the clue data, input the first prompt word into a large language model, and output key data corresponding to the clue data through the large language model; wherein the first prompt word is used to instruct the large language model to extract key data from the clue data according to the entity, the attribute and the association relationship; Performing statistical analysis on the key data using a statistical algorithm to obtain statistical analysis results; According to the statistical analysis result, the pre-constructed first analysis template is updated to obtain a second analysis template; wherein the first analysis template is determined according to the historical report data; The second analysis template is filled based on the key data to generate an analysis report.
2. The method according to claim 1, characterized in that: The first analysis template includes a plurality of template elements; the pre-constructed first analysis template is updated according to the statistical analysis result to obtain a second analysis template, including: A newly added template element is determined according to the statistical analysis result, and the newly added template element is added to the first analysis template to obtain the second analysis template.
3. The method according to claim 1, characterized in that The method for determining the core analysis elements includes: Acquire historical report data, and perform data cleaning on the historical report data; Determining an initial analysis template based on historical report data that has undergone data cleaning; wherein the initial analysis template includes a plurality of template elements; All template elements in the initial analysis template are extracted, all template elements are summarized and generalized, and the core analysis elements are determined.
4. The method according to claim 3, characterized in that Determining the initial analysis template based on the historical report data after data cleaning includes: Divide the cleaned historical report data into paragraphs to obtain multiple text paragraphs; Clustering algorithms are used to cluster multiple text paragraphs to obtain multiple clusters; For each text paragraph in each cluster, the core analysis information corresponding to the text paragraph in the cluster is extracted using a large language model; A second prompt word is constructed according to the core analysis information and the historical report data, the second prompt word is input into the large language model, and an initial analysis template corresponding to the historical report data is output through the large language model; wherein the second prompt word is used to instruct the large language model to replace the variable information in the historical report data with the corresponding template elements.
5. The method according to claim 3, characterized in that: Performing data cleaning on the historical report data, including: removing unnecessary characters and interfering characters from the historical report data; Standardizing punctuation marks in the historical report data; Natural language processing techniques are used to correct spelling errors in the historical reporting data.
6. The method according to claim 4, characterized in that For each text paragraph in each cluster, the core analysis information corresponding to the text paragraph in the cluster is extracted using a large language model, including: For each cluster, a plurality of text paragraphs are randomly selected from the cluster, a third prompt word is generated according to the plurality of text paragraphs, the third prompt word is input into the large language model, and the core analysis information is output through the large language model; wherein the third prompt word is used to instruct the large language model to summarize and analyze the plurality of text paragraphs to obtain the core analysis information.
7. The method according to claim 3, characterized in that The method for determining the first analysis template comprises: Determining new template elements based on the core analysis elements; The newly added template elements are added to the initial analysis template to obtain the first analysis template.
8. A data report intelligent generation device based on multi-dimensional feature mining, characterized in that: include: An acquisition module, configured to acquire clue data to be processed; A determination module, configured to determine entities and attributes to be extracted in the clue data according to the pre-extracted core analysis elements, and to establish association relationships between entities; wherein the core analysis elements are determined according to the historical report data; an extraction module, configured to construct a first prompt word according to the entity, the attribute, the association relationship and the clue data, input the first prompt word into a large language model, and output key data corresponding to the clue data through the large language model; wherein the first prompt word is used to instruct the large language model to extract key data from the clue data according to the entity, the attribute and the association relationship; An analysis module is configured to perform statistical analysis on the key data using a statistical algorithm to obtain a statistical analysis result; An updating module, configured to update a pre-built first analysis template according to the statistical analysis result to obtain a second analysis template; wherein the first analysis template is determined according to historical report data; A generation module is configured to fill in the second analysis template based on the key data to generate an analysis report.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Thematic report generation method and system
CN117931899A
Information extraction device and method based on large language model
CN118260390A