AI report automatic generation method based on big data
By identifying the original text type, introducing external memory modules and blocking context overlap and attention suppression, the adaptability of the encoder decoder model is enhanced, and the problems of excessive static rules, rough parameter adjustment granularity and poor processing of ultra-long text in the automatic generation of AI reports are solved, and the personalized configuration and information integrity of AI reports are achieved.
Patent Information
- Application Number
- CN202510505029.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology has too many static rules, rough parameter adjustment granularity, and lack of adaptability in the automatic generation of AI reports, resulting in misjudgment and information loss, especially when processing ultra-long text, resulting in information loss and redundancy.
By identifying the original text type, adjusting the length of AI report, introducing external memory modules and block context overlap and attention suppression, we will enhance the adaptability of the encoder decoder model, ensure information integrity, and achieve the extraction of block-level important information through word-level and sentence-level attention optimization.
It realizes the personalized configuration and information integrity of AI reports, solves the problems of information loss and redundancy in traditional methods when dealing with ultra-long texts, and improves the information coverage and simplicity of reports.
Smart Images

Figure CN120012730A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automatic report generation, and specifically refers to an automatic AI report generation method based on big data. Background Art
[0002] The automatic generation method of AI reports based on big data refers to a method of automatically generating reports based on AI technology.
[0003] Among the existing approximate solutions, for example, CN119322846AAI content summary generation method and device, this solution targets the technical problems of poor generalization ability of existing technical models, large consumption of computing resources, and insufficient multi-language and cross-cultural support. By cleaning the original data, evaluating the user creation time, interaction volume, related value and attention volume, extracting content-based screening vectors, screening out high-quality data, using the AI engine's Token library to convert the high-quality data into a Token list, checking the validity of the Token list, judging whether the quality of the input data meets the requirements for generating a summary, and dynamically adjusting the generation parameters of the AI model through the verification results to ensure that the generated Token list is valid and consistent with the topic, achieving the technical effect of improving the limitations of the traditional linear weighted method and reducing the impact of low-quality input. However, the existing technology has too many static rules and coarse parameter adjustment granularity, large differences in the ideal Token lengths of different original text types such as news and user comments, lack of adaptive ability, resulting in misjudgment, rough handling of edge cases, and the summary may lose important content.
[0004] In addition, for example, CN115455954A discloses a text summary generation method, system, computer device and storage medium. This solution aims to solve the technical problem that when summarizing a long text, the original content cannot be accurately expressed. A text summary generation model with an encoder-decoder structure is used, combined with an attention mechanism, to enhance the model's ability to capture key information in long texts, thereby achieving the technical effect of improving the efficiency of users in extracting core information from lengthy documents. However, the traditional sequence-to-sequence encoder-decoder structure cannot effectively process ultra-long texts, resulting in information loss, and the model repeatedly focuses on the same position of the original text, resulting in report redundancy. Short original texts may be misjudged as low quality, while lengthy but inefficient original texts may pass verification. Summary of the invention
[0005] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides a method for automatically generating AI reports based on big data. In view of the technical problems in the prior art, such as too many static rules and coarse granularity of parameter adjustment, large differences in ideal Token lengths of different original text types such as news and user comments, lack of adaptive ability, resulting in misjudgment, rough handling of edge cases, and the possible loss of important content in the summary, this solution predicts user interest sentences through eye movement features containing user reading habits, replaces manual rules, realizes personalized configuration of reports, and adjusts the ideal target AI report length by identifying the original text type, thereby enhancing the adaptive ability of the encoder-decoder model. The block context overlap and attention suppression method is used to ensure the integrity of the information contained in the report. In view of the technical problems that the traditional sequence-to-sequence encoder-decoder structure in the existing technology cannot effectively process ultra-long texts, resulting in information loss, the model repeatedly pays attention to the same position of the original text, resulting in report redundancy, and short original texts may be misjudged as low quality, while long but inefficient original texts may pass verification, this solution introduces an external memory module to divide the original text into blocks, and through word-level and sentence-level attention optimization, it realizes the extraction of important information at the block level, and finally obtains the entire original text report, which solves the common problem of repeated generation in generative reports and balances the information coverage and conciseness of the report in an end-to-end manner.
[0006] The technical solution adopted by the present invention is as follows: The present invention provides a method for automatically generating an AI report based on big data, the method comprising the following steps:
[0007] Step S1: Identify the type of original text, which is used to achieve AI report length adaptation. Specifically, a lightweight classifier is used to automatically identify the type of original text, and the ideal target AI report length is adjusted according to the identification result.
[0008] Step S2: original text segmentation is used to segment the original text into blocks to avoid information fragmentation. Specifically, based on the original text semantics, the similarity between sentences is calculated, the maximum length of blocks is limited, overlapping contexts between blocks are allowed, paragraphs are divided while maintaining semantic integrity, and semantically coherent paragraphs are merged into blocks.
[0009] Step S3: knowledge graph construction, specifically, extracting entities and relationships from the original text, constructing a knowledge graph, and using the knowledge graph as external memory;
[0010] Step S4: memory fusion, specifically, using the encoder-decoder model to generate a local AI report for each block;
[0011] Step S5: Global aggregation, specifically, aggregating the local AI reports of all blocks, describing the relationship between local AI reports through the graph attention network, treating the local AI reports as nodes, obtaining the importance scores of the nodes, and generating the final global AI report based on the importance scores.
[0012] Furthermore, in step S4, the memory fusion is specifically performed in the following steps:
[0013] Step S41: Eliminate low-density information, which is used to eliminate low-density information in the original text. Specifically, pre-process the input original text, build and use an AI pre-training model, remove low-density information in the original text, the low-density information includes low-correlation, low-interaction and noise data, collect the eye movement data set of the user when reading the original text, and extract gaze features from the eye movement data set. The gaze features include gaze duration features, gaze frequency features, reading order features, the ratio of gaze duration to the number of words, and the ratio of gaze duration to the number of characters. The gaze duration features include the total gaze duration of the sentence, the relative gaze duration, the longest gaze time, the shortest gaze time, the mean, median and standard deviation of the gaze duration. The gaze frequency features include the number of gazes and the relative number of gazes. The reading order features include the first gaze order, the last gaze order and the number of words that appear in the first 10 gazes.
[0014] Step S42: Personalized screening, used to quantify user interests, specifically, using a gradient boosting decision tree, taking the gaze feature as input, predicting the sentences that the user is interested in, and outputting a binary label of whether any sentence in the original text acts on the AI report;
[0015] Step S43: obtaining a sentence-level score, specifically, introducing a sentence extractor based on a sentence-level attention mechanism, quantitatively evaluating the comprehensive quality of each sentence in the original text, and calculating the sentence-level importance score of each sentence in combination with the sentences that the user is interested in;
[0016] Step S44: Dynamic suppression, used to avoid the generated AI report from containing repeated identical content, and obtain a block-level word and sentence representation vector;
[0017] Step S45: Balanced generation, used for semantic compression and original text rewriting, specifically, the encoder queries the external memory to obtain block-related entities and the relationship between entities, the external memory returns a triple encoding vector, and extracts structured linguistic features based on linguistics. The structured linguistic features include sentence position, number of digits, named entity tags, part-of-speech tags and word weights. The triple encoding vector, the structured linguistic features and the block-level sentence representation vector are spliced by dimension, and the decoder outputs a local AI report.
[0018] Furthermore, in step S44, the dynamic suppression is specifically performed in the following steps:
[0019] Step S441: sentence-word consistency, specifically, introducing a word extractor based on a word-level attention mechanism, obtaining word-level attention through the word extractor, setting a loss function between sentences and words, and forcibly maintaining the consistency between word-level attention and sentence-level importance scores. The consistency means that if any sentence is considered important by the sentence extractor, the words contained in the sentence are also considered important;
[0020] Step S442: Dynamic suppression, specifically, in the process of encoder-decoder model iteration, the word-level attention of all words in the historical iteration process is accumulated and updated to obtain the sum of historical attention distribution, and the attention degree of the encoder-decoder model to each word in the original text in the process of generating the AI report is recorded. If the word-level attention of any word is high, it means that the word has been focused on, and the attention degree of the encoder-decoder model to the word is suppressed and reduced, and the encoder-decoder model is forced to pay attention to the rest of the original text. According to the sum of historical attention distribution, a dynamic penalty term is added on the basis of the traditional attention mechanism, and the word-level attention of the word that is focused on in the current iteration is calculated and reduced. Finally, the block-level sentence representation vector is output according to the word-level attention. The formula used is as follows: ;
[0021] In the formula, express After iterations, the sum of historical attention distributions of all words in the original text is, Indicates At the iteration, the word-level attention of each word in the original text is Indicates the current iteration round. Represents the iteration index; ;
[0022] In the formula, Indicates the original The word The attention weight after iterations, represents the hyperbolic tangent function used for compression, Indicates the encoder The expression of a word, Indicates that the decoder is The state after iterations, Indicates The word The sum of historical attention distribution after iterations, express Dimensional parameter matrix.
[0023] The present invention provides a method for automatically generating an AI report based on big data. The beneficial effects achieved by the present invention using the above scheme are as follows:
[0024] (1) In view of the technical problems that the existing technology has too many static rules and coarse granularity of parameter adjustment, the ideal token lengths of different original text types such as news and user comments vary greatly, lack of adaptive ability, resulting in misjudgment, rough handling of edge cases, and the summary may lose important content, this solution predicts user interest sentences by including eye movement features of user reading habits, replaces manual rules, and realizes personalized configuration of reports. By identifying the original text type, the ideal target AI report length is adjusted, the adaptive ability of the encoder-decoder model is enhanced, and the integrity of the information contained in the report is guaranteed by block context overlap and attention suppression;
[0025] (2) In view of the technical problems that the traditional sequence-to-sequence encoder-decoder structure of the existing technology cannot effectively process very long texts, resulting in information loss, the model repeatedly focusing on the same position of the original text, resulting in report redundancy, and short original texts may be misjudged as low quality, while long but inefficient original texts may pass verification, this solution introduces an external memory module to divide the original text into blocks, and through word-level and sentence-level attention optimization, it realizes the extraction of important information at the block level, and finally obtains the overall report of the original text, solving the common problem of repeated generation in generative reports, and balancing the information coverage and conciseness of the report in an end-to-end manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A flowchart of a method for automatically generating an AI report based on big data provided by the present invention;
[0027] Figure 2 is a schematic diagram of step S4;
[0028] Figure 3 is a schematic diagram of step S44.
[0029] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0031] In the description of the present invention, it should be understood that terms such as “upper”, “lower”, “front”, “back”, “left”, “right”, “top”, “bottom”, “inside” and “outside” indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as limiting the present invention.
[0032] Example 1, see Figures 1 to 3 The present invention provides a method for automatically generating an AI report based on big data, the method comprising the following steps:
[0033] Step S1: Identify the type of original text, which is used to achieve AI report length adaptation. Specifically, a lightweight classifier is used to automatically identify the type of original text, and the ideal target AI report length is adjusted according to the identification result.
[0034] Step S2: original text segmentation is used to segment the original text into blocks to avoid information fragmentation. Specifically, based on the original text semantics, the similarity between sentences is calculated, the maximum length of blocks is limited, overlapping contexts between blocks are allowed, paragraphs are divided while maintaining semantic integrity, and semantically coherent paragraphs are merged into blocks.
[0035] Step S3: knowledge graph construction, specifically, extracting entities and relationships from the original text, constructing a knowledge graph, and using the knowledge graph as external memory;
[0036] Step S4: memory fusion, specifically, using the encoder-decoder model to generate a local AI report for each block;
[0037] Step S5: Global aggregation, specifically, aggregating the local AI reports of all blocks, describing the relationship between local AI reports through the graph attention network, treating the local AI reports as nodes, obtaining the importance scores of the nodes, and generating the final global AI report based on the importance scores.
[0038] Example 2, see Figure 1 and Figure 2 This embodiment is based on the above embodiment. In step S4, the memory fusion is specifically performed as follows:
[0039] Step S41: Eliminate low-density information, which is used to eliminate low-density information in the original text. Specifically, pre-process the input original text, build and use an AI pre-training model, remove low-density information in the original text, the low-density information includes low-correlation, low-interaction and noise data, collect the eye movement data set of the user when reading the original text, and extract gaze features from the eye movement data set. The gaze features include gaze duration features, gaze frequency features, reading order features, the ratio of gaze duration to the number of words, and the ratio of gaze duration to the number of characters. The gaze duration features include the total gaze duration of the sentence, the relative gaze duration, the longest gaze time, the shortest gaze time, the mean, median and standard deviation of the gaze duration. The gaze frequency features include the number of gazes and the relative number of gazes. The reading order features include the first gaze order, the last gaze order and the number of words that appear in the first 10 gazes.
[0040] Step S42: Personalized screening, used to quantify user interests, specifically, using a gradient boosting decision tree, taking the gaze feature as input, predicting the sentences that the user is interested in, and outputting a binary label of whether any sentence in the original text acts on the AI report;
[0041] Step S43: obtaining a sentence-level score, specifically, introducing a sentence extractor based on a sentence-level attention mechanism, quantitatively evaluating the comprehensive quality of each sentence in the original text, and calculating the sentence-level importance score of each sentence in combination with the sentences that the user is interested in;
[0042] Step S44: Dynamic suppression, used to avoid the generated AI report from containing repeated identical content, and obtain a block-level word and sentence representation vector;
[0043] Step S45: Balanced generation, used for semantic compression and original text rewriting, specifically, the encoder queries the external memory to obtain block-related entities and the relationship between entities, the external memory returns a triple encoding vector, and extracts structured linguistic features based on linguistics. The structured linguistic features include sentence position, number of digits, named entity tags, part-of-speech tags and word weights. The triple encoding vector, the structured linguistic features and the block-level sentence representation vector are spliced by dimension, and the decoder outputs a local AI report.
[0044] Example 3, see Figures 1 to 3 This embodiment is based on the above embodiment. In step S44, the dynamic suppression is performed in the following specific steps:
[0045] Step S441: sentence-word consistency, specifically, introducing a word extractor based on a word-level attention mechanism, obtaining word-level attention through the word extractor, setting a loss function between sentences and words, and forcibly maintaining the consistency between word-level attention and sentence-level importance scores. The consistency means that if any sentence is considered important by the sentence extractor, the words contained in the sentence are also considered important;
[0046] Step S442: Dynamic suppression, specifically, in the process of encoder-decoder model iteration, the word-level attention of all words in the historical iteration process is accumulated and updated to obtain the sum of historical attention distribution, and the attention degree of the encoder-decoder model to each word in the original text in the process of generating the AI report is recorded. If the word-level attention of any word is high, it means that the word has been focused on, and the attention degree of the encoder-decoder model to the word is suppressed and reduced, and the encoder-decoder model is forced to pay attention to the rest of the original text. According to the sum of historical attention distribution, a dynamic penalty term is added on the basis of the traditional attention mechanism, and the word-level attention of the word that is focused on in the current iteration is calculated and reduced. Finally, the block-level sentence representation vector is output according to the word-level attention. The formula used is as follows: ;
[0047] In the formula, express After iterations, the sum of historical attention distributions of all words in the original text is, Indicates At the iteration, the word-level attention of each word in the original text is Indicates the current iteration round. Represents the iteration index; ;
[0048] In the formula, Indicates the original The word The attention weight after iterations, represents the hyperbolic tangent function used for compression, Indicates the encoder The expression of a word, Indicates that the decoder is The state after iterations, Indicates The word The sum of historical attention distribution after iterations, express Dimensional parameter matrix.
[0049] Example 4, see Figures 1 to 3 This embodiment is based on the above embodiment. In step S2, the overlapping context is allowed to be retained between blocks, specifically including the following steps:
[0050] The content of block 1 is: "Apple's revenue increased by 12% in 2023. CEO Cook said this was due to the success of AI products." The content of block 2 is: "He also announced new investment plans." The sentence related to CEO Cook at the end of block 1 is overlapped to the head of block 2, and block 2 is updated to: "CEO Cook announced new investment plans."
[0051] Example 5, see Figures 1 to 3 ,This embodiment is based on the above embodiment. In step S1, the recognition results ,include news, comments and poems.
[0052] Example 6, see Figures 1 to 3 , this embodiment is based on the above embodiment, and in step S41, the AI pre-training model is BERT.
[0053] Embodiment 7, see Figures 1 to 3 This embodiment is based on the above embodiment. In step S3, the entities include companies, people and events.
[0054] Embodiment 8, see Figures 1 to 3 This embodiment is based on the above embodiment. In step S3, the graph database Neo4j is used to store the knowledge graph and an index is established to accelerate the query.
[0055] Embodiment 9, see Figures 1 to 3 ,This embodiment is based on the above embodiment. In step S2, the maximum length of the block ,is set to 100 Tokens.
[0056] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0057] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
[0058] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.
Claims
1. A method for automatically generating AI reports based on big data, characterized by: The method comprises the following steps: Step S1: Identify the type of original text, which is used to achieve AI report length adaptation. Specifically, a lightweight classifier is used to automatically identify the type of original text, and the ideal target AI report length is adjusted according to the identification result. Step S2: original text segmentation is used to segment the original text into blocks to avoid information fragmentation. Specifically, based on the original text semantics, the similarity between sentences is calculated, the maximum length of blocks is limited, overlapping contexts between blocks are allowed, paragraphs are divided while maintaining semantic integrity, and semantically coherent paragraphs are merged into blocks. Step S3: knowledge graph construction, specifically, extracting entities and relationships from the original text, constructing a knowledge graph, and using the knowledge graph as external memory; Step S4: memory fusion, specifically, using the encoder-decoder model to generate a local AI report for each block; Step S5: Global aggregation, specifically, aggregating the local AI reports of all blocks, describing the relationship between local AI reports through the graph attention network, treating the local AI reports as nodes, obtaining the importance scores of the nodes, and generating the final global AI report based on the importance scores.
2. The method for automatically generating AI reports based on big data according to claim 1, characterized in that: In step S4, the memory fusion is specifically performed as follows: Step S41: Eliminate low-density information, which is used to eliminate low-density information in the original text. Specifically, pre-process the input original text, build and use an AI pre-trained model, remove low-density information in the original text, collect the eye movement data set of the user when reading the original text, and extract gaze features from the eye movement data set; Step S42: Personalized screening, used to quantify user interests, specifically, using a gradient boosting decision tree, taking the gaze feature as input, predicting the sentences that the user is interested in, and outputting a binary label of whether any sentence in the original text acts on the AI report; Step S43: obtaining a sentence-level score, specifically, introducing a sentence extractor based on a sentence-level attention mechanism, quantitatively evaluating the comprehensive quality of each sentence in the original text, and calculating the sentence-level importance score of each sentence in combination with the sentences that the user is interested in; Step S44: Dynamic suppression, used to avoid the generated AI report from containing repeated identical content, and obtain a block-level word and sentence representation vector; Step S45: Balanced generation, used for semantic compression and original text rewriting, specifically, the encoder queries the external memory to obtain block-related entities and the relationship between entities, the external memory returns the triple encoding vector, extracts structured linguistic features based on linguistics, the triple encoding vector, the structured linguistic features and the block-level word representation vector are spliced by dimension, and the decoder outputs a local AI report.
3. The method for automatically generating AI reports based on big data according to claim 2, characterized in that: In step S44, the dynamic suppression is performed in the following specific steps: Step S441: sentence-word consistency, specifically, introducing a word extractor based on a word-level attention mechanism, obtaining word-level attention through the word extractor, setting a loss function between sentences and words, and forcibly maintaining the consistency between word-level attention and sentence-level importance scores. The consistency means that if any sentence is considered important by the sentence extractor, the words contained in the sentence are also considered important; Step S442: Dynamic suppression, specifically, in the process of encoder-decoder model iteration, the word-level attention of all words in the historical iteration process is accumulated and updated to obtain the sum of historical attention distribution, and the attention degree of the encoder-decoder model to each word of the original text in the process of generating AI report is recorded. If the word-level attention of any word is high, it means that the word has been focused on, and the attention degree of the encoder-decoder model to the word is suppressed and reduced, and the encoder-decoder model is forced to pay attention to the rest of the original text. According to the sum of historical attention distribution, a dynamic penalty term is added on the basis of the traditional attention mechanism, and the word-level attention of the words that are focused on in the current iteration is calculated and reduced. Finally, the block-level sentence representation vector is output according to the word-level attention.
Citation Information
Patent Citations
AI content abstract generation method and device
CN119322846A
Semantic extraction model and method for enhanced generation of case information retrieval
CN119621934A
Document abstract method based on domain knowledge and multi-granularity graph network
CN119623617A