Aggregate question and answer method, system and device based on embedded semantic index and hierarchical aggregation reasoning and medium
By employing embedded semantic indexing and hierarchical aggregation reasoning, dialogue logs are automatically anonymized and multi-layered semantically vectorized. Combined with a question-answering engine and RAG correction mechanism, the problems of cross-session aggregation reasoning and group identification are solved, enabling the extraction of common viewpoints from massive dialogues and providing reliable business decision-making basis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-01
Smart Images

Figure CN121958495A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of intelligent decision-making, and in particular to an aggregated question-answering method, system, device, and medium based on embedded semantic indexing and hierarchical aggregated reasoning. Background Technology
[0002] Large Language Models (LLMs) are currently widely used in chatbots, intelligent customer service, and virtual assistant systems, generating massive amounts of human-computer dialogue data. This data itself contains rich social and business insights, such as customer complaint hotspots, user attention trends, and potential risk events. However, existing technologies mainly suffer from the following problems: Lack of cross-conversation aggregation reasoning capabilities: Traditional dialogue models treat each round of interaction as an independent event, failing to integrate the semantic patterns of multiple users to form macro-level insights. For example, a customer repeatedly mentions that "the policy change process is complicated," but the system cannot summarize this common problem. Lack of semantic attribution and group identification mechanisms: Existing public opinion analysis or customer service monitoring only performs keyword statistics, failing to identify the differentiated concerns of different groups (such as elderly customers, VIP customers) at the semantic level. Summary of the Invention
[0003] This invention provides an aggregated question-answering method, system, computer device, and medium based on embedded semantic indexing and hierarchical aggregated reasoning, in order to solve the technical problems of lack of cross-session aggregated reasoning capabilities and lack of semantic attribution and group recognition mechanisms in the prior art.
[0004] Firstly, a question-answering method based on embedded semantic indexing and hierarchical aggregation reasoning is provided, including: Automated anonymization processing is performed on the original dialogue logs, wherein the original dialogue logs include at least one separable semantic block; Perform multi-layer semantic vectorization transformation on the desensitized semantic blocks; Integrating semantic vectors and a question-answering engine, the question-answering engine performs semantic aggregation reasoning based on set-level semantic vectors to generate semantic aggregation reasoning results; Based on the aggregated reasoning results of the question-answering engine, the semantic vectors are subjected to group difference analysis to obtain the initial group difference analysis results; The semantic aggregation reasoning results and the initial group difference analysis results are double-corrected through the embedded RAG correction mechanism, and the corrected results are generated. By combining the semantic aggregation reasoning results, the corrected group difference analysis results, and the correction process records, a standardized structured aggregation report is constructed.
[0005] Secondly, an aggregated question-answering system based on embedded semantic indexing and hierarchical aggregated reasoning is provided, including: The desensitization module is used to automatically desensitize the original dialogue logs, wherein the original dialogue logs include at least one separable semantic block; The vectorization module is used to perform multi-level semantic vectorization transformation on the desensitized semantic blocks; The aggregation module is used to integrate semantic vectors and the question answering engine. It performs semantic aggregation reasoning based on the semantic vectors at the set level by the question answering engine and generates semantic aggregation reasoning results. The analysis module is used to perform group difference analysis on semantic vectors based on the aggregated reasoning results of the question-answering engine to obtain the initial group difference analysis results; The correction module is used to perform dual correction on the semantic aggregation reasoning results and the initial group difference analysis results through the embedded RAG correction mechanism, and generate the correction results. The output module is used to combine the semantic aggregation reasoning results, the corrected group difference analysis results, and the correction process record to build a standardized structured aggregation report.
[0006] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described aggregate question-answering method based on embedded semantic indexing and hierarchical aggregate reasoning.
[0007] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described aggregate question-answering method based on embedded semantic indexing and hierarchical aggregate reasoning.
[0008] The above-mentioned solution, implemented by the aggregated question-answering method, system, computer device, and storage medium based on embedded semantic indexing and hierarchical aggregated reasoning, can achieve the following: Automated desensitization of the original dialogue log, where the original dialogue log includes at least one separable semantic block; multi-layer semantic vectorization transformation of the desensitized semantic block; integration of semantic vectors and a question-answering engine; semantic aggregated reasoning based on set-level semantic vectors by the question-answering engine to generate semantic aggregated reasoning results; group difference analysis of the semantic vectors based on the aggregated reasoning results of the question-answering engine to obtain initial group difference analysis results; dual correction of the semantic aggregated reasoning results and the initial group difference analysis results through an embedded RAG correction mechanism to generate a correction result; and construction of a standardized structured aggregated report by combining the semantic aggregated reasoning results, the corrected group difference analysis results, and the correction process record. Automated desensitization technology accurately masks sensitive information, meeting data security compliance requirements while preserving core semantics to the greatest extent. Simultaneously, semantic block concatenation and metadata appending transform fragmented dialogues into structured, interconnected analytical units, resolving the issues of "disorganized data sources and lack of unified processing benchmarks" in subsequent analysis. This provides clean, standardized, and usable foundational data support for the entire analysis process. Multi-layered semantic vectorization converts text into machine-computable numerical vectors, enabling comprehensive extraction and structured representation of semantic features. Combined with the construction of semantic vector indexes, this endows the system with efficient cross-semantic retrieval and similarity aggregation capabilities. Semantic aggregation reasoning at the set level extracts common viewpoints from massive dialogues, solving the problem of "fragmented single-line analysis and lack of global understanding." By appending source chains to aggregated answers, the corresponding dialogue clusters and key statements are clearly identified, addressing the issue of "poor traceability of aggregation results." By constructing standardized and structured aggregated reports, scattered aggregated inference results, group difference data, and correction records are integrated into visualized and interpretable modules (hotspot distribution, difference charts, traceability tables, etc.), which not only lowers the threshold for non-technical personnel to understand the analysis results, but also provides clear and reliable basis for business decisions. Attached Figure Description
[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic diagram of an application environment for an aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning, according to an embodiment of the present invention.
[0011] Figure 2This is a flowchart illustrating an aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning in one embodiment of the present invention.
[0012] Figure 3 This is a schematic diagram of the structure of an aggregated question-answering system based on embedded semantic indexing and hierarchical aggregated reasoning in one embodiment of the present invention.
[0013] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention.
[0014] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] The aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning provided in this invention can be applied to, for example... Figure 1In this application environment, the client automatically de-identifies the original dialogue logs, which include at least one separable semantic block. The de-identified semantic block undergoes multi-layer semantic vectorization transformation. Semantic vectors are integrated with a question-answering engine. Based on the question-answering engine's set-level semantic vectors, semantic aggregation reasoning is performed to generate semantic aggregation reasoning results. Based on the question-answering engine's aggregation reasoning results, group difference analysis is performed on the semantic vectors to obtain initial group difference analysis results. An embedded RAG correction mechanism performs dual correction on the semantic aggregation reasoning results and the initial group difference analysis results, generating a corrected result. Combining the semantic aggregation reasoning results, the corrected group difference analysis results, and the correction process record, a standardized structured aggregation report is constructed. Automated de-identification technology achieves precise masking of sensitive information, meeting data security compliance requirements while preserving core semantics to the greatest extent. Simultaneously, through semantic block concatenation and metadata appending, fragmented dialogues are transformed into structured, associative analytical units, solving the problem of "disorganized data sources and lack of unified processing benchmarks" in subsequent analysis. This provides clean, standardized, and usable basic data support for the entire analysis process. By transforming text into machine-computable numerical vectors through multi-layer semantic vectorization, comprehensive extraction and structured representation of semantic features are achieved. Combined with the construction of semantic vector indexes, the system gains efficient cross-semantic retrieval and similarity aggregation capabilities. Through set-level semantic aggregation reasoning, common viewpoints are extracted from massive amounts of dialogue, solving the problem of "fragmented single-item analysis and lack of global understanding." By attaching a source chain to the aggregated answers, the corresponding dialogue clusters and key statements are clearly identified, solving the problem of "poor traceability of aggregation results." Through the construction of standardized structured aggregation reports, scattered aggregation reasoning results, group difference data, and correction records are integrated into visualized and interpretable modules (hotspot distribution, difference graphs, traceability tables, etc.), lowering the understanding threshold for non-technical personnel and providing clear and reliable basis for business decisions. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0017] Please see Figure 2 As shown, Figure 2 A flowchart illustrating the aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning provided in this embodiment of the invention includes the following steps: S10: Perform automated desensitization processing on the original dialogue log, wherein the original dialogue log includes at least one separable semantic block; S20: Perform multi-level semantic vectorization transformation on the desensitized semantic blocks; S30: Integrates semantic vectors and a question-answering engine, performs semantic aggregation reasoning based on semantic vectors at the set level by the question-answering engine, and generates semantic aggregation reasoning results; S40: Based on the aggregated reasoning results of the question-answering engine, perform group difference analysis on the semantic vectors to obtain the initial group difference analysis results; S50: Through an embedded RAG correction mechanism, the semantic aggregation reasoning results and the initial group difference analysis results are double-corrected, and the correction results are generated. S60: Combine the semantic aggregation reasoning results, the corrected group difference analysis results, and the correction process record to construct a standardized structured aggregation report.
[0018] S10 to S60 constitute a complete technical chain of "dialogue log preprocessing - core analysis - result correction - report output". The overall logic follows the progressive relationship of "data purification → data structuring → in-depth analysis → result verification → delivery". Each step is closely linked, and the output of the previous step becomes the input of the next step. Finally, a closed-loop transformation from raw dialogue data to standardized analysis report is achieved, which efficiently and accurately extracts the group semantic features, hot information and difference patterns in the dialogue, and ensures the reliability and interpretability of the analysis results.
[0019] In S10, privacy risks (such as sensitive information like user names, phone numbers, and addresses) in the original dialogue logs are eliminated while ensuring data availability, providing a "clean and compliant" foundation for subsequent semantic analysis. The smallest processable unit (semantic block) of the original data is clearly defined to avoid data fragmentation issues in subsequent analysis. Automated de-identification employs Named Entity Recognition (NER) models and regular expression matching to accurately detect and replace / hide sensitive entities in the dialogue (such as personal identification information and trade secrets), ensuring that the de-identified data meets data security compliance requirements. The semantic block definition clearly defines the basic structure of the original dialogue logs—containing at least one divisible semantic block (a semantic block refers to a dialogue unit with complete semantics, which can be a single-turn dialogue or the result of concatenating the context of multiple turns of interaction by the same user), defining clear data units for subsequent vectorization processing.
[0020] S10 includes: detecting and replacing sensitive entities in the original dialogue logs through a named entity recognition model to complete data desensitization; concatenating multiple rounds of interactions belonging to the same user session after desensitization through context to form structured semantic blocks; and attaching timestamps and user profile tags to each semantic block to provide data support for subsequent group difference analysis.
[0021] In S10, privacy risks (such as sensitive information like user names, phone numbers, and addresses) in the original dialogue logs are eliminated to achieve compliant data purification; structured semantic blocks are formed by splicing context to ensure semantic integrity; key metadata is added to the semantic blocks to provide data support for subsequent group difference analysis and time series analysis, while clarifying the processable units of the original data to avoid data fragmentation problems in subsequent analysis.
[0022] Sensitive entity detection and replacement employs Named Entity Recognition (NER) models (such as BERT-NER, BiLSTM-CRF, etc.) combined with regular expression matching rules to accurately locate sensitive entities (including personal identity information, account information, trade secrets, etc.) in the original dialogue logs. The desensitization process is completed by replacing them with placeholders (such as "[name]" or "[phone number]") or by hiding them, ensuring that the desensitized data not only meets data security compliance requirements but also retains core semantic information without loss.
[0023] The construction of structured semantic blocks mainly targets the anonymized dialogue data. It divides the data into units of "user sessions" and splices the text of multiple interactions between the same user in the same interaction scenario in chronological order to form structured semantic blocks with complete semantic logic (for example, multiple dialogues between users from inquiring about product features to providing feedback on usage issues, which are spliced together to form a complete "requirement-feedback" semantic unit), thus avoiding the problem of semantic fragmentation in single-turn dialogues.
[0024] Metadata appended to each completed structured semantic block are precisely appended with two types of core metadata: one is a timestamp (recording the initiation / end time of the corresponding session, supporting subsequent time-series trend analysis); the other is user profile tags (such as the user's customer group, consumption level, historical interaction preferences, etc., providing a direct basis for group segmentation for subsequent S40 group difference analysis), realizing the association and binding of semantic data with analysis dimensions.
[0025] S10 Input: Raw dialogue log (including user interaction text, session association identifier, user basic attribute information, and interaction time record); S10 Output: ① Anonymized dialogue text (sensitive information has been processed); ② Structured semantic blocks (composed of multiple rounds of interaction in the same user session); ③ A collection of semantic blocks with metadata (each semantic block is accompanied by a timestamp and user profile tag).
[0026] The quality of desensitization directly affects the accuracy of subsequent semantic analysis (avoiding sensitive information interfering with semantic feature extraction), while the integrity of the semantic block determines the semantic integrity of the vectorization result. The "structured semantic block with metadata" output by S10 is the foundation for subsequent steps: the desensitized semantic content is the direct input to S20's "multi-layer semantic vectorization," ensuring the semantic accuracy of the vectorization result; structured concatenation ensures the integrity of the semantic vector, avoiding vectorization bias caused by missing semantics in single-turn dialogues; the added timestamp supports subsequent time-series correlation analysis, and user profile tags directly provide group segmentation dimensions for S40's "group difference analysis," significantly improving the analysis efficiency and accuracy of S40. The integrity of the semantic block determines the semantic integrity of the vectorization result.
[0027] S20 transforms textual semantic blocks that cannot be directly computed into machine-processable numerical semantic vectors. Simultaneously, it achieves comprehensive semantic feature coverage through multi-layered extraction (such as literal semantics, deep semantics, and contextual semantics), providing a structured semantic representation foundation for subsequent aggregation reasoning and group difference analysis. Based on pre-trained language models (such as BERT and RoBERTa), multi-layered semantic feature extraction is performed on each anonymized semantic block output by S10, generating a corresponding high-dimensional semantic vector: the lower-level semantics extracts basic features such as the literal meaning of the text and keyword associations; the higher-level semantics mines deep features such as sentiment, intent, and contextual logical connections; the final output integrates the multi-layered semantic features to form a unified-dimensional semantic vector, achieving the transformation from "textual semantics to numerical vectors." The input to S20 is the anonymized semantic block output by S10; the output of S20 is the multi-layered semantic vector (structured semantic representation) corresponding to each semantic block.
[0028] S20 includes: performing multi-level semantic vectorization processing on the desensitized semantic blocks to generate corresponding semantic vectors; and establishing a semantic vector index based on the semantic vectors, wherein the semantic vector index supports cross-semantic retrieval and similarity aggregation functions.
[0029] S20 transforms textual semantic blocks that cannot be directly computed into machine-processable numerical semantic vectors, while achieving comprehensive semantic feature coverage through multi-layer extraction. Based on this, a semantic vector index is established, providing efficient data support for subsequent aggregation question retrieval and similarity aggregation, ensuring the efficiency and accuracy of core analysis steps. Based on pre-trained language models (such as BERT, RoBERTa, etc.), multi-layer semantic feature extraction is performed on each de-identified semantic block output by S10, generating corresponding high-dimensional semantic vectors. Subsequently, a semantic vector index is established based on the generated semantic vectors. The specific operations are as follows: Multi-layer semantic vectorization processing: Extracting low-level semantic features such as literal meaning and keyword associations from the text, while simultaneously mining high-level semantic features such as sentiment, intent, and contextual logical connections, integrating them to form a unified-dimensional semantic vector, achieving the transformation from "textual semantics to numerical vectors"; Semantic vector index establishment: Building a semantic vector index library based on the generated semantic vectors. This index has two core functions—first, cross-semantic retrieval, which can quickly match vector data related to the target semantics; second, similarity aggregation, which can efficiently cluster similar semantic vectors, providing a data retrieval and aggregation foundation for subsequent aggregation reasoning. S20 Input: Semantic blocks with desensitized metadata output from S10; S20 Output: ① Multi-level semantic vectors corresponding to each semantic block; ② Semantic vector index library with cross-semantic retrieval and similarity aggregation functions.
[0030] S30 achieves aggregation through semantic vector similarity calculation. Relying on the reasoning capabilities of the question-answering engine, it performs "collection-level" aggregation analysis on structured semantic vectors, extracts common semantic features from massive semantic blocks, summarizes core viewpoints, and generates comprehensive semantic aggregation reasoning results, solving the problem of "fragmented analysis of single semantic blocks and lack of a global perspective".
[0031] The semantic vectors generated by S20 are technically integrated with the question-answering engine to build a link between "semantic vectors and the inference engine," allowing the question-answering engine to directly access the semantic vector data. Based on semantic vectors at the "set level" (i.e., batch semantic vectors, not single vectors), the question-answering engine uses algorithms such as similarity clustering and common feature extraction to summarize common expressions from the clustering results of similar semantic blocks, ultimately extracting semantic aggregation inference results (such as common needs of a certain type of user, core viewpoints of frequently discussed topics, etc.). S30 input: semantic vectors output by S20, and the integrated question-answering engine; S30 output: semantic aggregation inference results (with globally comprehensive semantic conclusions).
[0032] The semantic aggregation reasoning results of S30 form the basis for S40's "group difference analysis" (difference analysis needs to be based on the aggregated semantic features to compare the differences between different groups). S40 quantifies group differences by calculating the distance between vectors, and the semantic representation ability of the vectors directly determines the accuracy of the subsequent analysis results.
[0033] S30 includes: after inputting the aggregation question, the aggregation question is transformed into a structured query template through a semantic parsing model; based on the structured query template, the corresponding session set is retrieved from the semantic index library, and an aggregated answer is formed by clustering similar semantic blocks and summarizing common expressions. Each aggregated answer is accompanied by a source chain, which includes dialogue clusters and key statements.
[0034] S30 achieves end-to-end analysis from "aggregation question → structured parsing → precise retrieval → aggregation answer generation", efficiently locating the target session set from the semantic vector index library and extracting aggregation answers with commonalities; at the same time, it ensures the interpretability of the aggregation results through the source chain, solving the problem of "unclear source of aggregation answer".
[0035] After inputting the aggregation question, the semantic parsing model performs semantic decomposition and structured transformation of the question to generate a standardized structured query template (clearly defining core information such as query dimensions and target scope) to avoid ambiguity in natural language questions. Based on the structured query template, the corresponding session set is accurately retrieved from the semantic vector index library built by S20. Based on the retrieved session set, relevant semantic vectors are aggregated through a similar semantic block clustering algorithm, and common expressions are summarized from the clustering results to extract aggregated answers. Each aggregated answer is bound to a unique traceability chain, which must clearly define two core pieces of information: first, the main source dialogue cluster of the aggregated answer (i.e., the relevant session group formed by clustering); and second, the key statements supporting the answer, ensuring that the aggregation results are traceable and verifiable.
[0036] S30 Inputs: Aggregation question, semantic vector index library built in S20, semantic parsing model; S30 Outputs: ① Structured query template; ② Aggregated answer; ③ Source chain with accompanying dialogue clusters and key statements. The aggregated answer output by S30 forms the basis for further analysis in S40, while the structured query template provides direction for subsequent precise analysis; the source chain provides core material for the "Evidence Tracing Table" module of the S60 structured report, ensuring the report's interpretability; simultaneously, the aggregated answer is also an important target for correction by the S50RAG correction mechanism, and its accuracy directly affects the quality of subsequent analysis.
[0037] Building upon the S30's "global aggregation," this study further refines the analytical dimensions, focusing on semantic differences at the "group" level (such as differing viewpoints among different user groups on the same topic, and differences in needs and preferences among different customer groups). This enriches the hierarchical nature of the analytical results and provides more granular evidence for precise decision-making. Using the aggregated inference results output by S30 as a benchmark, and combining semantic vector data, the study performs quantitative analysis of the differences in semantic vectors corresponding to different groups: based on user profile tags, interaction scenarios, and other dimensions, users corresponding to semantic vectors are divided into different groups; through vector distance calculation, feature comparison, and other methods, the study analyzes the deviation between the semantic vectors of different groups and the aggregated inference results, identifying semantic differences between groups; the results output the initial group difference analysis results (such as "Group A pays more attention to product price, while Group B pays more attention to product quality").
[0038] S40 Input: Semantic aggregation reasoning results from S30, semantic vectors from S20, and user group segmentation dimensions; S40 Output: Initial group difference analysis results (unverified group difference conclusions).
[0039] S40 includes: calculating the group center embedding vector based on the desensitized semantic vector and the timestamp information of the semantic vector, and quantifying group differences through Euclidean distance or KL divergence; constructing the time series of semantic vectors, and mining emerging topics or potential risk signals by combining cumulative sum and control charts and using change point detection methods; performing "secondary generation" aggregation inference in combination with a large language model, integrating the quantitative results of group differences, the signals from change point detection queries, and cluster statistics and contextual retrieval information, and generating initial group difference analysis results with accompanying analytical basis and explanation.
[0040] The analysis delves deeper into two core dimensions: “group differences” and “temporal changes”. It quantifies the semantic differences among different groups and uncovers emerging topics and potential risk signals. Then, it integrates the multi-dimensional analysis results through a large-scale language model (LLM) to generate structured aggregated report materials with supporting evidence and explanations, providing core content support for the final deliverables. Based on the anonymized semantic vectors output by S20 and the timestamp information added by S10, multi-dimensional analysis is conducted and aggregation inference is performed through LLM. The specific operations are as follows: Using user profile tags as the basis for group segmentation, the central embedding vectors of each group are calculated based on the anonymized semantic vectors; using algorithms such as Euclidean distance or KL divergence, the differences between the central embedding vectors of different groups are calculated to achieve a quantitative representation of group differences; using semantic vectors as core data and combining them with the attached timestamp information, a semantic feature time series is constructed; the Cumulative Sum Control Chart (CUSUM) change point detection method is used to analyze the time series data, locate the abrupt change nodes in the time series, and then mine the emerging topics or potential risk signals corresponding to the abrupt change nodes; a large-scale language model is introduced to perform "secondary generation" aggregation inference, integrating three types of core information—quantitative results of group differences, signals mined by change point detection, statistical data of clustering of similar semantic blocks, and contextual retrieval information; based on the integrated information, a structured aggregation report is generated, and the report must include the analysis basis and explanation to ensure the comprehensibility of the content.
[0041] Inputs: Desensitized semantic vectors from S20, timestamps and user profile tags from S10, large language model, and CUSUM change point detection tool; Outputs: ① Quantification results of group differences; ② Emerging topics / potential risk signals; ③ Structured aggregated report materials with analysis basis and explanation.
[0042] The group difference quantification results, emerging topics / risk signals, and structured aggregated report materials output by S40 are all subject to correction by the S50 embedded RAG correction mechanism; the corrected results will be directly integrated into the standardized structured aggregated report of S60, and the depth and accuracy of its analysis determine the business value of the final report.
[0043] The initial group difference analysis results are another target of the S50 "double correction," with their accuracy verified through the RAG mechanism. Simultaneously, the corrected group difference analysis results will serve as a key module in the S60 report, providing detailed analytical content. This addresses potential issues such as semantic bias and conflicts with external knowledge in the S30 and S40 analysis results. Verification with authoritative external knowledge enhances the accuracy and credibility of the analysis results, while generating correction records to ensure traceability. Correction preparation involves introducing an embedded RAG (Retrieval Enhanced Generation) correction mechanism. The semantic aggregation reasoning results of S30 and the initial group difference analysis results of S40 are used as search queries. Based on these queries, relevant verification knowledge is retrieved from an authoritative external knowledge base. The retrieved external knowledge is compared with the two analysis results to determine if semantic conflicts exist. If conflicts are found, the corresponding analysis results are marked "uncertain," and a correction result is generated. The entire correction process (including the source of retrieved knowledge, points of conflict, and correction basis) is recorded, forming a correction process record.
[0044] S50 Inputs: Semantic aggregation reasoning results from S30, initial group difference analysis results from S40, embedded RAG correction mechanism, and external authoritative knowledge base; S50 Outputs: Corrected semantic aggregation reasoning results, corrected group difference analysis results, and correction process record.
[0045] S50 includes: introducing an embedded RAG correction mechanism, using the conclusion of the aggregated semantic vector as a retrieval query to retrieve authoritative knowledge from an external knowledge base; performing consistency verification between the retrieved authoritative knowledge and the conclusion of the aggregated semantic vector; if a conflict is found, the conclusion of the aggregated semantic vector is marked as uncertain and a two-way interpretation report is generated; if no conflict is found, the conclusion of the aggregated semantic vector is marked as certain and a two-way interpretation report is generated.
[0046] The S50 incorporates external authoritative knowledge through an embedded RAG correction mechanism to verify the consistency of the conclusions of aggregated semantic vectors. Regardless of whether there are conflicts, it generates a two-way interpretation report and marks the certainty of the conclusions, fully ensuring the accuracy, interpretability, and traceability of the correction results. An embedded RAG (Retrieval Enhancement Generation) correction mechanism is introduced, using the semantic aggregation reasoning results of S30 and the initial group difference analysis results of S40 as retrieval query terms. Based on the retrieval query terms, relevant verification knowledge is retrieved from an external authoritative knowledge base. The retrieved external knowledge is compared with the two analysis results to determine if there are any semantic conflicts. The embedded RAG (Retrieval Enhancement Generation) correction mechanism then uses the conclusion of the aggregated semantic vector as the core retrieval query term. Based on this query term, matching authoritative knowledge is retrieved from an external authoritative knowledge base to form the verification basis. The retrieved external authoritative knowledge is compared with the conclusion of the aggregated semantic vector point by point to determine if there are any semantic conflicts, information biases, etc. If a conflict is found, the conclusion of the aggregated semantic vector is marked "uncertain" and a two-way explanation report is generated (clarifying the conflict point, the source of the external authoritative knowledge, and the supporting basis for the aggregation conclusion). If no conflict is found, the conclusion of the aggregated semantic vector is marked "confirmed," and a two-way explanation report is also generated (clarifying the consistency point, the source of the external authoritative knowledge, and the supporting basis for the aggregation conclusion).
[0047] S50 Inputs: Conclusions of aggregated semantic vectors, embedded RAG correction mechanism, external authoritative knowledge base; S50 Outputs: ① Aggregated semantic vector conclusions marked "definite" or "uncertain"; ② Two-way explanatory report with conflict / consistency points, authoritative knowledge sources, and supporting evidence for the conclusions.
[0048] The "corrected analysis results" and "correction process record" output by S50 serve as correction materials for S60 to construct structured aggregated reports. The corrected results ensure the accuracy of the report content, while the correction process record supports the report's "interpretability" (such as evidence traceability). S60 systematically integrates the analysis results from preceding steps (aggregate inference results, group difference analysis results) and correction-related information to form standardized, visualized, and interpretable aggregated reports. This lowers the barrier to entry for using the analysis results and provides direct and clear basis for business decisions. The systematic integration of three core deliverables from preceding steps—semantic aggregation inference results, corrected group difference analysis results, and correction process records—through standardized module design and content organization, forms visualized, interpretable, and directly decision-supporting structured aggregated reports. This transforms technical analysis results into business value while lowering the barrier to entry for non-technical personnel.
[0049] The specific operation follows a core logic of "material integration - module construction - standardized output," proceeding in three steps: First, material sorting and matching involves categorizing and breaking down three core input materials: semantic aggregation reasoning results correspond to "hot topics and core viewpoints," corrected group difference analysis results correspond to "customer group difference comparison," and correction process records correspond to "conclusion basis and source explanation." Second, standardized module construction involves filling fixed modules with the decomposed materials based on preset report specifications to ensure a consistent report structure. Third, visualization and interpretation optimization involves graphically processing core data (such as difference data and topic distribution data) and adding concise interpretations to each core conclusion, while combining correction process records to enhance the credibility of the conclusions. The structured aggregation report should include at least a distribution of hot topics, a customer group semantic difference map, a line graph of key indicator trend changes, a detailed table of evidence traceability for aggregation conclusions, and a confidence distribution heatmap. Typical standardized modules include: presenting high-frequency topics and their distribution percentages in visual charts (such as pie charts and bar charts) based on semantic aggregation reasoning results, and annotating the core viewpoints corresponding to the topics; intuitively displaying the semantic preferences and viewpoint differences of different customer groups through comparative charts (such as radar charts and heat maps) based on the corrected group difference analysis results, and clarifying the quantitative indicators of differences; linking the correction process records and presenting the supporting basis (key dialogue statements, dialogue clusters), correction status (confirmed / uncertain annotations), and authoritative knowledge verification sources for each aggregation conclusion in tabular form; and adding modules such as trend change analysis and conclusion confidence assessment according to business needs to further enhance the comprehensiveness and reference value of the report.
[0050] S60 Input: Three types of core materials, namely the semantic aggregation reasoning results generated by S30, the corrected group difference analysis results output by S50, and the correction process record formed by S50; S60 Output: Standardized structured aggregation report (with core attributes of unified structure, visual presentation, traceable conclusions, and clear interpretation).
[0051] S60 is the final value delivery stage in the entire dialogue log analysis process. Its core value lies in transforming the previously scattered technical analysis results (inference results, discrepancy data, and correction records) into decision-making basis that can be directly applied to the business. At the same time, the standardized report format not only ensures the reusability and inheritance of analysis results, but also provides feedback for subsequent process optimization. By reviewing the application effects of the reports, the parameters and models of previous steps such as anonymization, vectorization, and correction can be iterated backward, forming a complete business closed loop of "analysis-application-optimization".
[0052] As can be seen, the above solution first automates the anonymization of the original dialogue logs, where each log contains at least one separable semantic block. The anonymized semantic block undergoes multi-layered semantic vectorization transformation. Semantic vectors are integrated with a question-answering engine, and semantic aggregation reasoning is performed based on the set-level semantic vectors from the question-answering engine to generate semantic aggregation reasoning results. Based on the aggregation reasoning results from the question-answering engine, group difference analysis is performed on the semantic vectors to obtain initial group difference analysis results. An embedded RAG correction mechanism performs dual correction on the semantic aggregation reasoning results and the initial group difference analysis results, generating a corrected result. Combining the semantic aggregation reasoning results, the corrected group difference analysis results, and the correction process record, a standardized structured aggregation report is constructed. Automated anonymization technology achieves precise masking of sensitive information, meeting data security compliance requirements while preserving core semantics to the greatest extent. Simultaneously, through semantic block splicing and metadata appending, fragmented dialogues are transformed into structured, associative analytical units, solving the problem of "disorganized data sources and lack of unified processing benchmarks" in subsequent analysis, providing clean, standardized, and usable basic data support for the entire analysis process. By transforming text into machine-computable numerical vectors through multi-layer semantic vectorization, comprehensive extraction and structured representation of semantic features are achieved. Combined with the construction of semantic vector indexes, the system gains efficient cross-semantic retrieval and similarity aggregation capabilities. Semantic aggregation reasoning at the set level extracts common viewpoints from massive amounts of dialogue, solving the problem of "fragmented single-item analysis and lack of global understanding." By attaching a source chain to the aggregated answers, the corresponding dialogue clusters and key statements are clearly identified, addressing the problem of "poor traceability of aggregation results." The construction of standardized structured aggregation reports integrates scattered aggregation reasoning results, group difference data, and correction records into visualized and interpretable modules (hotspot distribution, difference graphs, traceability tables, etc.), lowering the barrier to understanding for non-technical personnel and providing clear and reliable evidence for business decisions.
[0053] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0054] In one embodiment, an aggregate question-answering system based on embedded semantic indexing and hierarchical aggregation reasoning is provided. This system corresponds one-to-one with the aggregate question-answering methods based on embedded semantic indexing and hierarchical aggregation reasoning described in the above embodiments. Figure 3 As shown, this aggregated question-answering system based on embedded semantic indexing and hierarchical aggregation reasoning includes a desensitization module 101, a vectorization module 102, an aggregation module 103, an analysis module 104, a correction module 105, and an output module 106. Detailed descriptions of each functional module are as follows: The desensitization module 101 is used to automatically desensitize the original dialogue log, wherein the original dialogue log includes at least one separable semantic block. Vectorization module 102 is used to perform multi-level semantic vectorization transformation on the desensitized semantic blocks; The aggregation module 103 is used to integrate semantic vectors and question answering engine, and to perform semantic aggregation reasoning based on semantic vectors at the set level by question answering engine to generate semantic aggregation reasoning results; Analysis module 104 is used to perform group difference analysis on semantic vectors based on the aggregated reasoning results of the question answering engine to obtain initial group difference analysis results; The correction module 105 is used to perform dual correction on the semantic aggregation reasoning results and the initial group difference analysis results through the embedded RAG correction mechanism, and generate the correction results. Output module 106 is used to combine semantic aggregation reasoning results, corrected group difference analysis results, and correction process records to construct a standardized structured aggregation report.
[0055] In one embodiment, the desensitization module 101 is further configured to: The data desensitization process is completed by detecting and replacing sensitive entities in the original dialogue log using a named entity recognition model. Multiple rounds of interaction belonging to the same user session, after being de-identified, are concatenated using context to form structured semantic blocks; Each semantic block is appended with a timestamp and user profile tag to provide data support for subsequent group difference analysis.
[0056] In one embodiment, the vectorization module 102 is further configured to: The desensitized semantic blocks are subjected to multi-layer semantic vectorization processing to generate corresponding semantic vectors; A semantic vector index is built based on semantic vectors, which supports cross-semantic retrieval and similarity aggregation functions.
[0057] In one embodiment, the aggregation module 103 is further configured to: After inputting the aggregation question, the semantic parsing model transforms the aggregation question into a structured query template; Based on structured query templates, the system retrieves the corresponding session sets from the semantic index. Through clustering of similar semantic blocks and summarizing and refining common expressions, it forms aggregated answers. Each aggregated answer is accompanied by a source chain, which includes dialogue clusters and key statements.
[0058] In one embodiment, the analysis module 104 is further configured to: Based on the desensitized semantic vectors and the timestamp information of the semantic vectors, the group center embedding vector is calculated, and the group differences are quantified by Euclidean distance or KL divergence. Construct time series of semantic vectors, combine cumulative sums and control charts, and use change point detection methods to mine emerging topics or potential risk signals. By combining large-scale language models to perform "secondary generation" aggregation reasoning, integrating the results of group difference quantification, change point detection query signals, and cluster statistics and contextual retrieval information, initial group difference analysis results with accompanying analytical basis and explanation are generated.
[0059] In one embodiment, the correction module 105 is further configured to: An embedded RAG correction mechanism is introduced, which uses the conclusion of the aggregated semantic vector as the retrieval query to retrieve authoritative knowledge from an external knowledge base; The retrieved authoritative knowledge is consistent with the conclusions of the aggregated semantic vector; If a conflict is found between the two, the conclusion of the aggregated semantic vector is marked as uncertain, and a two-way interpretation report is generated; If no conflict is found between the two, the conclusion of the aggregated semantic vector is marked as confirmed, and a two-way interpretation report is generated.
[0060] This invention provides an aggregated question-answering system based on embedded semantic indexing and hierarchical aggregated reasoning. It automates the anonymization of original dialogue logs, where each log includes at least one separable semantic block. The anonymized semantic block undergoes multi-layered semantic vectorization transformation. Semantic vectors are integrated with a question-answering engine, which performs semantic aggregated reasoning based on set-level semantic vectors to generate a semantic aggregated reasoning result. Based on the aggregated reasoning result, the semantic vectors are subjected to group difference analysis to obtain an initial group difference analysis result. An embedded RAG correction mechanism performs dual correction on the semantic aggregated reasoning result and the initial group difference analysis result, generating a corrected result. Combining the semantic aggregated reasoning result, the corrected group difference analysis result, and the correction process record, a standardized structured aggregated report is constructed. Automated anonymization technology achieves precise masking of sensitive information, meeting data security compliance requirements while preserving core semantics to the greatest extent. Simultaneously, through semantic block concatenation and metadata appending, fragmented dialogues are transformed into structured, associative analytical units, solving the problem of "disorganized data sources and lack of unified processing benchmarks" in subsequent analysis. This provides clean, standardized, and usable basic data support for the entire analysis process. By transforming text into machine-computable numerical vectors through multi-layer semantic vectorization, comprehensive extraction and structured representation of semantic features are achieved. Combined with the construction of semantic vector indexes, the system gains efficient cross-semantic retrieval and similarity aggregation capabilities. Semantic aggregation reasoning at the set level extracts common viewpoints from massive amounts of dialogue, solving the problem of "fragmented single-item analysis and lack of global understanding." By attaching a source chain to the aggregated answers, the corresponding dialogue clusters and key statements are clearly identified, addressing the problem of "poor traceability of aggregation results." The construction of standardized structured aggregation reports integrates scattered aggregation reasoning results, group difference data, and correction records into visualized and interpretable modules (hotspot distribution, difference graphs, traceability tables, etc.), lowering the barrier to understanding for non-technical personnel and providing clear and reliable evidence for business decisions.
[0061] Specific limitations regarding the aggregated question-answering system based on embedded semantic indexing and hierarchical aggregation reasoning can be found in the limitations of the aggregated question-answering method based on embedded semantic indexing and hierarchical aggregation reasoning mentioned above, and will not be repeated here. Each module in the aforementioned aggregated question-answering system based on embedded semantic indexing and hierarchical aggregation reasoning can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0062] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements server-side functions or steps of an aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning.
[0063] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input system connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps of an aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning. In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Automated anonymization processing is performed on the original dialogue logs, wherein the original dialogue logs include at least one separable semantic block; Perform multi-layer semantic vectorization transformation on the desensitized semantic blocks; Integrating semantic vectors and a question-answering engine, the question-answering engine performs semantic aggregation reasoning based on set-level semantic vectors to generate semantic aggregation reasoning results; Based on the aggregated reasoning results of the question-answering engine, the semantic vectors are subjected to group difference analysis to obtain the initial group difference analysis results; The semantic aggregation reasoning results and the initial group difference analysis results are double-corrected through the embedded RAG correction mechanism, and the corrected results are generated. By combining the semantic aggregation reasoning results, the corrected group difference analysis results, and the correction process records, a standardized structured aggregation report is constructed.
[0064] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Automated anonymization processing is performed on the original dialogue logs, wherein the original dialogue logs include at least one separable semantic block; Perform multi-layer semantic vectorization transformation on the desensitized semantic blocks; Integrating semantic vectors and a question-answering engine, the question-answering engine performs semantic aggregation reasoning based on set-level semantic vectors to generate semantic aggregation reasoning results; Based on the aggregated reasoning results of the question-answering engine, the semantic vectors are subjected to group difference analysis to obtain the initial group difference analysis results; The semantic aggregation reasoning results and the initial group difference analysis results are double-corrected through the embedded RAG correction mechanism, and the corrected results are generated. By combining the semantic aggregation reasoning results, the corrected group difference analysis results, and the correction process records, a standardized structured aggregation report is constructed.
[0065] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0066] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0067] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.
[0068] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning, characterized in that, include: Automated anonymization processing is performed on the original dialogue logs, wherein the original dialogue logs include at least one separable semantic block; Perform multi-layer semantic vectorization transformation on the desensitized semantic blocks; Integrating semantic vectors and a question-answering engine, the question-answering engine performs semantic aggregation reasoning based on set-level semantic vectors to generate semantic aggregation reasoning results; Based on the aggregated reasoning results of the question-answering engine, the semantic vectors are subjected to group difference analysis to obtain the initial group difference analysis results; The semantic aggregation reasoning results and the initial group difference analysis results are double-corrected through the embedded RAG correction mechanism, and the corrected results are generated. By combining the semantic aggregation reasoning results, the corrected group difference analysis results, and the correction process records, a standardized structured aggregation report is constructed.
2. The aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning as described in claim 1, characterized in that, The step of automatically de-identifying the original dialogue logs, wherein the original dialogue logs include at least one separable semantic block, includes: The data desensitization process is completed by detecting and replacing sensitive entities in the original dialogue log using a named entity recognition model. Multiple rounds of interaction belonging to the same user session, after being de-identified, are concatenated using context to form structured semantic blocks; Each semantic block is appended with a timestamp and user profile tag to provide data support for subsequent group difference analysis.
3. The aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning as described in claim 2, characterized in that, The step of performing multi-layer semantic vectorization transformation on the desensitized semantic block includes: The desensitized semantic blocks are subjected to multi-layer semantic vectorization processing to generate corresponding semantic vectors; A semantic vector index is built based on semantic vectors, which supports cross-semantic retrieval and similarity aggregation functions.
4. The aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning as described in claim 3, characterized in that, The steps for integrating semantic vectors and the question-answering engine, and generating semantic aggregation reasoning results based on semantic vectors at the set level by the question-answering engine, include: After inputting the aggregation question, the semantic parsing model transforms the aggregation question into a structured query template; Based on structured query templates, the system retrieves the corresponding session sets from the semantic index. Through clustering of similar semantic blocks and summarizing and refining common expressions, it forms aggregated answers. Each aggregated answer is accompanied by a source chain, which includes dialogue clusters and key statements.
5. The aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning as described in claim 4, characterized in that, The step of performing group difference analysis on semantic vectors based on the aggregated reasoning results of the question-answering engine to obtain the initial group difference analysis results includes: Based on the desensitized semantic vectors and the timestamp information of the semantic vectors, the group center embedding vector is calculated, and the group differences are quantified by Euclidean distance or KL divergence. Construct time series of semantic vectors, combine cumulative sums and control charts, and use change point detection methods to mine emerging topics or potential risk signals. By combining large-scale language models for secondary generation of aggregate reasoning, and integrating group difference quantification results, change point detection query signals, cluster statistics and contextual retrieval information, initial group difference analysis results with accompanying analytical basis and explanation are generated.
6. The aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning as described in claim 5, characterized in that, The step of performing dual correction on the semantic aggregation reasoning results and the initial group difference analysis results through the embedded RAG correction mechanism, and generating the correction results, includes: An embedded RAG correction mechanism is introduced, which uses the conclusion of the aggregated semantic vector as the retrieval query to retrieve authoritative knowledge from an external knowledge base; The retrieved authoritative knowledge is consistent with the conclusions of the aggregated semantic vector; If a conflict is found between the two, the conclusion of the aggregated semantic vector is marked as uncertain, and a two-way interpretation report is generated; If no conflict is found between the two, the conclusion of the aggregated semantic vector is marked as confirmed, and a two-way interpretation report is generated.
7. The aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning as described in claim 6, characterized in that, The structured aggregation report includes at least the distribution of hot topics, a semantic difference map of customer groups, a line graph of trend changes in key indicators, a detailed table of evidence for aggregation conclusions, and a heat map of confidence distribution.
8. An aggregated question-answering system based on embedded semantic indexing and hierarchical aggregated reasoning, characterized in that, include: The desensitization module is used to automatically desensitize the original dialogue logs, wherein the original dialogue logs include at least one separable semantic block; The vectorization module is used to perform multi-level semantic vectorization transformation on the desensitized semantic blocks; The aggregation module is used to integrate semantic vectors and the question answering engine. It performs semantic aggregation reasoning based on the semantic vectors at the set level by the question answering engine and generates semantic aggregation reasoning results. The analysis module is used to perform group difference analysis on semantic vectors based on the aggregated reasoning results of the question-answering engine to obtain the initial group difference analysis results; The correction module is used to perform dual correction on the semantic aggregation reasoning results and the initial group difference analysis results through the embedded RAG correction mechanism, and generate the correction results. The output module is used to combine the semantic aggregation reasoning results, the corrected group difference analysis results, and the correction process record to build a standardized structured aggregation report.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the aggregated question-answering method based on embedded semantic indexing and hierarchical aggregated reasoning as described in any one of claims 1 to 7.