Intelligent data analysis method and system based on semantic interaction
Through intelligent data analysis methods based on semantic interaction, the limitations of traditional data analysis methods when processing complex data are solved, automatic semantic understanding and efficient data analysis are realized, technical threshold is lowered, and analysis depth and efficiency are improved.
Patent Information
- Application Number
- CN202510495160.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Traditional data analysis methods have limitations when dealing with large-scale, multi-dimensional and semantic complex data sets, making it difficult to automatically understand the semantic meaning of data, which increases the technical threshold and analysis cost of users.
Using an intelligent data analysis method based on semantic interaction, a corpus is built and a semantic interaction model is trained through multimodal data acquisition and preprocessing, a semantic interaction model is supported, and a semantic parser is converted into structured analysis tasks through a semantic parser is used, and an intuitive and easy-to-use interactive interface is built with low-code technology.
It significantly lowers the technical threshold for data analysis, improves analysis efficiency and depth, can automatically understand the semantic meaning of data, and provides more intelligent and personalized data analysis services, suitable for business intelligence, public opinion analysis and other fields.
Smart Images

Figure CN120011572A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing and analysis, and specifically to an intelligent data analysis method and system based on semantic interaction. Background Art
[0002] With the rapid development of big data technology, enterprises and research institutions have accumulated massive amounts of data resources in their daily operations and decision-making processes. These data resources are not only huge in quantity but also diverse in form, including structured data (such as database records), semi-structured data (such as log files, XML documents) and unstructured data (such as text, images, audio, etc.). However, how to quickly and accurately extract valuable information from such huge amounts of data has become an important challenge facing the current field of data analysis.
[0003] Traditional data analysis methods mainly rely on preset algorithm models and rules, and perform statistics, mining, and visualization on data to reveal the hidden laws and patterns behind the data. However, these methods often have many limitations when dealing with large-scale, multi-dimensional, and semantically complex data sets. For example, it is difficult to automatically understand the semantic meaning of the data, resulting in the analysis results not being able to accurately reflect the true intent of the data; at the same time, traditional methods usually require users to have a high technical background in order to effectively explore and query data, which undoubtedly increases the threshold and cost of data analysis.
[0004] In recent years, with the rapid development of natural language processing (NLP) and artificial intelligence technology, semantic interaction technology has gradually emerged. By simulating human language understanding and communication capabilities, semantic interaction technology enables computer systems to more accurately understand users' query intentions and contextual information, thereby providing more intelligent and personalized data analysis services. However, the application of semantic interaction technology in actual data analysis scenarios still faces many technical challenges. For example, how to efficiently build and maintain large-scale data semantic models to support complex data queries and understanding, how to design intuitive and easy-to-use interactive interfaces to lower users' technical barriers, and how to combine domain knowledge and machine learning algorithms to achieve intelligent data analysis and decision support. Summary of the invention
[0005] The technical task of the present invention is to provide an intelligent data analysis method and system based on semantic interaction to solve the limitations and insufficient analysis depth of traditional data analysis methods when processing large-scale, multi-dimensional and semantically complex data sets.
[0006] The technical task of the present invention is achieved in the following way: an intelligent data analysis method based on semantic interaction, the method is as follows: Multimodal data collection and preprocessing: Efficiently collect data from multi-source heterogeneous data sources, analyze the number of multi-source heterogeneous data sources to be collected, the data processing complexity coefficient, the unit core processing capacity, and the configuration parameters of a single physical machine, calculate the number of nodes in the data collection cluster, and perform distributed cleaning on the collected multimodal data by removing duplicate data, supplementing missing values, and replacing erroneous data. The cleaned data is labeled and classified to build a corpus; the annotation content covers the semantic category and contextual information of the data; Use the corpus to train the pre-trained language model and generate a semantic interaction model with intelligent semantic parsing capabilities; Analyze and process user input content based on the trained semantic interaction model, and finally output a data result set that meets user requirements; Use low-code technology to build a user interaction interface, collect user feedback through the user interaction interface, and optimize and adjust the semantic interaction model.
[0007] As a preferred method, the number of nodes of the data collection cluster is calculated as follows: Calculate the total number of virtual cores required, vCores, using the following formula: vCores = (total data volume × processing complexity coefficient) / unit core processing capacity; Calculate the total required memory book vMemorys, the formula is as follows: vMemorys = total data volume × memory consumption coefficient; The cluster size is calculated based on the total number of virtual cores vCores and the total memory vMemorys required, and then the number of nodes nums in the data collection cluster is calculated. The formula is as follows: nums = MAX (total required number of virtual cores vCore / number of virtual cores on a single node, total required memory vMemorys / memory on a single node).
[0008] As a preferred method, the pre-trained language model is trained using the corpus to generate a semantic interaction model with semantic intelligent parsing capabilities as follows: Focusing on the scenario requirements of data analysis, fine-tune the vocabulary embedding layer of the pre-trained language model according to the specific vocabulary of professional terms and industry terms in the data analysis scenario; The pre-trained language model is further trained using the annotated corpus, so that it can accurately understand natural language instructions in a specific field and improve the accuracy and adaptability of semantic understanding. The pre-trained language model is then trained with a large amount of annotated data, so that it can automatically learn the semantic features and contextual relationships of the language, obtain a semantic interaction model with semantic intelligent parsing capabilities, and improve the ability to understand complex semantics. The trained semantic interaction model is evaluated through the cross-validation method, and the semantic interaction model is optimized and adjusted according to the evaluation results. Among them, the optimization indicators focus on the accuracy, recall rate and F1 score of semantic understanding to ensure the performance of the semantic interaction model in the real application scenarios of intelligent analysis.
[0009] As a preferred method, the user input content is analyzed and processed based on the trained semantic interaction model, and the data result set that meets the user's requirements is finally output as follows: Based on the trained semantic interaction model (SI-Model), the natural language query input by the user is semantically parsed and converted into a semantic vector; Vector retrieval technology is used to match the parsed semantic vector with the semantic vector in the data set to quickly retrieve data related to the user's query semantics; Sort the search results according to the semantic matching degree, and optimize the search results based on the user's historical query records and preferences to improve the relevance and accuracy of the search results; Using the K-Means algorithm, the retrieved data is clustered and divided into multiple clusters according to business type; The Apriori algorithm is used to identify frequent co-occurrence relationships between data items. By mining association rules, potential associations and trends between data as well as patterns, trends and anomalies in the data are automatically identified, and finally a data result set that meets the user's ideas is output.
[0010] Preferably, the user interface has the following functions: ① Support users to interact through natural language input, voice input, drag and drop operation or visual charts; ② The intelligent analysis results of user semantics are presented to users in an intuitive and visual way in the form of charts and reports; ③ Provide intelligent prompts and automatic completion functions, record user query history and feedback information, and automatically adjust the interactive interface and recommended content.
[0011] An intelligent data analysis system based on semantic interaction, the system comprising: Data Processor is used to efficiently collect data from multiple heterogeneous data sources, dynamically adjust the cluster size to cope with different data volumes through a load-based dynamic expansion algorithm, and perform data cleaning to remove noise and duplicate data from the collected data, optimize data quality, annotate and classify the cleaned data, and build a corpus; The semantic interaction model trainer (SIMT) is used to train the pre-trained language model using the corpus, optimize the performance of the pre-trained language model through the multi-head attention mechanism and regularization technology, obtain the semantic interaction model, evaluate the semantic interaction model through the cross-validation method, and adjust the parameters in real time to improve the accuracy, recall rate and F1 score, so as to ensure that the semantic interaction model can accurately understand natural language instructions in practical applications; The semantic analyzer (SAD) is used to deeply analyze the user input content based on the trained semantic interaction model (SI-Model), extract the query intent, entity and relationship semantic information, and then convert the natural language into executable semantic instructions. It uses context-aware technology to ensure that the analysis results accurately reflect the user's true intention and provide precise semantic guidance for data retrieval and analysis; The Visual Interaction Layer (VIL) is used to use low-code technology to drag components, configure properties, and optimize layout and style through a visual designer to quickly build a simple and easy-to-use visual interface. It provides intuitive and easy-to-use visual operations and presents analysis results in the form of intuitive charts and reports, allowing users to explore data in depth through filtering and drilling interactive operations, making data analysis more convenient and efficient.
[0012] Preferably, the data acquisition processor comprises: Data collection engine (DEG), used to efficiently collect data from multi-source heterogeneous data sources, analyze the number of multi-source heterogeneous data sources to be collected, data processing complexity coefficient, unit core processing capacity and single physical machine configuration parameters, and calculate the number of nodes in the data collection cluster; The data cleaning engine (DCE) is used to perform distributed cleaning on the collected multimodal data by removing duplicate data, supplementing missing values, and replacing erroneous data, annotating and classifying the cleaned data, and building a corpus; the annotation content covers the semantic category and contextual information of the data.
[0013] Preferably, the number of nodes in the data collection cluster is calculated as follows: Calculate the total number of virtual cores required, vCores, using the following formula: vCores = (total data volume × processing complexity coefficient) / unit core processing capacity; Calculate the total required memory book vMemorys, the formula is as follows: vMemorys = total data volume × memory consumption coefficient; The cluster size is calculated based on the total number of virtual cores vCores and the total memory vMemorys required, and then the number of nodes nums in the data collection cluster is calculated. The formula is as follows: nums = MAX (total required number of virtual cores vCore / number of virtual cores on a single node, total required memory vMemorys / memory on a single node).
[0014] Preferably, the semantic interaction model trainer comprises: The fine-tuning module is used to fine-tune the vocabulary embedding layer of the pre-trained language model based on the specific vocabulary of professional terms and industry terms in the data analysis scenario around the scenario requirements of data analysis; The speech interaction model generation module is used to further train the pre-trained language model using the annotated corpus, so that the pre-trained language model can accurately understand the natural language instructions in a specific field and improve the accuracy and adaptability of semantic understanding; the pre-trained language model is then trained with a large amount of annotated data so that the pre-trained language model can automatically learn the semantic features and contextual relationships of the language, obtain a semantic interaction model with semantic intelligent parsing capabilities, and improve the ability to understand complex semantics; The model evaluation module is used to evaluate the trained semantic interaction model through the cross-validation method, and optimize and adjust the semantic interaction model according to the evaluation results; among them, the optimization indicators focus on the accuracy, recall rate and F1 score of semantic understanding to ensure the performance of the semantic interaction model in the real application scenarios of intelligent analysis.
[0015] Preferably, the semantic parser comprises: The conversion module is used to perform semantic analysis on the natural language query input by the user based on the trained semantic interaction model (SI-Model) and convert it into a semantic vector; The matching module is used to match the parsed semantic vector with the semantic vector in the data set using vector retrieval technology to quickly retrieve data related to the user's query semantics; The sorting module is used to sort the search results according to the semantic matching degree, and optimize the search results based on the user's historical query records and preferences to improve the relevance and accuracy of the search results; The clustering analysis module is used to perform cluster analysis on the retrieved data using the K-Means algorithm and divide the data into multiple clusters according to business types; The output module is used to use the Apriori algorithm to identify frequent co-occurrence relationships between data items. By mining association rules, it automatically identifies potential associations and trends between data as well as patterns, trends and anomalies in the data, and finally outputs a data result set that meets the user's ideas.
[0016] The intelligent data analysis method and system based on semantic interaction of the present invention have the following advantages: (i) The present invention realizes the automation of the whole process from natural language query to intelligent analysis by integrating natural language processing, knowledge graph and machine learning technology. Based on semantic understanding and intent recognition, it converts user queries into structured analysis tasks, and performs semantic expansion and reasoning through dynamic knowledge graph, automatically associates related data dimensions, and intelligently selects the optimal algorithm and executes distributed computing tasks based on the adaptive analysis engine, and finally presents it to data analysts in various forms such as visual charts, tables and reports; (ii) The present invention realizes bidirectional mapping of semantics and data, interactive analysis with context awareness, and visualization of results, significantly lowering the technical threshold for data analysis and improving analysis efficiency and depth. It can be widely used in business intelligence, public opinion analysis, command and decision-making, and other fields, providing users with an intelligent data analysis experience of "what you think is what you get"; (III) The present invention combines deep learning semantic understanding technology with domain knowledge graphs to achieve intelligent processing and analysis of multi-source heterogeneous data. By constructing four core technical modules, namely, data acquisition processor (DataProcessor), semantic interaction model trainer (SIMT), semantic parser (SAD), and visual interaction layer (VIL), it solves the pain points of traditional data analysis methods in processing large-scale, multi-dimensional, and semantically complex data sets, such as limitations and insufficient analysis depth, and provides a new generation of data analysis infrastructure for the construction of smart government affairs. (IV) The present invention can automatically understand the semantic meaning of data, support users to query and explore data through natural language and other intuitive methods, and combine domain knowledge and machine learning algorithms to achieve deep mining and intelligent analysis of data. It can not only significantly improve the efficiency and accuracy of data analysis and lower the technical threshold of users, but also provide more intelligent and personalized data decision support services for all walks of life; (V) In view of the challenges currently faced in the field of data analysis, especially when dealing with large-scale, multi-dimensional and semantically complex data sets, the limitations of traditional data analysis methods are becoming increasingly prominent. The advantages of the present invention are mainly reflected in the following aspects: ① Improve the intelligence level of data analysis: By introducing advanced natural language processing (NLP) technology and semantic understanding engines, the present invention aims to achieve accurate understanding of data semantics, including not only the basic meaning of data items, but also the associations between data, contextual information, etc., so as to provide users with more intelligent and personalized data analysis services; the improvement of the intelligence level will greatly enhance the depth and breadth of data analysis, helping users to dig out more valuable insights from massive data; ② Optimize user interaction experience: Traditional data analysis methods often require users to have a high technical background to effectively explore and query data. The present invention supports users to interact with the system in various forms such as natural language, drag and drop operations or visual charts by designing an intuitive and easy-to-use interactive data analysis interface. The humanized design will significantly reduce the technical threshold of users, allowing non-professionals to easily get started and enjoy the convenience brought by data analysis; ③ Enhance the accuracy and efficiency of data analysis: Combining domain knowledge and machine learning algorithms, the present invention aims to achieve deep mining and intelligent analysis of data. By automatically identifying patterns, trends and anomalies in data, it can provide users with advanced functions such as predictive analysis, association rule mining, and sentiment analysis. It will not only improve the accuracy of data analysis, but also significantly improve analysis efficiency, enabling users to make data-driven decisions faster. ④ Promote the innovation and development of data analysis technology: This invention not only solves the practical problems faced by the current data analysis field, but also provides new ideas and directions for the future development of data analysis technology; by integrating advanced technologies such as semantic interaction, natural language processing, machine learning and domain knowledge graphs, this invention is expected to lead technological innovation in the field of data analysis and promote the sustainable development of the entire industry; ⑤ Promote the maximum utilization of data value: In the data-driven era, the value of data lies in its ability to be effectively analyzed and utilized; this invention aims to help enterprises, research institutions, government departments and other industries better explore and utilize data resources by providing intelligent data analysis methods, thereby promoting business innovation, optimizing decision-making processes, improving operational efficiency, etc., and ultimately achieving the maximum utilization of data value; To sum up, the present invention solves the limitations of traditional data analysis methods when processing complex data, improves the intelligence level of data analysis, optimizes user interaction experience, enhances the accuracy and efficiency of data analysis, and at the same time promotes the innovation and development of data analysis technology, and promotes the maximum utilization of data value. It will provide more efficient and intelligent data analysis solutions for all walks of life, and help digital transformation and intelligent upgrading. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The present invention is further described below in conjunction with the accompanying drawings.
[0018] Attached Figure 1 This is a structural diagram of an intelligent data analysis system based on semantic interaction. DETAILED DESCRIPTION
[0019] The intelligent data analysis method and system based on semantic interaction of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments of the specification.
[0020] Embodiment 1: This embodiment provides an intelligent data analysis method based on semantic interaction, the method is as follows: S1. Multimodal data collection and preprocessing: Efficiently collect data from multi-source heterogeneous data sources, analyze the number of multi-source heterogeneous data sources to be collected, the data processing complexity coefficient, the unit core processing capacity, and the configuration parameters of a single physical machine, calculate the number of nodes in the data collection cluster, and perform distributed cleaning of the collected multimodal data by removing duplicate data, supplementing missing values, and replacing erroneous data. The cleaned data is labeled and classified to build a corpus; the annotation content covers the semantic category and contextual information of the data; multi-source heterogeneous data sources include business data, IoT data, system logs, monitoring data, user behavior data, etc. S2. Use the corpus to train the pre-trained language model to generate a semantic interaction model with intelligent semantic parsing capabilities; S3, analyze and process the user input content based on the trained semantic interaction model, and finally output a data result set that meets the user's requirements; S4. Use low-code technology to build a user interaction interface, collect user feedback through the user interaction interface, and optimize and adjust the semantic interaction model.
[0021] The number of nodes in the data collection cluster calculated in step S1 of this embodiment is specifically as follows: S101. Calculate the total number of virtual cores required (vCores). The formula is as follows: vCores = (total data volume × processing complexity coefficient) / unit core processing capacity; S102, calculate the total required memory book vMemorys, the formula is as follows: vMemorys = total data volume × memory consumption coefficient; S103, the cluster size is calculated according to the total number of virtual cores vCores and the total number of memory books vMemorys required, and then the number of nodes nums of the data collection cluster is calculated, and the formula is as follows: nums = MAX (total required number of virtual cores vCore / number of virtual cores on a single node, total required memory vMemorys / memory on a single node).
[0022] In step S2 of this embodiment, the pre-trained language model is trained using the corpus to generate a semantic interaction model with semantic intelligent parsing capability as follows: S201. Based on the scenario requirements of data analysis, fine-tune the vocabulary embedding layer of the pre-trained language model according to the specific vocabulary of professional terms and industry terms in the data analysis scenario; S202, using the annotated corpus, further training the pre-trained language model, so that the pre-trained language model can accurately understand the natural language instructions in a specific field, and improve the accuracy and adaptability of semantic understanding; then training the pre-trained language model with a large amount of annotated data, so that the pre-trained language model can automatically learn the semantic features and contextual relationships of the language, obtain a semantic interaction model with semantic intelligent parsing capabilities, and improve the ability to understand complex semantics; S203. Evaluate the trained semantic interaction model through cross-validation method, and optimize and adjust the semantic interaction model according to the evaluation results; among them, the optimization index focuses on the accuracy, recall rate and F1 score of semantic understanding to ensure the performance of the semantic interaction model in the real application scenario of intelligent analysis.
[0023] In step S3 of this embodiment, the user input content is analyzed and processed based on the trained semantic interaction model, and the data result set that meets the user's requirements is finally output as follows: S301, semantically analyzing the natural language query input by the user based on the trained semantic interaction model (SI-Model) and converting it into a semantic vector; S302, using vector retrieval technology to match the parsed semantic vector with the semantic vector in the data set, and quickly retrieve data related to the user's query semantics; S303, sorting the search results according to the semantic matching degree, optimizing the search results in combination with the user's historical query records and preferences, and improving the relevance and accuracy of the search results; S304, using the K-Means algorithm, performing cluster analysis on the retrieved data, and dividing the data into multiple clusters according to business types; S305. Use the Apriori algorithm to identify frequent co-occurrence relationships between data items, and automatically identify potential associations and trends between data as well as patterns, trends and anomalies in the data by mining association rules, and finally output a data result set that meets the user's ideas.
[0024] The user interaction interface in step S4 of this embodiment has the following functions: ① Support users to interact through natural language input, voice input, drag and drop operation or visual charts; ② The intelligent analysis results of user semantics are presented to users in an intuitive and visual way in the form of charts and reports; ③ Provide intelligent prompts and automatic completion functions, record user query history and feedback information, and automatically adjust the interactive interface and recommended content.
[0025] Embodiment 2: As attached Figure 1 As shown, this embodiment provides an intelligent data analysis system based on semantic interaction, the system comprising: Data Processor is used to efficiently collect data from multi-source heterogeneous data sources, dynamically adjust the cluster size to cope with different data volumes through load-based dynamic expansion algorithms, and perform data cleaning to remove noise and duplicate data from the collected data, optimize data quality, annotate and classify the cleaned data, and build a corpus. Multi-source heterogeneous data sources include business data, IoT data, system logs, monitoring data, user behavior data, etc. The semantic interaction model trainer (SIMT) is used to train the pre-trained language model using the corpus, optimize the performance of the pre-trained language model through the multi-head attention mechanism and regularization technology, obtain the semantic interaction model, evaluate the semantic interaction model through the cross-validation method, and adjust the parameters in real time to improve the accuracy, recall rate and F1 score, so as to ensure that the semantic interaction model can accurately understand natural language instructions in practical applications; The semantic analyzer (SAD) is used to deeply analyze the user input content based on the trained semantic interaction model (SI-Model), extract the query intent, entity and relationship semantic information, and then convert the natural language into executable semantic instructions. It uses context-aware technology to ensure that the analysis results accurately reflect the user's true intention and provide precise semantic guidance for data retrieval and analysis; The Visual Interaction Layer (VIL) is used to use low-code technology to drag components, configure properties, and optimize layout and style through a visual designer to quickly build a simple and easy-to-use visual interface. It provides intuitive and easy-to-use visual operations and presents analysis results in the form of intuitive charts and reports, allowing users to explore data in depth through filtering and drilling interactive operations, making data analysis more convenient and efficient.
[0026] The data acquisition processor in this embodiment includes: Data collection engine (DEG), used to efficiently collect data from multi-source heterogeneous data sources, analyze the number of multi-source heterogeneous data sources to be collected, data processing complexity coefficient, unit core processing capacity and single physical machine configuration parameters, and calculate the number of nodes in the data collection cluster; The data cleaning engine (DCE) is used to perform distributed cleaning on the collected multimodal data by removing duplicate data, supplementing missing values, and replacing erroneous data, annotating and classifying the cleaned data, and building a corpus; the annotation content covers the semantic category and contextual information of the data.
[0027] The number of nodes in the computing data collection cluster in this embodiment is as follows: Calculate the total number of virtual cores required, vCores, using the following formula: vCores = (total data volume × processing complexity coefficient) / unit core processing capacity; Calculate the total required memory book vMemorys, the formula is as follows: vMemorys = total data volume × memory consumption coefficient; The cluster size is calculated based on the total number of virtual cores vCores and the total memory vMemorys required, and then the number of nodes nums in the data collection cluster is calculated. The formula is as follows: nums = MAX (total required number of virtual cores vCore / number of virtual cores on a single node, total required memory vMemorys / memory on a single node).
[0028] The semantic interaction model trainer in this embodiment includes: The fine-tuning module is used to fine-tune the vocabulary embedding layer of the pre-trained language model based on the specific vocabulary of professional terms and industry terms in the data analysis scenario around the scenario requirements of data analysis; The speech interaction model generation module is used to further train the pre-trained language model using the annotated corpus, so that the pre-trained language model can accurately understand the natural language instructions in a specific field and improve the accuracy and adaptability of semantic understanding; the pre-trained language model is then trained with a large amount of annotated data so that the pre-trained language model can automatically learn the semantic features and contextual relationships of the language, obtain a semantic interaction model with semantic intelligent parsing capabilities, and improve the ability to understand complex semantics; The model evaluation module is used to evaluate the trained semantic interaction model through the cross-validation method, and optimize and adjust the semantic interaction model according to the evaluation results; among them, the optimization indicators focus on the accuracy, recall rate and F1 score of semantic understanding to ensure the performance of the semantic interaction model in the real application scenarios of intelligent analysis.
[0029] The semantic parser in this embodiment includes: The conversion module is used to perform semantic analysis on the natural language query input by the user based on the trained semantic interaction model (SI-Model) and convert it into a semantic vector; The matching module is used to match the parsed semantic vector with the semantic vector in the data set using vector retrieval technology to quickly retrieve data related to the user's query semantics; The sorting module is used to sort the search results according to the semantic matching degree, and optimize the search results based on the user's historical query records and preferences to improve the relevance and accuracy of the search results; The clustering analysis module is used to perform cluster analysis on the retrieved data using the K-Means algorithm and divide the data into multiple clusters according to business types; The output module is used to use the Apriori algorithm to identify frequent co-occurrence relationships between data items. By mining association rules, it automatically identifies potential associations and trends between data as well as patterns, trends and anomalies in the data, and finally outputs a data result set that meets the user's ideas.
[0030] The visualization interaction layer in this embodiment has the following functions: ① Support users to interact through natural language input, voice input, drag and drop operation or visual charts; ② The intelligent analysis results of user semantics are presented to users in an intuitive and visual way in the form of charts and reports; ③ Provide intelligent prompts and automatic completion functions, record user query history and feedback information, and automatically adjust the interactive interface and recommended content.
[0031] The working process of the system is as follows: Step 1: Build a data acquisition processor (DataProcessor) for data acquisition and preprocessing: By building a distributed data acquisition processor (DataProcessor), the acquisition and preprocessing of multimodal data can be realized; the data acquisition processor (DataProcessor) is mainly composed of two parts: the data acquisition engine (DEG) and the data cleaning engine (DCE); first, the data acquisition engine (DEG) is used to analyze the parameters such as the amount of data to be collected, the complexity coefficient of data processing, the unit core processing capacity and the configuration of a single physical machine, and the number of nodes in the data acquisition cluster is calculated. The first step is to calculate the total number of virtual cores (vCores) required: vCores = (total data volume × processing complexity coefficient) / unit core processing capacity, and the second step is to calculate the total required memory book (vMemorys) vMemorys = total data volume × memory consumption coefficient, and then the cluster size is calculated based on vCores and vMemorys. The number of nodes (nums) of the data acquisition cluster = MAX(total required vCore / single node vCore, Total required memory / single node memory); then, using the data cleaning engine (DCE), the collected large-scale data is distributed cleaned by removing duplicate data, supplementing missing values, replacing erroneous data, etc., and the cleaned data is labeled and classified to build a corpus; the annotation content needs to cover the semantic category and context information of the data; Step 2: Construct a semantic interaction model trainer (SIMT) to train the semantic interaction model: By constructing a semantic interaction model trainer (SIMT), the collected data is used to train the semantic interaction model and generate a semantic interaction model (SI-Model) with semantic intelligent parsing capabilities. First, based on the scenario requirements of data analysis, the vocabulary embedding layer of the pre-trained model is fine-tuned according to the specific vocabulary in the data analysis scenario (such as professional terms, industry terms, etc.). The model is further trained using annotated corpora to enable it to accurately understand natural language instructions in specific fields and improve the accuracy and adaptability of semantic understanding. Secondly, the model is trained with a large amount of annotated data so that it can automatically learn the semantic features and contextual relationships of the language and improve the ability to understand complex semantics. Thirdly, the trained model is evaluated through cross-validation, and the model is optimized and adjusted based on the evaluation results. The optimization indicators mainly focus on three aspects: accuracy, recall, and F1 score of semantic understanding to ensure the performance of the model in real application scenarios of intelligent analysis.
[0032] Step 3: Build a semantic parser (SAD) to perform semantic parsing and vectorized storage: By building a semantic parser (SAD), semantic parsing of user input content is achieved and converted into vectorized data; first, based on the trained semantic interaction model (SI-Model), semantic parsing of the natural language query input by the user is performed and converted into a semantic vector, and vector retrieval technology is used to match the parsed semantic vector with the semantic vector in the data set to quickly retrieve data related to the user's query semantics; secondly, the retrieval results are sorted according to the semantic matching degree, and the retrieval results are optimized in combination with the user's historical query records and preferences to improve the relevance and accuracy of the retrieval results; thirdly, the K-Means algorithm is used to perform cluster analysis on the retrieved data, and the data is divided into multiple clusters according to business types, and then the Apriori algorithm is used to identify the frequent co-occurrence relationship between data items, and by mining association rules, the potential associations and trends between data, as well as patterns, trends and anomalies in the data, are automatically identified, and finally a data result set that meets the user's ideas is output; Step 4: Build a visual interaction layer (VIL) to provide users with a visual interactive interface: Utilize low-code technology and use its visual designer to drag components, configure properties, and optimize layout and style to quickly build a simple and easy-to-use interface that supports users to interact with the system through natural language input, voice input, drag and drop operations, or visual charts. Through the above-mentioned intelligent analysis of user semantics, the analysis results are presented to users in intuitive visual ways such as charts, reports, and so on. Provide intelligent prompts and automatic completion functions, record user query history and feedback information, and automatically adjust the interactive interface and recommended content. Collect user feedback through the user interaction interface, and optimize and adjust the semantic interaction model and data analysis algorithm.
[0033] So far, after the above steps, combined with technologies such as machine learning and big models, the user's intelligent interactive experience in data analysis has been optimized, and the intelligence level of data analysis has been greatly improved. Through advanced natural language processing technology and semantic understanding engine, it can accurately understand the semantic meaning of data and provide users with more intelligent and personalized data analysis services, which greatly enhances the depth and breadth of data analysis and helps users to dig out more valuable insights from massive data.
[0034] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent data analysis method based on semantic interaction, characterized in that: The method is as follows: Multimodal data collection and preprocessing: Efficiently collect data from multi-source heterogeneous data sources, analyze the number of multi-source heterogeneous data sources to be collected, the data processing complexity coefficient, the unit core processing capacity, and the configuration parameters of a single physical machine, calculate the number of nodes in the data collection cluster, and perform distributed cleaning on the collected multimodal data by removing duplicate data, supplementing missing values, and replacing erroneous data. The cleaned data is annotated and classified to build a corpus; the annotation content covers the semantic category and contextual information of the data; Use the corpus to train the pre-trained language model and generate a semantic interaction model with intelligent semantic parsing capabilities; Analyze and process user input content based on the trained semantic interaction model, and finally output a data result set that meets user requirements; Use low-code technology to build a user interaction interface, collect user feedback through the user interaction interface, and optimize and adjust the semantic interaction model.
2. The intelligent data analysis method based on semantic interaction according to claim 1 is characterized in that: The number of nodes in the data collection cluster is calculated as follows: Calculate the total number of virtual cores required, vCores, using the following formula: vCores = (total data volume × processing complexity coefficient) / unit core processing capacity; Calculate the total required memory book vMemorys, the formula is as follows: vMemorys = total data volume × memory consumption coefficient; The cluster size is calculated based on the total number of virtual cores vCores and the total memory vMemorys required, and then the number of nodes nums in the data collection cluster is calculated. The formula is as follows: nums = MAX (total required number of virtual cores vCore / number of virtual cores on a single node, total required memory vMemorys / memory on a single node).
3. The intelligent data analysis method based on semantic interaction according to claim 1 is characterized in that: The pre-trained language model is trained using the corpus to generate a semantic interaction model with semantic intelligent parsing capabilities as follows: Focusing on the scenario requirements of data analysis, fine-tune the vocabulary embedding layer of the pre-trained language model according to the specific vocabulary of professional terms and industry terms in the data analysis scenario; The pre-trained language model is further trained using the annotated corpus, so that the pre-trained language model can accurately understand natural language instructions in a specific field; the pre-trained language model is then trained using annotated data, so that the pre-trained language model can automatically learn the semantic features and contextual relationships of the language, and obtain a semantic interaction model with semantic intelligent parsing capabilities; The trained semantic interaction model is evaluated through the cross-validation method, and the semantic interaction model is optimized and adjusted according to the evaluation results. Among them, the optimization indicators focus on the accuracy, recall rate and F1 score of semantic understanding to ensure the performance of the semantic interaction model in the real application scenarios of intelligent analysis.
4. The intelligent data analysis method based on semantic interaction according to claim 1 is characterized in that: The user input content is analyzed and processed based on the trained semantic interaction model, and the final output data result set that meets the user's requirements is as follows: Based on the trained semantic interaction model, the natural language query input by the user is semantically parsed and converted into a semantic vector; Vector retrieval technology is used to match the parsed semantic vector with the semantic vector in the data set to quickly retrieve data related to the user's query semantics; Sort the search results according to the semantic matching degree, and optimize the search results based on the user's historical query records and preferences to improve the relevance and accuracy of the search results; Using the K-Means algorithm, the retrieved data is clustered and divided into multiple clusters according to business types; The Apriori algorithm is used to identify frequent co-occurrence relationships between data items. By mining association rules, potential associations and trends between data as well as patterns, trends and anomalies in the data are automatically identified, and finally a data result set that meets the user's ideas is output.
5. The intelligent data analysis method based on semantic interaction according to any one of claims 1 to 4, characterized in that: The user interface has the following functions: ① Support users to interact through natural language input, voice input, drag and drop operation or visual charts; ② The intelligent analysis results of user semantics are presented to users in an intuitive and visual way in the form of charts and reports; ③ Provide intelligent prompts and automatic completion functions, record user query history and feedback information, and automatically adjust the interactive interface and recommended content.
6. An intelligent data analysis system based on semantic interaction, characterized in that: The system includes: The data acquisition processor is used to efficiently collect data from multiple heterogeneous data sources, dynamically adjust the cluster size to cope with different data volumes through a load-based dynamic expansion algorithm, and perform data cleaning to remove noise and duplicate data from the collected data, optimize data quality, annotate and classify the cleaned data, and build a corpus; The semantic interaction model trainer is used to train the pre-trained language model using the corpus, optimize the performance of the pre-trained language model through the multi-head attention mechanism and regularization technology, obtain the semantic interaction model, evaluate the semantic interaction model through the cross-validation method, and adjust the parameters in real time to improve the accuracy, recall rate and F1 score, so as to ensure that the semantic interaction model can accurately understand natural language instructions in practical applications; The semantic parser is used to deeply parse the user input content based on the trained semantic interaction model, extract the query intent, entity and relationship semantic information, and then convert the natural language into executable semantic instructions. The context-aware technology is used to ensure that the parsing results accurately reflect the user's true intent and provide precise semantic guidance for data retrieval and analysis. The visual interaction layer is used to use low-code technology to drag components, configure properties, and optimize layout and style through a visual designer to quickly build a simple and easy-to-use visual interface, provide intuitive and easy-to-use visual operations, and present analysis results in the form of intuitive charts and reports, allowing users to explore data in depth through filtering and drilling interactive operations, making data analysis more convenient and efficient.
7. The intelligent data analysis system based on semantic interaction according to claim 6 is characterized in that: The data acquisition processor includes: The data collection engine is used to efficiently collect data from multi-source heterogeneous data sources, analyze the number of multi-source heterogeneous data sources to be collected, the data processing complexity coefficient, the unit core processing capacity, and the configuration parameters of a single physical machine, and calculate the number of nodes in the data collection cluster; The data cleaning engine is used to perform distributed cleaning on the collected multimodal data by removing duplicate data, supplementing missing values, and replacing erroneous data, annotating and classifying the cleaned data, and building a corpus; the annotation content covers the semantic category and contextual information of the data.
8. The intelligent data analysis system based on semantic interaction according to claim 7 is characterized in that: The number of nodes in the data collection cluster is calculated as follows: Calculate the total number of virtual cores required, vCores, using the following formula: vCores = (total data volume × processing complexity coefficient) / unit core processing capacity; Calculate the total required memory book vMemorys, the formula is as follows: vMemorys = total data volume × memory consumption coefficient; The cluster size is calculated based on the total number of virtual cores vCores and the total memory vMemorys required, and then the number of nodes nums in the data collection cluster is calculated. The formula is as follows: nums = MAX (total required number of virtual cores vCore / number of virtual cores on a single node, total required memory vMemorys / memory on a single node).
9. The intelligent data analysis system based on semantic interaction according to claim 6, characterized in that: The semantic interaction model trainer includes: The fine-tuning module is used to fine-tune the vocabulary embedding layer of the pre-trained language model based on the specific vocabulary of professional terms and industry terms in the data analysis scenario around the scenario requirements of data analysis; The speech interaction model generation module is used to further train the pre-trained language model using the annotated corpus, so that the pre-trained language model can accurately understand the natural language instructions in a specific field; the pre-trained language model is then trained by annotated data, so that the pre-trained language model can automatically learn the semantic features and contextual relationships of the language, and obtain a semantic interaction model with semantic intelligent parsing capabilities; The model evaluation module is used to evaluate the trained semantic interaction model through the cross-validation method, and optimize and adjust the semantic interaction model according to the evaluation results; among them, the optimization indicators focus on the accuracy, recall rate and F1 score of semantic understanding to ensure the performance of the semantic interaction model in the real application scenarios of intelligent analysis.
10. The intelligent data analysis system based on semantic interaction according to claim 6, characterized in that: The semantic parser includes: The conversion module is used to perform semantic analysis on the natural language query input by the user based on the trained semantic interaction model and convert it into a semantic vector; The matching module is used to match the parsed semantic vector with the semantic vector in the data set using vector retrieval technology to quickly retrieve data related to the user's query semantics; The sorting module is used to sort the search results according to the semantic matching degree, and optimize the search results based on the user's historical query records and preferences to improve the relevance and accuracy of the search results; The clustering analysis module is used to perform cluster analysis on the retrieved data using the K-Means algorithm and divide the data into multiple clusters according to business types; The output module is used to use the Apriori algorithm to identify frequent co-occurrence relationships between data items. By mining association rules, it automatically identifies potential associations and trends between data as well as patterns, trends and anomalies in the data, and finally outputs a data result set that meets the user's ideas.
Citation Information
Patent Citations
Cluster resource evaluation method and device
CN115757059A
Intelligent question number large-screen display system and method based on multi-modal large model
CN119088820A
Large language model training-oriented cluster monitoring method and related device
CN119271505A
Personalized recommendation system based on semantic analysis
CN119579287A
Customer service processing method, device and equipment based on big data, storage medium and product
CN119669405A
Cited By
Enhanced application method and device of semantic analysis model, equipment and storage medium
CN120832889A
Language control interaction method and system based on AR and VR
CN120913568A
Questionnaire missing data filling method, system and equipment and medium
CN121457626A
A questionnaire missing data filling method, system, device and medium
CN121457626B