Data analysis and visualization method and system based on artificial intelligence

Through the knowledge graph and multiple rounds of cross-attention mechanisms, multi-modal data is processed, and the problem of insufficient modal correlation and feature fusion is solved, achieving efficient and comprehensive data fusion and entity alignment.

CN120296672APending Publication Date: 2025-07-11HANGZHOU CHUANGXIANG YUEDONG ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510460347.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing multimodal data processing technology fails to fully consider the correlation between modals and visual, auditory and other characteristics, resulting in large consumption of computing resources and limited fusion effect.

Method used

The knowledge graph is used to materialize multimodal data, use L2 norm normalization and multiple rounds of cross-attention mechanisms, combine the shared matrix and Sigmoid activation function, and enhance the embedding matrix through cross-modal correlation modeling and complementarity to form a complementary embedding matrix.

Benefits of technology

It improves the entity alignment accuracy and computing efficiency of modal data, reduces the demand for computing resources, and comprehensively integrates multimodal data features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296672A_ABST
    Figure CN120296672A_ABST
Patent Text Reader

Abstract

The invention provides a data analysis and visualization method based on artificial intelligence. The method comprises the following steps: acquiring multi-modal data; carrying out materialization on the modal data by using a knowledge graph; carrying out normalization by using an L2 norm; a multi-round cross attention mechanism is used, in each round, a certain mode serves as a query set, other modes serve as key value sets in sequence, and parameterized representation is conducted on the key value sets through a shared matrix; then, a Sigmoid activation function is used for normalization to obtain cross attention weight, and finally attention distribution is weighted and summed to obtain modal output; after multiple rounds of cross attention mechanism processing, each modal already considers complementarity with other modals, and the complementarity enhanced embedding vectors are spliced to form a complementarity embedding matrix. According to the technical scheme, the mode complementarity and correlation are comprehensively considered, the calculation efficiency is improved, the data features are fused in multiple dimensions, and the data fusion method is more comprehensive and deep.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data analysis, and particularly to a data analysis and visualization method and system based on artificial intelligence. Background Art

[0002] Multimodal data processing technology refers to the simultaneous processing of data from multiple sources, such as visual data, auditory data, text data, etc. These data can provide rich information about user behavior and environmental conditions, thereby improving the automation level and user experience of the data lighting management system. At the same time, the lighting control system is a very important part of the data lighting management system. It can automatically adjust the brightness and color of the lights according to the needs of users and environmental changes, thereby creating a comfortable lighting environment.

[0003] Solutions of the prior art: Existing multimodal data processing technologies usually use knowledge graphs to materialize modal data entities, and then use multi-round cross-attention mechanisms and cross-modal correlation modeling to fuse data of different modalities. A knowledge graph is a graph-structured data representation method that can structure entities and their relationships, making it convenient for machine learning models to understand and process. The multi-round cross-attention mechanism can make data of different modalities complement each other, thereby improving the effect of data fusion. Cross-modal correlation modeling can capture the correlation relationships between different modalities, thereby more comprehensively understanding and utilizing multimodal data.

[0004] Problems of the prior art: However, existing multimodal data processing technologies still have some problems in practical applications. First, existing technologies usually only consider the complementarity between modalities, without fully considering the correlation between modalities. Although the multi-round cross-attention mechanism can capture some complementarity between modalities, it cannot fully capture the correlation between modalities. Second, existing technologies usually require a large amount of computing resources when processing multimodal data, which is a great challenge for some devices with limited resources. Finally, existing technologies usually only consider the semantic information of data when fusing multimodal data, without fully considering other features of data, such as visual features and auditory features. These limitations have greatly restricted the practical application of existing multimodal data processing technologies.

[0005] The invention patent with the Chinese patent application number "2024111749525" and the patent name "An Automotive Lighting Management Method, Device, Storage Medium, and Equipment" proposes to judge the environment where the data is located through multi-modal perception data, which can effectively respond to changes in road conditions and improve driving safety. Additionally, a large model is used to perform in-depth semantic analysis on the perception data, enabling the generation of more intelligent lighting control strategies. However, due to the large amount of data processed, the response time from obtaining data information to controlling automotive lighting is relatively long, and it cannot quickly detect oncoming vehicles and quickly turn off the corresponding headlight modules. Summary of the Invention

[0006] The present disclosure provides a method and system for data analysis and visualization based on artificial intelligence to solve at least one technical problem existing in the prior art.

[0007] According to the first aspect of the present disclosure, a method for data analysis and visualization based on artificial intelligence is provided, which acquires multi-modal data of target data; entity-izes the modal data using a knowledge graph; performs normalization using the L2 norm;

[0008] Uses a multi-round cross-attention mechanism. In each round, first, a certain modality is used as the query set, and other modalities are used as the key-value sets in sequence, and a shared matrix is used to parameterize their representations;

[0009] Then, the cross-attention weights are normalized using the Sigmoid activation function, and finally, the attention distributions are weighted and summed to obtain the modal output;

[0010] After being processed by the multi-round cross-attention mechanism, each modality has considered the complementarity with other modalities. These embedding vectors enhanced by complementarity are concatenated to form a complementary embedding matrix.

[0011] In an implementable manner, entity-izing the modal data using a knowledge graph is specifically represented as each sub-modality M∈{r, a, v, n, g} of the entity ei, where r represents the relationship, a represents the attribute, v represents the vision, n represents the numerical value, and g represents the graph structure; performing normalization using the L2 norm. The purpose of L2 norm normalization is to scale the feature vector to unit length, and the formula is as follows:

[0012] Where is the feature vector of modality m of entity ei, and d is the dimension of the feature vector;

[0013] Using a multi-round cross-attention mechanism, in each round, first take a certain modality as the query set, and the other modalities are used as the key-value sets in turn. Then, use the shared matrices Wq, Wk, and Wv to parameterize them to obtain Q, K, and V. Next, use the Sigmoid activation function to normalize and obtain the cross-attention weights. Finally, perform a weighted sum of the attention distributions to obtain the modality output, which is expressed by the following formula:

[0014] Q = mW q

[0015] K = V = pW k

[0016]

[0017] where dk represents the dimension of the Q matrix, KT represents the transpose of the K matrix, σ represents the Sigmoid activation function, m, p ∈ M represent different modalities, represents the m-modal embedding of entity ei. The Wq matrix is used to map the features of the input modality into the query vector (Q). The vector query represents the target to be searched for. The Wk matrix is used to map the features of other modalities into the key vector (K). The key vector is used to compare with the query vector to determine the correlation between different modality features. The Wv matrix is used to map the features of other modalities into the value vector (V). The value vector contains the selected feature information and will be weighted and summed according to the matching degree between the query and the key. The functions of the above formulas are to map the features of modality m into the query vector Q through Wq, and map the features of modality p into the key vector K and value vector V through Wk and Wv. Here, m and p represent different modalities. For example, m can be the visual modality, and p can be the text modality;

[0018] After being processed by the multi-round cross-attention mechanism, each modality m has considered the complementarity with other modalities. Concatenate these embedding vectors enhanced by complementarity to form a complementary embedding matrix, which is expressed by the following formula:

[0019]

[0020] where, represents the concatenation operation, M is the set of all modalities, represents the m-modal embedding of entity ei;

[0021] After attention weighting, the embeddings of different modalities have a certain degree of complementarity. Each modality considers the contributions of the features of other modalities and incorporates them into its own embedding. Cross-modal correlation modeling is used to capture the connections and shared information between different modalities. When modeling modality complementarity, a correlation matrix S ∈ R^{N_m×N_m}, where N_m represents the number of modalities, is introduced to adjust the modality embeddings, so that each entity comprehensively considers the correlation relationship between modalities while taking into account the complementary relationship between modalities. After introducing cross-modal correlation modeling, the above formula is further updated as follows:

[0022]

[0023] where, · represents the dot product of vectors, represents the m-modal embedding of entity e_i, and S mp is the correlation score between modality m and modality p, is the embedding vector of entity e_i on modality p.

[0024] In one implementable manner, the multi-modal data at least includes text data, image data, time-series data, graph structure data, and time-series data.

[0025] In one implementable manner, the multi-modal data further includes

[0026] user behavior data: user operation path, interaction frequency;

[0027] environmental data: system running environment parameters, external API structure status;

[0028] business rule data: preset policies, compliance constraints;

[0029] anomaly detection data: outlier markers, risk scores.

[0030] In one implementable manner, when the following scenarios are detected, an adaptive analysis strategy is triggered:

[0031] The number of outliers in the data stream exceeds the threshold;

[0032] The periodic fluctuation of time-series data fails;

[0033] There are unlabeled features in the image data;

[0034] The business rules conflict with the real-time data.

[0035] In one implementable manner, when the data flow analysis shows that the complexity of the external branch path exceeds the preset value, the sub-module isolation or feature dimensionality reduction is started.

[0036] In one implementable manner, when the external environmental data exceeds the safe range, the data compression or cache optimization strategy is enabled.

[0037] In one implementable manner, when the user behavior analysis module detects the following actions, it generates real-time alerts: high-frequency abnormal operations; unauthorized data access requests; behavior patterns deviating from historical baselines.

[0038] In one implementable manner, the following scenarios trigger visual highlighting prompts:

[0039] Sudden changes in the trends of key metrics;

[0040] Conflicts in multimodal data consistency;

[0041] The confidence level of model prediction is lower than the threshold;

[0042] The system resource occupancy rate exceeds the limit.

[0043] According to the second aspect of the present disclosure, there is provided an artificial intelligence-based data analysis and visualization system, including an acquisition module: first acquiring multimodal data of a target object; entityifying the modal data using a knowledge graph; and normalizing using the L2 norm;

[0044] An output module: using a multi-round cross-attention mechanism, in each round, first taking a certain modality as the query set, and sequentially taking other modalities as the key-value sets, and parametrically representing them using a shared matrix;

[0045] An activation module: then normalizing using the Sigmoid activation function to obtain cross-attention weights, and finally weighted-summing the attention distributions to obtain the modal output;

[0046] A complementary module: after being processed by the multi-round cross-attention mechanism, each modality has considered the complementarity with other modalities, and these embedding vectors enhanced by complementarity are concatenated to form a complementary embedding matrix.

[0047] According to the third aspect of the present disclosure, there is provided an electronic device, including:

[0048] At least one processor; and a memory communicatively connected to the at least one processor; wherein,

[0049] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of the present disclosure.

[0050] According to the fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the method of the present disclosure.

[0051] Compared with the prior art, the artificial intelligence-based data analysis and visualization method of the present disclosure has the following beneficial effects:

[0052] Compared with the existing technology, the beneficial effects of the present technical solution are as follows:

[0053] 1. Comprehensive consideration of modality complementarity and correlation: In the process of multi-modal data processing, the present invention not only considers the complementarity between modalities, but also fully considers the correlation between modalities. By introducing a correlation matrix, the present invention can more comprehensively understand and utilize multi-modal data, thereby improving the accuracy of entity alignment. This is in sharp contrast to the existing technology that only considers modality complementarity, and the fusion method of the present invention is more comprehensive and accurate.

[0054] 2. Improvement of computational efficiency: The present invention adopts a shared matrix to share parameters between different modalities, which can not only improve the efficiency of the model, but also reduce the number of parameters of the model. This is a great advantage for some devices with limited resources. In contrast, the existing technology usually requires a large amount of computational resources when processing multi-modal data, and the present invention is more efficient.

[0055] 3. Multi-dimensional fusion of data features: When fusing multi-modal data, the present invention not only considers the semantic information of the data, but also fully considers other features of the data, such as visual features and auditory features. This enables the present invention to more comprehensively understand and utilize multi-modal data, thereby improving the accuracy of entity alignment. In contrast, the existing technology usually only considers the semantic information of the data when fusing multi-modal data, and the data fusion method of the present invention is more comprehensive and in-depth.

[0056] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings

[0057] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become easily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, wherein:

[0058] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.

[0059] Figure 1 Shows a schematic diagram of the implementation process of a method for data analysis and visualization based on artificial intelligence according to an embodiment of the present disclosure;

[0060] Figure 2 Shows a schematic diagram of the structure of a system for data analysis and visualization based on artificial intelligence according to an embodiment of the present disclosure; Detailed Description

[0061] To make the objectives, features, and advantages of the present disclosure more apparent and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present disclosure.

[0062] Reference Figure 1 , a flowchart of a data analysis and visualization method based on artificial intelligence, which obtains multimodal data of target data; entityifies the modal data using a knowledge graph; normalizes it using the L2 norm; uses a multi-round cross-attention mechanism, in each round, first takes a certain modality as the query set, and the other modalities are taken as the key-value sets in turn, and parametrically represents them using a shared matrix; then normalizes using the Sigmoid activation function to obtain the cross-attention weights, and finally sums the attention distributions weighted to obtain the modal output; after being processed by the multi-round cross-attention mechanism, each modality has considered the complementarity with other modalities, and these embedding vectors enhanced by complementarity are concatenated to form a complementary embedding matrix.

[0063] A knowledge graph is a structured semantic knowledge base used to describe concepts in the physical world and their interrelationships in symbolic form. Its basic unit of composition is the "entity-relationship-entity" triple, as well as entity and its related attribute-value pairs. Entities are interconnected through relationships to form a networked knowledge structure.

[0064] Entityifying the modal data using a knowledge graph is specifically represented as each sub-modality M∈{r,a,v,n,g} of entity ei, where r represents relationship, a represents attribute, v represents vision, n represents numerical value, and g represents graph structure; normalizing using the L2 norm means dividing a vector by its L2 norm so that the length (L2 norm) of the vector is equal to 1. This operation is also called "unitization" or "normalization". The purpose of L2 norm normalization is to scale the feature vector to unit length, and the formula is as follows:

[0065] where is the feature vector of modality m of entity ei, and d is the dimension of the feature vector. The normalized vector can avoid the scale difference between different modality features, making the feature vectors of different modalities comparable when calculating similarity. In multimodal learning, normalization helps the model better capture the complementarity and correlation between different modalities. Through this formula, we can normalize the feature vectors of different modalities to unit length to prepare for subsequent feature fusion and entity alignment.

[0066] Using a multi-round cross-attention mechanism, in each round, first take a certain modality as the query set, and the other modalities are used as the key-value sets in turn, and parameterize them using the shared matrices Wq, Wk, and Wv to obtain Q, K, and V. Then, use the Sigmoid activation function to normalize to obtain the cross-attention weights. Finally, perform weighted summation on the attention distribution to obtain the modality output, which is expressed by the following formula:

[0067] Q = mW q

[0068] K = V = pW k

[0069]

[0070] where dk represents the dimension of the Q matrix, KT represents the device of the K matrix, σ represents the Sigmoid activation function, m, p ∈ M represent different modalities, represents the m-modal embedding of entity ei. The Wq matrix is used to map the features of the input modality into the query vector (Q), and the vector query represents the target to be searched for. The Wk matrix is used to map the features of other modalities into the key vector (K), and the key vector is used to compare with the query vector to determine the correlation between different modality features. The Wv matrix is used to map the features of other modalities into the value vector (V), and the value vector contains the selected feature information, which will be weighted and summed according to the matching degree of the query and the key. The functions of the above formulas are to map the features of modality m into the query vector Q through Wq, and map the features of modality p into the key vector K and value vector V through Wk and Wv. Here, m and p represent different modalities. For example, m can be the visual modality, and p can be the text modality. In this way, the model can learn the complementary information between different modalities. The advantage of using a shared matrix is that it allows the model to share parameters between different modalities, which can improve the efficiency and generalization ability of the model while reducing the number of model parameters. In addition, the shared matrix can also help the model learn the common features across modalities, which is very useful for understanding and fusing multi-modal data.

[0071] After being processed by the multi-round cross-attention mechanism, each modality m has considered the complementarity with other modalities. Concatenate these embedding vectors enhanced by complementarity to form a complementary embedding matrix, which is expressed by the following formula:

[0072]

[0073] where, represents the concatenation operation, M is the set of all modalities, Denote the m-modal embedding of entity ei. After attention weighting, the embeddings of different modalities have a certain degree of complementarity. Each modality considers the contributions of the features of other modalities and adds them to its own embedding. Cross-modal correlation modeling is used to capture the connections and shared information between different modalities. When modeling modality complementarity, a correlation matrix S∈RNm×Nm is introduced, where Nm represents the number of modalities, to adjust the modality embeddings, so that each entity comprehensively considers the correlation relationship between modalities while taking into account the complementary relationship between modalities. After introducing cross-modal correlation modeling, the above formula is further updated as follows:

[0074]

[0075] where · represents the dot product of vectors, denote the m-modal embedding of entity ei, and S mp is the correlation score between modality m and modality p, is the embedding vector of entity ei on modality p. In this way, the model not only considers the complementarity between different modalities. For example, the visual modality may provide the physical characteristics of an entity, while the text modality provides the semantic description of the entity. The correlation matrix helps the model understand how these modalities complement each other. It also considers the correlation between modalities, so as to more comprehensively understand and fuse multi-modal data. For example, in video content, visual information (video frames) and audio information (dialogue, background music) are closely related. The correlation matrix can help the model identify and utilize these relationships. This fusion method helps to improve the accuracy of entity alignment because it synthesizes rich information from different modalities.

[0076] The multi-modal data includes but is not limited to the following types:

[0077] Text data:

[0078] Application scenarios: user comment analysis, log information parsing, contract text extraction.

[0079] Example: In an e-commerce platform, through natural language processing (NLP), user comments are analyzed to extract sentiment tendencies (such as "poor product quality"), and associated with sales data to generate product improvement reports.

[0080] Image data:

[0081] Application scenarios: medical image diagnosis, industrial quality inspection, satellite image analysis.

[0082] Example: In the medical field, a convolutional neural network (CNN) is used to identify lung nodules in CT images, automatically mark the positions and calculate the malignancy probability.

[0083] Time series data:

[0084] Application scenarios: Stock trading flow monitoring, equipment operation status prediction.

[0085] Example: In a factory, the time-series data of equipment vibration sensors is analyzed through an LSTM model to predict motor failures 24 hours in advance and trigger maintenance work orders.

[0086] Graph-structured data:

[0087] Application scenarios: Social network analysis, knowledge graph construction.

[0088] Example: In financial anti-fraud, a user transaction relationship graph is constructed to identify abnormal subgraphs (such as multi-account circular transfers) and mark potential money laundering behaviors.

[0089] In one embodiment, the multimodal data further includes:

[0090] User behavior data:

[0091] Examples: User click stream, page dwell time, function usage frequency.

[0092] Example: In an online education platform, analyze students' video viewing behaviors (such as pausing, replaying), and combine with quiz scores to recommend personalized learning paths.

[0093] Environmental data:

[0094] Examples: Server CPU temperature, API response latency, third-party service availability.

[0095] Example: In a cloud computing platform, monitor the cold start time of AWS Lambda functions and dynamically adjust the container preheating strategy to reduce latency.

[0096] Business rule data:

[0097] Examples: Compliance constraints (such as GDPR), promotion activity strategies.

[0098] Example: In a retail system, real-time verify whether the order price complies with the "full reduction" rule. If there is a conflict, trigger an alarm and freeze the abnormal order.

[0099] Anomaly detection data:

[0100] Examples: Outlier marking (such as transaction amount exceeding the threshold), risk score (0 - 100 points).

[0101] Example: In payment risk control, detect operations where the single transfer amount exceeds 3 times the user's historical average through the Isolation Forest algorithm, automatically intercept and notify manual review.

[0102] In one embodiment, an adaptive analysis strategy is triggered when the following scenarios are detected:

[0103] Outliers in the data stream exceed the threshold:

[0104] Example: In power grid monitoring, if the current value in a certain area exceeds the safety threshold (such as 1000 A) for 5 consecutive minutes, automatically switch to the backup line and push an alarm to the operation and maintenance personnel.

[0105] Periodic fluctuations in time-series data fail:

[0106] Example: In retail sales forecasting, if the sales volume during holidays does not increase according to the historical cycle (such as a 20% month-on-month decrease in sales volume during "Double Eleven"), trigger model retraining and adjust the inventory strategy.

[0107] Unlabeled features exist in the image data:

[0108] Example: In autonomous driving, if the camera captures an unlabeled temporary traffic sign (such as a construction area sign), suspend the autonomous driving function and prompt the driver to take over.

[0109] Business rules conflict with real-time data:

[0110] Example: In air traffic control, if the flight plan conflicts with the real-time weather data (such as a typhoon making the route unavailable), dynamically generate an alternate landing plan and synchronize it to the passenger APP.

[0111] In one embodiment, when the data flow analysis shows that the complexity of the external branch path exceeds the preset value (such as the number of node connections > 1000):

[0112] Example: In social network analysis, if the user relationship chain involves more than 1000 nodes (such as a star fan group), start subgraph isolation, only retain the core nodes (such as the fan group leader and active users), and use PCA dimensionality reduction technology to compress the feature dimensions to improve the calculation efficiency.

[0113] In one embodiment, when the external environmental data (such as network latency > 200 ms or storage load > 80%) exceeds the safe range:

[0114] Example: In a video live streaming platform, if the CDN node latency suddenly increases, enable H.265 encoding to compress the video stream and cache the non-key frames to the edge server to ensure smooth viewing for users.

[0115] In one embodiment, Right 8: Example of real-time warning of user behavior

[0116] Content:

[0117] The user behavior analysis module generates a real-time alarm when detecting the following actions:

[0118] High-frequency abnormal operations:

[0119] Example: In a banking system, if a user initiates 50 password reset requests within 1 hour, the risk control locks the account and sends a text message for verification.

[0120] Unauthorized data access request:

[0121] Example: In a corporate intranet, when it is detected that an employee attempts to access an unauthorized database (such as a financial statement), the IP is immediately blocked and an audit log is recorded.

[0122] Behavior pattern deviates from the historical baseline:

[0123] Example: In a smart home, if the user usually turns off the lights at 10 pm and suddenly turns on the lights remotely at 3 am one day, a security confirmation notice is pushed to the mobile phone.

[0124] In one embodiment, the following scenarios trigger a visual highlighting prompt:

[0125] Sudden change in the trend of key indicators:

[0126] Example: In a stock trading dashboard, if the rise and fall of a certain stock exceeds 10% within 5 minutes, the K-line chart is automatically marked in red and a news sentiment analysis pops up.

[0127] Multi-modal data consistency conflict:

[0128] Example: In a smart city monitoring system, if the traffic camera shows congestion at an intersection, but the GPS data indicates normal vehicle flow rate, the area on the map is highlighted and manual verification is prompted.

[0129] Model prediction confidence is lower than the threshold:

[0130] Example: In medical AI-assisted diagnosis, if the confidence in lung cancer prediction is <60%, a yellow warning bar is displayed on the report interface and expert consultation is recommended.

[0131] System resource occupancy rate exceeds the limit:

[0132] Example: In cloud server monitoring, if the CPU usage rate >95%, the corresponding area of the dashboard blinks red and the instance is automatically scaled out.

[0133] Reference Figure 2 , an artificial intelligence-based data analysis and visualization system, including a data acquisition module, a controller module, and an LED matrix light group. The data acquisition module includes:

[0134] Acquisition module 10: First, acquire multi-modal data of the target object; materialize the modal data using a knowledge graph; and perform normalization using the L2 norm.

[0135] Output module 20: Using a multi-round cross-attention mechanism, in each round, first take a certain modality as the query set, and the other modalities as the key-value sets in turn, and use a shared matrix to parameterize and represent them;

[0136] Activation module 30: Then use the Sigmoid activation function to normalize to obtain the cross-attention weights, and finally sum the attention distributions weighted to obtain the modality output;

[0137] Complementary module 40: After being processed by the multi-round cross-attention mechanism, each modality has considered the complementarity with other modalities. Concatenate these embedding vectors enhanced by complementarity to form a complementary embedding matrix.

[0138] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0139] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0140] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0141] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0142] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0143] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0144] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0145] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present disclosure can be achieved, and no limitation is imposed herein.

[0146] As described above, this is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. An artificial intelligence-based data analysis and visualization method, characterized in that first, multi-modal data of the target object is obtained; and the modal data is materialized using a knowledge graph; normalization is performed using the L2 norm; a multi-round cross-attention mechanism is used. In each round, first, a certain modality is used as the query set, and the other modalities are used as the key-value sets in turn, and a shared matrix is used to parameterize and represent them; then, the cross-attention weights are normalized using the Sigmoid activation function, and finally, the attention distributions are weighted and summed to obtain the modal output; after being processed by the multi-round cross-attention mechanism, each modality has considered the complementarity with other modalities, and these embedding vectors enhanced by complementarity are concatenated to form a complementary embedding matrix.

2. The data analysis and visualization method based on artificial intelligence according to claim 1, wherein Among them, using the knowledge graph to materialize the modal data is specifically represented as each sub-modality M ∈ {r, a, v, n, g} of the entity ei, where r represents the relationship, a represents the attribute, v represents the vision, n represents the numerical value, and g represents the graph structure; normalization is performed using the L2 norm, and the formula is as follows: , Among them, is the eigenvector of the modality m of the entity ei, and d is the dimension of the eigenvector; a multi-round cross-attention mechanism is used. In each round, first, a certain modality is used as the query set, and the other modalities are used as the key-value sets in turn, and the shared matrices Wq, Wk, and Wv are used to parameterize and represent them to obtain Q, K, and V. Then, the cross-attention weights are normalized using the Sigmoid activation function, and finally, the attention distributions are weighted and summed to obtain the modal output. The formula is expressed as follows: , , , where, dk represents the dimension of the Q matrix, and KT represents the device of the K matrix. denotes the Sigmoid activation function, m, p ∈ M represent different modalities. represents the m-modal embedding of entity ei. The Wq matrix is used to map the features of the input modality into a query vector (Q). The vector query represents the target to be searched for. The Wk matrix is used to map the features of other modalities into key vectors (K). The key vectors are used to compare with the query vector to determine the relevance between the features of different modalities. The Wv matrix is used to map the features of other modalities into value vectors (V). The value vectors contain the selected feature information and will be weighted and summed according to the matching degree between the query and the key. The functions of the above formulas are to map the features of modality m into the query vector Q through Wq, and map the features of modality p into the key vector K and value vector V through Wk and Wv respectively. Here, m and p represent different modalities. For example, m can be the visual modality, while p can be the text modality. after being processed by the multi-round cross-attention mechanism, each modality m has considered the complementarity with other modalities, and these embedding vectors enhanced by complementarity are concatenated to form a complementary embedding matrix. The formula is expressed as follows: , Among them, represents the splicing operation, M is the set of all modalities, represents the m-modal embedding of the entity ei; after attention weighting, the embeddings of different modalities have a certain degree of complementarity. Each modality has considered the contributions of the features of other modalities and added them to its own embedding. Cross-modal correlation modeling is used to capture the connections and shared information between different modalities. When modeling the modality complementarity, a correlation matrix S ∈ RNm × Nm is introduced, where Nm represents the number of modalities, to adjust the modality embeddings, so that each entity comprehensively considers the association relationship between modalities while considering the complementary relationship between modalities. After introducing cross-modal correlation modeling, the above formula is further updated as follows: , , Among them, represents the dot product of vectors, represents the m-modal embedding of entity ei, is the correlation score between modality m and modality p, is the embedding vector of entity ei on modality p.

3. The method for artificial intelligence-based data analysis and visualization according to claim 1 or 2, characterized in that The multi-modal data at least includes text data, image data, time-series data, graph structure data, and time-series data.

4. The data analysis and visualization method based on artificial intelligence according to claim 3, wherein The multi-modal data also includes user behavior data: user operation path, interaction frequency; environment data: system running environment parameters, external API structure status; business rule data: preset policies, compliance constraints; anomaly detection data: outlier markers, risk scores.

5. The data analysis and visualization method based on artificial intelligence according to claim 4, wherein When the following scenarios are detected, an adaptive analysis strategy is triggered: the number of outliers in the data stream exceeds the threshold; the periodic fluctuation of the time-series data fails; there are unlabeled features in the image data; the business rules conflict with the real-time data.

6. The method for data analysis and visualization based on artificial intelligence according to claim 5, wherein When the data flow analysis shows that the complexity of the external branch path exceeds the preset value, start sub-module isolation or feature dimensionality reduction.

7. The data analysis and visualization method based on artificial intelligence according to claim 6, characterized in that, When the external environment data exceeds the safe range, enable data compression or cache optimization strategies.

8. The method for data analysis and visualization based on artificial intelligence according to claim 7, wherein The user behavior analysis module generates real-time alerts when it detects the following actions: high-frequency abnormal operations; unauthorized data access requests; behavior patterns deviating from historical baselines.

9. The method for artificial intelligence-based data analysis and visualization according to any one of claims 5-8, characterized in that, The following scenarios trigger visual highlighting prompts: Sudden changes in the trends of key metrics; Multimodal data consistency conflicts; The confidence level of model predictions is lower than the threshold; The system resource occupancy rate exceeds the limit.

10. An artificial intelligence-based data analysis and visualization system, characterized in that, Including: Acquisition module: First, acquire the multimodal data of the target object; And materialize the modal data using a knowledge graph; perform normalization using the L2 norm; Output module: Use a multi-round cross-attention mechanism. In each round, first take a certain modality as the query set, and the other modalities are used as the key-value sets in turn, and use a shared matrix to parameterize and represent them; Activation module: Then use the Sigmoid activation function to normalize to obtain the cross-attention weights, and finally sum the attention distributions weighted to obtain the modal output; Complementary module: After being processed by the multi-round cross-attention mechanism, each modality has considered the complementarity with other modalities. Concatenate these embedding vectors enhanced by complementarity to form a complementary embedding matrix.