An AI Big Data Real-Time Processing and Analysis Method
Through feature extraction algorithms and meta-learning technology for different data types, combined with distributed computing and edge computing, the problem of inapplicability of feature extraction methods in the existing technology is solved, and efficient unified processing and analysis of AI big data is realized.
Patent Information
- Application Number
- CN202510071942.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
In the process of AI big data processing and analysis, feature extraction methods are not suitable for different data types, resulting in low processing efficiency and long time.
Feature extraction algorithms for different data types (text, image, audio, video, time series data), including TF-IDF, CNN, Mel frequency cepspectral coefficients and LSTM, are adopted, and different data types and distributions are automatically adapted through meta-learning technology, and unified processing and analysis is carried out by combining distributed computing frameworks and edge computing technology.
It realizes unified processing of different data types, improves data processing efficiency, shortens big data processing and analysis time, and improves the system's adaptability and intelligence level.
Smart Images

Figure CN119493823B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of AI big data processing, and specifically to a method for real-time processing and analysis of AI big data. Background Art
[0002] As an important product of the development of modern information technology, AI has brought a lot of convenience to people's lives. And realizing the evolution of business through AI big data services is an important assistance in the current social business process, especially in mining key data from a large amount of text and image data, so as to improve efficiency in people's work and life. A prior art method for real-time processing and analysis of AI big data with the publication number of CN 118245680 B, especially relates to the technical field of AI big data processing. This method includes using a data acquisition unit to acquire user demand data, and performing key feature extraction on the user demand data through a feature extraction unit; using a feature association unit to associate the user demand data after key feature extraction with an AI big database to extract associated data of the user demand data from the AI big database; using a feature recognition unit to perform feature recognition on the associated data and output the recognition result; using a data screening unit to screen the recognition result, and sending the screened associated data to the user terminal through a data analysis unit, and collecting user feedback data for analysis; using an adjustment unit to adjust the processing process based on the analysis result of the data analysis unit.
[0003] Although the foregoing technology solves the problem that the prior art cannot precisely control the process of processing and analyzing AI big data, feature extraction is one of the key steps in data processing in the foregoing technology. However, different data types require different feature extraction methods, and the foregoing technical solution does not mention specific feature extraction techniques, nor does it explain whether it can efficiently process large-scale data sets, which results in a longer time actually required for big data processing and analysis. Summary of the Invention
[0004] Technical Problems to be Solved
[0005] Aiming at the deficiencies of the prior art, the present invention provides a method for real-time processing and analysis of AI big data, which has the advantages of uniformly processing different data types to improve data processing efficiency and shortening the time for big data processing and analysis, and solves the above technical problems.
[0006] Technical Solutions
[0007] To achieve the above object, the present invention provides the following technical solutions: A method for real-time processing and analysis of AI big data, including the following steps:
[0008] Step 1. Data Acquisition and Preprocessing: Real-time data is acquired from multiple data sources: sensors, user behavior records, log files, and social media, supporting the collection of structured, semi-structured, and unstructured data. Data quality is improved through automated cleaning tools: outlier detection. Data formats from different sources: JSON, XML, text, and images are uniformly converted into a standard format, and large-scale data is compressed to reduce storage and transmission overhead;
[0009] Step 2. Feature Extraction and Intelligent Adaptation: For different data types: text, images, audio, video, and time-series data, the optimal feature extraction algorithms are selected, including:
[0010] Text data: Use the TF-IDF deep learning model to extract semantic features;
[0011] Image data: Use the convolutional neural network CNN to extract visual features;
[0012] Audio data: Use Mel-frequency cepstral coefficients to extract audio features;
[0013] Time-series data: Use LSTM to extract time-series features;
[0014] Adaptive Feature Extraction: Through meta-learning techniques, automatically adapt to different data types and distributions, and improve the generalization ability of feature extraction;
[0015] Feature Fusion: Fuse multi-modal features: text, image, and time-series features to generate a unified feature representation;
[0016] Step 3. Feature Association and Data Screening: Associate the extracted multi-modal features with the data in the AI large database, supporting cross-modal association analysis: text and image, audio and video associations. Use the distributed computing framework Spar to accelerate the association calculation of large-scale data. Based on the user demand data and the association analysis results, use the machine learning model: random forest, to intelligently screen the associated data;
[0017] Step 4. Data Analysis: Use the lightweight inference engine TensorRT to perform real-time analysis on the screened data, supporting multi-dimensional analysis of time, space, and behavior, generating visual analysis results. Use the user's real-time feedback data: click behavior, ratings, and comments, as input to optimize the analysis results. Based on user feedback and analysis results, dynamically adjust the parameters of the feature extraction algorithm and the screening strategy to improve the system's adaptive ability. Through online learning Online Learning technology, continuously update the feature extraction and analysis models to improve the system's intelligence level;
[0018] Step 5, System Architecture and Technical Support: Adopt the distributed computing framework Spark and GPU clusters to accelerate the processing and analysis of large-scale data, and use the distributed storage system HDFS and the caching mechanism Redis to support efficient data storage and access;
[0019] Step 6, Edge Computing Support: Conduct preliminary feature extraction and data screening on edge devices to reduce data transmission latency, and upload tasks that cannot be processed by edge devices to the cloud for further processing.
[0020] Preferably, in Step 1, outlier detection is performed based on the Z-score method, and the detection calculation is as follows:
[0021] ;
[0022] Where: is the data point, is the mean of the data, is the standard deviation of the data. If , then the data point is considered an outlier;
[0023] The calculation of the data compression rate in Step 1 is as follows:
[0024] ;
[0025] Where the original data size is the data size before compression, and the compressed data size is the data size after compression.
[0026] Preferably, the TF-IDF deep learning model in Step 2 is as follows:
[0027] ;
[0028] Where: is the term frequency, indicating the frequency of the word appearing in the document ;
[0029] The specific formula is:
[0030] Term frequency ;
[0031] Inverse document frequency ;
[0032] Where, is the total number of documents in the corpus.
[0033] Preferably, the convolutional neural network in Step 2 extracts visual features from images through convolutional operations, and the calculation is as follows:
[0034] ;
[0035] Wherein: is the pixel matrix of the input image, with a size of , is the height, is the width, is the number of channels; is the bias term;
[0036] In step 2, the Mel-frequency cepstral coefficients are used to extract the features of the audio data and are calculated as:
[0037] ;
[0038] Wherein: is the logarithmic energy of the th Mel band; is the index of the time frame.
[0039] Preferably, in step 2, the long short-term memory network (LSTM) is used to extract the features of the time series data and is calculated as:
[0040] Forget gate
[0041] ;
[0042] Wherein: is the output of the forget gate, a value between 0 and 1; is the weight matrix of the forget gate, used to learn the relationship between the input and the previous state; is the bias term of the forget gate, used to adjust the output of the weight matrix is the hidden state at time step , is the current input data at time step ;
[0043] Input gate
[0044] ;
[0045] Wherein: is the output of the input gate, a value between 0 and 1; is the weight matrix of the input gate, used to learn the relationship between the input and the previous state; is the bias term of the input gate, used to adjust the output of the weight matrix;
[0046] Cell state update
[0047] ;
[0048] ;
[0049] Wherein: is the candidate state, representing the potential value of the new information; is the weight matrix of the candidate state, used to learn the relationship between the input and the previous state; is the bias term of the candidate state, used to adjust the output of the weight matrix; represents the cell state information of the previous time step retained;
[0050] Output gate
[0051] ;
[0052] ;
[0053] Wherein: The activation function is the Sigmoid function, is the bias term, is the output of the output gate, which is a value between 0 and 1; is the weight matrix of the output gate, used to learn the relationship between the input and the previous state; is the bias term of the output gate, used to adjust the output of the weight matrix; is the cell state The activation value of, compresses the value of the cell state to the interval (-1, 1).
[0054] Preferably, the cross-modal association analysis in step three is used to find the association between text and image, audio and video. Given data of modality A and modality B, the cosine similarity is used to calculate the similarity or correlation between them. For two vectors and , the cosine similarity is defined as:
[0055] ;
[0056] Wherein: and are text data and image data respectively; is the vector and The cosine value of the included angle between; and are the vectors and The modulus of, that is, the length of the vector;
[0057] The Euclidean distance is used to measure the distance between feature vectors. The smaller the value, the more similar:
[0058] ;
[0059] Wherein: and is a vector and is the th element; is the square of the difference between two vectors in the th dimension. By squaring, the distance is ensured to be non - negative and the difference is amplified; is to sum up the squares of the differences in all dimensions to obtain a total squared difference; is the square root of the squared difference, that is, the Euclidean distance, which measures the straight - line distance between two vectors in Euclidean space. The smaller it is, the closer the two vectors are.
[0060] Preferably, the time - dimension analysis and calculation in step four are as follows: calculate the mean and variance of the data in the time dimension;
[0061] ;
[0062] where: is the mean value of the data within the time period ;
[0063] In the space - dimension analysis of step four, calculate the spatial distribution of the data, and use the K - means algorithm to cluster the spatial data:
[0064] ;
[0065] where: is the set of cluster centers; is the th cluster center.
[0066] Preferably, the acceleration calculation of the distributed computing framework Spark in step five is as follows:
[0067] Spark divides the data set into partitions:
[0068] ;
[0069] The calculation result of each partition is obtained through parallel computing:
[0070] ;
[0071] The final global result is the integration of the results of each partition:
[0072] .
[0073] Preferably, the fifth step further includes effectiveness calculation, calculating the processing time of each partition and speedup ratio :
[0074] ;
[0075] Wherein: is the single-node processing time, is the total time consumed for processing the th partition, is the number of nodes allocated to the th partition;
[0076] Preferably, the preliminary feature extraction calculation of the edge device is as follows: the edge device preliminarily processes the data, extracts key features, and reduces the data volume:
[0077] Let the original data be , where is the number of samples, is the number of features;
[0078] Feature extraction: Use the feature extraction algorithm to extract key features to obtain a new feature matrix
[0079] ;
[0080] Among them, is the feature extraction function;
[0081] Data transmission volume calculation: The original data transmission volume , where is the data size of each feature: bytes,
[0082] The data transmission volume after extraction ;
[0083] Data transmission volume reduction ratio: .
[0084] Compared with the prior art, the present invention provides an AI big data real-time processing and analysis method, which has the following beneficial effects:
[0085] 1. In the present invention, for different types of data, such as images, texts, sensor data, etc., the edge device first performs standardization processing to convert the data into a common format for subsequent unified processing. The image is converted into a specific resolution or color format, the text is converted into a unified encoding format, and the sensor data is converted into a specific data unit or range. Next, the edge device uses corresponding algorithms to extract key features from the standardized data. These features can represent the main information of the original data, thus reducing the amount of data to be processed. For image data, features such as edges, corners, or textures are extracted; for text data, features such as keywords or word frequencies are extracted. After feature extraction, the edge device filters the data according to preset rules or models, and only the data that meets specific conditions or contains important information is uploaded to the cloud for further processing, greatly reducing unnecessary data transmission and improving the overall processing efficiency, achieving the beneficial effect of unified processing of different data types and improving data processing efficiency.
[0086] 2. In the present invention, by performing standardization processing on different types of data on the edge device, the data is converted into a unified format for subsequent unified processing. This step can significantly reduce the complexity and redundancy of the data. Feature extraction is performed on the edge device to extract the key features of the data. These features can represent the main information of the original data, greatly reducing the amount of data to be transmitted. The feature data is filtered through preset rules or models, and only the data containing key information is uploaded to the cloud, further reducing the amount of data to be transmitted. Since the amount of data transmitted is greatly reduced, the occupancy of network bandwidth is also correspondingly reduced, thereby shortening the data transmission time. Reducing the amount of data transmission can also reduce network latency, enabling the data to reach the cloud faster. Preliminary data preprocessing, such as standardization, feature extraction, and filtering, is performed on the edge device to reduce the complexity of the data. Since the amount of data received by the cloud is greatly reduced, the computing burden on the cloud is also reduced, and the cloud can process the remaining data faster, improving the processing efficiency and achieving the beneficial effect of shortening the big data processing and analysis time. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 is a schematic flow diagram of the method of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0088] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0089] Please refer to Figure 1, an AI big data real-time processing and analysis method, including the following steps:
[0090] Step 1, data acquisition and preprocessing: Real-time data is acquired from multiple data sources: sensors, user behavior records, log files, and social media, supporting the collection of structured, semi-structured, and unstructured data. Through automated cleaning tools: outlier detection, data quality is improved. Different data formats from various sources: JSON, XML, text, and images are uniformly converted into a standard format, and large-scale data is compressed to reduce storage and transmission overhead;
[0091] Step 2, feature extraction and intelligent adaptation: For different data types: text, images, audio, video, and time series data, the optimal feature extraction algorithms are selected, including:
[0092] Text data: Use the TF-IDF deep learning model to extract semantic features;
[0093] Image data: Use the convolutional neural network CNN to extract visual features;
[0094] Audio data: Use Mel-frequency cepstral coefficients to extract audio features;
[0095] Time series data: Use LSTM to extract time series features;
[0096] Adaptive feature extraction: Through meta-learning techniques, automatically adapt to different data types and distributions, and improve the generalization ability of feature extraction;
[0097] Feature fusion: Integrate multi-modal features: text, image, and time series features to generate a unified feature representation;
[0098] Step 3, feature association and data screening: Associate the extracted multi-modal features with the data in the AI big database, supporting cross-modal association analysis: association between text and image, audio and video. Use the distributed computing framework Spar to accelerate the association calculation of large-scale data. Based on user demand data and association analysis results, use a machine learning model: random forest, to intelligently screen the associated data;
[0099] Step 4, data analysis: Use the lightweight inference engine TensorRT to perform real-time analysis on the screened data, supporting multi-dimensional analysis of time, space, and behavior, generating visual analysis results. Use the user's real-time feedback data: click behavior, rating, and comments, as input to optimize the analysis results. Based on user feedback and analysis results, dynamically adjust the parameters of the feature extraction algorithm and the screening strategy to improve the system's adaptive ability. Through online learning Online Learning technology, continuously update the feature extraction and analysis models to improve the system's intelligence level;
[0100] Step 5, System Architecture and Technical Support: Adopt the distributed computing framework Spark and GPU cluster to accelerate the processing and analysis of large-scale data, and use the distributed storage system HDFS and cache mechanism Redis to support efficient data storage and access;
[0101] Step 6, Edge Computing Support: Conduct preliminary feature extraction and data screening on edge devices to reduce data transmission latency, and upload tasks that cannot be processed by edge devices to the cloud for further processing.
[0102] For different types of data, such as images, text, sensor data, etc., edge devices will first perform standardization processing, which means converting the data into a common format for subsequent unified processing. For example, images are converted to a specific resolution or color format, text is converted to a unified encoding format, and sensor data is converted to a specific data unit or range.
[0103] Next, edge devices will use corresponding algorithms to extract key features from the standardized data. These features can represent the main information of the original data, thereby reducing the amount of data to be processed. For example, for image data, features such as edges, corners, or textures are extracted; for text data, features such as keywords or word frequencies are extracted.
[0104] After feature extraction, edge devices will screen the data according to preset rules or models, which means that only data that meets specific conditions or contains important information will be uploaded to the cloud for further processing. This screening greatly reduces unnecessary data transmission and improves the overall processing efficiency.
[0105] By performing data preprocessing and screening on edge devices, the amount of data that needs to be transmitted to the cloud is significantly reduced, which not only reduces the occupancy of network bandwidth but also reduces the latency of data transmission.
[0106] Since edge devices have already performed preliminary processing and screening on the data, the amount of data received by the cloud is greatly reduced, reducing the computing and storage burden on the cloud and enabling it to more efficiently process the remaining important data.
[0107] Through edge computing, the efficiency of the entire data processing system has been significantly improved. The rapid preprocessing and screening of edge devices reduce the latency of data transmission, while the processing burden on the cloud is also alleviated, enabling the entire system to respond and process data faster and improving the overall performance.
[0108] For different types of data, such as images, text, and sensor data, edge computing can achieve a unified processing flow. Through steps such as standardization, feature extraction, and data screening, this data is converted into a common and easily processable format, which not only simplifies the complexity of data processing but also improves the efficiency and accuracy of processing.
[0109] By performing preliminary data preprocessing and screening at the device edge, edge computing enables unified processing of different data types and significantly enhances data processing efficiency. This architecture not only reduces the amount of data transmitted and the processing burden on the cloud but also improves the response speed and processing capacity of the entire system. In practical applications, this method has significant advantages in handling large volumes of diverse data.
[0110] By standardizing different types of data on edge devices and converting the data into a unified format for subsequent unified processing, this step can significantly reduce the complexity and redundancy of the data.
[0111] Feature extraction is performed on edge devices to extract the key features of the data. These features can represent the main information of the original data and greatly reduce the amount of data that needs to be transmitted.
[0112] The feature data is screened through preset rules or models, and only the data containing key information is uploaded to the cloud. This step further reduces the amount of data that needs to be transmitted.
[0113] Reduce data transmission time: Since the amount of data transmitted is greatly reduced, the occupancy of network bandwidth is also correspondingly reduced, thereby shortening the data transmission time.
[0114] Reduce network latency: Reducing the amount of data transmitted also reduces network latency, enabling the data to reach the cloud faster.
[0115] Perform preliminary data preprocessing on edge devices, such as standardization, feature extraction, and screening, to reduce the complexity of the data, making the data received by the cloud more concise and organized.
[0116] Since the amount of data received by the cloud is greatly reduced, the computing burden on the cloud is also reduced, enabling the cloud to process the remaining data faster and improving the processing efficiency.
[0117] Reduce cloud processing time: The amount of data processed by the cloud is reduced, and computing resources are utilized more efficiently, thereby shortening the processing time.
[0118] Improve response speed: The cloud can respond to new data processing requests faster, improving the overall response speed of the system.
[0119] The preliminary processing and screening steps of the edge device are quickly completed locally without relying on cloud resources, which not only reduces the dependence on the network but also improves the real-time performance of data processing.
[0120] Multiple edge devices perform data preprocessing and screening simultaneously, achieving parallel processing, further improving the efficiency of the overall system and shortening the time for big data processing.
[0121] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it is understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for real-time processing and analysis of AI big data, characterized in that: The following steps are involved: Step 1: Data acquisition and preprocessing: Real-time data acquisition from multiple data sources: sensors, user behavior records, log files, and social media. Support the collection of structured, semi-structured, and unstructured data. Improve data quality through automated cleaning tools: outlier detection. Convert data formats from different sources: JSON, XML, text, and images into a standard format. Compress large-scale data to reduce storage and transmission costs. Step 2: Feature extraction and intelligent adaptation: Select the optimal feature extraction algorithm for different data types: text, image, audio, video and time series data, including: Text data: Use the TF-IDF deep learning model to extract semantic features; Image data: Use convolutional neural network (CNN) to extract visual features; Audio data: Use Mel-frequency cepstral coefficients to extract audio features; Time series data: Use LSTM to extract time series features; Adaptive feature extraction: Through meta-learning technology, it automatically adapts to different data types and distributions to improve the generalization ability of feature extraction; Feature fusion: Fuse multimodal features: text, image, and time series features to generate a unified feature representation; Step 3: Feature association and data screening: Associate the extracted multimodal features with the data in the AI big data database, support cross-modal association analysis: association between text and images, audio and video, use the distributed computing framework Spark to accelerate the association calculation of large-scale data, and use the machine learning model: random forest to intelligently screen the associated data based on user demand data and association analysis results; Step 4: Data analysis: Use the lightweight inference engine TensorRT to perform real-time analysis on the filtered data, support multi-dimensional analysis of time, space, and behavior, generate visual analysis results, and use users' real-time feedback data: click behavior, ratings, and comments, as input to optimize the analysis results. Based on user feedback and analysis results, dynamically adjust the parameters and screening strategies of the feature extraction algorithm to improve the system's adaptability. Through online learning technology, continuously update the feature extraction and analysis models to improve the intelligence level of the system. Step 5: System architecture and technical support: Use the distributed computing framework Spark and GPU cluster to accelerate the processing and analysis of large-scale data, and use the distributed storage system HDFS and cache mechanism Redis to support efficient data storage and access; Step 6: Edge computing support: Perform preliminary feature extraction and data screening on edge devices to reduce data transmission delays, and upload tasks that cannot be processed by edge devices to the cloud for further processing; The distributed computing framework Spark accelerates the computing in step 5 as follows: Spark converts the dataset Divide into Partitions: ; The calculation results for each partition Through parallel calculation, we get: ; Final global result It is the combination of the results of each partition: ; The step 5 also includes validity calculation, which calculates the processing time of each partition. and speedup : ; in: is the single node processing time, To process the The total time consumed by the partitions, To allocate to The number of nodes in a partition; The edge device preliminary feature extraction calculation is as follows: the edge device performs preliminary processing on the data, extracts key features, and reduces the amount of data: Assume the original data is ,in is the number of samples, is the characteristic number; Feature extraction: Extract using feature extraction algorithm key features, and obtain a new feature matrix ; in, is the feature extraction function; Data transfer volume calculation: Raw data transfer volume ,in The data size for each feature: bytes, Data transfer after extraction ; Data transmission reduction ratio: .
2. The AI big data real-time processing and analysis method according to claim 1, characterized in that: In the step 1, outlier detection is performed based on the Z-score method, and the detection calculation is: ; in: is a data point, is the mean of the data, is the standard deviation of the data, if , then the data point is considered to be an outlier; The data compression rate calculation in step 1 is: ; The original data size is the data size before compression, and the compressed data size is the data size after compression.
3. The AI big data real-time processing and analysis method according to claim 2 is characterized by: The TF-IDF deep learning model in step 2 is: ; in: is the word frequency, indicating the word In the documentation The frequency of occurrence in The specific formula is: Word frequency ; Inverse Document Frequency ; in, is the total number of documents in the corpus.
4. The AI big data real-time processing and analysis method according to claim 1, characterized in that: In step 2, the convolutional neural network extracts visual features from the image through convolution operation, which is calculated as: ; in: is the pixel matrix of the input image, with a size of , is the height, is the width, is the number of channels; is the bias term; In step 2, the Mel frequency cepstral coefficient is used to extract the features of the audio data and is calculated as: ; in: It is The logarithmic energy of the mel-bands; is the timeframe index.
5. The AI big data real-time processing and analysis method according to claim 1, characterized in that: In step 2, the long short-term memory network (LSTM) is used to extract the features of time series data, which is calculated as: Forget Gate ; in: is the output of the forget gate, a value between 0 and 1; is the weight matrix of the forget gate, which is used to learn the relationship between the input and the previous state; is the bias term of the forget gate, which is used to adjust the output of the weight matrix; is the time step The hidden state of is at the time step The current input data; Input Gate ; in: is the output of the input gate, a value between 0 and 1; is the weight matrix of the input gate, which is used to learn the relationship between the input and the previous state; is the bias term of the input gate, which is used to adjust the output of the weight matrix; Cell status update ; ; in: is a candidate state, representing the potential value of new information; is the weight matrix of the candidate state, which is used to learn the relationship between the input and the previous state; is the bias term of the candidate state, which is used to adjust the output of the weight matrix; Represents the cell state information of the previous time step that is retained; Output Gate ; ; in: The activation function is the Sigmoid function, is the bias term, is the output of the output gate, which is a value between 0 and 1; is the weight matrix of the output gate, which is used to learn the relationship between the input and the previous state; is the bias term of the output gate, which is used to adjust the output of the weight matrix; Is the cell state The activation value of , compresses the value of the cell state to the (-1, 1) interval.
6. The AI big data real-time processing and analysis method according to claim 1, characterized in that: The cross-modal association analysis in step 3 is used to find the association between text and image, audio and video. Suppose there are data of modality A and modality B, and the cosine similarity is used to calculate the similarity or correlation between them. For two vectors and , the cosine similarity is defined as: ; in: and They are text data and image data respectively; is a vector and The cosine of the angle between them; and Respectively, vector and The modulus of , which is the length of the vector; Use Euclidean distance to measure the distance between feature vectors. The smaller the distance, the more similar it is: ; in: and is a vector and No. elements; There are two vectors in The square of the difference in the dimension ensures that the distance is non-negative and amplifies the difference; It is to add the squares of the differences in all dimensions to get a total square difference; It is the square root of the square difference, that is, the Euclidean distance, which measures the straight-line distance between two vectors in Euclidean space. The smaller it is, the closer the two vectors are.
7. The AI big data real-time processing and analysis method according to claim 6, characterized in that: The time dimension analysis calculation in step 4 is as follows: calculating the mean and variance of the data in the time dimension; ; in: For time period The mean of the internal data; In step 4, the spatial dimension analysis calculates the spatial distribution of the data and clusters the spatial data using the K-means algorithm: ; in: is the cluster center set; For the Cluster centers.
Citation Information
Patent Citations
A real-time processing and analysis method for AI big data
CN118245680B
Drug information storage method based on distributed edge calculation and multi-modal data
CN118152481A