Multi-granularity extraction and enhancement method, device and equipment for multi-source heterogeneous data

By preprocessing multi-source heterogeneous data, multi-level feature extraction, cross-modal knowledge migration and dynamic fusion, the problems of in-depth mining of inter-feature correlation relationships and inter-modal knowledge sharing in multi-source heterogeneous data processing are solved, and efficient and accurate multi-task data analysis is achieved.

CN119513569BActive Publication Date: 2025-07-25ZHENGZHOU DIGITAL INTELLIGENCE TECH RES INST CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411550522.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-07-25
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

When processing multi-source heterogeneous data, the existing technology lacks in-depth exploration of the correlation relationship between features, the inter-modal knowledge transfer mechanism is imperfect, the feature fusion process is insufficient, the feature sharing and task correlation mechanism in multi-task processing is not sound, and it is difficult to achieve efficient utilization of resources.

Method used

By preprocessing multi-source heterogeneous data, a multi-modal data set and original feature parameters are generated, multi-level feature extraction is performed, feature association matrix and optimized feature representation are generated, cross-modal knowledge migration is performed, feature parameters are dynamically fused, multi-task segmentation and timing enhancement are performed, and visual display is finally performed.

Benefits of technology

It realizes efficient feature extraction and accurate analysis of multi-source heterogeneous data, enhances the knowledge sharing ability between modals, improves the efficiency and accuracy of data processing, and realizes multi-task collaborative processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119513569B_ABST
    Figure CN119513569B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and discloses a multi-granularity extraction and enhancement method, device and equipment for multi-source heterogeneous data. The method includes: performing multi-level feature extraction on a multi-modal data set and original feature parameters to obtain a multi-modal feature set; performing correlation analysis on the multi-modal feature set to obtain a feature correlation matrix and an optimized feature representation; performing cross-modal knowledge transfer processing on the feature correlation matrix and the optimized feature representation to obtain transfer feature parameters; performing dynamic fusion processing on the transfer feature parameters to obtain target fusion features; performing multi-task segmentation processing according to the target fusion features to obtain multiple sub-task sequences, and respectively performing temporal enhancement and multi-task fusion analysis on each sub-task sequence to obtain target extraction data, and performing visual display on the target extraction data. The present application improves the efficiency and accuracy of multi-granularity extraction and enhancement for multi-source heterogeneous data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular, to a multi-granularity extraction and enhancement method, device, and equipment for multi-source heterogeneous data. Background Art

[0002] With the rapid development of information technology, multi-source heterogeneous data has been widely used in various fields. These data come from diverse sources, including different modalities such as text, images, videos, and audio, and their data formats and structures are also different. Existing data processing technologies mainly focus on the processing of single-modal data. For example, text processing technologies focus on semantic analysis, image processing technologies emphasize visual feature extraction, video processing technologies focus on capturing dynamic information, and audio processing technologies concentrate on acoustic feature analysis. At the same time, some cross-modal data processing methods have emerged, attempting to establish association and mapping relationships between different-modal data.

[0003] However, the existing technologies still have the following deficiencies: First, the feature extraction for multi-source heterogeneous data is mostly carried out independently, lacking in-depth exploration of the association relationships between features; second, the knowledge transfer mechanism between different-modal data is imperfect, making it difficult to fully utilize the complementary information between modalities; third, the feature fusion process often adopts simple splicing or weighting methods, failing to fully consider the hierarchy and dynamics of features; finally, the feature sharing and task association mechanisms in the multi-task processing process are not sound enough, making it difficult to achieve efficient utilization of resources. Summary of the Invention

[0004] This application provides a multi-granularity extraction and enhancement method, device, and equipment for multi-source heterogeneous data, which is used to improve the efficiency and accuracy of multi-granularity extraction and enhancement of multi-source heterogeneous data.

[0005] In a first aspect, the present application provides a multi-granularity extraction and enhancement method for multi-source heterogeneous data. The multi-granularity extraction and enhancement method for multi-source heterogeneous data includes: preprocessing the multi-source heterogeneous data to obtain a multi-modal data set and original feature parameters, where the original feature parameters include: content feature values, distribution feature values, and quality feature values; performing multi-level feature extraction on the multi-modal data set and the original feature parameters to obtain a multi-modal feature set; performing correlation analysis on the multi-modal feature set to obtain a feature correlation matrix and an optimized feature representation; performing cross-modal knowledge transfer processing on the feature correlation matrix and the optimized feature representation to obtain transfer feature parameters; performing dynamic fusion processing on the transfer feature parameters to obtain a target fusion feature, where the target fusion feature includes: a low-level fusion feature, a middle-level fusion feature, and a high-level fusion feature; performing multi-task segmentation processing based on the target fusion feature to obtain a plurality of sub-task sequences, and respectively performing temporal enhancement and multi-task fusion analysis on each sub-task sequence to obtain target extraction data, and performing visual display on the target extraction data.

[0006] In a second aspect, the present application provides a multi-granularity extraction and enhancement device for multi-source heterogeneous data. The multi-granularity extraction and enhancement device for multi-source heterogeneous data includes:

[0007] A processing module for preprocessing the multi-source heterogeneous data to obtain a multi-modal data set and original feature parameters, where the original feature parameters include: content feature values, distribution feature values, and quality feature values;

[0008] An extraction module for performing multi-level feature extraction on the multi-modal data set and the original feature parameters to obtain a multi-modal feature set;

[0009] An analysis module for performing correlation analysis on the multi-modal feature set to obtain a feature correlation matrix and an optimized feature representation;

[0010] A transfer module for performing cross-modal knowledge transfer processing on the feature correlation matrix and the optimized feature representation to obtain transfer feature parameters;

[0011] A fusion module for performing dynamic fusion processing on the transfer feature parameters to obtain a target fusion feature, where the target fusion feature includes: a low-level fusion feature, a middle-level fusion feature, and a high-level fusion feature;

[0012] A segmentation module for performing multi-task segmentation processing based on the target fusion feature to obtain a plurality of sub-task sequences, and respectively performing temporal enhancement and multi-task fusion analysis on each sub-task sequence to obtain target extraction data, and performing visual display on the target extraction data.

[0013] A third aspect of the present application provides a computer device. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through a bus. When the machine-readable instructions are executed by the processor, the steps of the above-mentioned multi-granularity extraction and enhancement method for multi-source heterogeneous data are executed.

[0014] In the technical solution provided by the present application, by preprocessing multi-source heterogeneous data, a multi-modal data set and original feature parameters are obtained, effectively solving the problems of inconsistent data formats and non-standard feature expressions, and laying a foundation for subsequent feature extraction. Multi-level feature extraction is performed on the multi-modal data set and original feature parameters, realizing the mining of data features from different granularities and levels, and improving the integrity and accuracy of feature expression. In the correlation analysis stage, by generating a feature correlation matrix and optimizing feature representation, the internal connections between different features are fully mined, enhancing the feature expression ability. The cross-modal knowledge transfer processing link realizes knowledge sharing and transfer between different modalities by obtaining transfer feature parameters, improving the utilization efficiency of information between modalities. The target fusion features generated in the dynamic fusion processing stage include low-level fusion features, middle-level fusion features, and high-level fusion features, constructing a complete feature hierarchy and enhancing the richness of feature expression. Finally, multiple sub-task sequences are obtained through multi-task segmentation processing, and time-series enhancement and multi-task fusion analysis are performed, realizing the efficient decomposition and collaborative processing of tasks. At the same time, the target extraction data is visually displayed, improving the intuitiveness and interpretability of data analysis. The entire solution forms a complete multi-source heterogeneous data processing framework through the organic combination of multiple links, not only improving the accuracy and efficiency of feature extraction, but also enhancing the knowledge sharing ability between modalities, realizing the collaborative processing of multiple tasks, and ultimately achieving the unity of high efficiency, accuracy, and reliability in data analysis and processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic diagram of an embodiment of the multi-granularity extraction and enhancement method for multi-source heterogeneous data in an embodiment of the present application;

[0017] Figure 2 It is a schematic diagram of an embodiment of the multi-granularity extraction and enhancement device for multi-source heterogeneous data in an embodiment of the present application;

[0018] Figure 3 This is a schematic structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners

[0019] The embodiments of the present application provide a multi-granularity extraction and enhancement method, device, and equipment for multi-source heterogeneous data. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or equipment that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or equipment.

[0020] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 One embodiment of the multi-granularity extraction and enhancement method for multi-source heterogeneous data in the embodiments of the present application includes:

[0021] Step S101: Preprocess the multi-source heterogeneous data to obtain a multi-modal data set and original feature parameters. The original feature parameters include: content feature values, distribution feature values, and quality feature values;

[0022] Step S102: Perform multi-level feature extraction on the multi-modal data set and the original feature parameters to obtain a multi-modal feature set;

[0023] Step S103: Perform correlation analysis on the multi-modal feature set to obtain a feature correlation matrix and an optimized feature representation;

[0024] Step S104: Perform cross-modal knowledge transfer processing on the feature correlation matrix and the optimized feature representation to obtain transfer feature parameters;

[0025] Step S105: Perform dynamic fusion processing on the transfer feature parameters to obtain target fusion features. The target fusion features include: underlying fusion features, middle-level fusion features, and high-level fusion features;

[0026] Step S106: Perform multi-task segmentation processing based on the target fusion features to obtain multiple sub-task sequences, and perform temporal enhancement and multi-task fusion analysis on each sub-task sequence respectively to obtain target extraction data, and perform visual display on the target extraction data.

[0027] It can be understood that the execution entity of this application can be a multi-granularity extraction and enhancement device for multi-source heterogeneous data, or it can also be a terminal or a server, and specific details are not limited here. In the embodiments of this application, the server is taken as an example of the execution entity for illustration.

[0028] Specifically, preprocess the multi-source heterogeneous data. Among them, the multi-source heterogeneous data includes text, image, video, and audio data. There are a large number of invalid characters and noisy texts in the text data; the image data has problems such as blurring and uneven illumination; the video data has problems such as inconsistent frame rates and jitter; the audio data has problems such as noise interference and inconsistent sampling rates. In the preprocessing stage, through adaptive data cleaning, use wavelet transform to remove noise, tokenize and standardize the text; use histogram equalization to adjust the image contrast and median filtering to eliminate noise points; perform stability correction and frame rate standardization on the video; perform noise reduction and sampling rate unification on the audio. In the feature parameter extraction, the content feature value reflects the core information volume of the data and is obtained through means such as semantic analysis, image analysis, and time series analysis; the distribution feature value describes the statistical law of the data, including statistical quantities such as mean and variance; the quality feature value characterizes the integrity and reliability of the data.

[0029] In the multi-level feature extraction stage, adopt a hierarchical extraction strategy for different types of data. For text data, morphological features are extracted at the character level, semantic features are extracted at the word level, syntactic features are extracted at the sentence level, and topic features are extracted at the document level. For image data, color texture features are extracted at the pixel level, shape edge features are extracted at the region level, semantic features are extracted at the object level, and context features are extracted at the scene level. For video data, spatial features and time series features are extracted, including action features and scene change features. For audio data, time domain features, frequency domain features, and time-frequency joint features are extracted. In the correlation analysis, by constructing a feature correlation map, calculate the correlation between different features, and generate a feature correlation matrix. Adopt feature selection and optimization algorithms to screen out the most representative feature combinations to form an optimized feature representation. In the knowledge transfer process, transfer the knowledge of the information-rich modality to the information-sparse modality to enhance the feature expression ability.

[0030] Taking the quality control of an intelligent manufacturing production line as an example, the input data includes production parameter records, product appearance inspection images, production process videos, and equipment operation audio. In the parameter records, each production record contains 50 process parameters such as temperature, pressure, and speed. After removing outliers through data cleaning, the distribution characteristic values of each parameter are calculated. For product appearance inspection images, the contrast is increased by 35% and the clarity is increased by 40% through image enhancement, and features such as the color and texture of the product surface are extracted. The production process video records the complete assembly process. Through the key frame extraction algorithm, representative frames are extracted from the original video, and the timing characteristics of the assembly process are calculated. The equipment operation audio retains the effective signal of 100Hz - 5kHz through band-pass filtering, and the working state characteristics of the equipment are extracted.

[0031] After feature extraction and correlation analysis, it is found that the correlation coefficient between process parameters and product appearance features reaches 0.85, and the timing correlation between assembly action features and equipment operation audio reaches 0.78. Through knowledge transfer, the defect feature knowledge of appearance inspection is transferred to parameter analysis, improving the recognition accuracy of parameter anomalies. In the dynamic fusion stage, the underlying fusion features include basic data features such as temperature curves and pressure changes; the middle-level fusion features include state correlation features such as the correlation between parameter anomalies and appearance defects; the high-level fusion features include comprehensive evaluation features for the overall evaluation of product quality. Finally, the quality control task is decomposed into sub-tasks such as parameter monitoring, appearance inspection, and assembly verification to achieve full-process quality control. Through multi-dimensional data analysis and feature fusion, product quality problems are accurately identified and traced back to specific process links, providing data support for production optimization.

[0032] In the embodiments of the present application, by preprocessing multi-source heterogeneous data, a multi-modal data set and original feature parameters are obtained, effectively solving the problems of inconsistent data formats and non-standard feature expressions, and laying a foundation for subsequent feature extraction. Multi-level feature extraction is performed on the multi-modal data set and the original feature parameters, realizing the mining of data features from different granularities and levels, and improving the integrity and accuracy of feature expressions. In the correlation analysis stage, by generating a feature correlation matrix and optimizing feature representations, the internal relationships between different features are fully explored, enhancing the expression ability of features. The cross-modal knowledge transfer processing link realizes knowledge sharing and transfer between different modalities by obtaining transfer feature parameters, improving the utilization efficiency of information between modalities. The target fusion features generated in the dynamic fusion processing stage include low-level fusion features, middle-level fusion features, and high-level fusion features, constructing a complete feature hierarchy and enhancing the richness of feature expressions. Finally, multiple sub-task sequences are obtained through multi-task segmentation processing, and time-series enhancement and multi-task fusion analysis are performed, realizing the efficient decomposition and collaborative processing of tasks. At the same time, the target extraction data is visually displayed, improving the intuitiveness and interpretability of data analysis. Through the organic combination of multiple links, the entire solution forms a complete multi-source heterogeneous data processing framework, not only improving the accuracy and efficiency of feature extraction, but also enhancing the knowledge sharing ability between modalities, realizing the collaborative processing of multiple tasks, and ultimately achieving the unity of high efficiency, accuracy, and reliability in data analysis and processing.

[0033] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0034] (1) Perform data cleaning processing on the multi-source heterogeneous data to obtain an initial data marking result, and perform type recognition processing on the initial data marking result to obtain a data type identifier;

[0035] (2) Perform data analysis processing on the data type identifier to obtain a basic data sequence and basic feature data;

[0036] (3) Perform modal division processing on the basic data sequence to obtain a multi-modal data set, and perform content analysis processing on the basic feature data to obtain content feature values;

[0037] (4) Perform time-frequency decomposition processing on the basic feature data to obtain a multi-scale feature sequence, and perform distribution calculation processing on the multi-scale feature sequence to obtain distribution feature values;

[0038] (5) Perform quality assessment processing on the multi-scale feature sequence to obtain quality feature values, and perform combination processing on the content feature values, distribution feature values, and quality feature values to obtain original feature parameters.

[0039] Specifically, multi-source heterogeneous data refers to a data set from different sources and in different formats, including text, image, video, and audio data. Data cleaning and processing refers to the process of preprocessing the original data to remove outliers, noise, and redundant data. Specifically: for text data, special characters, duplicate content, and meaningless symbols are removed; for image data, blurred pixels, noise, and distorted areas are removed; for video data, jitter frames, blurred frames, and discontinuous frames are removed; for audio data, background noise and distorted segments are removed. The initial data marking result obtained after data cleaning contains basic attribute markings of the data, such as information on the data generation time, data format, data scale, etc. Type recognition processing is performed on these marking results to identify the specific type and characteristics of the data, generating a data type identifier, which contains information such as the modality type, format type, and structure type of the data.

[0040] Data analysis processing is performed on the data type identifier to analyze the internal characteristics and structure of the data. The basic data sequence refers to the time series of the original data after preliminary processing, and the basic feature data are the basic features extracted from these sequences, such as the word frequency feature of text, the color feature of images, the motion feature of videos, and the frequency feature of audio.

[0041] Modality division processing classifies and organizes the basic data sequences according to different data modalities (text, image, video, audio) to form a multi-modal data set. Content analysis processing is performed on the basic feature data to extract the core features of the data content, obtaining content feature values. The content feature values reflect the main content information of the data, such as the theme feature of text, the target feature of images, the scene feature of videos, and the speech feature of audio.

[0042] Time-frequency decomposition processing decomposes the basic feature data in the time and frequency dimensions to obtain multi-scale feature sequences. The multi-scale feature sequences contain feature information on different time scales and frequency scales. Distribution calculation processing is performed on these sequences to calculate the statistical distribution features of the features, such as mean, variance, skewness, kurtosis, etc., obtaining distribution feature values.

[0043] Quality assessment processing measures the quality of the multi-scale feature sequences, evaluating the integrity, accuracy, and reliability of the data, obtaining quality feature values. Finally, the content feature values (reflecting the data content), distribution feature values (reflecting the data distribution), and quality feature values (reflecting the data quality) are combined and processed to be fused into a complete set of original feature parameters.

[0044] Taking the intelligent manufacturing environment monitoring as an example, the input temperature sensor data (10 values per second) is first subjected to data cleaning to remove outliers outside the range of 20 - 80 °C, and the cleaned data is obtained and the sampling time and validity are marked. Type recognition determines that this is numerical time-series data, and data analysis generates a basic data sequence (temperature change sequence over time) and basic feature data (temperature change rate, temperature fluctuation range, etc.). Modal division classifies the temperature data as numerical sensing data, and content analysis extracts temperature change characteristic values. Time-frequency decomposition obtains temperature change characteristics at different time scales (hourly, daily), and distribution calculation obtains statistical characteristics of temperature. Quality assessment calculates the completeness rate and accuracy rate of the data, and finally combines to form the original feature parameters including temperature characteristics, distribution characteristics, and quality indicators.

[0045] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0046] (1) Perform hierarchical parsing processing on the multi-modal data set to obtain initial text data, initial image data, initial video data, and initial audio data, and perform feature matching processing on the initial text data, initial image data, initial video data, and initial audio data with the original feature parameters to obtain a basic feature sequence;

[0047] (2) Perform in-depth semantic processing on the text part in the basic feature sequence to obtain text-level features, and perform spatial structure processing on the image part in the basic feature sequence to obtain image-level features;

[0048] (3) Perform spatio-temporal change processing on the video part in the basic feature sequence to obtain video-level features, and perform spectral analysis processing on the audio part in the basic feature sequence to obtain audio-level features;

[0049] (4) Perform feature integration processing on the text-level features, image-level features, video-level features, and audio-level features to obtain a multi-level feature matrix, and perform feature screening processing on the multi-level feature matrix to obtain a target feature vector;

[0050] (5) Perform feature recombination processing on the target feature vector to obtain a multi-modal feature set.

[0051] Specifically, hierarchical parsing processing is a process of hierarchically splitting a multimodal dataset. Structured parsing of text data includes sentence splitting, word segmentation, and part-of-speech tagging to generate initial text data; pixel-level parsing of image data includes color space conversion, edge detection, and region segmentation to generate initial image data; frame-level parsing of video data includes frame extraction, scene segmentation, and motion detection to generate initial video data; sampling-level parsing of audio data includes time-domain analysis, frequency-domain conversion, and acoustic feature extraction to generate initial audio data. Feature matching processing is to correspond and map these initial data with the original feature parameters, where the original feature parameters include content feature values, distribution feature values, and quality feature values, and feature matching is performed by calculating a similarity matrix to obtain a basic feature sequence. Deep semantic processing is performed on the text part of the basic feature sequence, and natural language processing techniques are used to extract semantic information. First, a word vector space is constructed to calculate the semantic similarity between words, then the sentence structure features are extracted through syntactic analysis, and finally the document-level semantic features are extracted through topic modeling to form text hierarchical features including character-level, word-level, sentence-level, and document-level. Spatial structure processing is performed on the image part. First, multi-scale decomposition is carried out to extract color and texture features at the pixel level, shape and edge features at the region level, semantic features at the object level, and finally global features at the scene level to obtain image hierarchical features.

[0052] The spatio-temporal change processing of the video part includes the extraction of spatial features and time features. Visual features of each frame are extracted in the spatial dimension, and frame-by-frame changes are analyzed in the time dimension to capture motion information and scene transitions to obtain video hierarchical features. The spectral analysis processing of the audio part includes short-time Fourier transform, Mel-frequency cepstral coefficient extraction, pitch analysis, etc., to extract acoustic features from different frequency dimensions to obtain audio hierarchical features. Feature integration processing is to uniformly represent and organize the hierarchical features of different modalities. The features of different modalities are projected into the same feature space through feature mapping to construct a multi-level feature matrix. Feature screening is performed on this matrix, and principal component analysis and feature importance evaluation are used to select the most representative feature combination to obtain the target feature vector. Finally, through feature recombination processing, the target feature vector is reorganized into a multimodal feature set to achieve a structured representation of the features.

[0053] Taking intelligent manufacturing quality control as a specific example, the input data includes production process parameter texts (recording 50 parameters per minute), product appearance inspection images (2048×1536 resolution), production process videos (1080p, 30fps), and device operation audio (48kHz sampling rate). In the hierarchical parsing stage, word segmentation is performed on the process parameter texts to identify the process parameter names and values; multi-scale decomposition is performed on the appearance inspection images to construct a 4-layer image pyramid; key frames are extracted from the production videos at 2-second intervals; and the device audio is framed in units of 50ms.

[0054] In the feature matching stage, the process parameters are compared with the set ranges, the image features are matched with the standard templates, the video sequences are aligned with the standard process flows, and the audio features are compared with the acoustic features of normal device operation. Deep semantic processing extracts the variation patterns and abnormal modes of the process parameters, spatial structure processing identifies the surface defect features of the products, spatio-temporal variation processing captures the abnormal actions during the production process, and spectral analysis processing identifies the abnormal vibration features of the devices. Through feature integration, these features are organized into a multi-dimensional feature matrix, where each row of the matrix represents a time point and each column represents a feature dimension. Feature screening retains the feature combinations with the highest contribution, and finally generates a feature set that can comprehensively characterize the production quality status.

[0055] In a specific embodiment, the process of performing step S103 may specifically include the following steps:

[0056] (1) Perform semantic association processing on the text-level features to obtain a text association vector, and perform spatial association processing on the image-level features to obtain an image association vector;

[0057] (2) Perform temporal sequence association processing on the video-level features to obtain a video association vector, and perform spectral association processing on the audio-level features to obtain an audio association vector;

[0058] (3) Perform cross-modal matching processing on the text association vector and the image association vector to obtain a text-image association matrix, and perform cross-modal matching processing on the video association vector and the audio association vector to obtain a video-audio association matrix;

[0059] (4) Perform matrix fusion processing on the text-image association matrix and the video-audio association matrix to obtain a feature association matrix, and perform feature extraction processing on the feature association matrix to obtain an optimized feature representation.

[0060] Specifically, semantic association processing calculates the correlation of semantic information in the text-level features. For each text unit, first calculate the semantic relatedness between words, and convert the text into a vector representation in a high-dimensional space through word vectorization. Semantic association processing includes word-level association, sentence-level association, and document-level association, and finally generates a text association vector. For the spatial association processing of image-level features, by calculating the spatial dependence relationship between different regions of the image, including the feature similarity of adjacent regions, the positional relationship and structural relationship between regions, a spatial association graph of image features is constructed and converted into an image association vector. When performing temporal association processing on video-level features, it is necessary to analyze the variation law of the video frame sequence in the time dimension. Calculate the feature change between adjacent frames, extract the motion trajectory information, identify the scene transition points, construct the temporal dependence relationship, and form a video association vector. The spectral association processing of audio-level features focuses on the association relationship between different frequency components, and generates an audio association vector by analyzing the spectral energy distribution, band correlation, and harmonic structure.

[0061] Cross-modal matching processing is a key step in establishing the corresponding relationship between different modal data. Match the text association vector and the image association vector, calculate the corresponding relationship between semantic content and visual content, and construct a text-image association matrix reflecting the association strength between the two. Similarly, match the video association vector and the audio association vector, analyze the temporal corresponding relationship between visual changes and acoustic features, and obtain a video-audio association matrix. Matrix fusion processing integrates the text-image association matrix and the video-audio association matrix, and constructs a complete feature association matrix by calculating the mutual information between different modal features. Feature extraction processing then extracts the most representative feature combination from this matrix to form an optimized feature representation.

[0062] Taking the equipment monitoring in the intelligent manufacturing environment as an example, the input includes the text record of equipment operation parameters (recording 20 parameters per second), the surface temperature image of the equipment (resolution 1024×768), the equipment operation video (720p, 25fps), and the equipment vibration audio (sampling rate 96kHz). Through semantic association processing, key parameters such as "rotation speed", "temperature", and "vibration" in the parameter record are associated and analyzed, and the correlation coefficient between parameters is calculated. When the correlation coefficient between the rotation speed and vibration exceeds 0.8, it is marked as a strongly correlated parameter pair. Spatial association processing analyzes the hot spot areas of the temperature image, calculates the spatial autocorrelation coefficient of the hot spot distribution, and identifies the spatial distribution pattern of the temperature anomaly area. Temporal association processing analyzes the motion characteristics in the equipment operation video, calculates the motion vector every 25 frames (1 second), and constructs a motion feature sequence. Spectral association processing performs short-time Fourier transform on the vibration audio according to a 100ms window, analyzes the spectral correlation in the range of 0 - 20kHz, and identifies the characteristic frequency combination. Text-image matching establishes a correspondence between the temperature data in the parameter record and the temperature distribution of the thermal imaging map, and constructs a 16×16 association matrix. Video-audio matching generates a 32×32 association matrix by aligning the motion characteristics of the video and the vibration characteristics of the audio. Finally, through matrix fusion, the association information of all modalities is integrated into a 64×64 feature association matrix, and each element reflects the association strength between different features. Feature extraction retains the feature combinations with an association strength exceeding 0.7, and finally obtains an optimized feature representation that can comprehensively characterize the equipment operation state. It can not only discover the feature associations within a single modality, but also identify the cross-modal feature correspondence relationships.

[0063] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0064] (1) Perform matrix decomposition processing on the feature association matrix to obtain associated feature components, and perform probability distribution mapping processing on the associated feature components to obtain a feature distribution sequence;

[0065] (2) Perform entropy value calculation processing on the feature distribution sequence to obtain an information entropy matrix, and perform feature probability processing on the optimized feature representation to obtain a conditional probability vector;

[0066] (3) Perform mutual information calculation processing on the information entropy matrix and the conditional probability vector to obtain a knowledge transfer matrix, and perform feature weight assignment processing on the knowledge transfer matrix to obtain a weight coefficient sequence;

[0067] (4) Perform marginal probability calculation processing on the weight coefficient sequence to obtain a probability distribution map, and perform distribution consistency processing on the probability distribution map to obtain a transfer probability sequence;

[0068] (5) Process the migration probability sequence for feature correlation to obtain a correlation matrix, and perform threshold screening on the correlation matrix to obtain a screened feature sequence;

[0069] (6) Perform migration loss calculation on the screened feature sequence to obtain a loss function vector, and perform parameter optimization on the loss function vector and the weight coefficient sequence to obtain an optimized parameter group;

[0070] (7) Perform feature enhancement on the optimized parameter group to obtain an enhanced feature sequence, and perform parameter mapping on the enhanced feature sequence to obtain migration feature parameters, where the migration feature parameters include: feature importance parameters, structural similarity parameters, and modal conversion parameters.

[0071] Specifically, perform matrix decomposition on the feature correlation matrix, and use the singular value decomposition method to decompose the feature correlation matrix into feature vectors and eigenvalues. The associated feature components represent the basic components of different feature dimensions, and each component reflects a main direction in the feature space. The probability distribution mapping process maps these associated feature components to the probability space, and generates a feature distribution sequence by calculating the distribution density function of each feature component. Calculate the entropy value of the feature distribution sequence. The information entropy reflects the degree of uncertainty of the feature. For each feature dimension, calculate the information entropy of its probability distribution to form an information entropy matrix. At the same time, perform feature probability processing on the optimized feature representation, calculate the probability distribution of each feature under given conditions, and form a conditional probability vector. The mutual information calculation process measures the information dependence degree between different features, and constructs a knowledge migration matrix by calculating the mutual information between the information entropy matrix and the conditional probability vector. Feature weight assignment assigns importance weights to different features according to the size of the mutual information to generate a weight coefficient sequence.

[0072] The marginal probability calculation process calculates the marginal distribution of each feature in the weight coefficient sequence to obtain a probability distribution graph. The distribution consistency process adjusts the distribution parameters by comparing the probability distribution similarities of different features, making the feature distributions of the source domain and the target domain tend to be consistent, and forming a transfer probability sequence. The feature correlation process calculates the correlation coefficients between the features in the transfer probability sequence and constructs a correlation matrix. The threshold screening process sets a correlation threshold and retains the feature pairs with high correlation to obtain a screened feature sequence. The transfer loss calculation process evaluates the information loss in the knowledge transfer process, calculates the distance metric between the source domain features and the target domain features, and forms a loss function vector. The parameter optimization process adjusts the transfer parameters by minimizing the loss function while considering the weight coefficient sequence to obtain an optimized parameter set. The feature enhancement process enhances the original features using the optimized parameter set to improve the expression ability of the features and generates an enhanced feature sequence. The parameter mapping process maps the enhanced feature sequence to a standard feature space to obtain transfer feature parameters, including feature importance parameters, structural similarity parameters, and modality conversion parameters.

[0073] Taking the equipment status monitoring in an intelligent factory as an example, the input data includes vibration sensor data (sampling rate 10 kHz), temperature sensor data (sampling rate 1 Hz), current sensor data (sampling rate 1 kHz), and acoustic sensor data (sampling rate 48 kHz). First, perform singular value decomposition on a 64×64 feature correlation matrix, and retain the eigenvectors corresponding to the first 10 largest eigenvalues as the associated feature components. Map these feature components to a probability distribution through kernel density estimation to obtain a 10-dimensional feature distribution sequence. Calculate the information entropy for each feature distribution and construct a 10×10 information entropy matrix. Calculate the conditional probability through Bayesian estimation to obtain a 64-dimensional conditional probability vector. Mutual information calculation shows that the mutual information value between the vibration feature and the acoustic feature is 0.85, and the mutual information value between the temperature feature and the current feature is 0.72. Based on this, construct a knowledge transfer matrix and assign feature weights.

[0074] Calculations show that the probability distributions of the vibration signal in the 20 - 100 Hz frequency band and the acoustic signal in the 1 - 5 kHz frequency band are highly similar (similarity 0.92), and the migration parameters are adjusted accordingly. Feature correlation analysis shows that the correlation coefficient between temperature change and current fluctuation is 0.78, and a threshold value of 0.7 is set for feature screening. Migration loss calculation shows that when 90% of the information is retained, the feature dimension can be reduced from 64 dimensions to 32 dimensions. Finally, the migration feature parameters obtained through parameter optimization include: feature importance parameters (reflecting the weights of the vibration - acoustic feature pair 1.5 and the temperature - current feature pair 1.2), structural similarity parameters (spectrum structure similarity 0.92, time - series structure similarity 0.85), and modal conversion parameters (mapping coefficient from vibration to acoustics 0.8, mapping coefficient from temperature to current 0.75). These parameters effectively guide the knowledge migration process between different modalities, achieving the optimization and enhancement of data representation.

[0075] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0076] (1) Perform feature grouping processing on the feature importance parameters to obtain a bottom - layer feature group, and perform edge feature extraction processing on the bottom - layer feature group to obtain bottom - layer fusion features;

[0077] (2) Perform semantic association processing on the structural similarity parameters to obtain a middle - layer feature group, and perform content feature association processing on the middle - layer feature group to obtain middle - layer fusion features;

[0078] (3) Perform high - dimensional mapping processing on the modal conversion parameters to obtain a high - layer feature group, and perform context feature association processing on the high - layer feature group to obtain high - layer fusion features;

[0079] (4) Perform feature hierarchical division processing on the bottom - layer fusion features, middle - layer fusion features, and high - layer fusion features to obtain a multi - level feature sequence, and perform inter - layer relationship analysis processing on the multi - level feature sequence to obtain an inter - layer feature vector;

[0080] (5) Perform adaptive adjustment processing on the inter - layer feature vector to obtain an adjustment parameter matrix, and perform feature combination processing on the adjustment parameter matrix and the multi - level feature sequence to obtain target fusion features.

[0081] Specifically, the feature importance parameters are grouped. The feature importance parameters are numerical indicators reflecting the contribution degrees of various features to the target task. Through clustering analysis, features with similar importance levels are grouped to form underlying feature groups. The underlying feature groups contain basic data features, such as the word frequency features of text, the pixel features of images, the frame features of videos, and the waveform features of audio. Edge feature extraction processing extracts local significant features from the underlying feature groups, including the keyword features of text, the edge contour features of images, the motion boundary features of videos, and the spectral edge features of audio. These features together constitute the underlying fusion features. The structural similarity parameters describe the structural correspondence relationships between different features. Through semantic association processing, features with similar semantic structures are combined together to form middle-level feature groups. The middle-level feature groups contain more advanced semantic information, such as the syntactic structure of text, the object structure of images, the scene structure of videos, and the harmonic structure of audio. Content feature association processing analyzes the semantic connections between these structural features, extracts co-occurrence patterns and dependency relationships, and generates middle-level fusion features.

[0082] The modality conversion parameters describe the conversion rules between different modality data. Through high-dimensional mapping processing, these parameters are projected into a high-dimensional feature space to form high-level feature groups. The high-level feature groups contain semantic concepts at the abstract level, such as the theme features of text, the scene semantics of images, the event features of videos, and the semantic features of audio. Context feature association processing analyzes the semantic associations of these high-level features in different scenarios, constructs the context dependency relationships of the features, and obtains high-level fusion features. Feature hierarchy division processing organizes the underlying fusion features, middle-level fusion features, and high-level fusion features according to the semantic levels to construct a multi-level feature sequence. Inter-layer relationship analysis processing studies the dependency relationships between features at different levels, including the bottom-up feature abstraction relationship and the top-down feature refinement relationship, and generates inter-layer feature vectors. Adaptive adjustment processing dynamically adjusts the weights of features at different levels according to the importance of the features and the task requirements to form an adjustment parameter matrix. Feature combination processing performs weighted fusion on the adjustment parameter matrix and the multi-level feature sequence to obtain the final target fusion feature.

[0083] Taking intelligent video analysis as a specific example, the input includes video caption text, visual image sequences, action video clips, and scene audio. The feature importance parameters show that the key action words in the text (importance weight 0.8), the human targets in the images (importance weight 0.9), the action trajectories in the videos (importance weight 0.85), and the speech parts in the audio (importance weight 0.75) have relatively high importance, and these features are grouped into the underlying feature groups. Through edge feature extraction, the action description words in the text, the human outlines in the images (edge points with a confidence greater than 0.9), the motion boundaries in the videos (regions with an optical flow gradient greater than 0.5), and the speech segments in the audio (segments with a signal-to-noise ratio greater than 15 dB) are identified. In the middle-level feature processing, the structural similarity parameters indicate that the structural similarity between the action description words and the motion trajectories is 0.82, and the structural similarity between the human outlines and the speech features is 0.78. Through semantic association analysis, these features are organized into an action-speech correspondence graph, and each node contains a feature descriptor and an association strength. The content feature association processing finds that when the action amplitude increases, the speech intensity also increases accordingly, and the correlation coefficient reaches 0.85.

[0084] In the high-level feature processing, the modality conversion parameters guide the feature mapping between different modalities, such as mapping the visual action features to the speech feature space (mapping accuracy 0.88), and mapping the text semantics to the visual scene space (mapping accuracy 0.85). The context feature association analysis shows that in different scenarios, the combination patterns of the action-speech features have significant differences, and the scene correlation reaches 0.92. The feature hierarchical division organizes these features into a three-layer structure: the bottom layer contains basic features (64-dimensional feature vectors), the middle layer contains semantic features (32-dimensional feature vectors), and the high layer contains scene features (16-dimensional feature vectors). The inter-layer relationship analysis shows that the feature abstraction accuracy from the bottom layer to the middle layer is 0.85, and the feature abstraction accuracy from the middle layer to the high layer is 0.82. Through adaptive adjustment, the bottom layer feature weight is set to 0.3, the middle layer feature weight is set to 0.4, and the high layer feature weight is set to 0.3, generating a 112-dimensional target fusion feature, realizing the effective fusion of multi-modal data.

[0085] In a specific embodiment, the process of executing step S106 may specifically include the following steps:

[0086] (1) Perform density clustering processing on the target fusion feature to obtain a feature density map, and perform task boundary detection processing on the feature density map to obtain an initial task sequence;

[0087] (2) Perform task relevance calculation processing on the initial task sequence to obtain a relevance vector, and perform hierarchical clustering processing on the relevance vector to obtain multiple subtask sequences;

[0088] (3) Perform temporal feature extraction processing on multiple sub-task sequences to obtain a temporal feature matrix, and perform autocorrelation analysis processing on the temporal feature matrix to obtain a periodic feature vector;

[0089] (4) Perform feature spectrum analysis processing on the periodic feature vector to obtain a spectrum distribution diagram, and perform principal component extraction processing on the spectrum distribution diagram to obtain a principal component sequence;

[0090] (5) Perform temporal correlation degree calculation processing on the principal component sequence to obtain a correlation feature group, and perform sub-task feature mapping processing on the correlation feature group to obtain a mapping matrix set;

[0091] (6) Perform multi-task feature fusion processing on the mapping matrix set to obtain a fusion feature vector, and perform feature reconstruction processing on the fusion feature vector to obtain target extraction data, where the target extraction data includes: temporal feature data, correlation feature data, and fusion feature data;

[0092] (7) Perform dimensionality reduction projection processing on the temporal feature data to obtain a two-dimensional feature map, and perform hierarchical layout processing on the correlation feature data to obtain a layout structure diagram;

[0093] (8) Perform dynamic rendering processing on the fusion feature data to obtain an interactive view, and perform scene synthesis processing on the two-dimensional feature map, layout structure diagram, and interactive view to obtain a visualization interface.

[0094] Specifically, perform density clustering processing on the target fusion features. Density clustering processing is a method of clustering based on the density distribution of data points in the feature space. The core is to calculate the local density of each data point and the distance to high-density points. By setting a density threshold (the local density is greater than 1.5 times the average density) and a distance threshold (the distance is less than 0.5 times the average distance in the feature space), the density center points are identified to generate a feature density map. Task boundary detection processing determines the boundaries of different task categories by analyzing the density jump regions in the density map to form an initial task sequence. Task relevance calculation is to quantitatively analyze the association degree between each task in the initial task sequence. The cosine similarity is used to calculate the similarity between task feature vectors to generate a relevance vector. Hierarchical clustering processing clusters the relevance vector in a top-down manner, setting a clustering threshold (the relevance is greater than 0.7), and combining related tasks into sub-task sequences. Each sub-task sequence represents a set of sub-tasks with similar features and goals.

[0095] The time-series feature extraction process analyzes the variation patterns of the subtask sequence in the time dimension. The sequence is segmented by a sliding window (window size of 30 time points), and the statistical features, trend features, and fluctuation features of each segment are extracted to construct a time-series feature matrix. The autocorrelation analysis process calculates the correlation coefficients of the feature sequence at different time lags to identify the periodic patterns of the sequence and obtain a periodic feature vector. The feature spectrum analysis process transforms the periodic feature vector into the frequency domain and calculates the spectrum distribution diagram through the fast Fourier transform. The principal component extraction process performs principal component analysis on the spectrum distribution diagram and retains the principal components with a cumulative contribution rate exceeding 90% to form a principal component sequence. The time-series correlation degree calculation process analyzes the correlation of the principal component sequence in the time dimension, calculates the correlation coefficients at different time scales, and generates a correlation feature group. The subtask feature mapping process maps the correlation feature group to the feature spaces of each subtask to construct a set of mapping matrices.

[0096] The multi-task feature fusion process performs weighted combination on the features in the set of mapping matrices. The weights are determined based on task importance and feature significance to generate a fused feature vector. The feature reconstruction process converts the fused feature vector into an interpretable data form to obtain the target extraction data containing time-series feature data, correlation feature data, and fused feature data.

[0097] Taking the production line monitoring of an intelligent factory as an example, the input data includes continuous monitoring records of equipment operation parameters (100 parameters per second), quality inspection data (10 indicators per minute), and environmental monitoring data (1 set per 10 seconds). Density clustering analysis finds three main task density centers, corresponding to equipment maintenance, quality control, and environmental management tasks respectively. Task correlation analysis shows that the correlation between equipment maintenance and quality control is 0.85, and the correlation between quality control and environmental management is 0.72. Based on this, the tasks are decomposed into 8 subtask sequences. Time-series feature analysis shows that the equipment parameters have a periodic fluctuation every 8 hours (corresponding to the production shift), the quality indicators have a fluctuation every 2 hours (corresponding to the product batch), and the environmental parameters show a daily and nightly variation pattern every 24 hours. Spectrum analysis extracts the frequency components corresponding to these three main periods, and principal component analysis retains 5 main components, explaining 92% of the data variance.

[0098] Correlation analysis found significant correlations between device temperature and product quality, and between environmental humidity and product performance (correlation coefficients are 0.82 and 0.75 respectively). Feature mapping converts these correlation relationships into a 64×64 mapping matrix, and generates a 128-dimensional feature vector through feature fusion. Perform t-SNE dimensionality reduction on the time-series feature data to generate a 2D feature scatter plot; construct a hierarchical force-directed graph for the correlation feature data to display the correlation relationships between features; perform dynamic rendering on the fused feature data to generate an interactive heat map. Integrate these three views into a unified visualization interface to achieve intuitive display and interactive analysis of multi-dimensional data.

[0099] The above described the multi-granularity extraction and enhancement method for multi-source heterogeneous data in the embodiments of the present application. Next, the multi-granularity extraction and enhancement device for multi-source heterogeneous data in the embodiments of the present application will be described. Please refer to Figure 2 One embodiment of the multi-granularity extraction and enhancement device for multi-source heterogeneous data in the embodiments of the present application includes:

[0100] A processing module 201, configured to preprocess multi-source heterogeneous data to obtain a multi-modal data set and original feature parameters, where the original feature parameters include: content feature values, distribution feature values, and quality feature values;

[0101] An extraction module 202, configured to perform multi-level feature extraction on the multi-modal data set and the original feature parameters to obtain a multi-modal feature set;

[0102] An analysis module 203, configured to perform correlation analysis on the multi-modal feature set to obtain a feature correlation matrix and an optimized feature representation;

[0103] A migration module 204, configured to perform cross-modal knowledge migration processing on the feature correlation matrix and the optimized feature representation to obtain migration feature parameters;

[0104] A fusion module 205, configured to perform dynamic fusion processing on the migration feature parameters to obtain target fusion features, where the target fusion features include: underlying fusion features, middle-level fusion features, and high-level fusion features;

[0105] A segmentation module 206, configured to perform multi-task segmentation processing according to the target fusion features to obtain multiple sub-task sequences, and perform time-series enhancement and multi-task fusion analysis on each sub-task sequence respectively to obtain target extraction data, and perform visual display on the target extraction data.

[0106] Through the collaborative cooperation of the above-mentioned various components, by preprocessing multi-source heterogeneous data, a multi-modal dataset and original feature parameters are obtained, effectively solving the problems of inconsistent data formats and non-standard feature expressions, laying a foundation for subsequent feature extraction. Multi-level feature extraction is performed on the multi-modal dataset and original feature parameters, realizing the mining of data features from different granularities and levels, and improving the integrity and accuracy of feature expressions. In the correlation analysis stage, by generating a feature correlation matrix and optimizing feature representations, the internal relationships between different features are fully explored, enhancing the expressive power of features. The cross-modal knowledge transfer processing link realizes knowledge sharing and transfer between different modalities by obtaining transfer feature parameters, improving the utilization efficiency of information between modalities. The target fusion features generated in the dynamic fusion processing stage include low-level fusion features, middle-level fusion features, and high-level fusion features, constructing a complete feature hierarchy and enhancing the richness of feature expressions. Finally, multiple sub-task sequences are obtained through multi-task segmentation processing, and temporal enhancement and multi-task fusion analysis are performed, realizing the efficient decomposition and collaborative processing of tasks. At the same time, the target extraction data is visually displayed, improving the intuitiveness and interpretability of data analysis. Through the organic combination of multiple links, the entire solution forms a complete multi-source heterogeneous data processing framework, not only improving the accuracy and efficiency of feature extraction, but also enhancing the knowledge sharing ability between modalities, realizing the collaborative processing of multi-tasks, and ultimately achieving the unity of high efficiency, accuracy, and reliability in data analysis and processing.

[0107] Based on the same inventive concept, the embodiments of the present application also provide an electronic device. Referring to Figure 3 As shown, it is a schematic structural diagram of the electronic device 300 provided by the embodiments of the present application, including a processor 301, a memory 302, and a bus 303. Among them, the memory 302 is used to store execution instructions, including an internal memory 3021 and an external memory 3022; the internal memory 3021 here is also called the main memory, which is used to temporarily store the operation data in the processor 301 and the data exchanged with the external memory 3022 such as a hard disk. The processor 301 exchanges data with the external memory 3022 through the internal memory 3021. When the electronic device 300 runs, the processor 301 communicates with the memory 302 through the bus 303.

[0108] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described system, system, and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0109] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0110] As described above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application.

Claims

1. A multi-granularity extraction and enhancement method for multi-source heterogeneous data, characterized in that, The multi-granularity extraction and enhancement method for multi-source heterogeneous data includes: preprocessing the multi-source heterogeneous data to obtain a multi-modal data set and original feature parameters, where the original feature parameters include: content feature values, distribution feature values, and quality feature values; Performing multi-level feature extraction on the multi-modal data set and the original feature parameters to obtain a multi-modal feature set; Performing correlation analysis on the multi-modal feature set to obtain a feature correlation matrix and an optimized feature representation; Performing cross-modal knowledge transfer processing on the feature correlation matrix and the optimized feature representation to obtain transfer feature parameters, specifically including: performing matrix decomposition processing on the feature correlation matrix to obtain associated feature components, and performing probability distribution mapping processing on the associated feature components to obtain a feature distribution sequence; performing entropy value calculation processing on the feature distribution sequence to obtain an information entropy matrix, and performing feature probability processing on the optimized feature representation to obtain a conditional probability vector; performing mutual information calculation processing on the information entropy matrix and the conditional probability vector to obtain a knowledge transfer matrix, and performing feature weight assignment processing on the knowledge transfer matrix to obtain a weight coefficient sequence; performing marginal probability calculation processing on the weight coefficient sequence to obtain a probability distribution graph, and performing distribution consistency processing on the probability distribution graph to obtain a transfer probability sequence; performing feature correlation processing on the transfer probability sequence to obtain a correlation matrix, and performing threshold screening processing on the correlation matrix to obtain a screened feature sequence; performing transfer loss calculation processing on the screened feature sequence to obtain a loss function vector, and performing parameter optimization processing on the loss function vector and the weight coefficient sequence to obtain an optimized parameter group; performing feature enhancement processing on the optimized parameter group to obtain an enhanced feature sequence, and performing parameter mapping processing on the enhanced feature sequence to obtain transfer feature parameters, where the transfer feature parameters include: feature importance parameters, structural similarity parameters, and modality conversion parameters; Performing dynamic fusion processing on the transfer feature parameters to obtain a target fusion feature, where the target fusion feature includes: a low-level fusion feature, a middle-level fusion feature, and a high-level fusion feature; Performing multi-task segmentation processing according to the target fusion feature to obtain multiple sub-task sequences, and respectively performing temporal enhancement and multi-task fusion analysis on each sub-task sequence to obtain target extraction data, and performing visual display on the target extraction data.

2. The multi-granularity extraction and enhancement method for multi-source heterogeneous data according to claim 1, characterized in that The preprocessing of the multi-source heterogeneous data to obtain a multi-modal data set and original feature parameters, where the original feature parameters include: content feature values, distribution feature values, and quality feature values, includes: performing data cleaning processing on the multi-source heterogeneous data to obtain an initial data marking result, and performing type recognition processing on the initial data marking result to obtain a data type identifier; Performing data analysis processing on the data type identifier to obtain a basic data sequence and basic feature data; Performing modality division processing on the basic data sequence to obtain a multi-modal data set, and performing content analysis processing on the basic feature data to obtain content feature values; Perform time-frequency decomposition processing on the basic feature data to obtain a multi-scale feature sequence, and perform distribution calculation processing on the multi-scale feature sequence to obtain distribution feature values; Perform quality assessment processing on the multi-scale feature sequence to obtain quality feature values, and perform combination processing on the content feature values, the distribution feature values, and the quality feature values to obtain original feature parameters.

3. The multi-granularity extraction and enhancement method for multi-source heterogeneous data according to claim 1, wherein Perform multi-level feature extraction on the multi-modal data set and the original feature parameters to obtain a multi-modal feature set, including: perform hierarchical parsing processing on the multi-modal data set to obtain initial text data, initial image data, initial video data, and initial audio data, and perform feature matching processing on the initial text data, the initial image data, the initial video data, the initial audio data, and the original feature parameters to obtain a basic feature sequence; Perform deep semantic processing on the text part of the basic feature sequence to obtain text-level features, and perform spatial structure processing on the image part of the basic feature sequence to obtain image-level features; Perform spatio-temporal change processing on the video part of the basic feature sequence to obtain video-level features, and perform spectral analysis processing on the audio part of the basic feature sequence to obtain audio-level features; Perform feature integration processing on the text-level features, the image-level features, the video-level features, and the audio-level features to obtain a multi-level feature matrix, and perform feature screening processing on the multi-level feature matrix to obtain a target feature vector; Perform feature recombination processing on the target feature vector to obtain a multi-modal feature set.

4. The multi-granularity extraction and enhancement method for multi-source heterogeneous data according to claim 3, characterized in that Perform correlation analysis on the multi-modal feature set to obtain a feature correlation matrix and an optimized feature representation, including: perform semantic correlation processing on the text-level features to obtain a text correlation vector, and perform spatial correlation processing on the image-level features to obtain an image correlation vector; Perform temporal correlation processing on the video-level features to obtain a video correlation vector, and perform spectral correlation processing on the audio-level features to obtain an audio correlation vector; Perform cross-modal matching processing on the text correlation vector and the image correlation vector to obtain a text-image correlation matrix, and perform cross-modal matching processing on the video correlation vector and the audio correlation vector to obtain a video-audio correlation matrix; Perform matrix fusion processing on the text-image correlation matrix and the video-audio correlation matrix to obtain a feature correlation matrix, and perform feature extraction processing on the feature correlation matrix to obtain an optimized feature representation.

5. The multi-granularity extraction and enhancement method for multi-source heterogeneous data according to claim 1, wherein, Perform dynamic fusion processing on the migration feature parameters to obtain a target fusion feature, where the target fusion feature includes: a low-level fusion feature, a middle-level fusion feature, and a high-level fusion feature, including: perform feature grouping processing on the feature importance parameters to obtain a low-level feature group, and perform edge feature extraction processing on the low-level feature group to obtain a low-level fusion feature; Perform semantic correlation processing on the structural similarity parameters to obtain a middle-level feature group, and perform content feature correlation processing on the middle-level feature group to obtain a middle-level fusion feature; Perform high-dimensional mapping processing on the modal conversion parameters to obtain a high-level feature group, and perform context feature association processing on the high-level feature group to obtain high-level fusion features; Perform feature level division processing on the low-level fusion features, the middle-level fusion features, and the high-level fusion features to obtain a multi-level feature sequence, and perform inter-layer relationship analysis processing on the multi-level feature sequence to obtain an inter-layer feature vector; Perform adaptive adjustment processing on the inter-layer feature vector to obtain an adjustment parameter matrix, and perform feature combination processing on the adjustment parameter matrix and the multi-level feature sequence to obtain target fusion features.

6. The multi-granularity extraction and enhancement method for multi-source heterogeneous data according to claim 1, characterized in that Perform multi-task segmentation processing according to the target fusion features to obtain multiple sub-task sequences, and perform temporal enhancement and multi-task fusion analysis on each sub-task sequence respectively to obtain target extraction data, and perform visual display on the target extraction data, including: performing density clustering processing on the target fusion features to obtain a feature density map, and performing task boundary detection processing on the feature density map to obtain an initial task sequence; Perform task relevance calculation processing on the initial task sequence to obtain a relevance vector, and perform hierarchical clustering processing on the relevance vector to obtain multiple sub-task sequences; Perform temporal feature extraction processing on the multiple sub-task sequences to obtain a temporal feature matrix, and perform autocorrelation analysis processing on the temporal feature matrix to obtain a periodic feature vector; Perform feature spectrum analysis processing on the periodic feature vector to obtain a spectrum distribution map, and perform principal component extraction processing on the spectrum distribution map to obtain a principal component sequence; Perform temporal correlation degree calculation processing on the principal component sequence to obtain a correlation feature group, and perform sub-task feature mapping processing on the correlation feature group to obtain a mapping matrix set; Perform multi-task feature fusion processing on the mapping matrix set to obtain a fusion feature vector, and perform feature reconstruction processing on the fusion feature vector to obtain target extraction data, where the target extraction data includes: temporal feature data, correlation feature data, and fusion feature data; Perform dimensionality reduction projection processing on the temporal feature data to obtain a two-dimensional feature map, and perform hierarchical layout processing on the correlation feature data to obtain a layout structure diagram; Perform dynamic rendering processing on the fusion feature data to obtain an interactive view, and perform scene synthesis processing on the two-dimensional feature map, the layout structure diagram, and the interactive view to obtain a visual interface.

7. A multi-granularity extraction and enhancement device for multi-source heterogeneous data, which is used to implement the multi-granularity extraction and enhancement method for multi-source heterogeneous data as described in any one of claims 1 to 6, characterized in that The multi-granularity extraction and enhancement device for multi-source heterogeneous data includes: a processing module for preprocessing multi-source heterogeneous data to obtain a multi-modal data set and original feature parameters, where the original feature parameters include: content feature values, distribution feature values, and quality feature values; An extraction module for performing multi-level feature extraction on the multi-modal data set and the original feature parameters to obtain a multi-modal feature set; An analysis module for performing correlation analysis on the multi-modal feature set to obtain a feature correlation matrix and an optimized feature representation; A migration module for performing cross-modal knowledge migration processing on the feature association matrix and the optimized feature representation to obtain migration feature parameters; A fusion module for performing dynamic fusion processing on the migration feature parameters to obtain target fusion features, where the target fusion features include: low-level fusion features, middle-level fusion features, and high-level fusion features; A segmentation module for performing multi-task segmentation processing based on the target fusion features to obtain multiple sub-task sequences, and respectively performing temporal enhancement and multi-task fusion analysis on each sub-task sequence to obtain target extraction data, and visually displaying the target extraction data.

8. A computer device, characterized in that, Comprising: A processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the multi-granularity extraction and enhancement method for multi-source heterogeneous data as described in any one of claims 1 to 6 are executed.

Citation Information

Patent Citations

  • Multi-modal data fusion method and device based on tensor and mutual information

    CN116975776A

  • Underwater multi-modal target segmentation method based on mask complementary cross-layer fusion

    CN117975000A