Multimodal large model cross-modal feature alignment optimization method

By decomposing and inverse mapping the multimodal data of industrial equipment, a time-frequency co-occurrence feature tensor and an inverse mapping topological weight matrix are generated. This solves the problem of long-tail modal feature data being submerged in a unified feature space, realizes deep coupling and alignment between mainstream modalities and long-tail modalities, and improves the generalization ability and semantic consistency of cross-modal feature alignment.

CN121365239BActive Publication Date: 2026-02-17ZHONGSHU (XIAMEN) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511935311.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-02-17
Estimated Expiration
2045-12-22

AI Technical Summary

Technical Problem

Existing technologies cannot effectively capture the core information of long-tail modal feature data when processing multimodal data from industrial monitoring equipment. This results in the long-tail modal feature data being overwhelmed by mainstream modal data in a unified feature space, reducing the generalization ability and semantic consistency of cross-modal feature alignment.

Method used

By decomposing the multimodal data of industrial equipment, a time-frequency co-occurrence feature tensor and an inverse mapping topological weight matrix are generated. Combined with fractal-enhanced long-tailed feature manifolds, cross-modal pre-cooperative feature vectors are generated, achieving deep coupling and alignment between mainstream modes and long-tailed modes.

Benefits of technology

It improves the generalization ability and semantic consistency of data processing when aligning cross-modal features, strengthens the stability and structural regularity of feature representation, and improves the accuracy and collaborative processing effect of multimodal data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365239B_ABST
    Figure CN121365239B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses a multimodal large model cross-modal feature alignment optimization method, which comprises the following steps: acquiring multimodal data of an industrial device, the multimodal data being divided into long-tail minority modal data and mainstream modal data, performing decomposition processing on the long-tail minority modal data, and generating a time-frequency co-occurrence feature tensor; the long-tail minority modal data is a vibration signal of the industrial device; through the generated multimodal alignment feature vector, the marginalization problem that long-tail modal feature data is submerged by mainstream modal data in a unified feature space is avoided; the generalization ability and semantic consistency of data processing during cross-modal feature alignment are improved; the collaborative processing precision of industrial multimodal data is optimized; meanwhile, the stability, structural regularity and core recognition degree of feature expression are strengthened; the accuracy and structural level of multimodal data alignment are ensured; and the multimodal data processing effect in an industrial monitoring scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically to a method for cross-modal feature alignment optimization of multimodal large models. Background Technology

[0002] At the data processing level, cross-modal feature alignment optimization of large multimodal models mainly focuses on batch pre-training and mapping of feature data from mainstream modalities such as text, image, and speech. A unified feature data space is constructed through iterative training of large-scale mainstream modal data. For feature data from niche modalities such as EEG signals and industrial equipment vibration signals, only a shallow projection layer is used for simple dimensional transformation and mapping.

[0003] However, the above data processing methods still have the following shortcomings when processing data from industrial monitoring equipment: When it is necessary to process multimodal data such as production video frames, equipment log text, and equipment vibration signals simultaneously, the existing processing methods for aligning long-tailed, non-mainstream modal data such as vibration signals usually convert the vibration signals into frequency domain feature data through fast Fourier transform, and then use only a shallow linear projection layer or a simple fully connected network to perform direct dimensional transformation and mapping processing. This results in the inability to capture the core information in the frequency domain data, which in turn causes the long-tailed modal feature data to be submerged by the mainstream modal data in the unified feature space, resulting in the marginalization of mapping processing. This leads to poor data processing results and reduces the generalization ability and semantic consistency of data processing when aligning cross-modal features. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method for cross-modal feature alignment optimization of multimodal large models, thus solving the aforementioned problems.

[0005] The above-mentioned technical objective of the present invention is achieved through the following technical solution:

[0006] Multimodal large model cross-modal feature alignment optimization methods include:

[0007] Step S1: Obtain multimodal data of industrial equipment. The multimodal data is divided into long-tail niche modal data and mainstream modal data. The long-tail niche modal data is decomposed and processed to generate time-frequency co-occurrence feature tensors. The long-tail niche modal data is the vibration signal of industrial equipment.

[0008] Step S2: Perform modal decoupling and adaptation inverse mapping on the mainstream modal data, calculate the inverse mapping weights from the mainstream modality to the long-tail modality space, and generate the inverse mapping topology weight matrix. The mainstream modal data includes: production videos of industrial equipment and log text of industrial equipment.

[0009] Step S3: Analyze the time-frequency co-occurrence feature tensor to generate a fractal-enhanced long-tailed feature manifold;

[0010] Step S4: Alignment analysis is performed on the inverse mapping topological weight matrix and the fractal-enhanced long-tailed feature manifold to generate cross-modal pre-cooperative feature vectors;

[0011] Step S5: Calculate the cross-modal pre-cooperative feature vector to generate a multimodal aligned feature vector.

[0012] Furthermore, the long-tail niche modality data is decomposed to generate a time-frequency co-occurrence feature tensor, including:

[0013] The dynamic trajectory of the preprocessed long-tail niche modal data in high-dimensional phase space is analyzed to obtain the trajectory topology tensor.

[0014] Analyze the structural consistency of the trajectory topology tensor under changing conditions and generate a structural stability scalar.

[0015] Furthermore, the long-tail niche modality data is decomposed to generate a time-frequency co-occurrence feature tensor, which also includes:

[0016] Stability calculations are performed on the trajectory topology tensor and the structural stability scalar to generate a stable eigenspectrum.

[0017] The eigenmodes in the stable eigenspectrum are arranged and coupled to generate a time-frequency co-occurrence feature tensor.

[0018] Furthermore, modal decoupling and adaptation inverse mapping are performed on the mainstream modal data, the inverse mapping weights from the mainstream modes to the long-tail modal space are calculated, and the inverse mapping topological weight matrix is ​​generated, including:

[0019] Analyze each frame sequence of video and each semantic segment of text in the preprocessed mainstream modal data to generate semantic quantum states;

[0020] For semantic quantum states of different modes, calculate the correlation strength and action at a distance between them to generate modal entanglement entropy.

[0021] Furthermore, modal decoupling and adaptation inverse mapping are performed on the mainstream modal data, the inverse mapping weights from the mainstream modes to the long-tail modal space are calculated, and an inverse mapping topological weight matrix is ​​generated. This also includes:

[0022] Using the long-tailed mode space as the base space and the mainstream mode space as the fiber, a fiber bundle is constructed. Based on the modal entanglement entropy, the connection form from the fiber to the base space is calculated, and the fiber bundle connection vector is generated.

[0023] The stability of the fiber bundle connection vector in the entire feature space is analyzed, and the inverse mapping topological weight matrix is ​​generated.

[0024] Furthermore, the time-frequency co-occurrence feature tensor is analyzed to generate a fractal-enhanced long-tailed feature manifold, including:

[0025] The time-frequency co-occurrence feature tensor is calculated to generate morphological gene sequences;

[0026] Causal analysis of morphogenetic sequences was performed to analyze the drastic changes in trait expression when basic genes were interfered with, and causal emergence intensity was generated.

[0027] After filtering the causal emergence intensity, the self-organization, self-replication and adaptive processes of the characteristic structure are analyzed to generate an evolutionarily stable configuration.

[0028] Furthermore, the analysis of the time-frequency co-occurrence feature tensor generates a fractal-enhanced long-tailed feature manifold, which also includes:

[0029] In an evolutionarily stable configuration, identify the region of trajectory convergence in its phase space, i.e., the strange attractor, calculate the distribution and topological structure of the strange attractor, and generate a cloud of characteristic attractors.

[0030] Using the feature-attracting subcloud as anchor points and constraints, reverse reconstruction is performed to generate a fractal-enhanced long-tailed feature manifold.

[0031] Furthermore, alignment analysis is performed on the inverse mapping topological weight matrix and the fractal-enhanced long-tailed feature manifold to generate cross-modal pre-cooperative feature vectors, including:

[0032] The structure tensor is generated by calculating the inverse mapping topological weight matrix and the local curvature information of the fractal-enhanced long-tailed feature manifold.

[0033] Using the structural tensor as the transformation core, the mainstream modal features are reconstructed on the manifold under constraints to generate a cooperative manifold.

[0034] Furthermore, alignment analysis is performed on the inverse mapping topological weight matrix and the fractal-enhanced long-tailed feature manifold to generate cross-modal pre-cooperative feature vectors, which also includes:

[0035] Calculate the information density of each point on the co-manifold and the multi-scale fractal dimension of the corresponding point on the fractal-enhanced long-tailed feature manifold to generate a resonant focusing kernel;

[0036] By applying the resonant focusing kernel to the cooperative manifold, cross-modal pre-cooperative feature vectors are generated through local feature aggregation and global pooling of the resonant focusing kernel.

[0037] Furthermore, the cross-modal pre-cooperative feature vectors are calculated to generate multimodal aligned feature vectors, including:

[0038] Multi-scale stability analysis is performed on cross-modal pre-cooperative eigenvectors to calculate their structure preservation ability in the feature space and generate a stability distribution vector.

[0039] Based on the stability distribution vector, selective enhancement and structured recombination of cross-modal pre-cooperative feature vectors are performed to generate feature crystallization nuclei;

[0040] The feature crystallization kernel is dimensionally regularized and optimized to generate multimodal aligned feature vectors.

[0041] In summary, the present invention has the following main beneficial effects:

[0042] By analyzing long-tailed niche modal data such as industrial equipment vibration signals, a time-frequency co-occurrence feature tensor is generated to deeply mine the core dynamic information of vibration signals in the frequency domain, overcoming the limitation of traditional fast Fourier transform + shallow linear projection in capturing deep features. Then, mainstream modal data such as production videos and log texts are analyzed to generate an inverse mapping topological weight matrix, accurately understand the nonlinear relationship between mainstream and long-tailed modes, and generate a fractal-enhanced long-tailed feature manifold to strengthen the structural integrity and recognizability of long-tailed features.

[0043] By generating cross-modal pre-cooperative feature vectors, deep coupling of bimodal features is achieved, and multimodal aligned feature vectors are obtained. This avoids the marginalization problem of long-tailed modal feature data being submerged by mainstream modal data in a unified feature space, improves the generalization ability and semantic consistency of data processing during cross-modal feature alignment, optimizes the accuracy of collaborative processing of industrial multimodal data, and strengthens the stability, structural regularity and core recognition of feature representation. This ensures the accuracy and structure level of multimodal data alignment and improves the multimodal data processing effect in industrial monitoring scenarios. Attached Figure Description

[0044] Figure 1 This is a flowchart illustrating the steps of the multimodal large model cross-modal feature alignment optimization method of the present invention. Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] refer to Figure 1 Multimodal large model cross-modal feature alignment optimization methods include:

[0047] Step S1: Obtain multimodal data of industrial equipment. The multimodal data is divided into long-tail niche modal data and mainstream modal data. The long-tail niche modal data is decomposed and processed to generate time-frequency co-occurrence feature tensors. The long-tail niche modal data is the vibration signal of industrial equipment.

[0048] Step S2: Perform modal decoupling and adaptation inverse mapping on the mainstream modal data, calculate the inverse mapping weights from the mainstream modality to the long-tail modality space, and generate the inverse mapping topology weight matrix. The mainstream modal data includes: production videos of industrial equipment and log text of industrial equipment.

[0049] Step S3: Analyze the time-frequency co-occurrence feature tensor to generate a fractal-enhanced long-tailed feature manifold;

[0050] Step S4: Alignment analysis is performed on the inverse mapping topological weight matrix and the fractal-enhanced long-tailed feature manifold to generate cross-modal pre-cooperative feature vectors;

[0051] Step S5: Calculate the cross-modal pre-cooperative feature vector to generate a multimodal aligned feature vector.

[0052] In one embodiment, the long-tail niche modality data is decomposed to generate a time-frequency co-occurrence feature tensor, including:

[0053] The dynamic trajectory of the preprocessed long-tailed niche modal data in the high-dimensional phase space is analyzed to obtain the trajectory topology tensor. Specifically, this includes: reconstructing the phase space of the preprocessed vibration signal, which is standardized vibration data that has undergone noise reduction, detrending and normalization.

[0054] The permutation entropy value of the preprocessed vibration signal at different delay times is calculated to form a dynamic information entropy sequence. The dynamic information entropy sequence records the process of signal complexity changing with the increase of delay time. By identifying the delay time corresponding to the first peak of the permutation entropy value, it is used as the optimal delay parameter.

[0055] The neighborhood preservation optimization algorithm is used to analyze the ability of a signal to preserve its geometric structure under different embedding dimensions. Specifically, when the signal is upgraded from the current dimension to a higher dimension, the proportion of the original neighborhood relationship between its phase space trajectory points is correctly preserved, and the neighborhood preservation rate sequence under different dimensions is obtained. When the neighborhood preservation rate exceeds 95% for the first time, the corresponding dimension is the optimal embedding dimension.

[0056] The coordinates of each phase space point are obtained. Based on the reconstructed phase space, the continuous trajectory is divided into multiple overlapping trajectory segments. Each segment contains a fixed number of phase space points. The local geometric features of each trajectory segment are calculated, including the trajectory curvature change rate and motion consistency index.

[0057] Specifically, for each internal trajectory point in the trajectory segment (excluding the start and end points of the trajectory segment), its local tangent vector is calculated. The local tangent vector is the vector pointing from the current trajectory point to the next adjacent trajectory point. For each internal trajectory point, the angle between its own tangent vector and the tangent vector of the next point is calculated, which is the angle between the continuous tangent vectors. The mean and standard deviation of the angle between the continuous tangent vectors of all internal trajectory points in the trajectory segment are calculated, and the mean and standard deviation are multiplied to obtain the curvature change intensity value. The curvature change intensity value is divided by the actual length of the trajectory segment and then multiplied by 100% to obtain the trajectory curvature change rate.

[0058] The vector pointing from the start point to the end point of a trajectory segment is taken as the overall motion direction of that segment. Next, the consistency of local motion directions is calculated. For each trajectory point in the segment, the cosine of the angle between its local tangent vector and the overall motion direction vector of the segment is calculated. The mean of the cosines of the angles between the directions of all trajectory points within the segment is taken as the direction consistency score. The standard deviation of the motion velocities (Euclidean distances between adjacent points) of all trajectory points within the segment is taken as the velocity consistency coefficient. The direction consistency score is multiplied by the velocity consistency coefficient to obtain the motion consistency index.

[0059] Based on local geometric features, a trajectory association topology network is constructed, and the dynamic similarity between different trajectory segments is calculated. Specifically, the absolute value of the difference between the trajectory curvature change rate of one trajectory segment and the trajectory curvature change rate of another trajectory segment is calculated to obtain the geometric morphology difference; the absolute value of the difference between the motion consistency index of one trajectory segment and the motion consistency index of another trajectory segment is calculated to obtain the motion pattern difference; the Euclidean distance between the geometric center point of one trajectory segment and the geometric center point of another trajectory segment in the reconstructed phase space is calculated, and the reciprocal of the Euclidean distance is normalized to the 0-1 interval to obtain the spatial proximity correlation; the spatial proximity correlation plus (1 minus the geometric morphology difference) plus (1 minus the motion pattern difference) is used to obtain the dynamic similarity.

[0060] Establish connections between trajectory segments whose dynamic similarity is higher than the average of all dynamic similarities to form an associated topology network of trajectory segments. In the associated topology network, each node represents a trajectory segment, and the edges represent the dynamic similarity between segments.

[0061] Three core parameters are calculated from the trajectory association topology network: network connectivity, node centrality distribution value, and path transmission efficiency.

[0062] The local connectivity density is obtained by calculating the ratio of the actual number of connections to the maximum number of connections for each trajectory segment. The local connectivity density of all trajectory segments is then weighted, averaged, and normalized to the 0-1 interval to obtain the network connectivity. The weight in the weighted average is the trajectory length of each trajectory segment in phase space.

[0063] Using the geometric center of the segment in phase space as the sphere center and a radius of 5% of the overall size of phase space, a multidimensional spherical neighborhood is constructed. The total number of trajectory points falling into this multidimensional spherical neighborhood is counted, including trajectory points of the current trajectory segment and other trajectory segments. The total number of trajectory points is divided by the n-dimensional volume of the spherical neighborhood (where n is the optimal embedding dimension) to obtain the trajectory point density value per unit volume. If the trajectory points contained in the spherical neighborhood come from more than 3 different trajectory segments, the trajectory point density value is multiplied by 1.2 to obtain the density of neighboring trajectory points for each trajectory segment; otherwise, the trajectory point density value is directly used as the density of neighboring trajectory points. The number of direct connections between the trajectory segment and other trajectory segments is obtained. The density of neighboring trajectory points and the number of direct connections are normalized to the 0-1 interval and then multiplied to obtain the centrality score of each trajectory segment.

[0064] The centrality score of each trajectory segment is passed to its directly connected neighbor segments. The top 20% of trajectory segments with the highest centrality scores are selected, and the mean and variance of the centrality scores of these key trajectory segments are calculated. The mean plus the variance is used to obtain the node centrality distribution value.

[0065] Calculate the shortest path length between any two connected trajectory segments, take the reciprocal of the shortest path length as the transmission efficiency of the node pair, exclude node pairs with a length less than the average trajectory segment length, and multiply the average transmission efficiency of all node pairs by 100% to obtain the path transmission efficiency.

[0066] The network connectivity, node centrality distribution value, and path propagation efficiency are aligned and superimposed according to their geometric positions in the original phase space to obtain the trajectory topology tensor.

[0067] The structural consistency of the trajectory topology tensor under change is analyzed, and a structural stability scalar is generated. Specifically, this includes: calculating the impact of the numerical changes of the three parameters in the trajectory topology tensor on the overall structure, generating independent Gaussian white noise with an amplitude of 5% of each parameter value, and superimposing it on the corresponding parameters to generate 10 sets of perturbation variants, each set of perturbation variants maintaining the original relationship of the three parameters.

[0068] For each perturbation variant, calculate the differences between its current network connectivity, node centrality distribution value, and path transit efficiency and the original element, and obtain the differences in network connectivity, node centrality distribution, and path transit efficiency, respectively.

[0069] The weighted difference value is obtained by multiplying the network connectivity difference by 0.4, the node centrality distribution difference by 0.35, and the path transmission efficiency difference by 0.25. The mean of the weighted difference values ​​of the 10 perturbation variants is calculated and normalized to the 0-1 interval, which is the structural stability scalar.

[0070] In one embodiment, the decomposition of long-tail niche modality data to generate a time-frequency co-occurrence feature tensor further includes:

[0071] The stability of the trajectory topology tensor and the structural stability scalar is calculated to generate a stable eigenspectrum. Specifically, this involves: normalizing the network connectivity, node centrality distribution value, and path propagation efficiency in the trajectory topology tensor to the 0-1 interval; multiplying these three elements to obtain a fusion value; using 10% of the total number of trajectory segments as a sliding window, calculating the local maxima and minima of the fusion value, filtering out valid extrema with a difference greater than 0.05, and excluding isolated points with no adjacent extrema in the surrounding 5 windows.

[0072] Each effective extremum and its corresponding phase space geometric coordinates are taken as an eigenmode. The amplitude of the eigenmode is the fusion value multiplied by the structural stability scalar. All eigenmodes are arranged in order of optimal embedding dimension to form a stable eigenspectrum.

[0073] The eigenmodes in the stable eigenspectrum are arranged and coupled to generate a time-frequency co-occurrence feature tensor. Specifically, this involves: sorting the eigenmodes in the stable eigenspectrum from the radial coordinate distance from the origin in phase space in ascending order to form an eigenmode sequence; for two adjacent eigenmodes in the eigenmode sequence, calculating the ratio of the amplitude of the latter eigenmode to the amplitude of the former eigenmode, multiplying the ratio by the reciprocal of the Euclidean distance between the corresponding spatial coordinates of the two modes to obtain the initial coupling coefficient; multiplying the initial coupling coefficient by an exponential function with the natural constant e as the base and the initial coupling coefficient as the exponent to obtain the coupling coefficient of adjacent modes; and arranging all the coupling coefficients of adjacent modes in order to form the time-frequency co-occurrence feature tensor.

[0074] By deeply mining the core dynamic information of long-tail niche modal data such as industrial equipment vibration signals, this method effectively solves the problem that traditional shallow mapping cannot capture core information in the frequency domain. The generated time-frequency co-occurrence feature tensor fully preserves the structural stability and semantic association of long-tail modalities, preventing such data from being submerged by mainstream modalities in a unified feature space. This improves the generalization ability and semantic consistency when aligning cross-modal features, enhances the collaborative processing effect of production video frames, equipment log text, and vibration signals, improves the marginalization problem of long-tail modal mapping, and ensures the accuracy of multimodal data processing.

[0075] In one embodiment, modal decoupling and adaptation inverse mapping are performed on the mainstream modal data, the inverse mapping weights from the mainstream mode to the long-tail modal space are calculated, and an inverse mapping topology weight matrix is ​​generated, including:

[0076] The process involves analyzing each frame sequence of video and each semantic segment of text in the preprocessed mainstream modal data to generate semantic quantum states. Specifically, for video frame sequences, the output feature map of the last convolutional layer is extracted using a pre-trained deep convolutional network. This output feature map is then globally averaged in the spatial dimension to obtain a 512-dimensional visual feature vector.

[0077] For a text paragraph, all word vectors of the last hidden layer are extracted by a language model pre-trained based on a bidirectional attention mechanism. This language model consists of an input embedding layer, at least 6 bidirectional Transformer encoder layers, and an output layer. The mean and standard deviation of these word vectors are calculated one by one along the sequence dimension. The mean vector and the standard deviation vector are concatenated and then reduced to 512 dimensions through a fully connected layer to obtain the semantic feature vector.

[0078] The visual feature vector and the semantic feature vector are normalized using the L2 norm. The normalized visual feature vector is then transposed into a column vector, and the semantic feature vector is used as a row vector. The tensor product of the two is calculated to obtain a joint matrix of 512 rows and 512 columns.

[0079] Calculate the square root of the sum of squares of all elements in the joint matrix, divide the square root by each element in the matrix, and make the Frobenius norm of the joint matrix equal to 1, thus completing the normalization process of the matrix.

[0080] The first 64 elements are extracted along the main diagonal of the normalized joint matrix. At the same time, the first 63 and the first 62 elements on the first and second secondary diagonals parallel to the main diagonal are extracted. These 189 elements are concatenated into a 189-dimensional vector in the order of main diagonal, first diagonal, and second diagonal to obtain the semantic quantum state of the video modality and the semantic quantum state of the text modality, respectively.

[0081] For semantic quantum states of different modalities, calculate the correlation strength and action at a distance between them to generate modal entanglement entropy. Specifically, this includes performing a tensor product operation on the 189-dimensional semantic quantum state of the video modality and the 189-dimensional semantic quantum state of the text modality to obtain a joint density matrix of 189 rows and 189 columns, where the row index of the joint density matrix corresponds to the state of the video modality and the column index corresponds to the state of the text modality.

[0082] The joint density matrix is ​​divided into a block matrix structure of 189 rows and 189 columns. The sum of the diagonal elements of each block matrix is ​​calculated to form a video reduced density matrix of order 189. The text reduced density matrix can be obtained by performing the same operation as above. The video reduced density matrix and the text reduced density matrix are multiplied together, and the sum of all diagonal elements of the product matrix is ​​calculated to obtain the correlation strength factor.

[0083] Keeping the block structure of the joint density matrix unchanged, the block matrix corresponding to each text state index is transposed, and the eigenvalues ​​of the transposed matrix are calculated. Negative eigenvalues ​​less than -0.01 are selected, and the absolute values ​​of these negative eigenvalues ​​are added together to obtain the action-at-distance strength. The association strength factor is multiplied by the action-at-distance strength, and the natural logarithm of the product is multiplied by -2 to obtain the modal entanglement entropy representing the nonlinear dependency between the two modes.

[0084] In one embodiment, modal decoupling and adaptation inverse mapping are performed on the mainstream modal data, the inverse mapping weights from the mainstream mode to the long-tail modal space are calculated, and an inverse mapping topology weight matrix is ​​generated. The method further includes:

[0085] Using the long-tailed mode space as the base space and the mainstream mode space as the fibers, a fiber bundle is constructed. Based on the modal entanglement entropy, the connection form from the fibers to the base space is calculated, and the fiber bundle connection vector is generated. Specifically, the following steps are taken: a base space coordinate system is constructed using the time-frequency co-occurrence feature tensor of the long-tailed mode, and a fiber coordinate system is constructed using the semantic quantum state of the video mode in the mainstream mode. Adjacent sampling points are selected from the base space, and the numerical changes of each dimension of the 189-dimensional semantic quantum state of the video mode between adjacent sampling points are calculated to obtain the fiber coordinate increment. The Euclidean distance between adjacent sampling points in the time-frequency co-occurrence feature tensor is also calculated, and this Euclidean distance is used as the base space coordinate increment.

[0086] Multiplying the fiber coordinate increment of each dimension by the square root of the modal entanglement entropy and dividing by the base space coordinate increment yields 189 basic connection components. Performing a hyperbolic tangent function nonlinear transformation on each basic connection component ensures that the output value range is between -1 and 1, thus obtaining the single-dimensional connection coefficient. Arranging the 189 single-dimensional connection coefficients in dimensional order together forms a 189-dimensional vector, which is the fiber bundle connection vector.

[0087] The stability of the fiber bundle connection vector in the entire feature space is analyzed, and the inverse mapping topological weight matrix is ​​generated. Specifically, 35 sampling center points are randomly selected in the long-tailed modal feature space, and a local spherical neighborhood is constructed with a radius of 8% of the overall diameter of the feature space.

[0088] Within each spherical neighborhood, the standard deviation of all elements in the fiber bundle connection vector is calculated. Then, three independent random perturbation vectors are generated, each with 189 dimensions, and the value of each element is randomly sampled from a Gaussian distribution with a mean of zero and a standard deviation of 0.07. These three perturbation vectors are then added to the fiber bundle connection vector to obtain three different perturbation variants. The cosine similarity between the fiber bundle connection vector and these three perturbation variants is calculated to obtain three similarities. The standard deviation of these three similarities is then calculated, which is the local stability index of the spherical neighborhood.

[0089] Calculate the mean of all 35 local stability indices, multiply the reciprocal of the mean by the single-dimensional connection coefficient of each dimension in the fiber bundle connection vector, and combine the results to obtain a 189-dimensional weight vector. Use the elements in the weight vector as diagonal elements to construct a 189-row, 189-column diagonal matrix, which is the inverse mapping topological weight matrix.

[0090] By extracting visual features from video frames and semantic features from text to generate semantic quantum states, and then calculating modal entanglement entropy through a joint density matrix, the nonlinear dependency between video and text modalities is captured, and an inverse mapping topological weight matrix is ​​generated to achieve accurate inverse mapping from the mainstream modality to the long-tail modality space. This breaks the limitation of dimensional transformation in shallow projection layers, strengthens the feature association between mainstream and long-tail modalities, avoids long-tail modality data being submerged in a unified feature space, and improves the semantic consistency and generalization ability of cross-modal feature alignment.

[0091] In one embodiment, the time-frequency co-occurrence feature tensor is analyzed to generate a fractal-enhanced long-tailed feature manifold, including:

[0092] The time-frequency co-occurrence feature tensor is calculated to generate morphological gene sequences. Specifically, the time-frequency co-occurrence feature tensor is evenly divided into 12 continuous segments along the time dimension. For each segment, its data is divided into 10 sub-intervals. The ratio of the range to the standard deviation of each sub-interval is calculated, and the natural logarithmic mean of all ratios is used as the Hurst exponent.

[0093] Six scales (2 to the power of -1, 2 to the power of -2, 2 to the power of -3, 2 to the power of -4, 2 to the power of -5, and 2 to the power of -6, i.e., scales 0.5, 0.25, 0.125, 0.0625, 0.03125, and 0.015625) were used to perform grid coverage on the fragment data: For each scale, the data space was divided into grids with side lengths equal to the current scale, the number of non-empty grids containing data points was counted, and linear regression was performed on the logarithm of the scale (base 2) and the logarithm of the number of non-empty grids (base 2). The least squares method was used to fit a straight line, and the absolute value of the slope of the fitted line is the box dimension.

[0094] Calculate the sum of squares of all data points in the segment, and then calculate the ratio of this sum of squares to the total sum of squares of all elements in the entire time-frequency co-occurrence feature tensor, which is the energy proportion;

[0095] The morphological gene value of the fragment is obtained by multiplying the Hurst index, box dimension, and energy percentage. The morphological gene values ​​of the 12 fragments are arranged in chronological order to form a morphological gene sequence.

[0096] Causal analysis is performed on morphogenetic sequences to analyze the drastic changes in eigenvalue expression when the basic gene is intervened, generating causal emergence intensity. Specifically, this includes: taking the first morphogenetic gene value in the morphogenetic sequence as the basic gene and forcibly setting its value to zero to complete the intervention; calculating the Euclidean distance between the intervened morphogenetic sequence and the original morphogenetic sequence to obtain the initial difference value; calculating the standard deviation of all 12 element values ​​of the original morphogenetic sequence; dividing the initial difference value by the standard deviation to obtain the normalized difference ratio; and multiplying the normalized difference ratio by the square root of the natural constant e to obtain the causal emergence intensity.

[0097] After screening the causal emergence intensity, the self-organization, self-replication and self-adaptation processes of the characteristic structure are analyzed to generate an evolutionarily stable configuration. Specifically, this includes: screening effective sequences with a causal emergence intensity greater than 0.5 from all morphological gene sequences; for each effective sequence, calculating the variance of the difference between adjacent elements of the effective sequence; and multiplying the variance by the reciprocal of the natural constant e to obtain the self-organization index value.

[0098] The effective sequence is divided into three consecutive segments in sequence to ensure that each segment contains the same number of morphological gene values. The square of the Pearson correlation coefficient between the first and third segments is calculated to obtain the self-replication index value.

[0099] Calculate the mean and standard deviation of all morphogenetic values ​​of the effective sequence, and construct an ideal Gaussian distribution with the same mean and standard deviation. Divide the numerical range of the effective sequence into 10 equal-width intervals, and count the frequency of the morphogenetic values ​​of the effective sequence falling into each interval as the empirical distribution. Calculate the KL divergence between the empirical distribution and the ideal Gaussian distribution, and take the negative natural logarithm of the KL divergence, that is, calculate the negative natural logarithm (logarithm with the natural constant e as the base), to obtain the adaptive index value.

[0100] The weights of the self-organization index, self-replication index, and adaptive index are set to 0.4, 0.3, and 0.3 respectively. The self-organization index, self-replication index, and adaptive index are multiplied by their respective weights and then summed. The sum is then compressed to the interval [-1, 1] using the hyperbolic tangent function to obtain the scalar stable value. The self-organization index, self-replication index, and adaptive index corresponding to each effective sequence are used as x, y, and z axis coordinates, respectively, forming a point in three-dimensional space, which is the configuration point. Combining all configuration points yields the evolutionarily stable configuration, and the scalar stable value serves as the stability score of the configuration point.

[0101] In one embodiment, analyzing the time-frequency co-occurrence feature tensor to generate a fractal-enhanced long-tailed feature manifold further includes:

[0102] In an evolutionarily stable configuration, the region where the trajectory converges in its phase space is identified, i.e., the strange attractor. The distribution and topological structure of the strange attractor are calculated, and a characteristic attractor cloud is generated. Specifically, in the three-dimensional phase space formed by the evolutionarily stable configuration, the mean of the Euclidean distance between all configuration points is calculated, and 15% of the mean is used as the radius of the spherical neighborhood. Then, a spherical neighborhood is constructed with each configuration point as the center of the sphere.

[0103] For each configuration point in space, the number of configuration points contained in its spherical neighborhood is divided by the volume of the sphere to obtain the local density; the average distance from the configuration point to its 5 nearest neighbor configuration points is calculated and used as the local clustering index; based on the configuration points corresponding to each time segment, configuration points of 12 consecutive adjacent time segments are extracted, and the standard deviation of the coordinate changes of these configuration points in the three dimensions of x, y, and z is calculated. The arithmetic mean of the three standard deviations is calculated to obtain the trajectory fluctuation index.

[0104] The attractor strength at a point is obtained by multiplying the local density, the inverse of the local aggregation degree, and the trajectory fluctuation index. Configuration points with attractor strength greater than the median attractor strength of all points are identified as singular attractors. The three-dimensional coordinates of the singular attractors are combined with the corresponding attractor strengths to form a characteristic attractor cloud.

[0105] Using the feature attractor cloud as anchor points and constraints, a reverse reconstruction is performed to generate a fractal-enhanced long-tailed feature manifold. Specifically, this involves: using the 3D coordinates of each singular attractor in the feature attractor cloud as an anchor point, and the corresponding attractor strength as a constraint weight; constructing a local reconstruction region centered on each anchor point in the long-tailed modal feature space, with the region radius being 20% ​​of the attractor strength of the anchor point; for each configuration point within the region, calculating its weighted Euclidean distance to all anchor points, with the weight being the attractor strength of the corresponding anchor point; based on these weighted Euclidean distances, using a fractal interpolation function (where the interpolation coefficients are set to the square root of the reciprocal of the weighted Euclidean distance) to calculate the enhancement value of each configuration point; and connecting the enhancement values ​​of all configuration points according to the topology of the original phase space to form a fractal-enhanced long-tailed feature manifold.

[0106] By generating morphological gene sequences using Hurst exponent, box dimension, and energy proportion, the deep fractal characteristics of time-frequency co-occurrence feature tensors are explored. The response patterns of feature interventions are captured by combining causal emergence intensity. Furthermore, evolutionarily stable configurations are constructed through self-organization, self-replication, and adaptive indices. The fractal-enhanced long-tail feature manifold fully preserves the nonlinear structure and dynamic evolution information of the long-tail mode, breaking through the limitations of traditional shallow processing in capturing core information and preventing it from being submerged by mainstream modes in a unified feature space. This improves the generalization ability and semantic consistency of cross-modal feature alignment, and enhances the stability and recognizability of feature expression.

[0107] In one embodiment, alignment analysis is performed on the inverse mapping topological weight matrix and the fractal-enhanced long-tailed feature manifold to generate cross-modal pre-cooperative feature vectors, including:

[0108] The structure tensor is generated by calculating the inverse mapping topological weight matrix and the local curvature information of the fractal-enhanced long-tailed manifold. Specifically, this involves: taking each configuration point on the fractal-enhanced long-tailed manifold as the center, calculating the eigenvalues ​​of the covariance matrix of the rate of change of the tangent vectors of all configuration points in the spherical neighborhood of each configuration point, and dividing the largest eigenvalue by the smallest eigenvalue to obtain the local curvature value; and then weighting and fusing the diagonal elements of the inverse mapping topological weight matrix with the local curvature values. Specifically, for each local curvature value, it is multiplied by the element of the corresponding dimension in the inverse mapping topological weight matrix (where, if the number of local curvature values ​​exceeds 169 dimensions, it is mapped to 169 dimensions through linear interpolation), and then multiplied by the square root of the natural constant e for scaling to obtain the weighted curvature value. All weighted curvature values ​​are arranged into a 169-dimensional vector according to the topological order of the manifold and reconstructed into a 13-row, 13-column matrix, which is the structure tensor.

[0109] Using the structural tensor as the core of the transformation, a constrained reconstruction of the mainstream modal features is performed on the manifold to generate a cooperative manifold. Specifically, the 13x13 structural tensor is expanded into a 169-dimensional vector in row-major order. The 189-dimensional semantic quantum state vector of the video modality in the mainstream modality is compressed to 169 dimensions using a piecewise linear interpolation method, where the interpolation nodes are evenly distributed on the 189-dimensional vector, generating a total of 169 interpolation points. Then, the interpolated 169-dimensional feature vector is multiplied element-wise with the 169-dimensional vector of the flattened structural tensor to obtain a 169-dimensional intermediate vector. Finally, the intermediate vector is nonlinearly transformed using the hyperbolic tangent function to output a 169-dimensional vector with a value range of [-1, 1], which is the cooperative manifold.

[0110] In one embodiment, the alignment analysis of the inverse mapping topological weight matrix and the fractal-enhanced long-tailed feature manifold to generate cross-modal pre-cooperative feature vectors further includes:

[0111] The information density of each point on the cooperative manifold and the multi-scale fractal dimension of the corresponding point on the fractal-enhanced long-tailed feature manifold are calculated to generate the resonance focusing kernel. Specifically, this involves: using a fixed orthogonal projection matrix constructed from a 64th-order Haar wavelet basis to map the 169-dimensional vector of the cooperative manifold to a 256-dimensional space; on the fractal-enhanced long-tailed feature manifold, with each configuration point as the center, calculating the direction vectors of all other configuration points in its spherical neighborhood relative to the center point, and calculating the arithmetic mean of the angles between each pair of these direction vectors to obtain the local directional divergence of that configuration point; multiplying each element of the projected 256-dimensional vector by the reciprocal of the corresponding local directional divergence, and then multiplying by (169 divided by the square root of 256), and standardizing the calculation result to make its Euclidean norm equal to 1. The resulting 256-dimensional vector is the resonance focusing kernel.

[0112] Applying a resonant focusing kernel to a cooperative manifold, cross-modal pre-cooperative feature vectors are generated through local feature aggregation and global pooling of the resonant focusing kernel. Specifically, this involves: extending the 169-dimensional vector of the cooperative manifold to 256 dimensions via linear interpolation to match the dimension of the resonant focusing kernel; performing element-wise multiplication of the extended vector with the resonant focusing kernel; and performing local feature aggregation on the product result: using every 16 consecutive dimensions as an aggregation unit, calculating the sum of squares of all element values ​​within each aggregation unit, and then multiplying each sum of squares by... The reciprocal of pi (π) yields 16 aggregated values. Multi-scale global pooling is then applied to these 16 aggregated values: they are divided into four groups of four values ​​each. The product of the arithmetic mean and standard deviation of each group is calculated, and this product is transformed using the hyperbolic tangent function and multiplied by 0.3 to obtain four primary pooling values. These four primary pooling values ​​are then mapped to 128 dimensions through a fully connected layer, and non-linearly activated using the Sigmoid function to obtain a 128-dimensional cross-modal pre-cooperative feature vector.

[0113] By fusing the inverse mapping topological weight matrix with the local curvature information of the fractal-enhanced long-tailed feature manifold to generate a structural tensor, the mainstream modal features are reconstructed using this tensor as the core constraint to obtain a collaborative manifold. Then, cross-modal pre-collaborative feature vectors are generated, and deep coupling between mainstream and long-tailed modal features is achieved. The core information of the long-tailed modality and the semantic association with the mainstream modality are fully preserved, solving the problem of long-tailed modality marginalization caused by traditional shallow processing. This improves the generalization ability and semantic consistency of cross-modal feature alignment, optimizes the collaborative processing accuracy of industrial production video frames, equipment log text and vibration signals, and ensures the alignment accuracy of multimodal data.

[0114] In one embodiment, the calculation of cross-modal pre-cooperative feature vectors to generate multimodal aligned feature vectors includes:

[0115] Multi-scale stability analysis is performed on cross-modal pre-cooperative eigenvectors to calculate their structure preservation ability in the feature space and generate a stability distribution vector. Specifically, the cross-modal pre-cooperative eigenvectors are divided into eight consecutive segments in sequence, each containing sixteen dimensions. Four analysis scales (0.8, 0.4, 0.2, 0.1) are used. For each segment at each scale, the ratio of the interquartile range to the median is calculated to obtain the scale sensitivity. The sensitivity values ​​of each consecutive segment at the four scales are multiplied consecutively to obtain the stability value. The stability values ​​of the eight consecutive segments are arranged in the original order to form an 8-dimensional stability distribution vector.

[0116] Based on the stability distribution vector, selective enhancement and structured recombination of cross-modal pre-cooperative feature vectors are performed to generate feature crystallization kernels. Specifically, this includes: multiplying each element value in the stability distribution vector by π, taking the cosine function value, and then multiplying by a coefficient of 0.5 to obtain the enhancement factor; dividing the cross-modal pre-cooperative feature vector into eight continuous 16-dimensional feature segments, where each element in the stability distribution vector corresponds to the stable value of a feature segment.

[0117] For each feature segment, its 16 dimensional elements are multiplied by the corresponding enhancement factor to enhance the feature segment. From each enhanced 16-dimensional feature segment, the top 4 elements with the largest values ​​are selected, and the arithmetic mean of these 4 elements is calculated to obtain the representative value of the feature segment.

[0118] The representative values ​​of the eight feature segments are arranged in order to form an 8-dimensional vector. The Euclidean norm of this vector is scaled to 5% of the overall diameter of the feature space. The output is a standardized 8-dimensional vector, which is the feature crystal nucleus.

[0119] The feature kernel is dimensionally regularized and optimized to generate a multimodal aligned feature vector. Specifically, this involves: calculating the average value of all elements in the feature kernel; subtracting the average value from each element value to obtain a difference vector; dividing each element in the difference vector by the standard deviation of the first four elements to complete the standardization; arranging the standardized vector in ascending order and dividing it into four equal segments; and arranging the medians of each segment in the original order to form a four-dimensional vector, which is the multimodal aligned feature vector.

[0120] By generating a stability distribution vector and combining it with an enhancement factor to selectively enhance and structurally reorganize the cross-modal pre-cooperative feature vector, a multimodal aligned feature vector is obtained. This accurately preserves the correlation between the core frequency domain information of the long-tailed mode and the semantics of the mainstream mode, solves the problem of long-tailed mode marginalization caused by traditional shallow processing, improves the generalization ability and semantic consistency of cross-modal feature alignment, enhances the stability and core recognition of feature expression, improves the accuracy of collaborative processing of industrial production video frames, equipment log text and vibration signals, and ensures the structured and accurate alignment of multimodal data.

[0121] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multi-modal large model cross-modal feature alignment optimization method, characterized in that, The method comprises the following steps: Step S1, obtaining multi-modal data of the industrial equipment, the multi-modal data being divided into long-tail small modal data and mainstream modal data, decomposing the long-tail small modal data, and generating a time-frequency co-occurrence feature tensor, comprising: analyzing the dynamic trajectory of the preprocessed long-tail small modal data in a high-dimensional phase space to obtain a trajectory topological tensor; analyzing the structural consistency of the trajectory topological tensor under changes to generate a structural stability scalar; performing stability calculation on the trajectory topological tensor and the structural stability scalar to generate a stable eigen spectrum; arranging and coupling each eigen mode in the stable eigen spectrum to generate a time-frequency co-occurrence feature tensor; the long-tail small modal data is a vibration signal of the industrial equipment; Step S2, performing modal decoupling and adaptive inverse mapping on the mainstream modal data, calculating inverse mapping weights of the mainstream modal data to the long-tail modal space, and generating an inverse mapping topological weight matrix, the mainstream modal data comprising: production videos of the industrial equipment and log texts of the industrial equipment; Step S3, analyzing the time-frequency co-occurrence feature tensor to generate a fractal-enhanced long-tail feature manifold, comprising: calculating the time-frequency co-occurrence feature tensor to generate a morphological gene sequence, specifically comprising: dividing the time-frequency co-occurrence feature tensor into 12 continuous segments, and calculating the Hurst index, box dimension and energy proportion of each segment, multiplying the Hurst index, box dimension and energy proportion to obtain the morphological gene value of the segment, arranging the morphological gene values of the 12 segments in time sequence to form a morphological gene sequence; performing causal analysis on the morphological gene sequence to analyze the degree of change in feature expression when the basic gene is intervened to generate a causal emergence intensity; after screening the causal emergence intensity, analyzing the self-organization, self-replication and self-adaptation process of the feature structure to generate an evolutionary stable configuration; in the evolutionary stable configuration, identifying the region of trajectory convergence in the phase space, i.e. a strange attractor, calculating the distribution and topological structure of the strange attractor to generate a feature attractor cloud; using the feature attractor cloud as an anchor and a constraint to perform reverse reconstruction to generate a fractal-enhanced long-tail feature manifold; Step S4, aligning and analyzing the inverse mapping topological weight matrix and the fractal-enhanced long-tail feature manifold to generate a cross-modal pre-cooperative feature vector; Step S5, calculating the cross-modal pre-cooperative feature vector to generate a multi-modal alignment feature vector.

2. The multi-modal large model cross-modal feature alignment optimization method of claim 1, wherein, The method comprises the following steps: performing modal decoupling and adaptive inverse mapping on the mainstream modal data, calculating inverse mapping weights of the mainstream modal data to the long-tail modal space, and generating an inverse mapping topological weight matrix, comprising: analyzing each frame sequence of the video and each semantic paragraph of the text in the preprocessed mainstream modal data to generate a semantic quantum state, specifically comprising: extracting a visual feature vector from the video frame sequence by using a deep convolutional network, extracting a semantic feature vector from the text paragraph by using a language model, and analyzing and processing the visual feature vector and the semantic feature vector to respectively obtain a semantic quantum state of the video modal and a semantic quantum state of the text modal; for the semantic quantum states of different modalities, calculating the correlation strength and the super-distance effect between them to generate a modal entanglement entropy.

3. The multi-modal large model cross-modal feature alignment optimization method of claim 2, wherein, The mainstream modal data is decoupled and adapted inverse mapping, the inverse mapping weight of the mainstream modal to the long tail modal space is calculated, the inverse mapping topological weight matrix is generated, and the method further comprises: The long tail modal space is taken as the base space, and the mainstream modal space is taken as the fiber to construct the fiber bundle. Based on the modal entanglement entropy, the connection form from the fiber to the base space is calculated to generate the fiber bundle connection vector. The stability of the fiber bundle connection vector in the whole feature space is analyzed to generate the inverse mapping topological weight matrix.

4. The multi-modal large model cross-modal feature alignment optimization method of claim 3, wherein, The inverse mapping topological weight matrix and the fractal enhanced long tail feature manifold are aligned and analyzed to generate the cross-modal pre-collaborative feature vector, which comprises: The inverse mapping topological weight matrix and the local curvature information of the fractal enhanced long tail feature manifold are calculated to generate the structure tensor. The structure tensor is taken as the transformation core to perform the constraint reconstruction of the mainstream modal feature on the manifold to generate the collaborative manifold.

5. The multi-modal large model cross-modal feature alignment optimization method of claim 4, wherein, The inverse mapping topological weight matrix and the fractal enhanced long tail feature manifold are aligned and analyzed to generate the cross-modal pre-collaborative feature vector, which further comprises: The information density of each point on the collaborative manifold and the multi-scale fractal dimension of the corresponding point of the fractal enhanced long tail feature manifold are calculated to generate the resonance focusing kernel. The resonance focusing kernel is applied to the collaborative manifold to generate the cross-modal pre-collaborative feature vector through local feature aggregation and global pooling of the resonance focusing kernel.

6. The multi-modal large model cross-modal feature alignment optimization method of claim 5, wherein, The cross-modal pre-collaborative feature vector is calculated to generate the multi-modal alignment feature vector, which comprises: The multi-scale stability of the cross-modal pre-collaborative feature vector is analyzed to calculate the structure retention ability thereof in the feature space to generate the stability distribution vector. Based on the stability distribution vector, the cross-modal pre-collaborative feature vector is selectively strengthened and structured to generate the feature crystallization core. The feature crystallization core is dimensionally normalized and optimized to generate the multi-modal alignment feature vector.

Citation Information

Patent Citations

  • Multimodal zero-order learning classification method and device for hyperspectral remote sensing image

    CN117746097A

  • Multi-mode communication signal intelligent identification method based on neural network

    CN120910800A