Multi-modal large model cross-modal feature alignment optimization method
By decomposing and inverse mapping the multimodal data of industrial equipment, a time-frequency co-occurrence feature tensor and an inverse mapping topological weight matrix are generated, which solves the problem of long-tailed modal data being submerged in a unified feature space and achieves more efficient cross-modal feature alignment and data processing.
Patent Information
- Application Number
- CN202511935311.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2045-12-22
AI Technical Summary
Existing technologies cannot effectively capture core frequency domain information when processing multimodal data from industrial monitoring equipment, especially long-tailed modal data such as vibration signals, resulting in poor data processing performance and insufficient generalization ability and semantic consistency.
By decomposing the multimodal data of industrial equipment, a time-frequency co-occurrence feature tensor is generated, the inverse mapping topological weight matrix and fractal-enhanced long-tailed feature manifold are calculated, and cross-modal pre-cooperative feature vectors are generated to achieve deep coupling between the mainstream mode and the long-tailed mode.
It improves the generalization ability and semantic consistency of cross-modal feature alignment, strengthens the stability and structural regularity of feature representation, and improves the accuracy and collaborative processing effect of multimodal data processing.
Smart Images

Figure CN121365239A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a multi-modal large model cross-modal feature alignment optimization method. BACKGROUND
[0002] At present, the cross-modal feature alignment optimization of the multi-modal large model in the data processing level mainly carries out batch pre-training and mapping processing on the feature data of mainstream modalities such as text, image and voice, and constructs a unified feature data space through the iterative training of large-scale mainstream modal data; for the feature data of long-tail and small modalities such as electroencephalogram signals and industrial equipment vibration signals, only a shallow projection layer is used for simple dimension conversion and mapping processing.
[0003] However, the above data processing method still has the following defects when processing the data of industrial monitoring equipment: when multiple modal data such as production video frames, device log texts and device vibration signals need to be processed at the same time, the existing processing method usually converts the vibration signal into frequency domain feature data through fast Fourier transform, and then only uses a shallow linear projection layer or a simple fully connected network to perform direct dimension conversion and mapping processing, which leads to the inability to capture the core information in the frequency domain data, and further leads to the problem that the long-tail modality feature data is submerged in the unified feature space by the mainstream modality data, resulting in edge problem in mapping processing, poor data processing effect, and reduced generalization ability and semantic consistency of data processing in cross-modal feature alignment. SUMMARY
[0004] In view of the defects of the prior art, the present application provides a multi-modal large model cross-modal feature alignment optimization method, which solves the above problems.
[0005] The above technical purpose of the present application is realized by the following technical scheme: The multi-modal large model cross-modal feature alignment optimization method comprises: Step S1, acquiring multi-modal data of an industrial equipment, the multi-modal data being divided into long-tail small modality data and mainstream modality data, decomposing the long-tail small modality data to generate a time-frequency co-occurrence feature tensor, the long-tail small modality data being a vibration signal of the industrial equipment; Step S2, modality decoupling and adaptive inverse mapping are performed on the mainstream modality data, inverse mapping weights of the mainstream modality to the long-tail modality space are calculated, an inverse mapping topological weight matrix is generated, and the mainstream modality data includes production video of the industrial equipment and log text of the industrial equipment; Step S3, analyzing the time-frequency co-occurrence feature tensor to generate a fractal enhanced long-tail feature manifold; Step S4, aligning analysis is performed on the inverse mapping topological weight matrix and the fractal enhanced long tail feature manifold to generate a cross-modal pre-cooperative feature vector; Step S5, the cross-modal pre-cooperative feature vector is calculated to generate a multi-modal alignment feature vector.
[0006] Further, the long-tail niche modal data is decomposed to generate a time-frequency co-occurrence feature tensor, including: The dynamic trajectory of the pre-processed long-tail niche modal data in the high-dimensional phase space is analyzed to obtain a trajectory topological tensor; The structural consistency of the trajectory topological tensor under changes is analyzed to generate a structural stability scalar.
[0007] Further, the long-tail niche modal data is decomposed to generate a time-frequency co-occurrence feature tensor, including: The trajectory topological tensor and the structural stability scalar are calculated to generate a stable eigen spectrum; Each eigen mode in the stable eigen spectrum is arranged and coupled to generate a time-frequency co-occurrence feature tensor.
[0008] Further, the mainstream modal data is modal decoupled and adapted to inverse mapping, the inverse mapping weight of the mainstream modal to the long-tail modal space is calculated, and an inverse mapping topological weight matrix is generated, including: The semantic quantum state is generated by analyzing each frame sequence of the video and each semantic paragraph of the text in the pre-processed mainstream modal data; The correlation strength and the hyperdistance effect between the semantic quantum states of different modalities are calculated to generate modal entanglement entropy.
[0009] Further, the mainstream modal data is modal decoupled and adapted to inverse mapping, the inverse mapping weight of the mainstream modal to the long-tail modal space is calculated, and an inverse mapping topological weight matrix is generated, including: The long-tail modal space is taken as the base space and the mainstream modal space is taken as the fiber to construct a fiber bundle. Based on the modal entanglement entropy, the connection form from the fiber to the base space is calculated to generate a fiber bundle connection vector; The stability of the fiber bundle connection vector in the entire feature space is analyzed to generate an inverse mapping topological weight matrix.
[0010] Further, the time-frequency co-occurrence feature tensor is analyzed to generate a fractal enhanced long tail feature manifold, including: The time-frequency co-occurrence feature tensor is calculated to generate a morphological gene sequence; The morphological gene sequence is causally analyzed to analyze the degree of change in feature expression when the basic gene is intervened to generate causal emergence intensity; After screening the causality emergence intensity, the self-organization, self-replication and self-adaptation process of the characteristic structure is analyzed, and an evolutionary stable configuration is generated.
[0011] Further, the time-frequency co-occurrence feature tensor is analyzed to generate a fractal enhanced long tail feature manifold, which further includes: In the evolutionary stable configuration, the area where the trajectory converges in its phase space, i.e. the strange attractor, is identified, the distribution and topological structure of the strange attractor are calculated, and a characteristic attractor cloud is generated. The characteristic attractor cloud is used as an anchor point and constraint for reverse reconstruction to generate a fractal enhanced long tail feature manifold.
[0012] Further, the inverse mapping topological weight matrix and the fractal enhanced long tail feature manifold are aligned and analyzed to generate a cross-modal pre-cooperative feature vector, including: The structure tensor is generated by calculating the local curvature information of the inverse mapping topological weight matrix and the fractal enhanced long tail feature manifold. The structure tensor is used as a transformation core to perform constraint reconstruction on the main flow mode feature on the manifold to generate a cooperative manifold.
[0013] Further, the inverse mapping topological weight matrix and the fractal enhanced long tail feature manifold are aligned and analyzed to generate a cross-modal pre-cooperative feature vector, including: The resonance focusing kernel is generated by calculating the information density of each point on the cooperative manifold and the multi-scale fractal dimension of the corresponding point on the fractal enhanced long tail feature manifold. The resonance focusing kernel acts on the cooperative manifold, and through local feature aggregation and global pooling of the resonance focusing kernel, a cross-modal pre-cooperative feature vector is generated.
[0014] Further, the cross-modal pre-cooperative feature vector is calculated to generate a multi-modal alignment feature vector, including: The multi-scale stability of the cross-modal pre-cooperative feature vector is analyzed, and its structure retention ability in the feature space is calculated to generate a stability distribution vector. Based on the stability distribution vector, the cross-modal pre-cooperative feature vector is selectively strengthened and structurally reorganized to generate a feature crystallization nucleus. The feature crystallization nucleus is dimensionally normalized and optimized to generate a multi-modal alignment feature vector.
[0015] In summary, the present application mainly has the following advantages: By analyzing the long-tail small-mode data such as industrial equipment vibration signal, a time-frequency co-occurrence feature tensor is generated, the frequency domain core dynamic information of the vibration signal is deeply mined, and the limitation that the traditional fast Fourier transform + shallow linear projection cannot capture deep features is broken through; then the mainstream modal data such as production video and log text are analyzed, an inverse mapping topological weight matrix is generated, the nonlinear association between mainstream and long-tail modal is accurately understood, and a fractal enhanced long-tail feature manifold is generated to strengthen the structural integrity and recognition of the long-tail feature.
[0016] By generating a cross-modal pre-collaborative feature vector, deep coupling of double-modal features is realized, and a multi-modal alignment feature vector is obtained, avoiding the marginalization problem that long-tail modal feature data is submerged by mainstream modal data in a unified feature space, improving the generalization ability and semantic consistency of data processing during cross-modal feature alignment, optimizing the collaborative processing precision of industrial multi-modal data, and at the same time strengthening the stability, structural regularity and core recognition of feature expression, ensuring the accuracy and structural level of multi-modal data alignment, and improving the multi-modal data processing effect in industrial monitoring scene. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is a multi-modal large model cross-modal feature alignment optimization method step diagram of the present application. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] Reference Figure 1 , the multi-modal large model cross-modal feature alignment optimization method comprises: Step S1, obtaining multi-modal data of industrial equipment, the multi-modal data is divided into long-tail small-mode data and mainstream modal data, the long-tail small-mode data is decomposed and processed to generate a time-frequency co-occurrence feature tensor, and the long-tail small-mode data is an industrial equipment vibration signal; Step S2, modality decoupling and inverse mapping of the mainstream modal data are performed, the inverse mapping weight of the mainstream modal to the long-tail modal space is calculated, an inverse mapping topological weight matrix is generated, and the mainstream modal data includes production video of industrial equipment and log text of industrial equipment; Step S3, analyzing the time-frequency co-occurrence feature tensor to generate a fractal enhanced long-tail feature manifold; Step S4, aligning and analyzing the inverse mapping topological weight matrix and the fractal enhanced long-tail feature manifold to generate a cross-modal pre-collaborative feature vector; Step S5, the cross-modal pre-coordination feature vector is calculated to generate a multi-modal alignment feature vector.
[0020] In one case of the embodiment, the long-tail niche modal data is decomposed to generate a time-frequency co-occurrence feature tensor, including: The dynamic trajectory of the pre-processed long-tail niche modal data in the high-dimensional phase space is analyzed to obtain a trajectory topological tensor, specifically including: reconstructing the phase space of the pre-processed vibration signal, the pre-processed signal being standardized vibration data after noise reduction, detrending and normalization processing; The permutation entropy value of the pre-processed vibration signal at different delay times is calculated to form a dynamic information entropy sequence, which records the process of signal complexity changing with the increase of delay time. The delay time corresponding to the first peak of the permutation entropy value is identified as the best delay parameter. The neighbor preserving optimization algorithm is used to analyze the geometric structure preserving ability of the signal at different embedding dimensions, specifically: calculating the proportion of the original neighborhood relationship between the trajectory points in the phase space that can be correctly preserved when the signal is promoted from the current dimension to a higher dimension, to obtain a neighborhood preserving rate sequence at different dimensions. When the neighborhood preserving rate first exceeds 95% of the neighborhood preserving rate, the corresponding dimension is the optimal embedding dimension. The coordinates of each phase space point are obtained, and the continuous trajectory is divided into multiple overlapping trajectory segments on the basis of the reconstructed phase space, each segment containing a fixed number of phase space points. The local geometric features of each trajectory segment are calculated, including the trajectory curvature change rate and the motion consistency index. Among them, for each internal trajectory point in the trajectory segment, excluding the starting point and the ending point of the trajectory segment, the local tangent vector is calculated, which is the vector from the current trajectory point to the next adjacent trajectory point. For each internal trajectory point, the angle between its own tangent vector and the tangent vector of the next point is calculated, which is the continuous tangent vector angle. The mean and standard deviation of the continuous tangent vector angle values of all internal trajectory points in the trajectory segment are calculated, and the mean is multiplied by the standard deviation to obtain the curvature change intensity value. The curvature change intensity value is divided by the actual length of the trajectory segment and multiplied by 100% to obtain the trajectory curvature change rate. The vector from the starting point to the ending point of the trajectory segment is taken as the overall motion direction of the segment. Then, the local motion direction consistency is calculated. For each trajectory point in the segment, the cosine value of the angle between its local tangent vector and the overall motion direction vector of the segment is calculated. The mean of the direction cosine values of all trajectory points in the trajectory segment is taken as the direction consistency score. The standard deviation of the motion speed (Euclidean distance between adjacent points) of all trajectory points in the trajectory segment is taken as the speed consistency coefficient. The direction consistency score is multiplied by the speed consistency coefficient to obtain the motion consistency index. Based on the local geometric features, a trajectory correlation topology network is constructed, and the dynamic similarity between different trajectory segments is calculated. Specifically, the absolute value of the difference between the trajectory curvature change rate of the trajectory segment and the trajectory curvature change rate of another trajectory segment is calculated to obtain the geometric shape difference degree; the absolute value of the difference between the motion consistency index of the trajectory segment and the motion consistency index of another trajectory segment is calculated to obtain the motion mode difference degree; the Euclidean distance between the geometric center point of the trajectory segment and the geometric center point of another trajectory segment in the reconstructed phase space is calculated, and the reciprocal of the Euclidean distance is normalized to the interval of 0-1 to obtain the spatial proximity correlation degree; the spatial proximity correlation degree + (1 minus the geometric shape difference degree) + (1 minus the motion mode difference degree) is obtained to obtain the dynamic similarity; The trajectory segments with a dynamic similarity higher than the average of all dynamic similarities are connected to form a correlation topology network of trajectory segments. In the correlation topology network, each node represents a trajectory segment, and the edge represents the dynamic similarity between the segments. Three core parameters are calculated from the trajectory correlation topology network. The three core parameters are network connectivity, node centrality distribution value, and path transmission efficiency. Among them, the ratio of the actual connection number of each trajectory segment to the maximum connection number is calculated to obtain the local connection density. After weighted averaging of the local connection densities of all trajectory segments, the local connection densities are normalized to the interval of 0-1 to obtain the network connectivity. The weight in the weighted averaging is the trajectory length of each trajectory segment in the phase space. A multi-dimensional spherical neighborhood is constructed with the geometric center point of the segment in the phase space as the center and 5% of the overall size of the phase space as the radius. The total number of trajectory points falling into the multi-dimensional spherical neighborhood is counted, including the trajectory points of the current trajectory segment and other trajectory segments. The trajectory point density value in the unit volume is obtained by dividing the total number of trajectory points by the n-dimensional volume of the spherical neighborhood (n is the optimal embedding dimension). If the trajectory points in the spherical neighborhood come from more than 3 different trajectory segments, the trajectory point density value is multiplied by 1.2 to obtain the adjacent trajectory point density of each trajectory segment. Otherwise, the trajectory point density value is directly used as the adjacent trajectory point density. The number of direct connections between the trajectory segment and other trajectory segments is obtained. The adjacent trajectory point density and the number of direct connections are normalized to the interval of 0-1 and then multiplied to obtain the centrality score of each trajectory segment. The centrality score of each trajectory segment is transmitted to its directly connected neighbor segment. The top 20% of trajectory segments with the highest centrality scores are selected, and the mean and variance of the centrality scores of these key trajectory segments are calculated. The mean + variance is obtained to obtain the node centrality distribution value. Calculate the shortest path length between any two connected trajectory segments, take the reciprocal of the shortest path length as the transmission efficiency of the node pair, exclude node pairs less than the average trajectory segment length, multiply the average of the transmission efficiency of all node pairs by 100%, and obtain the path transmission efficiency; Align and superimpose the network connectivity, node centrality distribution value and path transmission efficiency according to their geometric positions in the original phase space to obtain the trajectory topology tensor.
[0021] Analyze the structural consistency of the trajectory topology tensor under changes to generate a structural stability scalar, specifically including: calculating the influence of the numerical changes of the three parameters in the trajectory topology tensor on the overall structure, generating independent Gaussian white noise with an amplitude of 5% of each parameter value, superimposing it on the corresponding parameter to produce 10 groups of perturbed variants, and each group of perturbed variants maintains the original relationship of the three parameters; For each perturbed variant, calculate the difference between its current network connectivity, node centrality distribution value and path transmission efficiency and the original elements, respectively, to obtain the network connectivity difference value, node centrality distribution difference value and path transmission efficiency difference value; Set the network connectivity difference value multiplied by 0.4 plus the node centrality distribution difference value multiplied by 0.35 plus the path transmission efficiency difference value multiplied by 0.25 to obtain the weighted difference value, calculate the average of the weighted difference values of the 10 perturbed variants, and normalize the average to the 0-1 interval to obtain the structural stability scalar.
[0022] In one case of the embodiment, the long-tail niche modal data is decomposed and processed to generate a time-frequency co-occurrence feature tensor, which further includes: Perform stability calculation on the trajectory topology tensor and the structural stability scalar to generate a stable eigen spectrum, specifically including: normalizing the network connectivity, node centrality distribution value and path transmission efficiency in the trajectory topology tensor to the 0-1 interval; and multiplying the network connectivity, node centrality distribution value and path transmission efficiency to obtain a fusion value; take 10% of the total number of trajectory segments as a sliding window, calculate the local maximum and minimum of the fusion value, filter the effective extreme values with a difference greater than 0.05 between the extreme values, and exclude isolated points without adjacent extreme values in the surrounding 5 windows; Take each effective extreme value and its corresponding phase space geometric coordinates as an eigenmode, and the amplitude of the eigenmode is the fusion value multiplied by the structural stability scalar; arrange all eigenmodes in the order of the best embedding dimension to form a stable eigen spectrum.
[0023] The stable intrinsic modes in the stable intrinsic spectrum are arranged and coupled to generate a time-frequency co-occurrence feature tensor, specifically including: according to the distance of each intrinsic mode in the stable intrinsic spectrum from the origin in the radial coordinate of the phase space, sorting from small to large to form an intrinsic mode sequence, for two adjacent intrinsic modes in the intrinsic mode sequence, calculating the ratio of the amplitude of the latter intrinsic mode to the amplitude of the former intrinsic mode, multiplying the ratio by the reciprocal of the Euclidean distance between the two mode corresponding spatial coordinate points to obtain the initial coupling coefficient; multiplying the initial coupling coefficient by the exponential function value of the initial coupling coefficient with the natural constant e as the base to obtain the adjacent mode coupling coefficient, arranging all the adjacent mode coupling coefficients in order to form a time-frequency co-occurrence feature tensor.
[0024] By deeply mining the core dynamic information of long-tail small-mode data such as industrial equipment vibration signals, the defect that traditional shallow mapping cannot capture the core information in the frequency domain is effectively solved. The time-frequency co-occurrence feature tensor generated by the method retains the structural stability and semantic association of the long-tail mode, avoids the long-tail mode being submerged by the mainstream mode in the unified feature space, improves the generalization ability and semantic consistency when aligning cross-modal features, strengthens the collaborative processing effect of production video frames, device log texts and vibration signals, improves the marginalization problem of long-tail mode mapping, and ensures the accuracy of multi-modal data processing.
[0025] In one case of the embodiment, the mainstream mode data is decoupled and adapted inverse mapping, the inverse mapping weight of the mainstream mode to the long-tail mode space is calculated, and an inverse mapping topology weight matrix is generated, including: The preprocessed video frame sequence and text semantic paragraph in the mainstream mode data are analyzed to generate a semantic quantum state, specifically including: for the video frame sequence, using a pre-trained deep convolutional network to extract the output feature map of the last convolutional layer, performing global average pooling on the output feature map in the spatial dimension to obtain a 512-dimensional visual feature vector; For the text paragraph, all word vectors of the last hidden layer are extracted by a language model pre-trained based on a bidirectional attention mechanism, wherein the language model is composed of an input embedding layer, at least 6 bidirectional Transformer encoder layers and an output layer; the mean and standard deviation of each word vector along the sequence dimension are calculated, and the mean vector and the standard deviation vector are concatenated and then reduced to 512 dimensions by a fully connected layer to obtain a semantic feature vector; The visual feature vector and the semantic feature vector are respectively subjected to L2 norm normalization processing, and the normalized visual feature vector is transposed into a column vector, and the semantic feature vector is taken as a row vector, and the tensor product of the two is calculated to obtain a 512x512 joint matrix; The square root of the sum of squares of all elements in the joint matrix is calculated, and the square root is divided by each element in the matrix, so that the Frobenius norm of the joint matrix is equal to 1, and the normalization of the matrix is completed; The first 64 elements along the main diagonal of the normalized joint matrix are extracted, and the first 63 and 62 elements along the first and second sub-diagonals parallel to the main diagonal are extracted, and the 189 elements are spliced into a 189-dimensional vector in the order of the main diagonal, the first diagonal and the second diagonal to obtain the semantic quantum state of the video mode and the semantic quantum state of the text mode.
[0026] For the semantic quantum states of different modes, the correlation strength and the hyperdistance effect between them are calculated to generate modal entanglement entropy, which specifically includes: performing a tensor product operation on the 189-dimensional semantic quantum state of the video mode and the 189-dimensional semantic quantum state of the text mode to obtain a joint density matrix of 189 rows and 189 columns, wherein the row index of the joint density matrix corresponds to the state of the video mode, and the column index corresponds to the state of the text mode; The joint density matrix is divided into a block matrix structure of 189 rows and 189 columns, and the sum of the diagonal elements of each block matrix is calculated to form a 189-order video reduced density matrix. In the same way, the text reduced density matrix can be obtained. The video reduced density matrix and the text reduced density matrix are multiplied, and the sum of all diagonal elements of the product matrix is calculated to obtain the correlation strength factor; The block structure of the joint density matrix is kept unchanged, the transpose operation is performed on the block matrix corresponding to each text state index, and the eigenvalues of the transposed matrix are calculated, and the negative eigenvalues less than-0.01 are screened out. The absolute values of these negative eigenvalues are added to obtain the hyperdistance effect strength. The correlation strength factor is multiplied by the hyperdistance effect strength, and the natural logarithm of the product is multiplied by-2 to obtain the modal entanglement entropy representing the nonlinear dependence between the two modes.
[0027] In one case of the embodiment, the modal decoupling and adaptive inverse mapping of the mainstream modal data are performed, the inverse mapping weight of the mainstream modal space to the long-tail modal space is calculated, and an inverse mapping topological weight matrix is generated, which further includes: The long-tail modal space is taken as the base space, and the mainstream modal space is taken as the fiber to construct a fiber bundle. Based on the modal entanglement entropy, the connection form from the fiber to the base space is calculated to generate a fiber bundle connection vector, which specifically includes: constructing a base space coordinate system with the time-frequency co-occurrence feature tensor of the long-tail modal, constructing a fiber coordinate system with the video mode semantic quantum state in the mainstream modal, selecting adjacent sampling points from the base space, calculating the numerical change of each dimension of the 189-dimensional semantic quantum state of the video mode between the adjacent sampling points to obtain the fiber coordinate increment, and calculating the Euclidean distance between the adjacent sampling points in the time-frequency co-occurrence feature tensor as the base space coordinate increment. After multiplying the fiber coordinate increment of each dimension with the square root of the modal entanglement entropy and dividing by the base space coordinate increment, 189 basic connection components are obtained. After hyperbolic tangent function nonlinear transformation is performed on each basic connection component, the output value range is ensured to be between-1 and 1, that is, a single-dimensional connection coefficient is obtained. The 189 single-dimensional connection coefficients are arranged in dimension order to form a 189-dimensional vector, that is, the fiber bundle connection vector.
[0028] The stability of the fiber bundle connection vector in the entire feature space is analyzed to generate an inverse mapping topological weight matrix, specifically including: 35 sampling center points are randomly selected in the long tail modal feature space, and a local spherical neighborhood is constructed with a radius of 8% of the overall diameter of the feature space; In each spherical neighborhood, the standard deviation of all elements in the fiber bundle connection vector is calculated, and then three independent random perturbation vectors are generated, each of which is 189-dimensional, and each element in the vector is randomly sampled from a Gaussian distribution with a mean of zero and a standard deviation of 0.07 standard deviation. The three perturbation vectors are added to the fiber bundle connection vector respectively, thereby obtaining three different perturbed variants. The cosine similarity between the fiber bundle connection vector and the three perturbed variants is calculated respectively to obtain three similarities, and the standard deviation of the three similarities is calculated, which is the local stability indicator of the spherical neighborhood. The mean of all 35 local stability indicators is calculated, the reciprocal of the mean is multiplied by the single-dimensional connection coefficients of each dimension in the fiber bundle connection vector, and the calculation results are combined to obtain a 189-dimensional weight vector. The elements in the weight vector are used as diagonal elements to construct a 189x189 diagonal matrix, which is the inverse mapping topological weight matrix.
[0029] By extracting video frame visual features and text semantic features to generate semantic quantum states, and then calculating modal entanglement entropy by joint density matrix to capture the nonlinear dependence between video and text modal, and generating an inverse mapping topological weight matrix, precise inverse mapping of mainstream modal to long tail modal space is realized, breaking the dimensional conversion limitation of the shallow projection layer, strengthening the feature association of mainstream and long tail modal, avoiding the long tail modal data from being submerged in the unified feature space, and improving the semantic consistency and generalization ability of cross-modal feature alignment.
[0030] In one case of the embodiment, the time-frequency co-occurrence feature tensor is analyzed to generate a fractal enhanced long tail feature manifold, including: The time-frequency co-occurrence feature tensor is calculated to generate a morphological gene sequence, specifically including: dividing the time-frequency co-occurrence feature tensor into 12 continuous segments along the time dimension, dividing the data of each segment into 10 subintervals, calculating the ratio of the range to the standard deviation of each subinterval, and taking the natural logarithm average of all ratios as the Hurst index. Grid covering the segment data with 6 scales (2 raised to the power of -1, 2 raised to the power of -2, 2 raised to the power of -3, 2 raised to the power of -4, 2 raised to the power of -5, 2 raised to the power of -6, i.e. scales 0.5, 0.25, 0.125, 0.0625, 0.03125, 0.015625): for each scale, divide the data space into grids with the side length of the current scale, count the number of non-empty grids containing data points, linearly regress the logarithm (base 2) of the scale and the logarithm (base 2) of the number of non-empty grids, and use the least squares method to fit a straight line, and the absolute value of the slope of the fitted line is the box dimension; Calculate the sum of squares of all data points in the segment, then calculate the ratio of the sum of squares to the total sum of squares of the elements in the entire time-frequency co-occurrence feature tensor, which is the energy proportion; Multiply the Hurst index, box dimension and energy proportion to get the morphological gene value of the segment, arrange the morphological gene values of the 12 segments in chronological order to form a morphological gene sequence.
[0031] Perform causal analysis on the morphological gene sequence to analyze the degree of change in feature expression when the basic gene is intervened, and generate a causal emergence intensity, which specifically includes: taking the first morphological gene value in the morphological gene sequence as the basic gene, and forcibly setting its value to zero to complete the intervention operation; calculate the Euclidean distance between the intervened morphological gene sequence and the original morphological gene sequence to get the initial difference value; calculate the standard deviation of all 12 elements in the original morphological gene sequence; divide the initial difference value by the standard deviation to get the normalized difference ratio; multiply the normalized difference ratio by the square root of the natural constant e to get the causal emergence intensity.
[0032] After screening the causal emergence intensity, analyze the self-organization, self-replication and self-adaptation process of the feature structure to generate an evolutionary stable configuration, which specifically includes: selecting effective sequences with a causal emergence intensity greater than 0.5 from all morphological gene sequences, and for each effective sequence, calculating the variance of the difference between adjacent elements in the effective sequence, and multiplying the variance by the inverse of the natural constant e to get a self-organization index value; Divide each effective sequence into three consecutive subsegments in order to ensure that each subsegment contains the same number of morphological gene values, and calculate the square of the Pearson correlation coefficient between the first subsegment and the third subsegment to get a self-replication index value; The mean and standard deviation of all morphological gene values of the effective sequence are calculated, an ideal Gaussian distribution with the same mean and standard deviation is constructed, the numerical range of the effective sequence is divided into 10 equal-width intervals, the frequency of the morphological gene value of the effective sequence falling into each interval is counted as an empirical distribution, the KL divergence between the empirical distribution and the ideal Gaussian distribution is calculated, and the negative natural logarithm of the KL divergence is calculated, that is, the natural constant e is taken as the base, to obtain the adaptive index value; The weight of the self-organization index value is set to 0.4, the weight of the self-replication index value is set to 0.3, and the weight of the adaptive index value is set to 0.3. The self-organization index value, the self-replication index value and the adaptive index value are multiplied by the corresponding weights and added, and the added result is compressed to the [-1, 1] interval by the hyperbolic tangent function to obtain the scalar stability value. The self-organization index value, the self-replication index value and the adaptive index value corresponding to each effective sequence are respectively taken as the x, y and z axis coordinates to form a point in a three-dimensional space, that is, a configuration point. All configuration points are combined to form an evolutionary stable configuration, and the scalar stability value is taken as the stability score of the configuration point.
[0033] In one case of the embodiment, the time-frequency co-occurrence feature tensor is analyzed to generate a fractal enhanced long tail feature manifold, which further includes: In the evolutionary stable configuration, the region where the trajectory converges in the phase space, that is, the strange attractor, is identified, the distribution and topological structure of the strange attractor are calculated, and a characteristic attractor cloud is generated. Specifically, in the three-dimensional phase space formed by the evolutionary stable configuration, the mean of the Euclidean distance between all configuration points is calculated, 15% of the mean is taken as the spherical neighborhood radius, and then a spherical neighborhood is constructed with each configuration point as the center of the sphere. For each configuration point in the space, the number of configuration points contained in the spherical neighborhood of the configuration point is divided by the volume of the sphere to obtain the local density. The average distance from the configuration point to its five nearest neighbors is calculated as a local aggregation index. Based on the configuration points corresponding to each time segment, the configuration points of the last 12 adjacent time segments are extracted, the standard deviations of the coordinate changes of these configuration points in the x, y and z dimensions are calculated, the arithmetic mean of the three standard deviations is calculated to obtain a trajectory fluctuation index. The local density, the reciprocal of the local aggregation degree and the trajectory fluctuation index are multiplied to obtain the attractor strength of the point. The configuration points with an attractor strength greater than the median of all point attractor strengths are taken as strange attractors. The three-dimensional coordinates of the strange attractors and the corresponding attractor strengths are combined to form a characteristic attractor cloud.
[0034] The feature attractor cloud is taken as an anchor point and a constraint for reverse reconstruction to generate a fractal enhanced long-tail feature manifold, specifically including: taking the three-dimensional coordinates of each singular attractor in the feature attractor cloud as an anchor point, and taking the corresponding attractor strength as a constraint weight, in the long-tail modal feature space, a local reconstruction region is constructed with each anchor point as the center, and the region radius is 20% of the attractor strength of the anchor point; for each configuration point in the region, the weighted Euclidean distance of each configuration point to all anchor points is calculated, and the weight is the attractor strength of the corresponding anchor point; based on the weighted Euclidean distance, the enhanced value of each configuration point is calculated using a fractal interpolation function (where the interpolation coefficient is set as the square root of the reciprocal of the weighted Euclidean distance), and the enhanced values of all configuration points are connected according to the topological structure of the original phase space to form a fractal enhanced long-tail feature manifold.
[0035] The Hurst index, the box dimension and the energy proportion are used to generate a morphological gene sequence, the deep fractal characteristics of the time-frequency co-occurrence feature tensor are mined, the response law of the feature intervention is captured in combination with the causal emergence strength, the self-organization, self-replication and self-adaptation indexes are used to construct an evolution stable configuration, the fractal enhanced long-tail feature manifold completely retains the nonlinear structure and dynamic evolution information of the long-tail mode, breaks through the core information capture limitation of traditional shallow processing, avoids the long-tail mode being submerged by the mainstream mode in a unified feature space, improves the generalization ability and semantic consistency of the cross-modal feature alignment, and strengthens the stability and recognition degree of the feature expression.
[0036] In one case of the embodiment, the inverse mapping topological weight matrix and the fractal enhanced long-tail feature manifold are subjected to alignment analysis to generate a cross-modal pre-coordinated feature vector, including: The local curvature information of the inverse mapping topological weight matrix and the fractal enhanced long-tail feature manifold is calculated to generate a structure tensor, specifically including: taking each configuration point on the fractal enhanced long-tail feature manifold as the center, for the spherical neighborhood of each configuration point, the covariance matrix eigenvalue of the change rate of the tangent vector of all configuration points in the neighborhood is calculated, and the maximum eigenvalue is divided by the minimum eigenvalue to obtain the local curvature value; the main diagonal elements of the inverse mapping topological weight matrix and the local curvature value are fused by weighting, specifically: for each local curvature value, it is multiplied by the element in the corresponding dimension of the inverse mapping topological weight matrix (wherein, if the number of local curvature values exceeds 169 dimensions, it is mapped to 169 dimensions through linear interpolation), and then multiplied by the square root of the natural constant e for scaling to obtain a weighted curvature value, all weighted curvature values are arranged in the form of a 169-dimensional vector according to the topological order of the manifold, and are reconstructed into a 13-row and 13-column matrix, that is, the structure tensor.
[0037] The structural tensor is taken as a transformation core to perform constraint reconstruction on the main stream modal feature on a manifold to generate a collaborative manifold, specifically including: expanding the 13-row 13-column structural tensor as a row priority sequence into a 169-dimensional vector, and compressing the 189-dimensional semantic quantum state vector of the video modal in the main stream modal into a 169-dimensional vector through a piecewise linear interpolation method, wherein interpolation nodes are uniformly distributed on the 189-dimensional vector, and 169 interpolation points are generated; then, the 169-dimensional feature vector after interpolation is multiplied element by element with the 169-dimensional vector of the flattened structural tensor to obtain a 169-dimensional intermediate vector; finally, the intermediate vector is nonlinearly transformed through a hyperbolic tangent function to output a 169-dimensional vector with a value range of [-1, 1], which is the collaborative manifold.
[0038] In one case of the embodiment, the inverse mapping topological weight matrix is aligned and analyzed with the fractal enhanced long-tail feature manifold to generate a cross-modal pre-collaborative feature vector, further including: The information density of each point on the collaborative manifold and the multi-scale fractal dimension of the corresponding point of the fractal enhanced long-tail feature manifold are calculated to generate a resonance focusing kernel, specifically including: using a fixed orthogonal projection matrix constructed by a 64-order Haar wavelet basis to map the 169-dimensional vector of the collaborative manifold to a 256-dimensional space, and on the fractal enhanced long-tail feature manifold, taking each configuration point as the center, calculating the arithmetic mean of the included angles between the direction vectors of all other configuration points in the spherical neighborhood of the center point relative to the center point to obtain the local directional divergence of the configuration point; each element of the projected 256-dimensional vector is multiplied by the reciprocal of the corresponding local directional divergence, then multiplied by (169 divided by the square root of 256), and the calculation result is standardized to make its Euclidean norm equal to 1, and the obtained 256-dimensional vector is the resonance focusing kernel.
[0039] The resonance focusing kernel is applied to the collaborative manifold to generate a cross-modal pre-collaborative feature vector through local feature aggregation and global pooling of the resonance focusing kernel, specifically including: expanding the 169-dimensional vector of the collaborative manifold to 256 dimensions through linear interpolation to make it consistent with the dimension of the resonance focusing kernel, multiplying the expanded vector and the resonance focusing kernel element by element, and performing local feature aggregation on the product: taking every 16 consecutive dimensions as an aggregation unit, calculating the sum of squares of all element values in each aggregation unit, then multiplying each sum of squares by the reciprocal of pi, to obtain 16 aggregation values; performing multi-scale global pooling on the 16 aggregation values: dividing the 16 aggregation values into 4 groups in order, each group having 4 values, calculating the product of the arithmetic mean and the standard deviation of each group, and multiplying the product by 0.3 after hyperbolic tangent function transformation to obtain 4 primary pooling values; then, the 4 primary pooling values are mapped to 128 dimensions through a fully connected layer and nonlinearly activated through a Sigmoid function to obtain a 128-dimensional cross-modal pre-collaborative feature vector.
[0040] The structural tensor is generated by fusing the inverse mapping topological weight matrix with the local curvature information of the fractal enhanced long-tail feature manifold, and the main flow mode feature is reconstructed by taking the structural tensor as the core constraint to obtain a collaborative manifold, and then a cross-modal pre-collaborative feature vector is generated, and the deep coupling of the main flow and long-tail mode features is realized, the core information of the long-tail mode and the semantic association of the main flow mode are completely preserved, the problem of long-tail mode marginalization caused by traditional shallow processing is solved, the generalization ability and semantic consistency of cross-modal feature alignment are improved, the collaborative processing precision of industrial production video frames, device log texts and vibration signals is optimized, and the alignment accuracy of multi-modal data is ensured.
[0041] In one case of the embodiment, the cross-modal pre-collaborative feature vector is calculated to generate a multi-modal alignment feature vector, which includes: The multi-scale stability of the cross-modal pre-collaborative feature vector is analyzed, and its structure preservation ability in the feature space is calculated to generate a stability distribution vector, which specifically includes: dividing the cross-modal pre-collaborative feature vector into eight consecutive segments in order, each segment containing sixteen dimensions, using four analysis scales (0.8, 0.4, 0.2, 0.1), calculating the ratio of the interquartile range to the median of each segment under each scale, obtaining the scale sensitivity, multiplying the sensitivity values of each consecutive segment under the four scales in sequence, obtaining the stability value, and arranging the stability values of the eight consecutive segments in the original order to form an 8-dimensional stability distribution vector.
[0042] Based on the stability distribution vector, the cross-modal pre-collaborative feature vector is selectively strengthened and structured to generate a feature crystallization nucleus, which specifically includes: multiplying each element value in the stability distribution vector by the cosine function value after taking the remainder of π, and then multiplying it by the coefficient 0.5 to obtain the strengthening factor; dividing the cross-modal pre-collaborative feature vector into eight consecutive 16-dimensional feature segments, wherein each element in the stability distribution vector corresponds to the stability value of a feature segment. For each feature segment, multiply the 16-dimensional elements it contains by the corresponding strengthening factor of the segment to complete the strengthening of the feature segment; from each 16-dimensional feature segment after strengthening, select the top 4 elements with the largest values, calculate the arithmetic mean of the 4 elements, and obtain the representative value of the feature segment. Arrange the representative values of the eight feature segments in order to form an 8-dimensional vector, and scale the Euclidean norm of the vector to 5% of the overall diameter of the feature space, output the standardized eight-dimensional vector, which is the feature crystallization nucleus.
[0043] The dimension of the characteristic crystal nucleus is regularized and optimized to generate a multi-modal alignment feature vector, specifically including: calculating the average value of all elements in the characteristic crystal nucleus, subtracting the average value from each element value to obtain a difference vector, dividing each element in the difference vector by the standard deviation of the first four elements to complete standardization; the sorted vector is arranged in ascending order and divided into four sections, and the median of each section is arranged in the original order to form a four-dimensional vector, which is a multi-modal alignment feature vector.
[0044] By generating a stability distribution vector, the selective reinforcement and structural recombination of the cross-modal pre-coordination feature vector are combined with the reinforcement factor to obtain a multi-modal alignment feature vector, which accurately preserves the association of the long-tail modal core frequency domain information and the mainstream modal semantics, solves the long-tail modal marginalization problem caused by traditional shallow processing, improves the generalization ability and semantic consistency of cross-modal feature alignment, strengthens the stability and core recognition of feature expression, and improves the collaborative processing precision of industrial production video frames, device log texts and vibration signals, and guarantees the structure and accuracy of multi-modal data alignment.
[0045] Although embodiments of the present application have been shown and described, it is to be understood that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the present application, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multi-modal large model cross-modal feature alignment optimization method, characterized in that, The method comprises the following steps: Step S1, obtaining multi-modal data of the industrial equipment, the multi-modal data being divided into long-tail small modal data and mainstream modal data, decomposing the long-tail small modal data, and generating a time-frequency co-occurrence feature tensor, comprising: analyzing the dynamic trajectory of the preprocessed long-tail small modal data in a high-dimensional phase space to obtain a trajectory topological tensor; analyzing the structural consistency of the trajectory topological tensor under changes to generate a structural stability scalar; performing stability calculation on the trajectory topological tensor and the structural stability scalar to generate a stable eigen spectrum; arranging and coupling each eigen mode in the stable eigen spectrum to generate a time-frequency co-occurrence feature tensor; the long-tail small modal data is a vibration signal of the industrial equipment; Step S2, performing modal decoupling and adaptive inverse mapping on the mainstream modal data, calculating inverse mapping weights of the mainstream modal data to the long-tail modal space, and generating an inverse mapping topological weight matrix, the mainstream modal data comprising: production videos of the industrial equipment and log texts of the industrial equipment; Step S3, analyzing the time-frequency co-occurrence feature tensor to generate a fractal-enhanced long-tail feature manifold, comprising: calculating the time-frequency co-occurrence feature tensor to generate a morphological gene sequence, specifically comprising: dividing the time-frequency co-occurrence feature tensor into 12 continuous segments, and calculating the Hurst index, box dimension and energy proportion of each segment, multiplying the Hurst index, box dimension and energy proportion to obtain the morphological gene value of the segment, arranging the morphological gene values of the 12 segments in time sequence to form a morphological gene sequence; performing causal analysis on the morphological gene sequence to analyze the degree of change in feature expression when the basic gene is intervened to generate a causal emergence intensity; after screening the causal emergence intensity, analyzing the self-organization, self-replication and self-adaptation process of the feature structure to generate an evolutionary stable configuration; in the evolutionary stable configuration, identifying the region of trajectory convergence in the phase space, i.e. a strange attractor, calculating the distribution and topological structure of the strange attractor to generate a feature attractor cloud; using the feature attractor cloud as an anchor and a constraint to perform reverse reconstruction to generate a fractal-enhanced long-tail feature manifold; Step S4, aligning and analyzing the inverse mapping topological weight matrix and the fractal-enhanced long-tail feature manifold to generate a cross-modal pre-cooperative feature vector; Step S5, calculating the cross-modal pre-cooperative feature vector to generate a multi-modal alignment feature vector.
2. The multi-modal large model cross-modal feature alignment optimization method of claim 1, wherein, The method comprises the following steps: performing modal decoupling and adaptive inverse mapping on the mainstream modal data, calculating inverse mapping weights of the mainstream modal data to the long-tail modal space, and generating an inverse mapping topological weight matrix, comprising: analyzing each frame sequence of the video and each semantic paragraph of the text in the preprocessed mainstream modal data to generate a semantic quantum state, specifically comprising: extracting a visual feature vector from the video frame sequence by using a deep convolutional network, extracting a semantic feature vector from the text paragraph by using a language model, and analyzing and processing the visual feature vector and the semantic feature vector to respectively obtain a semantic quantum state of the video modal and a semantic quantum state of the text modal; for the semantic quantum states of different modalities, calculating the correlation strength and the super-distance effect between them to generate a modal entanglement entropy.
3. The multi-modal large model cross-modal feature alignment optimization method of claim 2, wherein, The mainstream modal data is decoupled and adapted inverse mapping, the inverse mapping weight of the mainstream modal to the long tail modal space is calculated, the inverse mapping topological weight matrix is generated, and the method further comprises: The long tail modal space is taken as the base space, and the mainstream modal space is taken as the fiber to construct the fiber bundle. Based on the modal entanglement entropy, the connection form from the fiber to the base space is calculated to generate the fiber bundle connection vector. The stability of the fiber bundle connection vector in the whole feature space is analyzed to generate the inverse mapping topological weight matrix.
4. The multi-modal large model cross-modal feature alignment optimization method of claim 3, wherein, The inverse mapping topological weight matrix and the fractal enhanced long tail feature manifold are aligned and analyzed to generate the cross-modal pre-collaborative feature vector, which comprises: The inverse mapping topological weight matrix and the local curvature information of the fractal enhanced long tail feature manifold are calculated to generate the structure tensor. The structure tensor is taken as the transformation core to perform the constraint reconstruction of the mainstream modal feature on the manifold to generate the collaborative manifold.
5. The multi-modal large model cross-modal feature alignment optimization method of claim 4, wherein, The inverse mapping topological weight matrix and the fractal enhanced long tail feature manifold are aligned and analyzed to generate the cross-modal pre-collaborative feature vector, which further comprises: The information density of each point on the collaborative manifold and the multi-scale fractal dimension of the corresponding point of the fractal enhanced long tail feature manifold are calculated to generate the resonance focusing kernel. The resonance focusing kernel is applied to the collaborative manifold to generate the cross-modal pre-collaborative feature vector through local feature aggregation and global pooling of the resonance focusing kernel.
6. The multi-modal large model cross-modal feature alignment optimization method of claim 5, wherein, The cross-modal pre-collaborative feature vector is calculated to generate the multi-modal alignment feature vector, which comprises: The multi-scale stability of the cross-modal pre-collaborative feature vector is analyzed to calculate the structure retention ability thereof in the feature space to generate the stability distribution vector. Based on the stability distribution vector, the cross-modal pre-collaborative feature vector is selectively strengthened and structured to generate the feature crystallization core. The feature crystallization core is dimensionally normalized and optimized to generate the multi-modal alignment feature vector.
Citation Information
Patent Citations
Multimodal zero-order learning classification method and device for hyperspectral remote sensing image
CN117746097A
Multi-mode communication signal intelligent identification method based on neural network
CN120910800A
Hydraulic engineering construction monitoring data management system and method
CN121168824A
Damage identification method for cantilever beam based on multifractal spectrum of multi-scale reconstructed attractor
US20230358631A1