Multi-modal data fusion and analysis method for tobacco industry

By constructing a five-dimensional mode multimodal data fusion method, the problem that multimodal data in the tobacco industry cannot be uniformly represented is solved, and the deep interaction and dynamic adaptation of multimodal data are achieved, which improves the accuracy and efficiency of the tobacco production process.

CN120354341APending Publication Date: 2025-07-22SHANDONG INSPUR DIGITAL BUSINESS TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510391062.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The multimodal data in the tobacco industry cannot be expressed uniformly. Traditional data analysis methods cannot handle the multidimensional relationship between space-time and business in the tobacco industry chain, and insufficient dynamic adaptability, resulting in low cross-link synergy efficiency and lack of feature interactions, and the high-order coupling relationship between tobacco leaf chemical components and production process parameters cannot be captured.

Method used

A five-dimensional mode multimodal data fusion method is constructed, and the tensor rank is dynamically adjusted through the cross-modal attention mechanism and tensor decomposition technology, so as to realize the unified representation and deep interaction of multimodal data, and establish an end-to-end process decision-making and execution system.

Benefits of technology

It realizes unified representation and deep interaction of multimodal data in the entire tobacco industry chain, improves the sensitivity of abnormal detection in emergencies, dynamically adjusts the model to adapt to the dynamic changes in tobacco production, and improves the accuracy and efficiency of the production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354341A_ABST
    Figure CN120354341A_ABST
Patent Text Reader

Abstract

The invention provides a tobacco industry-oriented multi-modal data fusion and analysis method, and belongs to the technical field of crossing of data mining and industrial digitization. A tensor multi-modal data representation method is applied to the tobacco industry, and a five-dimensional space-time business tensor is defined; a cross-modal attention mechanism is dynamically combined with tensor Tucker decomposition, and interactive modeling of cross-modal heterogeneous data is realized through optimization of attention weights; the rank of each mode is dynamically adjusted according to the information entropy, and the model capacity and the calculation efficiency are balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the cross - technical field of data mining and industrial digitization, and particularly to a multi - modal data fusion and analysis method for the tobacco industry. Background Art

[0002] The tobacco industry generally uses traditional matrix modeling methods to process multi - source data (such as time - series data of tobacco field sensors and image data of production lines). Existing federated learning frameworks rely on low - dimensional feature representations, such as two - dimensional data fusion systems.

[0003] Technical Defects

[0004] Serious data islands: Data from tobacco fields, factories, and logistics systems are difficult to be uniformly characterized due to dimensional and modal differences, resulting in low cross - link collaborative efficiency.

[0005] Lack of feature interaction: Traditional matrix decomposition methods cannot capture the high - order coupling relationship between tobacco leaf chemical components (such as nicotine content) and production process parameters.

[0006] Insufficient dynamic adaptation: The reconstruction error of the fixed - dimension model increases significantly when dealing with sudden working conditions (such as equipment anomalies and tobacco leaf mildew).

[0007] Root Causes of the Problems

[0008] Traditional Euclidean space modeling methods can only express two - dimensional linear relationships and cannot handle the spatio - temporal - business multi - dimensional associations in the tobacco industry chain.

[0009] Existing data fusion technologies lack cross - modal attention mechanisms, resulting in insufficient correlation between multi - spectral tobacco leaf image data and chemical inspection indicators.

[0010] CN119599481A discloses a multi - index collaborative optimization method. Excessive model simplification: Polynomial fitting cannot represent the non - linear coupling effect of path complexity and weather factors (such as the dynamic detour cost when a highway is closed due to heavy rain);. Single - source data: Only relying on logistics work order text data and ignoring the decision - making value of sensor real - time data (such as truck tire pressure and cold chain temperature); 3. Lack of dynamic adaptation: The index weights are fixed (such as the weather factor α4 is preset to 0.15) and cannot be automatically adjusted according to sudden road conditions (such as traffic accidents).

[0011] CN119600337A discloses a method and device for identifying the maturity of tobacco leaves based on the Lasso - Boruta - gcforest algorithm. 1. Single data modality: The user's solution only relies on image information for maturity discrimination, without fusing multi - modality data such as tobacco chemical indicators (such as nicotine content, total sugar content), environmental parameters (temperature, humidity, light intensity), etc., thus unable to capture the co - variation law of tobacco chemical - physical characteristics (such as the dose - effect relationship between sugar - alkali ratio and leaf surface texture).

[0012] 2. Lack of dynamic adaptability: The hyperparameters of the GCF model in the user's solution are statically optimized once through the genetic algorithm, without establishing a dynamic adjustment mechanism. When the tobacco field environment suddenly changes (such as sudden drought causing premature leaf senescence), the model needs to be retrained (increasing the time consumption) and cannot respond in real - time.

[0013] 3. Insufficient interpretability: The GCF model in the user's solution has a black - box structure, unable to trace the feature contribution degree and decision basis, resulting in agronomic experts being unable to verify whether the "over - ripe judgment" is related to key chemical indicators (such as starch accumulation), and lacking judicial credibility in the quality dispute scenario.

[0014] The prior art fails to solve the two core problems of asymmetric dimension representation and dynamic working condition adaptation. The root causes are as follows:

[0015] 1. Multi - modality data in the tobacco industry cannot be uniformly represented

[0016] 2. The mathematical constraints of traditional data analysis methods do not match the characteristics of tobacco data.

[0017] 3. The static model architecture cannot adapt to the dynamic data flow characteristics of the tobacco field - factory - logistics scenario. Summary of the Invention

[0018] To solve the above - mentioned technical problems, the present invention provides a multi - modality data fusion and analysis method for the tobacco industry. By constructing a multi - modality fusion tensor and performing cross - modality decomposition, and analyzing the obtained results, the bottleneck of the prior art can be broken through.

[0019] The technical solution of the present invention is as follows:

[0020] A multi - modality data fusion and analysis method for the tobacco industry,

[0021] First, construct a fifth - order tensor representing global data by integrating five - dimensional modalities of image, chemistry, environment, time, and space:

[0022] Secondly, deploy five groups of cross - modality attention correlation matrices, and analyze the dynamic interaction mechanism between different modalities through tensor decomposition technology:

[0023] Then, implement the core tensor rank adaptive optimization driven by the joint five-modal information entropy:

[0024] Finally, construct an end-to-end process decision-making and execution system.

[0025] Furthermore,

[0026] Through timestamp synchronization and spatial coordinate mapping, unify the multi-source heterogeneous data collected during the tobacco leaf baking process to the spatio-temporal reference framework, forming a continuous five-dimensional tensor structure, where each dimension respectively carries image features, chemical components, environmental parameters, time process, and spatial position information, ensuring that data units at any time and any position contain complete cross-modal attribute descriptions.

[0027] Furthermore,

[0028] Deploy five groups of cross-modal attention correlation matrices, including

[0029] Establish a two-way mapping relationship between visual appearance changes and internal component evolution between the image and chemical modalities;

[0030] Quantify the non-linear response of the material conversion rate to temperature and humidity conditions between the chemical and environmental modalities;

[0031] Capture the immediate impact of sudden physical fluctuations on the leaf surface morphology between the environmental and image modalities;

[0032] Model the long-term regulation of the cumulative effect of baking process stages on components between the time and chemical modalities;

[0033] Reveal the differences in environmental parameter distributions caused by the heterogeneity of the baking room area between the spatial and environmental modalities;

[0034] Each group of attention matrices dynamically screens cross-modal correlation features and suppresses noise interference through an autonomously learned weight assignment mechanism, generating a core tensor representing the key interaction patterns.

[0035] Furthermore,

[0036] Independently evaluate the texture confusion degree of the image modality, the index volatility of the chemical modality, the parameter deviation degree of the environmental modality, the stage mutation degree of the time modality, and the regional balance degree of the spatial modality, and quantify the data complexity of each dimension through information entropy; according to the comparison results of the entropy value thresholds, expand the core tensor rank in the high-complexity modality direction to enhance the feature representation ability, and compress the rank in the low-complexity modality direction to improve the calculation efficiency.

[0037] Furthermore,

[0038] Input the optimized core tensor into the depth decision network to output the control instruction set for the baking equipment, including the temperature adjustment range, the set value of the ventilation intensity, the leaf turning frequency parameter, and the abnormal condition handling strategy; circularly inject the newly collected data stream into the five-dimensional tensor construction module to trigger a new round of cross-modal analysis and model optimization, and achieve full-process adaptive iteration.

[0039] Furthermore,

[0040] Introduce cross-modal attention during the decomposition process to dynamically correct the factor matrix. The attention mechanism actively retrieves the key-value pairs, i.e., the relevant information in the chemical + environmental modality, by querying the image modality, and calculates the association weights to dynamically focus on the key modality combinations.

[0041] Including:

[0042] (1) Inter-modal attention calculation:

[0043] Query image modality: Q = U (1) W Q

[0044] where U (1) is the factor matrix of the image modality, and W Q is the weight matrix. This step maps the features of the image modality (such as leaf color and texture) to the query space for actively retrieving the associated information of other modalities.

[0045] Key-value pairs (chemical + environmental): K = [U (2) ; U (3) W K , V = [U (2) ; U (3) W V

[0046] where [U (2) ; U (3) is the vertical concatenation of the chemical and environmental modalities. This step jointly encodes the chemical indicators (passive responses) and environmental parameters (dynamic interferences) as key-value pairs to establish a cross-modal association library.

[0047] Association weights:

[0048]

[0049] where i, j, k, l, m traverse the image, chemical, environmental, spatial, and temporal dimensions respectively, aiming to calculate the dynamic association strength of the image feature i with the chemical indicator j, environmental parameter k, spatial position l, and time m. Among them, Softma is a non-linear function that converts a real vector into a probability distribution and satisfies: h is a scaling factor to prevent the numerator from being too large (usually taking the attention dimension value)

[0050] (2) Factor matrix correction:

[0051] Mathematical expression

[0052]

[0053] Among them, λ is the learning rate, which controls the attention correction intensity and takes values from 0.01 to 0.1. is the Hadamard product, that is, element-wise multiplication. According to the cross-modal correlation weight a ijklm , the chemical, environmental, spatial, and temporal information is fused into the image feature representation.

[0054] (3) Core tensor correction:

[0055] Through the corrected factor matrix, the core tensor needs to be updated by the projection reconstruction error:

[0056]

[0057] Among them, (·) + represents the pseudo-inverse. This step ensures that the decomposition result is compatible with the corrected factor matrix;

[0058] Final output:

[0059] The corrected core tensor Stores the correlation pattern after cross-modal attention correction;

[0060] The updated factor matrix Contains the feature representation with dynamic attention weights.

[0061] Further,

[0062] Dynamic core tensor adjustment

[0063] Dynamically adjust the rank r of each modality according to the data complexity m , balance the model capacity and computational efficiency, and perform incremental updates on the tensor model.

[0064] Includes

[0065] (1) Entropy value monitoring

[0066] First, calculate the information entropy for each modality of the core tensor. The formula is as follows:

[0067]

[0068] Among them, represents the s-th slice of the core tensor in the m-th modality, represents the Frobenius norm, which is the square root of the sum of the squares of all elements in the tensor, and p sis the energy proportion of the sth slice, H m The higher the value, the higher the data complexity of the mode, which requires a higher rank modeling; on the contrary, the complexity is low and the rank can be compressed;

[0069] (2) Rank Adaptive Adjustment

[0070] According to the information entropy calculated in the first step, the rank of the core tensor is adjusted. The following is the adjustment rule: If H m >1, the modal complexity is considered to be high, and the rank needs to be expanded if the complexity is high. Otherwise, the modal complexity is considered to be low, and the rank needs to be compressed.

[0071] Extended conditions mean high complexity:

[0072] H m >nH th ,but

[0073] The compression condition is low complexity:

[0074] H m <nH th ,but

[0075] Among them, n is the preset multiple, H th is the preset threshold.

[0076] The beneficial effects of the present invention are

[0077] 1. Realize unified representation and deep interaction of multimodal data of the entire tobacco industry chain through high-order tensor modeling, breaking through the dimensional limitations of traditional matrix methods

[0078] 2. Dynamically adjust the tensor decomposition rank and core dimension to significantly improve the sensitivity of anomaly detection under sudden conditions

[0079] 3. Construct a cross-modal attention mechanism to effectively explore the nonlinear correlation between tobacco leaf physical properties (such as texture density) and chemical indicators (such as sugar-alkali ratio) and other indicators. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 It is a schematic diagram of the workflow of the present invention;

[0081] Figure 2 This is a diagram of Tucker decomposition. DETAILED DESCRIPTION

[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0083] The present invention proposes a multi-modal data fusion and analysis method for the tobacco industry to achieve closed-loop optimization from data representation to decision support. The following is a complete description of the technical process ( Figure 1 ):

[0084] First, construct a fifth-order tensor that fuses five modalities: image, chemistry, environment, time, and space to represent global data:

[0085] Through timestamp synchronization and spatial coordinate mapping, the multi-source heterogeneous data collected during the tobacco leaf baking process is uniformly aligned to the spatio-temporal reference framework to form a continuous fifth-order tensor structure. Each dimension respectively carries image features, chemical components, environmental parameters, time progress, and spatial position information, ensuring that the data units at any time and any position contain complete cross-modal attribute descriptions.

[0086] Second, deploy five groups of cross-modal attention correlation matrices to analyze the dynamic interaction mechanism between different modalities through tensor decomposition technology:

[0087] Establish a two-way mapping relationship between visual appearance changes and internal component evolution between the image and chemistry modalities;

[0088] Quantify the non-linear response of the material conversion rate to temperature and humidity conditions between the chemistry and environment modalities;

[0089] Capture the immediate impact of sudden physical fluctuations on the leaf surface morphology between the environment and image modalities;

[0090] Model the long-term regulation of the baking process stage evolution on the component accumulation effect between the time and chemistry modalities;

[0091] Reveal the differences in environmental parameter distributions caused by the heterogeneity of the baking room area between the space and environment modalities.

[0092] Each group of attention matrices dynamically screens cross-modal correlation features and suppresses noise interference through a self-learning weight assignment mechanism to generate a core tensor representing key interaction patterns.

[0093] Then, implement the core tensor rank adaptive optimization jointly driven by the five-modal information entropy:

[0094] Independently evaluate the texture chaos of the image modality, the index volatility of the chemical modality, the parameter deviation of the environmental modality, the stage mutation of the time modality, and the regional equilibrium of the spatial modality, and quantify the data complexity of each dimension through information entropy; according to the comparison result of the entropy value threshold, expand the core tensor rank in the high-complexity modality direction to enhance the feature representation ability, and compress the rank in the low-complexity modality direction to improve the calculation efficiency.

[0095] Finally, construct an end-to-end process decision-making and execution system:

[0096] Input the optimized core tensor into the deep decision-making network to output a baking equipment control instruction set, including the temperature adjustment range, ventilation intensity setting value, leaf turning frequency parameter, and abnormal condition handling strategy; inject the newly collected data stream in real time into the five-dimensional tensor construction module in a loop to trigger a new round of cross-modal analysis and model optimization, and achieve full-process adaptive iteration.

[0097] Among them, the tobacco multi-modal tensor construction method, the combination of tensor decomposition and cross-modal attention mechanism, and the introduction of information entropy for core tensor rank adjustment are the first creations of the present invention. The following are the technical details of the present invention.

[0098] 1. Tobacco multi-modal tensor construction method

[0099] In the analysis of tobacco leaf maturity, data usually comes from multiple different sources and types, such as:

[0100] · Visual data: color photos of tobacco leaves, multi-spectral images (spectral reflectance of different bands)

[0101] · Chemical data: chlorophyll content (SPAD value), nicotine concentration, total sugar content, etc. detected in the laboratory

[0102] · Environmental data: real-time monitoring values of temperature, humidity, light intensity, etc. in the planting area

[0103] · Spatial data: specific position coordinates of the plant in the tobacco field

[0104] · Time data: specific time points of data collection (such as 8 am, 3 pm)

[0105] These data of different "modalities" are like books in different languages. Traditional methods can only read each one separately and cannot discover cross-language related knowledge.

[0106] Traditional technologies (such as Excel tables, two-dimensional databases) can only store data in a simple structure of "rows and columns". This two-dimensional structure has two major problems:

[0107] 1. Single dimension: unable to record multi-dimensional information such as image features and spatial positions simultaneously

[0108] 2. Loss of relationship: For example, it is impossible to directly express the complex association of "how high temperature affects the chlorophyll and image features of tobacco leaves at different positions".

[0109] Therefore, consider introducing the tensor, a multi-dimensional data structure, to solve the above problems. A tensor can be understood as a "multi-dimensional data table". For example, a five-dimensional tensor is like a super cube with five coordinate axes, and each "cell" stores multi-modal data under specific conditions:

[0110] · Dimension 1: Image features (such as leaf color, texture)

[0111] · Dimension 2: Chemical indicators (such as SPAD value, sugar content)

[0112] · Dimension 3: Environmental parameters (temperature, humidity)

[0113] · Dimension 4: Spatial location (tobacco field grid coordinates)

[0114] · Dimension 5: Time series (acquisition time point)

[0115] With this structure, we can analyze the "correlation between the image features and chemical indicators of a certain tobacco leaf at 28°C in grid area A3 at 8:00" in a unified framework. At the same time, tensors have specialized operations and decomposition methods, which can be used for data dimensionality reduction and feature extraction to better analyze multi-modal data.

[0116] Specific construction method:

[0117] Step 1: Data standardization, converting data with different dimensions into a unified scale:

[0118] · Image data: Extract image feature vectors using a convolutional neural network

[0119] · Chemical data: Normalize to the 0-1 interval (e.g., SPAD value 32.5 → 0.72)

[0120] · Environmental data: Record directly according to physical units (e.g., temperature = 28.5°C)

[0121] · Spatial data: Divide the tobacco field into 1m × 1m grids and number them with coordinates (e.g., grid A3)

[0122] · Time data: Convert to minutes (e.g., 8:00 → 480 minutes)

[0123] Step 2: Tensor filling

[0124] Construct a five-dimensional tensor T, and each of its elements T ijklm represents:

[0125] · i: The i-th type of image feature (e.g., i = 1 represents color, i = 2 represents texture...)

[0126] · j: The j-th chemical index (e.g., j = 1 represents SPAD value, j = 2 represents sugar content...)

[0127] · k: The k-th environmental parameter (e.g., k = 1 represents temperature, k = 2 represents humidity...)

[0128] · l: The l-th spatial grid

[0129] · m: The m-th time point

[0130] For example:

[0131] t1,3,2,5,10 = 0.85 means:

[0132] · Image feature 1 (such as red channel intensity)

[0133] · Chemical index 3 (such as nicotine content)

[0134] · Environmental parameter 2 (such as humidity)

[0135] · Spatial grid 5

[0136] · Time point 10 (such as 9:30)

[0137] The value under this condition is 0.85 (normalized humidity value)

[0138] Breakthrough of the present invention:

[0139] By constructing a five-dimensional tensor, the following problems are solved:

[0140] Unified storage: All modal data are integrated into a single structure, eliminating the need for manual comparison.

[0141] Full-dimensional analysis: The relationships of any dimensional combination can be examined at once (such as the three-dimensional slice of "temperature - time - space").

[0142] Efficient calculation: Using a tensor operation library (such as Tensorly) for direct batch processing to improve the calculation speed.

[0143] 2. Cross-modal attention decomposition algorithm

[0144] This method realizes the adaptive fusion of multi-modal data by dynamically adjusting the tensor decomposition structure and injecting an attention mechanism.

[0145] Input: A five-dimensional tensor of tobacco multi-modal data, and perform Tucker decomposition on the tobacco fusion tensor to extract low-dimensional features ( Figure 2 ):

[0146] T = G core × 1U (1) × 2U (2) × … × 5U (5)

[0147] · Core tensor: Core tensor (Compression ratio > 80%)

[0148] · Factor matrix:

[0149] (3 - 1) Attention mechanism injection

[0150] Introduce cross - modal attention during the decomposition process to dynamically correct the factor matrix. The attention mechanism actively retrieves relevant information in the key - value (chemical + environmental modality) by querying (image modality), calculates the correlation weight, and dynamically focuses on the key modality combination. For example, in a high - temperature environment, the model will strengthen the correlation weight between leaf color and chlorophyll content, so as to more accurately reflect the change of tobacco leaf maturity.

[0151] 1) Inter - modal attention calculation:

[0152] · Query (image modality): Q = U (1) W Q

[0153] where U (1) is the factor matrix of the image modality, and W Q is the weight matrix. This step maps the features of the image modality (such as leaf color and texture) to the query space for actively retrieving the associated information of other modalities.

[0154] · Key - value (chemical + environment): K = [U (2) ; U (3) W K , V = [U (2) ; U (3) W V

[0155] where [U (2) ; U (3) is the vertical concatenation of the chemical and environmental modalities. This step jointly encodes the chemical indicators (passive response) and environmental parameters (dynamic interference) as key - value pairs to establish a cross - modal association library.

[0156] · Correlation weight:

[0157]

[0158] (i, j, k, l, m traverse the image, chemistry, environment, space, and time dimensions respectively), with the aim of calculating the dynamic correlation strength of image feature i with chemical index j, environmental parameter k, spatial position l, and time m. Among them, Softmax is a non-linear function that converts a real-valued vector into a probability distribution and satisfies: h is the scaling factor to prevent the molecule from being too large (usually take the attention dimension value)

[0159] 2) Factor matrix correction:

[0160] · Mathematical expression

[0161]

[0162] Among them, λ is the learning rate (controlling the attention correction strength, usually taking 0.01 - 0.1), is the Hadamard product (element-wise multiplication). The purpose of this step is to fuse chemical, environmental, spatial, and time information into the image feature representation according to the cross-modal correlation weight a ijklm , and the chemical, environmental, spatial, and time information is fused into the image feature representation.

[0163] The above is the correction of the environmental factor matrix. This method can be used to correct the chemical, environmental, spatial, and time matrices in turn.

[0164] 3) Core tensor correction:

[0165] Through the corrected factor matrix, the core tensor needs to be updated by the projection reconstruction error:

[0166]

[0167] Among them, (·) + represents the pseudo-inverse. This step ensures that the decomposition result is compatible with the corrected factor matrix.

[0168] Final output:

[0169] The corrected core tensor Stores the correlation pattern after cross-modal attention correction.

[0170] The updated factor matrix Contains the feature representation with dynamic attention weights.

[0171] 3. Dynamic core tensor adjustment

[0172] This method dynamically adjusts the rank r of each modality according to the data complexity m , balances the model capacity and computational efficiency, and performs incremental updates on the tensor model.

[0173] (1) Entropy value monitoring

[0174] First, calculate the information entropy for each mode of the core tensor. The formula is as follows:

[0175]

[0176] Among them, represents the s-th slice of the core tensor in the m-th mode, represents the Frobenius norm (the square root of the sum of the squares of all elements in the tensor), and p s is the energy proportion of the s-th slice, and H m The higher the value, the higher the data complexity of this mode (the energy distribution is dispersed), and higher-rank modeling is required; conversely, the complexity is low and the rank can be compressed.

[0177] (2) Rank adaptive adjustment

[0178] The purpose is to adjust the rank of the core tensor according to the information entropy calculated in the first step. The following are the adjustment rules:

[0179] If H m > 1, it is considered that the complexity of this mode is high. For high complexity, the rank needs to be expanded; conversely, it is considered that the complexity of this mode is low and the rank needs to be compressed.

[0180] Expansion condition (high complexity):

[0181] H m > nH th , then

[0182] Compression condition (low complexity):

[0183] H m < nH th , then

[0184] Among them, n is a preset multiple, and H th is a preset threshold (usually 1, but it needs to be adjusted in combination with engineering experience).

[0185] The output result of the present invention is a multi-modal dynamic association feature and a generalizable core model. Its core value lies in transforming high-dimensional heterogeneous agricultural data into an intelligent foundation that can directly drive decision-making. Compared with traditional methods, this technical system realizes the closed-loop transition of multi-modal data from "static analysis" to "dynamic decision-making" while ensuring the light weight and interpretability of the model, providing a practical technical paradigm for the precision and intelligence of agricultural production.

[0186] The above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A multi-modal data fusion and analysis method for the tobacco industry, characterized in that: First, construct a fifth-order tensor that fuses five modalities of image, chemistry, environment, time, and space to represent global data: Second, deploy five groups of cross-modal attention correlation matrices, and analyze the dynamic interaction mechanism between different modalities through tensor decomposition technology: Third, implement core tensor rank adaptive optimization jointly driven by five-modal information entropy: Finally, construct an end-to-end process decision-making and execution system.

2. The method according to claim 1, characterized in that: Through timestamp synchronization and spatial coordinate mapping, multi-source heterogeneous data collected during the tobacco leaf baking process are uniformly aligned to the spatio-temporal reference framework to form a continuous five-dimensional tensor structure, where each dimension respectively carries image features, chemical components, environmental parameters, time process, and spatial position information, ensuring that data units at any time and any position contain complete cross-modal attribute descriptions.

3. The method according to claim 1, characterized in that: Deploy five groups of cross-modal attention correlation matrices, including Establish a two-way mapping relationship between visual appearance changes and internal component evolution between the image and chemistry modalities; Quantify the non-linear response of the substance conversion rate to temperature and humidity conditions between the chemistry and environment modalities; Capture the immediate impact of sudden physical fluctuations on the leaf surface morphology between the environment and image modalities; Model the long-term regulation of the component accumulation effect by the evolution of the baking process stage between the time and chemistry modalities; Reveal the difference in environmental parameter distribution caused by the heterogeneity of the baking room area between the space and environment modalities; Each group of attention matrices dynamically screens cross-modal correlation features and suppresses noise interference through a self-learning weight assignment mechanism, and generates a core tensor representing the key interaction mode.

4. The method according to claim 3, characterized in that: Independently evaluate the texture chaos of the image modality, the index volatility of the chemistry modality, the parameter deviation of the environment modality, the stage mutation of the time modality, and the regional balance of the space modality, and quantify the data complexity of each dimension through information entropy; according to the comparison result of the entropy value threshold, expand the core tensor rank in the high-complexity modality direction to enhance the feature representation ability, and compress the rank in the low-complexity modality direction to improve the calculation efficiency.

5. The method according to claim 4, characterized in that: Input the optimized core tensor into a deep decision network, and output a baking equipment control instruction set, including temperature adjustment amplitude, ventilation intensity setting value, leaf turning frequency parameter, and abnormal working condition handling strategy; inject the newly collected data stream into the five-dimensional tensor construction module in a loop, trigger a new round of cross-modal analysis and model optimization, and achieve full-process adaptive iteration.

6. The method according to claim 5, characterized in that: Introduce cross-modal attention during the decomposition process to dynamically correct the factor matrix. The attention mechanism actively retrieves key values, that is, relevant information in the chemistry + environment modality, by querying the image modality, and calculates the correlation weight to dynamically focus on the key modality combination.

7. The method according to claim 6, characterized in that: Includes: (1) Cross-modal attention calculation: Query image modality: Q = U (1) W Q Among them, U (1) is the factor matrix of the image modality, and W Q is the weight matrix that maps the features of the image modality to the query space and is used to actively retrieve the associated information of other modalities; Key value, i.e., chemistry + environment: K = [U (2) ; U (3) W K , V = [U (2) ; U (3) W V Among them, [U (2) ; U (3) is the vertical splicing of the chemical and environmental modalities. In this step, chemical indicators and environmental parameters (dynamic interferences) are jointly encoded into key-value pairs to establish a cross-modal association library; Correlation weight: Among them, i, j, k, l, and m traverse the image, chemistry, environment, space, and time dimensions respectively, aiming to calculate the dynamic correlation strength of image feature i with chemical index j, environmental parameter k, spatial position l, and time m. Among them, Softma is a non-linear function that converts a real number vector into a probability distribution and satisfies: h is a scaling factor to prevent the molecule from being too large; (2) Factor matrix correction: Mathematical expression Among them, λ is the learning rate, which controls the attention correction intensity and takes values from 0.01 to 0.

1. ο is the Hadamard product, that is, element-wise multiplication. According to the cross-modal correlation weight a ijklm , the chemical, environmental, spatial, and temporal information is fused into the image feature representation. (3) Core tensor correction: With the corrected factor matrix, the core tensor needs to be updated by the projection reconstruction error: where (·) + denotes the pseudo-inverse, and this step ensures that the decomposition result is compatible with the corrected factor matrix; Final output: Corrected core tensor Store the corrected association pattern of cross-modal attention; Updated factor matrix Feature representation containing dynamic attention weights.

8. The method according to claim 7, wherein Dynamic core tensor adjustment Dynamically adjust the rank r of each modality according to data complexity m , balance the model capacity and computational efficiency, and perform incremental updates on the tensor model.

9. The method according to claim 8, wherein including (1) Entropy value monitoring First, calculate the information entropy for each mode of the core tensor, and the formula is as follows: Among them, represents the s-th slice of the core tensor in the m-th mode, represents the Frobenius norm, which is the square root of the sum of the squares of all elements in the tensor, and p s is the energy proportion of the s-th slice, and H m The higher the value, the higher the data complexity of the mode, and higher-rank modeling is required; on the contrary, the complexity is low and the rank can be compressed. (2) Rank adaptive adjustment According to the information entropy calculated in the first step, adjust the rank of the core tensor. The following are the adjustment rules: If H m > 1, it is considered that the complexity of this mode is high. For high complexity, the rank needs to be expanded. Otherwise, it is considered that the complexity of this mode is low and the rank needs to be compressed; Expansion condition, i.e., high complexity: H m >nH th , then Compression condition, i.e., low complexity: H m <nH th then where n is a preset multiple, and H th is a preset threshold value.

Citation Information

Patent Citations

  • Multi-index collaborative optimization method and device, medium and equipment

    CN119599481A

  • Tobacco harvesting maturity identification method and equipment based on Lasso-Boruta-gcfrest algorithm

    CN119600337A

Cited By

  • Material dynamic regulation and control system and control method

    CN120949721A

  • Premium rate calculation method and system based on multi-source data analysis

    CN121169452A