E-commerce product visual and tactile fusion model training method based on deep learning
By using deep learning for modal relevance assessment and dynamic weight adjustment, combined with visual feature enhancement and information loss control, the problem of static weights in visual and tactile fusion methods is solved, improving the accuracy and robustness of e-commerce product recognition and enhancing the ability to recognize complex products.
Patent Information
- Application Number
- CN202511766459.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-02-27
AI Technical Summary
Existing visual and tactile fusion methods face problems in e-commerce product recognition, such as large differences in modal feature representation, static weight allocation, and inflexible adjustment, resulting in insufficient recognition accuracy, especially when recognizing complex products, and failing to fully learn key details.
A deep learning-based modal relevance assessment technique is adopted. Modal weights are dynamically adjusted through attention mechanism and context dependency analysis. Combined with visual feature enhancement and information loss control, priority ranking and cross-modal interaction assessment, the ratio of visual and tactile data is dynamically adjusted to generate a calibrated fusion perception vector.
It improves the model's recognition accuracy and robustness for complex products, enhances its sensitivity to geometric and textural details, reduces interference from irrelevant information, and improves the overall performance of intelligent recognition of e-commerce products.
Smart Images

Figure CN121580018A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent perception and multi-modal information technology, and in particular to a deep learning-based e-commerce product visual-haptic fusion model training method. BACKGROUND
[0002] In the current field of e-commerce product intelligent recognition and perception technology, multi-modal information fusion has become an important direction to improve product recognition accuracy and understanding ability. Traditional e-commerce image recognition usually relies on visual data alone, but as the complexity of product types increases, the difference in material quality increases, and users have higher demands on product texture understanding, relying solely on visual modalities is difficult to fully depict the real attributes of goods. Therefore, introducing haptic data (such as pressure texture, material feedback, etc.) and visual data for joint modeling has become an important trend to improve the intelligent recognition ability of e-commerce products.
[0003] However, existing visual and haptic fusion methods still face many challenges in practical applications. Firstly, the feature representation forms of different modalities differ greatly, and how to accurately evaluate the correlation between visual features and haptic features is a key difficulty that affects the quality of fusion. Secondly, in the process of multi-level information processing, the importance of each modality changes dynamically with the task stage, but existing methods usually use fixed weights or static fusion mechanisms, which are difficult to adjust flexibly according to the structural differences of e-commerce products, texture complexity, and specific recognition targets. In addition, in the training process of deep learning models, if there is no reasonable modality weight distribution mechanism, it is easy to cause uneven utilization of information between modalities. For example, when identifying small accessories or goods with complex edge structures on e-commerce platforms, visual data needs to highlight geometric features such as shape and position, while haptic data emphasizes material texture and pressure feedback. If the model cannot dynamically adjust the information proportion of visual and haptic data according to task requirements, it may not be able to fully learn key details during training, thereby affecting the final recognition accuracy. At the same time, as the scale of e-commerce data continues to expand, how to reduce the interference of excessive irrelevant information in the multi-modal fusion process, avoid the excessive dominance of haptic or visual modalities in model training, and ensure the effective interaction of multi-level features, has become a further challenge faced by current technology development. SUMMARY
[0004] The purpose of the present application is to provide a deep learning-based e-commerce product visual-haptic fusion model training method to solve the problems existing in the prior art.
[0005] To achieve the above object, the application provides the following technical scheme: a method for training an e-commerce product visual-haptic fusion model based on deep learning, comprising S1, collecting visual data and haptic data of an e-commerce product through a sensor, extracting shape and position features for the visual data, extracting pressure and texture features for the haptic data, using a deep learning-based modal correlation evaluation technique to analyze the mutual information metric between the two, and obtaining a preliminary multi-modal feature combination; S2, using an attention mechanism to perform dynamic weight calculation based on the preliminary multi-modal feature combination, combining context-dependent analysis and time series technology to evaluate the contribution of visual and haptic data in the current level of processing, and determining the initial weight distribution of each modality; S3, if the initial weight distribution shows that the contribution of visual data is lower than that of haptic data, then the resource allocation proportion of visual data is increased through a weight proportion adjustment technique, and the shape and position information is released using a visual feature enhancement method to obtain a preliminary fused perception vector; S4, using a fusion priority sorting technique to place visual data in a dominant position for the preliminary fused perception vector, balancing the auxiliary information of haptic data through an information loss control method, and generating a calibrated fusion perception vector; S5, based on the calibrated fusion perception vector, combining task demand matching technology to analyze the current operation's demand for e-commerce product object edge recognition, evaluating the interaction between visual and haptic data through cross-modal interaction quantity, and determining whether the visual weight proportion needs to be adjusted.
[0006] Preferably, S1 includes obtaining visual data of an e-commerce product through an image acquisition device, extracting shape and position features for the visual data using an edge detection tool to obtain a shape and position feature set; obtaining haptic data of an e-commerce product through a pressure sensor, extracting pressure and texture features for the haptic data using a texture analysis tool to obtain a pressure and texture feature set; using a deep neural network tool to calculate the mutual information metric between the shape and position feature set and the pressure and texture feature set, if the mutual information value is higher than a preset threshold, it is determined that the two have modal correlation, and a preliminary multi-modal feature combination data is obtained; using a feature fusion tool to weight and integrate the shape and position features and the pressure and texture features for the preliminary multi-modal feature combination data to obtain the final multi-modal feature representation.
[0007] Preferably, S2 comprises extracting a current level of feature representation from visual data and tactile data according to a preliminary multi-modal feature combination, performing standardization processing by using a data preprocessing tool to obtain a unified feature representation set; analyzing the historical change trend of visual data and tactile data by using a time series processing tool for the unified feature representation set, obtaining the performance difference of each modality at different time points, and determining the time correlation data of each modality; if the performance difference of a certain modality in the time correlation data is higher than a preset threshold, adjusting the dynamic weight of the modality by using an attention calculation tool, analyzing its role in combination with a context dependence tool, and obtaining an adjusted weight distribution result; according to the adjusted weight distribution result, using a data integration tool to comprehensively evaluate the contribution of visual data and tactile data, judging the initial weight distribution of each modality in the current level, and obtaining the final modality weight distribution data.
[0008] Preferably, S3 comprises obtaining a contribution degree comparison result of visual data and tactile data according to the initial weight distribution data, determining the allocation of visual data in the current level; if the contribution degree of visual data is lower than that of tactile data, dynamically improving the resource allocation ratio of visual data by using a proportion adjustment tool, and obtaining an adjusted allocation scheme; for the adjusted allocation scheme, using a feature extraction tool to separate shape position information from visual data, combining an enhancement tool to strengthen details, and determining an enhanced feature set; integrating the enhanced feature set and tactile data by using a data fusion tool to generate preliminary perception vector data, judging whether it meets a preset standard, and obtaining a final fusion result.
[0009] Preferably, S4 comprises adjusting the priority of visual data by using a sorting tool according to the preliminary fused perception vector, obtaining an adjusted data level distribution, and determining its dominant position; for the adjusted data level distribution, extracting auxiliary information of tactile data by using a balancing tool to obtain a key feature set, judging whether it meets a preset threshold, and obtaining a feature balanced information combination; if the feature balanced information combination meets the preset threshold, integrating the dominant information of visual data and the auxiliary information of tactile data by using an integration tool to obtain the integrated data set, and determining the preliminary structure of the fusion vector; processing the integrated data set by using a calibration tool to obtain the final fusion perception vector, judging whether it meets a preset balance standard, and obtaining the calibrated perception vector data.
[0010] Preferably, S5 comprises extracting feature distribution of visual data and tactile data by using analytical tool according to calibrated fusion perception vector, obtaining hierarchical structure of feature distribution, and determining dominant component in hierarchical structure; analyzing task requirement of object edge recognition by using comparison tool through hierarchical structure of feature distribution, obtaining feature weight distribution matched with requirement, judging whether distribution meets preset threshold value or not, and obtaining preliminary matching result; if preliminary matching result meets preset threshold value, calculating cross-modal interaction amount of visual data and tactile data by using evaluation tool, obtaining influence range of interaction amount, and determining saliency of influence range; according to saliency of influence range, dynamically modifying weight proportion of visual data by using adjustment tool, obtaining modified weight configuration, judging whether configuration meets edge recognition requirement or not, and obtaining final allocation scheme.
[0011] Preferably, S6 further comprises if conductivity of visual data after weight proportion is insufficient to capture details of object edge, re-adjusting allocation proportion of visual data and tactile data by using dynamic proportion feedback technology, enhancing processing depth of shape position information, and determining optimized multi-modal perception combination, specifically comprising according to requirement of visual detail capture, performing feature enhancement processing on original visual data by using image processing tool, obtaining feature information related to edge shape from original visual data, and judging whether feature information meets preset threshold value or not; if feature information does not meet preset threshold value, supplementing tactile information to original visual data by using data fusion tool, obtaining enhanced mixed data distribution, and determining contribution degree of mixed data distribution.
[0012] Preferably, S6 further comprises for mixed data distribution, dynamically modifying allocation proportion of visual data and tactile data by using proportion adjustment tool, obtaining new data interaction balance state, and judging whether state meets detail resolution requirement or not; according to data interaction balance state, performing optimization processing on multi-modal information fusion by using information integration tool, obtaining final perception combination configuration, and determining adaptability of configuration to edge recognition.
[0013] Preferably, S7 further comprises according to optimized multi-modal perception combination, removing irrelevant information by using data redundancy filtering technology, and simultaneously updating fusion process by priority division through hierarchical processing, generating fusion perception vector with improved e-commerce product recognition precision, specifically comprising according to requirement of multi-modal data processing, performing preprocessing on original input information by using data cleaning tool, extracting relevant features from original input information, and obtaining data set after preliminary screening; for data set after preliminary screening, processing redundant content by using information filtering tool, obtaining simplified information distribution, and determining integrity of information distribution.
[0014] Preferably, step S7 further includes, if the simplified information distribution reaches a preset threshold, using a priority sorting tool to adjust the information distribution hierarchy, obtaining the sorted information hierarchy, and judging the matching degree of the information hierarchy; based on the sorted information hierarchy, using a feature integration tool to uniformly process the information hierarchy, obtaining the final fusion perception vector, and determining the applicability of the fusion perception vector.
[0015] As can be seen from the above technical solution, the present invention has the following beneficial effects: This deep learning-based training method for a visual-tactile fusion model of e-commerce products utilizes attention mechanisms, context dependency analysis, and time-series methods to obtain the dynamic contribution of each modality at different processing levels. This effectively avoids the shortcomings of traditional methods, such as static weights that cannot be adjusted according to task changes. By adjusting weight ratios, the method increases the proportion of visual resources when visual contributions are insufficient, and combines visual feature enhancement methods to strengthen shape and position information, enabling the model to maintain higher robustness in recognizing small objects or complex edge structures. Furthermore, this invention introduces fusion priority ranking and information loss control techniques, ensuring that visual data maintains a more stable dominant position in the primary task while ensuring that tactile data plays a role in supplementing details. The invention achieves maximum auxiliary effect by evaluating the interaction between vision and touch through cross-modal interaction and dynamically determining whether to increase the visual weight, thus solving the problem that existing models cannot flexibly adapt to the complexity of different product structures. When the model's ability to capture minute edge features is insufficient, a dynamic proportional feedback mechanism is used to redistribute the data ratio between vision and touch, improving the model's bidirectional sensitivity to geometric and texture details. Finally, the invention generates optimized fusion perception vectors through data redundancy filtering and hierarchical processing priority update mechanisms, effectively reducing interference from irrelevant information, enhancing the training efficiency and recognition accuracy of deep learning models in large-scale e-commerce scenarios, and significantly improving the overall performance of intelligent recognition and attribute understanding of e-commerce products. Attached Figure Description
[0016] Figure 1 This is a flowchart of the training method for the visual-tactile fusion model of e-commerce products based on deep learning, as described in this invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] like Figure 1As shown, the present application provides a technical solution: an e-commerce product visual and tactile fusion model training method based on deep learning, comprising S1, collecting visual data and tactile data of e-commerce products through sensors, extracting shape position features for visual data, extracting pressure texture features for tactile data, using modal correlation evaluation technology based on deep learning to analyze the mutual information measure between the two, and obtaining a preliminary multi-modal feature combination; S2, according to the preliminary multi-modal feature combination, using attention mechanism to calculate dynamic weight, combining context dependence analysis and time series technology to evaluate the contribution of visual and tactile data in the current level processing, and determine the initial weight distribution of each modal; S3, if the initial weight distribution shows that the contribution of visual data is lower than that of tactile data, then increase the resource allocation proportion of visual data through weight proportion adjustment technology, and release shape position information using visual feature enhancement method, to obtain a preliminary fused perception vector; S4, for the preliminary fused perception vector, using fusion priority sorting technology to place visual data in a dominant position, balancing the auxiliary information of tactile data through information loss control method, and generating a calibrated fusion perception vector; S5, according to the calibrated fusion perception vector, combining task demand matching technology to analyze the demand of current operation on e-commerce product object edge recognition, evaluating the interaction influence of visual and tactile data through cross-modal interaction quantity, and judging whether it is necessary to adjust the proportion of visual weight; S6, if the conductivity of visual data after weight proportion is not enough to capture the details of object edge, then readjust the allocation proportion of visual data and tactile data through dynamic proportion feedback technology, enhance the processing depth of shape position information, and determine the optimized multi-modal perception combination; S7, according to the optimized multi-modal perception combination, using data redundancy filtering technology to remove irrelevant information, and updating the fusion process through hierarchical processing priority division, to generate a fusion perception vector with improved e-commerce product recognition accuracy.
[0019] In the above scheme, the shape, position, and pressure distribution, texture changes, and other original data of the e-commerce product are obtained through visual and tactile sensors. A deep convolutional neural network is used to analyze the spatial geometry of the visual data, and a tactile coding network is used to extract the local response features of the pressure texture. The system calculates the mutual information based on the modal correlation evaluation technology to identify the feature correlation between the visual and tactile modalities on the same product target, and generates a preliminary multi-modal feature combination accordingly. Subsequently, the system dynamically adjusts the contribution of different modalities at different processing stages by combining attention mechanisms, and identifies the dynamic relationship between features and operation steps through context-dependent analysis and time series analysis techniques to determine the initial weight distribution. When the visual contribution is low, the system calls the weight ratio adjustment module to increase the resource occupation of visual features, and uses visual feature enhancement methods to highlight the details of shape and position, so that the preliminary fused perception vector can better express the spatial characteristics. After generating the preliminary fusion vector, the system uses fusion priority sorting technology to prioritize visual information as the dominant channel, while balancing the tactile modality through information loss control to provide necessary but not excessive detail supplementation, and forms the calibrated fusion perception vector. The system then determines the specific visual intensity required for object edge recognition in product identification based on task demand matching technology, and evaluates the coupling of vision and touch in specific tasks through cross-modal interaction. When the visual guidance is insufficient to capture edge details, the system uses dynamic proportion feedback technology to update the modal allocation ratio again, thereby enhancing the participation of the visual modality in shape and position reasoning and ultimately obtaining an optimized multi-modal combination. After redundancy filtering and hierarchical processing priority updating, the output fusion perception vector can maintain effective information while reducing cross-modal conflicts, improving the recognition accuracy of e-commerce products.
[0020] The above embodiments build an effective visual-tactile fusion processing mechanism, enabling the system to automatically complete dynamic weight adjustment in scenarios where multi-modal data correlation is weak or visual information is insufficient, improving the ability to capture visual key content such as product shape, edge, and texture. The modal mutual information metric and attention mechanism are used to accurately evaluate the contribution of modalities, improving the rationality and reliability of feature fusion. The introduction of visual feature enhancement, fusion priority sorting, and information loss control helps to balance the expressiveness and robustness of the fusion vector, reducing feature redundancy and directional bias. Task demand matching and cross-modal interaction evaluation improve the adaptability of the model in specific recognition tasks, ensuring that the model can automatically amplify the role of key modalities according to recognition requirements. Dynamic proportion feedback technology and hierarchical redundancy filtering are used to further optimize the fusion results, making the final generated fusion perception vector more accurate and generalizable in various e-commerce product recognition tasks, significantly improving the stability, recognition efficiency, and overall performance of the model.
[0021] S1 includes acquiring visual data of the e-commerce product through an image acquisition device, extracting shape and position features from the visual data using an edge detection tool to obtain a shape position feature set; acquiring haptic data of the e-commerce product through a pressure sensor, extracting pressure texture features from the haptic data using a texture analysis tool to obtain a pressure texture feature set; calculating mutual information between the shape position feature set and the pressure texture feature set using a deep neural network tool, if the mutual information value is higher than a preset threshold, it is judged that the two have modal correlation, and a preliminary multi-modal feature combination data is obtained; for the preliminary multi-modal feature combination data, the shape position features and the pressure texture features are integrated by a feature fusion tool, and the final multi-modal feature representation is obtained.
[0022] In the process of acquiring visual data, the image acquisition device is fixedly installed on the shooting support, the camera lens is kept facing the e-commerce product during shooting, and the shooting distance is fixed at about 30-50 cm to ensure the picture clarity. After the visual image is collected, the system reads the image pixel by pixel to form a two-dimensional matrix composed of pixel brightness values. When the edge detection tool scans the matrix, it calculates the brightness value difference between each pixel and its adjacent pixels. The specific calculation method is as follows: after the system reads the brightness value of the target pixel, it reads the brightness values of the right and lower pixels, and calculates the difference between the two and the brightness value of the target pixel. In order to determine whether a pixel belongs to an edge point, the system sets the edge detection threshold to 20, which is derived from a statistical experiment on 100 e-commerce product sample images. In the experiment, the threshold was gradually adjusted from 10, 15, 20, 25 and the contour line clarity was observed. When the threshold is 20, the stability of edge recognition is the highest, so it is determined as the final threshold. When the brightness difference between a pixel and its adjacent pixels is greater than 20, the system clearly marks the pixel as an edge point. The system then connects all the pixels marked as edge points, and confirms whether the edge points are continuous in sequence from left to right and top to bottom. If two edge points differ by 1 pixel or less in coordinate position, they are determined to be part of the same contour line. Finally, the system records the coordinate values of all continuous edge points, with each coordinate point consisting of three elements: its horizontal position value, vertical position value, and gradient intensity with adjacent points. The system arranges all edge points in scanning order to form a complete shape position feature set and stores it in the feature cache area.
[0023] Subsequently, the tactile data is acquired by collecting the pressure information on the surface of the e-commerce product point by point through the pressure sensor. The pressure sensor is a fixed-area matrix sensing surface. During one contact process, the sensor records the pressure value of each sampling point at a frequency of 100 times per second. After sampling is completed, the system arranges the continuous pressure values into a pressure sequence. After the texture analysis tool receives the sequence, it performs denoising processing. The denoising method is to take the average of every three adjacent pressure sampling values by using a 3-point sliding window, so that the fluctuation of the pressure sequence is smoother. Subsequently, the system calculates the pressure difference between each two consecutive sampling points and records these difference values. In order to determine whether a certain segment of tactile signal contains texture characteristics, the system calculates the standard deviation of a number of consecutive pressure difference values. When the standard deviation is greater than 0.3, the system considers that the pressure change has regular fluctuations, and thus determines that there is a pressure texture feature at this position. The determination method of the standard deviation threshold 0.3 is as follows: when collecting tactile data on 50 types of e-commerce products with different surface materials, the standard deviation of the pressure difference is gradually calculated. In many experiments, when the standard deviation reaches 0.3, the consistency of manual judgment and system judgment of whether the texture exists is the highest, so the threshold is determined to be 0.3. The system records all positions judged as texture points, and combines the pressure intensity, pressure peak position and pressure change trend direction corresponding to each texture point into a pressure texture feature, and finally forms a pressure texture feature set.
[0024] After the visual and tactile feature extraction is completed, the system calls a deep neural network tool to calculate the mutual information metric between the two types of features. The calculation process is as follows: the system inputs the shape position feature set into the first feature channel in order, and inputs the pressure texture feature set into the second feature channel in the same order. Each channel is composed of multiple fixed numerical ratio feature compression steps, and each compression step reduces the numerical range of the input features to a comparable range by a fixed ratio. The system counts the frequency of occurrence of the features at the end of each channel to obtain the statistical distribution of the visual features and the statistical distribution of the tactile features, respectively. Then the system selects the visual features and tactile features at the same index position to form feature pairs, and counts the joint occurrence frequency of these feature pairs. The system stores the visual feature distribution, tactile feature distribution and joint distribution in the cache respectively, and then calculates the distribution difference in the same batch. In order to realize the mutual information metric, the system compares the difference between the high and low values of the joint distribution and the high and low values of the individual distribution for each feature pair. After the system repeats the above comparison process for a large number of feature pairs, it accumulates a mutual information value. The higher the mutual information value, the closer the visual features and tactile features are in numerical distribution, and the stronger the correlation between the features. The system sets the mutual information threshold to 0.6. The method of determining this threshold is as follows: the system uses 200 sets of visual and tactile feature samples for experiments, and respectively statistics their mutual information values, and gradually tests in the range of 0.4 to 0.8. Whenever the mutual information value is higher than 0.6, the probability that the visual and tactile feature pairs correspond to the same product structure is always higher than 90%, so the final mutual information threshold is determined to be 0.6. When the mutual information value is higher than 0.6, the system determines that there is a correlation between the shape position feature set and the pressure texture feature set, that is, the two are consistent in describing the same e-commerce product structure, and at this time the preliminary multi-modal feature combination data is generated. If the mutual information value is lower than 0.6, it will not enter the fusion step.
[0025] When the preliminary multi-modal feature combination data is obtained, the feature fusion tool is started to weight and integrate the shape position features and the pressure texture features. The steps of the weighted integration are as follows: the system first sets the initial weight of the shape position features to 0.5 and sets the initial weight of the pressure texture features to 0.5. Then the system adjusts the weights according to the contribution proportion of each feature in the mutual information metric. For example, the system statistics that the contribution proportion of the shape position features in the formation process of the overall mutual information value is 60%, and the contribution proportion of the pressure texture features is 40%, so the system adjusts the weight of the visual features to 0.6 and adjusts the weight of the tactile features to 0.4. The determination process of the weight is: the system statistics the frequency of each type of feature appearing in the joint distribution, and takes the frequency proportion as the contribution degree, and then takes the contribution degree proportion as the new weight value. The system performs a weighted calculation at each feature index position, multiplies the visual feature value by the visual weight, multiplies the tactile feature value by the tactile weight, and then adds them to obtain a single output value of the fused features. The system calculates all the features one by one to finally form a complete multi-modal feature representation. The multi-modal feature representation is arranged in a fixed order, each element contains the fused spatial position feature strength and the pressure texture response value, and its structure is completely compatible with the subsequent model training interface and can be directly used for model training.
[0026] S2 includes extracting the current level feature representation from the visual data and the tactile data according to the preliminary multi-modal feature combination, performing standardization processing by using the data preprocessing tool to obtain a unified feature representation set; for the unified feature representation set, analyzing the historical change trend of the visual data and the tactile data by using the time series processing tool to obtain the performance difference of each modality at different time points and determine the time correlation data of each modality; if the performance difference of a certain modality in the time correlation data is higher than a preset threshold, adjusting the dynamic weight of the modality by using the attention calculation tool, analyzing its effect by using the context dependence tool, and obtaining the adjusted weight distribution result; according to the adjusted weight distribution result, comprehensively evaluating the contribution degrees of the visual data and the tactile data by using the data integration tool, judging the initial weight distribution of each modality in the current level, and obtaining the final modality weight distribution data.
[0027] In the embodiment, firstly, according to the preliminary multi-modal feature combination, the feature representation used in the current level is extracted from the visual data and the tactile data respectively, in particular, the system reads the visual index position recorded in the preliminary multi-modal feature combination, then extracts the shape change value, the contour intensity value and the position offset of the corresponding position from the visual data, and arranges these values in the order of the index to form the visual current level feature sequence; at the same time, the system reads the pressure sampling value, the pressure change trend value and the pressure fluctuation amplitude of the same index position in the tactile data, and arranges them into the tactile current level feature sequence. Since the visual sequence and the tactile sequence are different in the range of values, the distribution mode and the change rhythm in the original form, which will interfere with the subsequent analysis, therefore, the system then calls the data preprocessing tool to standardize the two types of feature sequences. The execution mode of the standardization step is that the system first scans all the values of the visual feature sequence, records the maximum value and the minimum value, then scans all the values of the tactile feature sequence, records the maximum value and the minimum value, and calculates the scaling proportion interval according to the four values. When the maximum value and the minimum value of the visual feature sequence are greatly different, for example, the maximum value is 80 and the minimum value is 5, the system will reduce each value in the visual feature sequence by 5 and then scale it down by the proportion, so that the converted value is between 0 and 1; similarly, when the maximum value of the tactile feature sequence is 3 and the minimum value is 0.2, the system will reduce each tactile feature value by 0.2 and then scale it down by the proportion, so that it also falls within the range of 0 to 1. The proportion of standardization is formed by the difference between the maximum value and the minimum value, which is a fixed and determined scaling value, and does not need human intervention. After the above processing is completed, the system combines the standardized visual feature sequence and the tactile feature sequence to form a unified feature representation set, in which each visual feature and tactile feature is converted into a value in the same range, which can be directly compared and analyzed.
[0028] After the unified feature representation set is generated, the system calls a time series processing tool to analyze the historical change trend of the visual features and the tactile features in the previous layer to the current layer. The specific implementation manner is as follows: the system first arranges the visual feature sequence in chronological order into a historical change record, for example, records the value of the visual feature at time point 1, the value at time point 2, the value at time point 3, and so on; then the system calculates the value difference between each two consecutive time points, for example, the visual feature increases by 0.1 or decreases by 0.05 at time point 2 compared with time point 1, and records all the change values to generate a change trend sequence of the visual features. Similarly, the system performs the same steps on the tactile feature sequence, that is, records all the tactile values in chronological order, and then calculates the value change between adjacent time points, and records these changes as a tactile change trend sequence. Then the system respectively calculates the average change amplitude, the maximum change amplitude and the minimum change amplitude of the two trend sequences. In order to determine the performance difference between the visual modality and the tactile modality in the time dimension, the system performs the following steps: at each time point, the system reads the visual change amplitude and the tactile change amplitude and directly compares them, for example, the visual change amplitude at a certain time point is 0.15, and the tactile change amplitude is 0.04, then the system records the difference value of 0.11 that the visual is greater than the tactile at this time point; if the visual change is 0.05 and the tactile change is 0.20, the system records the difference value of 0.15 that the tactile is greater than the visual. The system accumulates all the difference values at the time points to form complete time correlation data, which is used to represent the stability and dominance of the visual modality or the tactile modality in the change amplitude at multiple time points.
[0029] In order to determine whether the modality weight needs to be adjusted, the system presets the performance difference threshold value as 0.2, and the specific threshold value determination manner is as follows: in the experimental stage, the system selects 100 groups of real sampling data of visual and tactile, adjusts the performance difference threshold value as 0.1, 0.15, 0.2, 0.25, 0.3 respectively, and observes the consistency between the weight adjustment result and the manual annotation of the modality importance judgment under each threshold value. After multiple experiments, when the threshold value is set as 0.2, the consistency between the automatic weight adjustment and the manual judgment reaches more than 90%, and this threshold value can adapt to the tactile structure difference of different types of e-commerce products, so 0.2 is taken as the finally determined performance difference threshold value. When the difference value of a time point or multiple time points in the time correlation data exceeds 0.2, the system determines that the modality has a significantly different performance in the current level and needs to be dynamically adjusted. If the visual change amplitude is greater than the tactile change amplitude at multiple time points and the difference exceeds 0.2, the system determines that the visual modality should increase the weight; if the tactile performance is stronger, the tactile modality weight is increased.
[0030] When the need for weight adjustment is detected, the attention calculation tool will be called to perform dynamic weight update. The specific way of attention update is: the system reads the difference value, for example, when the difference is 0.3, the system increases the weight of the dominant modality by 0.1 according to the fixed ratio, while reducing the weight of the non-dominant modality by 0.1 to keep the total weight sum to 1. If the difference is 0.25, the system increases the weight of the dominant modality by 0.08, and reduces the weight of the other modality by the same 0.08. If the difference is just close to the threshold 0.2, the system will increase the weight of the dominant modality by 0.05, and reduce the weight of the other modality by 0.05. All the change amplitudes are within the determined numerical range, ensuring the stable execution of the system. After the weight adjustment is completed, the system inputs the adjusted weight distribution result into the context-dependent tool for analysis to determine whether further optimization is needed in combination with historical level trends. The specific steps of context-dependent analysis are: the system reads the historical weight records of the past 5 levels, for example, and calculates the average weight of vision and touch in the past 5 levels. If the vision weight has always been greater than the touch weight in the historical levels, for example, the average value of vision is 0.62 and the average value of touch is 0.38, the system will increase the current vision weight by a stability compensation value of 0.03, and at the same time reduce the touch weight by the same 0.03, to ensure the continuity of weight change. If the touch modality is stronger in the historical levels, the same logic compensation value adjustment is performed on the touch modality, for example, by 0.03. After the context-dependent tool completes the calculation, it outputs a weight distribution result that has been context-corrected.
[0031] Then the data integration tool is called to comprehensively evaluate the contribution of visual data and tactile data according to the weight distribution result. The process of comprehensive evaluation is: the system reads the visual value and tactile value of each index position in the unified feature representation set, multiplies each visual value by the context-corrected vision weight, multiplies each tactile value by the context-corrected touch weight, and then compares each pair of weighted visual value and weighted tactile value. For example, if the visual weighted value is 0.48 and the tactile weighted value is 0.32, the vision contribution is greater; if the visual weighted value is 0.20 and the tactile weighted value is 0.45, the tactile contribution is greater. The system repeats this comparison process for all index positions, and calculates the vision modality contribution ratio and the touch modality contribution ratio. When the vision contribution ratio exceeds the touch contribution ratio, for example, the vision contribution ratio is 0.58 and the touch contribution ratio is 0.42, the system sets the vision modality as the larger initial weight value of the current level, for example, 0.6, and sets the touch modality as the smaller value, for example, 0.4; otherwise, the opposite setting is performed. Finally, the system takes the initial weight as the final modality weight distribution data of the current level.
[0032] S3 comprises obtaining a contribution degree comparison result of the visual data and the tactile data according to the initial weight distribution data, and determining an allocation situation of the visual data in the current level; if the contribution degree of the visual data is lower than that of the tactile data, dynamically improving a resource allocation proportion of the visual data by a proportion adjustment tool to obtain an adjusted allocation scheme; for the adjusted allocation scheme, separating shape position information from the visual data by a feature extraction tool, and combining an enhancement tool to strengthen details to determine a strengthened feature set; and integrating the strengthened feature set and the tactile data by a data fusion tool to generate preliminary perception vector data, judging whether the preliminary perception vector data meets a preset standard, and obtaining a final fusion result.
[0033] In the embodiment, first, the visual features and the tactile features of each index position in the unified feature representation set are weighted calculated according to the initial weight distribution data generated in the previous step. The visual feature value multiplied by the visual weight becomes the visual weighted contribution value, and the tactile feature value multiplied by the tactile weight becomes the tactile weighted contribution value. The system traverses all index positions one by one, compares the size of each pair of visual weighted contribution value and tactile weighted contribution value, accumulates the number of times that the visual weighted contribution value is greater than the tactile weighted contribution value as the visual contribution count, and accumulates the number of times that the tactile weighted contribution value is greater than the visual weighted contribution value as the tactile contribution count. Then, the two counts are respectively divided by the total number of indexes to obtain the visual contribution proportion and the tactile contribution proportion. If the visual contribution proportion is lower than the tactile contribution proportion, for example, the visual contribution proportion is 0.42 and the tactile contribution proportion is 0.58, the system confirms that the actual contribution degree of the visual data in this level is low. At this time, the system calls the proportion adjustment tool to perform a dynamic improvement operation according to the contribution difference value. The contribution difference value is calculated by subtracting the visual contribution proportion from the tactile contribution proportion, for example, the difference value in the foregoing example is 0.16, which determines the magnitude of the system to improve the visual weight. According to the fixed weight adjustment rule, the system corresponds the difference value to a preset proportion, for example, when the difference value is between 0.10 and 0.20, the system improves the visual weight by 0.1 and reduces the tactile weight by 0.1; if the difference value exceeds 0.20, for example, reaches 0.25, the system improves the visual weight by 0.15 and reduces the tactile weight by 0.15; if the difference value is small, for example, only 0.05, the system improves the visual weight by 0.03 and reduces the tactile weight by 0.03. All adjustment values are obtained by contribution proportion statistics of a large number of training samples, which ensures that the weight improvement magnitude is sufficient to correct the modal contribution deviation without causing system instability. The system records the adjusted visual weight and tactile weight as a new allocation scheme for the visual information strengthening process in the current level.
[0034] Subsequently, according to the adjusted allocation scheme, the shape position information is re-separated from the visual data through a feature extraction tool. Specifically, the system reads all shape boundary points, contour change points, coordinate positions, and boundary brightness gradients in the original index order of the visual data, and generates a new shape position information sequence one by one. In order to enhance the expression ability of the visual modal, the system calls an enhancement tool to perform a detail enhancement operation. The execution process of the detail enhancement includes three data processing steps. The first step is boundary continuity scanning processing. The system checks the distance between two consecutive points along the order of all visual boundary points. If the distance between two consecutive points exceeds 1 pixel, the system automatically inserts one or more intermediate points between the two points to make the boundary more continuous. The second step is local gradient enhancement processing. The system reads the boundary brightness gradient change value point by point and compares it with the preset gradient highlight threshold value. The threshold value is statistically set to 0.3, which is the best discrimination standard obtained by analyzing a large number of product edge samples. When the gradient change value of a certain point exceeds 0.3, the system increases the gradient value of the point by 0.1, so that it is more easily identified in the fusion process. If the gradient change value of a certain point is lower than the set minimum gradient intensity 0.05, the system increases the gradient value of the point to 0.05, ensuring that all boundary points have identifiable minimum intensity. The third step is shape smoothing and stabilizing processing. The system scans all enhanced shape points in order to determine whether there is a sudden jump, such as a sudden difference in brightness between a certain point and the previous two points. If such a jump occurs, the system adjusts the brightness value of the point to the average value of the brightness values of the previous and next points, making the entire boundary structure more stable. After the above enhancement process, the system combines all the enhanced shape position information in order to form a final enhanced feature set.
[0035] After the generation of the set of enhanced features, the system calls the data fusion tool to perform the integration of the visual enhancement features and the haptic features. The integration process is as follows: the system traverses all index positions, reads the visual enhancement feature values and the haptic feature values, multiplies the visual enhancement feature values by the latest visual weight, multiplies the haptic feature values by the latest haptic weight, and then adds the two values to form the fusion value of the current index position. The system repeats this operation for all indexes to finally generate the preliminary perception vector data. Subsequently, the system performs standard compliance detection on the perception vector data item by item. The determination method of the preset standard is to evaluate the minimum weight proportion of visual enhancement information required in the fusion result through a large number of training samples. When the proportion of visual enhancement part is found to be less than 0.45 in the statistical process, the model recognition accuracy decreases obviously, so the system sets the proportion of visual enhancement weighting contribution not less than 0.45 as the preset standard. When the system detects that the proportion of visual enhancement weighting contribution in the perception vector is less than 0.45, it is judged as not meeting the standard, and returns to the weight adjustment process; if it is detected that all index positions meet the preset standard, for example, the visual contribution proportion is in the interval of 0.48 to 0.55, the system judges that the perception vector is valid, records it as the final fusion result, and outputs the fusion result for the next step processing. The fusion result is a vector composed of a plurality of fusion values arranged in order, and each fusion value represents the final perception information after the joint action of visual and haptic.
[0036] S4 includes adjusting the priority of visual data according to the preliminary fused perception vector, obtaining the adjusted data level distribution, and determining its dominance; for the adjusted data level distribution, extracting the auxiliary information of the haptic data through the balancing tool to obtain the key feature set, judging whether it meets the preset threshold, and obtaining the feature balanced information combination; if the feature balanced information combination meets the preset threshold, the dominant information of the visual data and the auxiliary information of the haptic data are fused by the integration tool to obtain the integrated data set, and the preliminary structure of the fusion vector is determined; the integrated data set is processed by the calibration tool to obtain the final fusion perception vector, and whether the preset balance standard is reached is judged to obtain the calibrated perception vector data.
[0037] In the embodiment, firstly, the visual weighting values and the tactile weighting values of each index position are read according to the preliminary fused perception vector, and the priority of the visual data is adjusted by using a sorting tool. The specific processing manner of the sorting tool is that the system reorders all the visual weighting values in the preliminary fused perception vector in the order from the maximum to the minimum, and records the position of each visual weighting value after the ordering, for example, the visual weighting value of 0.72 is recorded as the ordering position 1, the visual weighting value of 0.66 is recorded as the ordering position 2, and so on. The system forms the hierarchical distribution of the visual data according to the above ordering result, and calculates the mean of the weighting values of the front positions in the visual ordering sequence, for example, the visual weighting values in the front 10% are taken to be averaged, when the mean is greater than 0.60, the system determines that the visual data has a dominant position in the current hierarchy, and the limit 0.60 is from the division point at which the visual dominant effect is obviously improved when 300 groups of training samples are statistically processed, and thus is determined as the dominant judgment parameter. After the visual hierarchical distribution is determined, the system calls a balancing tool to extract auxiliary information of the tactile data. The specific manner is that the system reads the tactile weighting value of the corresponding index position in the visual priority list after the ordering one by one, and reads the pressure change amount, the pressure fluctuation amplitude and the pressure peak position corresponding to the index position in the tactile data, and sequentially groups these tactile features to form a tactile key feature set. Then the system statistically processes the tactile key feature set, reads all the pressure change amount values, all the pressure fluctuation amplitude values and all the pressure peak position features in the set one by one, and divides the sum of the values by the number of elements in the set to obtain the mean of the tactile key features. The system compares the mean with a preset threshold, and the determination manner of the preset threshold is to statistically process the pressure change mean of the tactile information which can keep auxiliary effectiveness in the fusion process in a large number of training samples, and after analyzing a plurality of batches of data, when the mean of the tactile key features is greater than or equal to 0.30, the tactile mode can provide stable supplement under the visual dominant mode, and thus the system sets 0.30 as the tactile auxiliary effectiveness threshold. If the mean of the tactile key features is greater than or equal to 0.30, the system determines that the tactile data has the necessary condition to participate in the fusion and generates the information combination after the feature balancing. If it is less than 0.30, the system will record the current insufficient tactile information and return a prompt for optimizing the tactile sampling, but will not continue to enter the fusion process. If the information combination after the feature balancing meets the above threshold requirement, the system then calls an integration tool to fuse the visual dominant information and the tactile auxiliary information.The specific execution steps of the integration tool are as follows: the system reads the visual ranking value of each index position one by one, multiplies the visual ranking value by the visual dominance weight, for example, the determination method of the dominance weight is the proportional difference between the visual ranking mean value and the tactile key feature mean value, when the visual ranking mean value reaches 0.60 and the tactile key feature mean value is 0.30, the visual dominance weight is set to 0.65, and the tactile auxiliary weight is set to 0.35, the sum of the two is 1 and the proportional determination method is clear. The system then multiplies the tactile feature value by the tactile auxiliary weight, and adds the above two to obtain the integration value of the current index, and the system repeats this process for all indexes to form the integrated data set. After the integrated data set is generated, the system calls the calibration tool to calibrate it. The execution method of the calibration tool is as follows: the system first sets the balance standard of the visual and tactile fusion proportion, the determination process of the balance standard is to conduct error statistics on the fusion process of a large number of training models, it is found that when the visual weighting proportion is between 0.45 to 0.60 and the tactile weighting proportion is between 0.40 to 0.55, the recognition accuracy of fusion is the highest, therefore the interval is determined as the final balance standard. The calibration tool reads each integration value in the integrated data set one by one, and reverses the visual weighting proportion and the tactile weighting proportion from it, for example, the visual weighting proportion is 0.62 and exceeds 0.60, then the system performs visual proportion reduction operation on the index position, reduces the visual weighting value by 0.03; if the visual weighting proportion is lower than 0.45, for example, only 0.42, the system increases the visual weighting value by 0.03 to return it to the standard range. If the tactile weighting proportion exceeds 0.55 or is lower than 0.40, the system increases or decreases the tactile weighting value in the same way to return it to the standard interval. The calibration tool repeats the above adjustment steps for each index until the visual weighting proportion and the tactile weighting proportion of all index positions are kept within the above interval. When the calibration is completed, the system combines the calibration results of all indexes into the final fusion perception vector, and judges whether all indexes are within the balance standard range, if all indexes meet the conditions that the visual is not lower than 0.45 and not higher than 0.60, and the tactile is not lower than 0.40 and not higher than 0.55, the system marks the vector as the calibrated perception vector data, and as the final output result of this step.
[0038] S5 comprises extracting feature distribution of visual data and tactile data by using an analysis tool according to the calibrated fusion perception vector, obtaining a hierarchical structure of the feature distribution, and determining a dominant component in the hierarchical structure; through the hierarchical structure of the feature distribution, analyzing task requirements of object edge recognition by using a comparison tool, obtaining a feature weight distribution matched with the requirements, judging whether the distribution meets a preset threshold, and obtaining a preliminary matching result; if the preliminary matching result meets the preset threshold, calculating a cross-modal interaction amount of the visual data and the tactile data by using an evaluation tool, obtaining an influence range of the interaction amount, and determining a saliency degree of the influence range; according to the saliency degree of the influence range, dynamically correcting a weight proportion of the visual data by using an adjustment tool, obtaining a corrected weight configuration, judging whether the configuration meets the edge recognition requirements, and obtaining a final allocation scheme.
[0039] In the embodiment, firstly, according to the calibrated fusion perception vector, the visual weighting value and the tactile weighting value of each index position are read item by item, and the analysis tool is called to extract the feature distribution of the visual data and the tactile data in the fusion vector. The specific processing mode of the analysis tool is: the system reorders all the visual weighting values in descending order, and records the index position of each visual feature point in the original fusion vector at the same time; then the system sorts the tactile weighting values in descending order in the same way, and records the original index position. The system combines the visual sorting sequence and the tactile sorting sequence to form a multi-modal hierarchical structure arranged according to the weight size, which represents the hierarchical level of the visual feature point or the tactile feature point with the highest contribution degree from top to bottom. After the structure is generated, the system calculates the proportion of visual features and the proportion of tactile features in each level, for example, the number of times of visual features and the number of times of tactile features in the top 20% of the level are counted, and the two are divided by the total number of elements in the level to obtain the visual dominance and the tactile dominance. When the visual dominance is greater than the tactile dominance, for example, the visual dominance is 0.70 and the tactile dominance is 0.30, the system determines that the visual modality is the dominant component in the current hierarchical structure, and the determination standard is obtained by statistical analysis of 200 training samples. When the visual dominance is greater than 0.60, it contributes most to the overall recognition task, so 0.60 is selected as the dominant determination limit. After determining the dominant component, the system uses the comparison tool to analyze the hierarchical structure layer by layer to determine whether the current fusion vector meets the feature requirements of the object edge recognition task. The execution mode of the comparison tool is: the system reads the proportion of visual features and tactile features in each layer of the hierarchical structure one by one, and then judges whether it matches according to the feature requirements of the object edge recognition task. The edge recognition task requirement is determined by training statistics to be that the weighting proportion of visual features should not be less than 0.55 and the proportion of tactile features should not be higher than 0.45, because edge recognition relies on the shape change value, brightness gradient change value and spatial positioning information of the visual modality, and the tactile is mainly used to provide supplementary details, so the minimum proportion of visual features is 0.55 as the threshold. The comparison tool compares the proportion of visual features in the hierarchical structure with the threshold, when the proportion of visual features in a certain level structure is 0.60 or 0.65, it is determined that the structure meets the task requirement, if the proportion of visual features in a certain level is only 0.48, it is determined that it does not meet the requirement. The system summarizes the comparison results of all levels, if most of the levels (more than 70%) meet the minimum proportion requirement of visual features, the system generates a preliminary matching result and determines that it meets the preset threshold. If the preliminary matching result meets the preset threshold, the system will call the evaluation tool to continue calculating the cross-modal interaction amount of visual data and tactile data.The specific execution method of the evaluation tool is that the system reads the absolute difference between the visual weighting value and the tactile weighting value in index order, for example, the visual value is 0.62 and the tactile value is 0.38, then the difference is 0.24; the system adds all the difference values and divides by the total number of indexes to get the cross-modal average difference, which is the cross-modal interaction. The larger the interaction, the more obvious the difference between the two modalities at the index position, and the stronger the interaction. The influence range of the interaction is determined as follows: the system calculates the proportion of the interaction values greater than 0.20 according to the distribution of all interaction values, if the interaction of more than 50% of the indexes is greater than 0.20, it means that the influence range is significant; the value of 0.20 comes from the distribution analysis of a large amount of training data, it is found that when the visual and tactile difference is stable more than 0.20, the pressure difference provided by the tactile in the edge area has a significant effect on the visual boundary enhancement, therefore 0.20 is set as the cross-modal interaction significant threshold. The system determines whether the visual weight needs to be dynamically corrected according to the significant degree of the influence range. If the influence range is significant, for example, more than 50% of the index interaction is greater than 0.20, the system calls the adjustment tool to perform dynamic correction of the weight. The processing method of the adjustment tool is that the system first determines the correction amplitude according to the average value of the interaction, when the average interaction is 0.25, the system increases the visual weight by 0.05; when the average interaction is 0.30, the system increases the visual weight by 0.08; if the average interaction is more than 0.35, the system increases the visual weight by 0.10, and decreases the tactile weight by the same value to ensure that the total weight is unchanged. If the average interaction is only 0.15, the system keeps the original weight unchanged. After the weight correction is completed, the system performs edge recognition requirement test on the corrected visual weight and tactile weight, that is, it judges whether the proportion of the visual after correction reaches the minimum requirement of 0.55. If the proportion of the visual after correction reaches or exceeds 0.55, for example, 0.58 or 0.60, the system determines that the corrected weight can meet the edge recognition task requirement, and the corrected weight configuration is determined as the final allocation scheme; if the proportion of the visual after correction is still less than 0.55, for example, less than 0.50, the system returns to the process of increasing the visual weight again until the task requirement is met. Finally, the system outputs the final allocation scheme that meets the edge recognition requirement.
[0040] S6 includes according to the demand of visual detail capture, using image processing tools to original visual data feature enhancement processing, from which to obtain the edge shape related feature information, judge whether the feature information meets the preset threshold; if the feature information does not reach the preset threshold, the tactile information is supplemented to the original visual data through the data fusion tool, the enhanced mixed data distribution is obtained, and the contribution degree of the mixed data distribution is determined; for the mixed data distribution, the allocation proportion of visual data and tactile data is dynamically corrected by using the proportion adjustment tool, the new data interaction balance state is obtained, and whether the state meets the detail resolution requirement is judged; according to the data interaction balance state, the multi-modal information fusion is optimized by using the information integration tool, the final perception combination configuration is obtained, and the adaptability of the configuration to edge recognition is determined.
[0041] In the embodiment, first, according to the requirement of visual detail capture, the original visual data is called for image processing tool for feature enhancement processing, the execution mode of the processing is that the system scans the original visual image pixel by pixel, reads the brightness value of all pixels and the brightness difference of adjacent pixels, and calculates the local contrast of each pixel, that is, the average value of the brightness difference between the current pixel and its four adjacent pixels above and below, such as the contrast of a certain pixel is 0.12 or 0.15, which indicates that there is certain detail structure; the system calculates the local contrast of all pixels, and counts the average contrast, the highest and the lowest contrast of the whole image, to form a feature information set related to edge shape. When the system judges the average value of the statistical visual edge features and the preset threshold, if the average value reaches the preset threshold, for example, reaches 0.25 or more, the system considers that the visual features can meet the requirement of detail capture; the threshold 0.25 is derived from the contrast statistical analysis of 200 e-commerce product images, when the average contrast of the image reaches 0.25 or more, the edge curve, corner and texture mutation can be clearly coded, so it is established as the minimum visual detail requirement of this step. If the visual feature information does not reach the preset threshold, for example, the average contrast is only 0.18, the system calls the data fusion tool to supplement the tactile data to the original visual data for enhancement. The specific execution mode of the tactile supplement is that the system finds the corresponding tactile sampling point for each visual pixel position, and adds these tactile values to the brightness change of the visual pixel in proportion, for example, when the brightness gradient of a certain pixel is 0.10 and the tactile pressure gradient is 0.08, the system increases the visual gradient by a fixed proportion, for example, by 0.04, so that the mixed gradient reaches 0.14, thereby forming the enhanced mixed data distribution. Then the system repeats the supplement process for all pixels to form the mixed data distribution of the whole image, and performs contribution analysis on the mixed data distribution, the contribution analysis mode is that the system calculates the difference between the mixed gradient and the original gradient pixel by pixel, and then divides the cumulative difference by the number of pixels to obtain the mixed contribution, for example, the contribution is 0.06 or 0.08. When the mixed contribution reaches the visual enhancement standard, for example, reaches not less than 0.05, the system considers that the tactile supplement plays a correct role. Then the system calls the proportion adjustment tool for the mixed data distribution to dynamically correct the allocation proportion of visual data and tactile data, the correction mode is that the visual weight increase and the tactile weight decrease are determined according to the mixed contribution, for example, when the contribution is 0.06, the system increases the visual weight by 0.05 and decreases the tactile weight by 0.05; if the contribution is 0.10, the system increases the visual weight by 0.08 and decreases the tactile weight by 0.08; all the adjustment proportions are derived from the training sample statistics to ensure that they are within a reasonable range.When the weight correction is completed, the system checks the new data interaction balance state. The checking method is that the system calculates the average value of the enhancement gradient of all mixed pixel points. If the average value reaches the minimum requirement of detail resolution of 0.25, it is determined that the corrected balance state is valid. If the average value is still insufficient, for example, only 0.22, the system increases the visual weight again, continues to perform balance correction according to the fixed increment of 0.03, until the average gradient reaches or exceeds 0.25. When the data interaction balance state meets the detail resolution requirement, the system calls the information integration tool to further optimize the multi-modal information fusion processing. The optimization processing method is that the system integrates all enhanced visual information and tactile information by index, multiplies the final visual enhancement gradient of each pixel by the visual weight, multiplies the tactile auxiliary gradient by the tactile weight, and then adds the two to generate a fusion gradient value. The system repeats the above calculation for the entire image to obtain the final perception combination configuration. The system then makes an adaptive judgment according to the edge response intensity of the entire configuration. The judgment method is that the average value of all fusion gradients and the highest value and the lowest value of the edge region gradient are calculated. If the average value is not less than 0.25 and the difference between the highest gradient and the lowest gradient in the edge region is not less than 0.15, it means that the final configuration can clearly distinguish the shape change, boundary details, texture mutation and special-shaped contour of the e-commerce product, so it is determined that the configuration has good adaptability to edge recognition, and it is output as the final perception combination configuration.
[0042] S7 includes according to the requirement of multi-modal data processing, using data cleaning tool to preprocess the original input information, extracting relevant features to obtain the data set after preliminary screening; for the data set after preliminary screening, the redundant content is processed by information filtering tool to obtain the simplified information distribution, and the integrity of the information distribution is determined; if the simplified information distribution reaches the preset threshold, the priority sorting tool is used to adjust the level of the information distribution to obtain the sorted information level, and the matching degree of the information level is judged; according to the sorted information level, the feature integration tool is used to uniformly process the information level to obtain the final fusion perception vector, and the applicability of the fusion perception vector is determined.
[0043] In the embodiment, first, according to the requirements of multi-modal data processing, the original input information is preprocessed by calling a data cleaning tool. The execution mode of the data cleaning tool is as follows: the system reads the brightness value, gradient value and spatial position value of each pixel in the visual data one by one, and performs validity detection on each pixel. When the brightness value of a certain pixel is zero, exceeds 255, or the gradient value is negative, it is marked as an invalid data point, and the average of the brightness values of the four pixels above and below the pixel is replaced to ensure that the replacement result is within the true brightness range. At the same time, the system reads the pressure intensity value, pressure rise, pressure drop and local pressure texture change of each touch data point. If a certain pressure value is lower than the physical lower limit of the sensor 0.01 or higher than the upper limit of the sensor 1.00, the touch point is marked as invalid data, and the average pressure value of the previous and next two touch points is replaced to make it consistent with the true measurement range of the touch sensor. After all the invalid data is replaced, the system matches each visual pixel with its corresponding touch point according to the mapping relationship between the visual sampling index and the touch sampling index, so as to obtain visual brightness gradient, visual contour mutation value, touch pressure peak value, touch pressure difference value and other information at each index position, and generate a preliminary screened data set from these information. Then the system calls an information filtering tool to perform a redundant deletion operation on the preliminary screened data set. The processing mode is as follows: the system reads all visual features in the screened set one by one, calculates the difference between the gradient values of two adjacent visual features, and determines that they are repeated if the difference is less than 0.02. The touch features are processed in the same way. If the adjacent touch pressure difference is less than 0.01, it is also considered as a repeated feature. These two thresholds are determined according to the statistical rules of 300 real e-commerce product data. Experiments show that when the visual difference is less than 0.02 or the touch difference is less than 0.01, it does not contribute to the multi-modal fusion process, so it is fixed as the redundant judgment threshold. The system removes all visual features and touch features that are judged to be redundant, and obtains a simplified information distribution. After the simplified information distribution is generated, the system performs integrity detection on it. The detection method is as follows: the sum of the number of remaining visual features and the number of touch features is calculated, and the retention ratio is obtained by dividing the total number of preliminary screened data. If the retention ratio is not less than 0.70, it is considered that the simplified information distribution has sufficient integrity. The 0.70 threshold is determined according to the model training result. When the retention ratio is less than 0.70, the model reasoning accuracy decreases significantly, so this ratio is determined as the minimum integrity standard. If the simplified information distribution reaches the integrity threshold, the system continues to call a priority sorting tool to adjust the hierarchy of the simplified information distribution. The specific processing mode of the sorting tool is as follows: the system sorts the visual features according to the gradient value from large to small. The larger the gradient value, the more obvious the image edge, so it is arranged in the front. The touch features are sorted according to the pressure peak value and pressure fluctuation amplitude from large to small. The higher the pressure peak value, the greater the touch difference, which is also arranged in the front level.Subsequently, the system cross-merges the visual sorting sequence and the tactile sorting sequence, and generates an information hierarchy in numerical size order. After generating the information hierarchy, the system judges its matching degree. The matching degree judgment method is: the system reads the number of visual features in the information hierarchy level by level, and divides the number by the total number of features in the level to obtain the visual proportion. Then the visual proportion is compared with the edge recognition minimum demand threshold 0.55. The visual proportion 0.55 comes from the analysis results of a large number of edge recognition data. Experiments show that when the visual proportion is not less than 0.55, the model has stable edge recognition ability, so it is fixed as the minimum proportion threshold. The system performs the above comparison on all levels, and counts the proportion of levels whose visual proportion reaches or exceeds 0.55. If the proportion exceeds 0.60, it means that the sorted information hierarchy has a high matching degree with the task demand, and the required visual and tactile distribution mode meets the model requirements. After the matching degree meets the requirements, the system calls the feature integration tool to perform unified fusion processing on the information hierarchy. The fusion processing method is: the system reads each visual feature in the level structure one by one, multiplies the gradient value of the visual feature by the visual integration weight, for example, the visual integration weight is 0.60, which is determined by the proportion of the dominant role of vision in the edge task. At the same time, for each tactile feature, multiply the pressure change amount by the tactile integration weight, for example, the tactile integration weight is 0.40, which is determined according to the contribution proportion of tactile in the texture change compensation process. The system adds the visual weighted value and the tactile weighted value of each index position to obtain the fused single-point feature value, and then combines all the index points in order to form the final fused perception vector. After the generation of the fused perception vector, the system judges its applicability. The applicability judgment method is: the system calculates the average gradient and the lowest gradient of all gradient values of the fused vector. If the average gradient is not less than 0.25 and the lowest gradient is not less than 0.10, it proves that the fused perception vector can not only capture the overall trend of the edge contour, but also retain the necessary local detail changes, so it is determined to be applicable to subsequent multi-modal recognition tasks, and it is output as the final perception vector.
[0044] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and changes can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A method for training an e-commerce product visual-haptic fusion model based on deep learning, characterized in that, The method comprises the following steps: S1, collecting visual data and tactile data of e-commerce products through sensors, extracting shape and position features from visual data, extracting pressure texture features from tactile data, using deep learning-based modal correlation evaluation technology to analyze the mutual information between the two, and obtaining a preliminary multi-modal feature combination; S2, according to the preliminary multi-modal feature combination, using attention mechanism to calculate dynamic weight, combining context dependence analysis and time series technology to evaluate the contribution of visual and tactile data in the current level processing, and determining the initial weight distribution of each modal; S3, if the initial weight distribution shows that the contribution of visual data is lower than that of tactile data, then increase the resource allocation proportion of visual data through weight proportion adjustment technology, and release shape position information by using visual feature enhancement method, to obtain a preliminary fused perception vector; S4, for the preliminary fused perception vector, using fusion priority sorting technology to place visual data in a dominant position, balancing the auxiliary information of tactile data through information loss control method, and generating a calibrated fusion perception vector; S5, according to the calibrated fusion perception vector, combining task demand matching technology to analyze the demand of current operation on e-commerce product object edge recognition, evaluating the interaction influence of visual and tactile data through cross-modal interaction quantity, and judging whether it is necessary to adjust the weight proportion of visual data.
2. The deep learning-based e-commerce product visual-haptic fusion model training method of claim 1, wherein: The S1 comprises: acquiring visual data of e-commerce products through an image acquisition device, extracting shape and position features from the visual data using edge detection tools to obtain a shape position feature set; acquiring tactile data of e-commerce products through a pressure sensor, extracting pressure texture features from the tactile data using texture analysis tools to obtain a pressure texture feature set; according to the shape position feature set and the pressure texture feature set, using a deep neural network tool to calculate the mutual information between the two, if the mutual information value is higher than a preset threshold, it is judged that the two have modal correlation, and a preliminary multi-modal feature combination data is obtained; for the preliminary multi-modal feature combination data, the shape position features and the pressure texture features are integrated by a feature fusion tool, and the final multi-modal feature representation is obtained.
3. The deep learning-based e-commerce product visual-haptic fusion model training method of claim 1, wherein: The S2 comprises: according to the preliminary multi-modal feature combination, extracting the current level feature representation from the visual data and the tactile data, using a data preprocessing tool for standardization processing to obtain a unified feature representation set; for the unified feature representation set, using a time series processing tool to analyze the historical change trend of visual data and tactile data, obtaining the performance difference of each modal at different time points, and determining the time correlation data of each modal; if the performance difference of a certain modal in the time correlation data is higher than a preset threshold, the dynamic weight of the modal is adjusted by an attention calculation tool, and the effect is analyzed by a context dependence tool to obtain an adjusted weight distribution result; according to the adjusted weight distribution result, using a data integration tool to comprehensively evaluate the contribution of visual data and tactile data, judging the initial weight distribution of each modal in the current level, and obtaining the final modal weight distribution data.
4. The deep learning-based e-commerce product visual-haptic fusion model training method of claim 1, wherein: The S3 comprises: According to the initial weight distribution data, the contribution degree contrast result of the visual data and the tactile data is obtained, and the allocation of the visual data in the current level is determined; If the contribution degree of the visual data is lower than that of the tactile data, the resource allocation proportion of the visual data is dynamically improved through the proportional adjustment tool to obtain an adjusted allocation scheme; According to the adjusted allocation scheme, the shape position information is separated from the visual data by using a feature extraction tool, and the details are strengthened by combining an enhancement tool to determine a strengthened feature set; The strengthened feature set and the tactile data are integrated by using a data fusion tool to generate preliminary perception vector data, and whether the preliminary perception vector data meets a preset standard is determined to obtain a final fusion result.
5. The deep learning-based e-commerce product visual-haptic fusion model training method of claim 1, wherein: The S4 includes: According to the preliminary fused perception vector, the priority of the visual data is adjusted by using a sorting tool to obtain an adjusted data level distribution, and the dominant position is determined; According to the adjusted data level distribution, auxiliary information of the tactile data is extracted by using a balancing tool to obtain a key feature set, and whether the key feature set meets a preset threshold is determined to obtain a feature-balanced information combination; If the feature-balanced information combination meets the preset threshold, the dominant information of the visual data and the auxiliary information of the tactile data are fused by using an integration tool to obtain the integrated data set, and a preliminary structure of the fusion vector is determined; The integrated data set is processed by using a calibration tool to obtain a final fusion perception vector, and whether the final fusion perception vector meets a preset balance standard is determined to obtain a calibrated perception vector data.
6. The deep learning-based e-commerce product visual-haptic fusion model training method of claim 1, wherein: The S5 includes: According to the calibrated fusion perception vector, the feature distribution of the visual data and the tactile data is extracted by using an analysis tool to obtain a hierarchical structure of the feature distribution, and the dominant component in the hierarchical structure is determined; According to the hierarchical structure of the feature distribution, a demand matching feature weight distribution is obtained by using a comparison tool to analyze the task demand of object edge recognition, and whether the distribution meets a preset threshold is determined to obtain a preliminary matching result; If the preliminary matching result meets the preset threshold, the cross-modal interaction amount of the visual data and the tactile data is calculated by using an evaluation tool to obtain an influence range of the interaction amount, and the significance of the influence range is determined; According to the significance of the influence range, the weight proportion of the visual data is dynamically corrected by using an adjustment tool to obtain a corrected weight configuration, and whether the configuration meets the edge recognition demand is determined to obtain a final allocation scheme.
7. The deep learning-based e-commerce product visual-haptic fusion model training method of claim 1, wherein, The S6 also includes: if the conductivity of the visual data after the weight proportion is insufficient to capture the details of the object edge, the allocation proportion of the visual data and the tactile data is readjusted by using a dynamic proportional feedback technology to enhance the processing depth of the shape position information, and an optimized multi-modal perception combination is determined, which specifically includes: According to the demand of visual detail capture, the original visual data is processed by using an image processing tool to enhance the features, and the feature information related to the edge shape is obtained, and whether the feature information meets a preset threshold is determined; If the feature information does not meet the preset threshold, the tactile information is supplemented to the original visual data by using a data fusion tool to obtain an enhanced mixed data distribution, and the contribution degree of the mixed data distribution is determined.
8. The deep learning-based e-commerce product visual-haptic fusion model training method of claim 7, wherein: The S6 also includes: For the mixed data distribution, the allocation ratio of visual data and tactile data is dynamically corrected by using a proportion adjustment tool to obtain a new data interaction balance state, and it is judged whether the state meets the detail resolution requirement; According to the data interaction balance state, the multi-modal information fusion is optimized by using an information integration tool to obtain the final perception combination configuration and determine the adaptability of the configuration to edge recognition.
9. The deep learning-based e-commerce product visual-haptic fusion model training method of claim 7, wherein, It also includes S7, according to the optimized multi-modal perception combination, using data redundancy filtering technology to remove irrelevant information, and through hierarchical priority division update fusion process, generating a fusion perception vector with improved e-commerce product recognition accuracy, specifically including: According to the demand of multi-modal data processing, the original input information is preprocessed by using a data cleaning tool to extract relevant features and obtain a preliminary screened data set; For the preliminary screened data set, the redundant content is processed by using an information filtering tool to obtain a simplified information distribution and determine the integrity of the information distribution.
10. The deep learning-based e-commerce product visual-haptic fusion model training method of claim 9, wherein: The S7 also includes: If the simplified information distribution reaches a preset threshold, a priority sorting tool is used to adjust the information distribution in layers to obtain a sorted information level and judge the matching degree of the information level; According to the sorted information level, the information level is uniformly processed by using a feature integration tool to obtain a final fusion perception vector and determine the applicability of the fusion perception vector.