Emergency rescue scene-oriented scalable cross-modal tactile signal generation method
Through the collaboration between edge servers and robotic equipment, deep learning and contrast learning methods are used to generate reliable tactile signals in emergency environments, solving the problem of rescue robots lacking tactile signals, and improving the operator's immersive experience and the practicality of the robot.
Patent Information
- Application Number
- CN202510213216.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The lack of tactile signals in the emergency environment by existing rescue robots, which makes it impossible for operators to fully understand the properties of objects, affecting fine operations. At the same time, the existing cross-modal haptic generation scheme is difficult to implement in an environment with large bandwidth fluctuations.
Through collaboration between edge servers and robotic devices, visual and audio features are extracted using residual neural networks and contrast learning methods, and cross-modal fusion is carried out through multi-headed attention blocks to achieve scalable semantic coding and tactile signal generation.
Provide reliable tactile feedback in a dynamically fluctuating emergency environment, enhancing the practicality of the robot and enhancing the immersive experience of human operators.
Smart Images

Figure CN120143975A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of adaptive tactile generation, and more particularly to a scalable cross-modal tactile signal generation method for emergency rescue scenarios. Background Art
[0002] The development of edge computing and artificial intelligence has promoted the integration of robotics and the Internet of Things (IoT), thus giving rise to the emergence of the Internet of Robotic Things (IoRT). This integration has made robots more intelligent and significantly expanded their applications in various fields, especially in the field of emergency rescue. Collaboration between humans and robots can significantly improve the rescue success rate and the survival rate of victims. In addition, robots can also be used to enter disaster areas that are inaccessible or dangerous for humans.
[0003] Most existing rescue robots use audio and visual signals to collect environmental information. The operator receives this information and controls the robot to perform rescue operations. Unfortunately, these robots rely only on audio-visual signals and cannot provide the operator with a comprehensive understanding of the object's properties, which is crucial for fine operations in emergency rescue. As a solution, many studies have shown that combining tactile signals can allow humans to experience virtual touch, which can reflect the physical properties of objects (such as hardness, elasticity, and viscosity), thereby facilitating robot operation. There are mainly two reasons why current rescue robots lack tactile signals: one is that real tactile signals are difficult to obtain, and the other is that bandwidth fluctuations are frequent and large in emergency environments.
[0004] On the one hand, due to the limitations of tactile sensors, tactile signals are not always available. Different from audio-visual signals, obtaining tactile signals is complex and time-consuming. The main method of obtaining real tactile signals is to use accelerometers or force sensors to capture the signals generated when sliding over the surface of an object, and these signals reflect the texture characteristics of the object. However, these sensors are not common in most existing rescue robots. Therefore, there is an urgent need to provide tactile signals to humans without directly collecting tactile signals.
[0005] On the other hand, the frequent and significant bandwidth fluctuations in emergency environments make it difficult to implement existing cross-modal tactile generation schemes because they lack scalability. Robots can easily obtain visual and audio signals through cameras and microphones, and researchers have conducted some studies on cross-modal tactile generation. For example, in Patent 1, "A Fine-Grained Tactile Signal Reconstruction Method Aided by Audio-Visual", cross-modal generation of touch is achieved through the encoding of audio-visual signals. However, it does not involve the transmission link and is merely a reconstruction algorithm at the receiving end. Patent 2, "A Source-Channel Joint Encoding and Decoding Method Based on a Cross-Modal Communication System", designs a cross-modal source codec to achieve fusion encoding and decoding of vision and touch and can realize the generation from vision to touch. However, it must rely on complete semantic information for reconstruction and is not suitable for emergency communication environments because dynamic bandwidth fluctuations can lead to semantic loss. Patent 3, "An Audio-Video Aided Tactile Signal Reconstruction Method Based on Cloud-Edge Collaboration", uses a cloud-edge collaboration strategy to achieve the generation of audio-visual signals into tactile signals. However, it directly transmits and reconstructs audio-visual signals, which poses a great challenge to network bandwidth and wastes computing power resources at the device end. Therefore, it is crucial to develop a highly adaptable cross-modal tactile generation scheme that can generate tactile feedback based on the received audio-visual information and work effectively even under poor network conditions. Summary of the Invention
[0006] To solve the above problems, the present invention discloses a scalable cross-modal tactile signal generation method for emergency rescue scenarios. By exploring the collaboration between edge servers and robot devices, the association between multi-modal signals is mined to provide reliable tactile feedback in a dynamically fluctuating emergency environment, enhance the practicality of the robot, and further improve the immersive experience of human operators.
[0007] To achieve the above object, the present invention is realized by the following technical solutions:
[0008] The present invention is a scalable cross-modal tactile signal generation method for emergency rescue scenarios, including the following steps:
[0009] Step (1): Use a residual neural network as a feature encoder to extract low-dimensional features from visual and audio data. However, since features of different modalities are maintained in their respective spaces, it is difficult to perform effective cross-modal interaction. Therefore, a contrastive learning method is used to align the extracted visual and audio features, alleviate modal differences, improve model performance, and provide support for the next step of feature fusion;
[0010] Step (2): Design an effective cross-modal fusion method that can exchange information between modalities and integrate key information beneficial to haptic generation in a complementary manner. This method first normalizes multi-modal features, calculates intermediate representations of different modalities, then uses a multi-head attention block to extract cross-modal interaction information, and finally merges the fused feature representations through element-wise summation to obtain fused features;
[0011] Step (3): Use the fused features extracted in the previous step to achieve scalable semantic encoding to obtain hierarchical semantic representations. At the same time, through the network situation awareness module, gain a comprehensive understanding of the current network environment, and select the optimal transmission strategy based on the network bandwidth;
[0012] Step (4): After receiving the semantic representation transmitted from the device side, the edge side performs semantic decoding and cross-modal haptic signal generation. According to different semantic granularities, the edge side adaptively adopts corresponding generation strategies, namely coarse-grained cross-modal retrieval, fine-grained cross-modal retrieval, and cross-modal generation;
[0013] Step (5): Use an optimization method to train the entire model, and finally obtain the optimal model parameters for the subsequent test phase;
[0014] Step (6): Input paired audio and visual signals into the optimal network model, which will extract the features of the audio and visual signals, perform fusion, and then execute the optimal haptic signal recovery strategy according to the current network state.
[0015] A further improvement of the present invention lies in that step (1) includes the following steps:
[0016] (1-1): To generate more powerful features and alleviate the problem of model overfitting caused by the increase in network depth, the present invention uses ResNet as a unified single-modal feature encoder instead of VGG. Specifically, input the visual image v i and the audio spectrogram a i into ResNet-18 to extract low-dimensional feature vectors f V and f A ;
[0017] (1-2): To alleviate the modality differences, use contrastive learning to align the features of different modalities, and use a similarity function to evaluate the similarity between the features extracted from different modalities. Contrastive learning aims to improve the similarity score between paired visual and audio signals, and vice versa. To construct paired audio-visual signals, positive samples are defined as features from different modalities but belonging to the same category, while negative samples are features from different modalities and categories.
[0018] The similarity function is defined as follows:
[0019]
[0020] Among them and represent the linear projection that maps features to a low-dimensional representation, V represents video features, and A represents audio features; the value range of S is from 0 to 1, and the higher the value, the greater the similarity between features.
[0021] For each visual and audio feature, calculate the similarity between vision-to-audio and audio-to-vision, and the formula is as follows:
[0022]
[0023] where B is the batch size, λ represents the temperature parameter, V b and b respectively represent the video features and audio features under the b-th sample;
[0024] Finally, the contrastive loss is defined as cross-entropy and can be expressed as:
[0025]
[0026] where and represent that the probability of the true similarity and the positive sample pair is 1, while the probability of the negative sample pair is 0.
[0027] A further improvement of the present invention lies in that: step (2) includes the following steps:
[0028] (2-1) Use two Transformer blocks for feature fusion, and each block consists of a multi-head attention block (MHA), a normalization layer (NL), and a multi-layer perceptron (MLP). The key factor for promoting cross-modal information exchange is to utilize MHA. In traditional Transformer, the MHA block captures the intrinsic context relationship in images or texts. In contrast, MHA is used to extract the semantic association between visual and audio features. The specific steps are as follows:
[0029] Step 211: First, normalize the input multi-modal features to obtain f V-norm and f A-norm .
[0030] Step 212: Then, use linear projection to calculate the intermediate representations of different modalities: query (Q), key (K), and value (V).
[0031] Step 213: After that, obtain the cross-modal interaction information f V-iter and f A-iter through MHA. The process is described as follows:
[0032]
[0033]
[0034] Among them, Q a , K v , V v , Q v , K a , V a are intermediate representations of visual and audio features, and d represents the dimension.
[0035] Step 214: Combine the outputs of the two Transformer blocks through element-wise summation to finally obtain f V-A . The whole process can be described as:
[0036]
[0037] A further improvement of the present invention lies in that step (3) includes the following steps:
[0038] (3-1) Use a neural network to extract a concise representation of the fused features to achieve scalable semantic encoding, that is, represent different semantic granularities through feature maps of different depths.
[0039] Step 311: The fused feature representation with complete semantics is denoted as f V-A , while the shallow and deep feature maps extracted from the fused features through the neural network are denoted as f V-A-s and f V-A-d respectively. The feature f V-A will be used for direct haptic generation, while f V-A-s and f V-A-d will be used for fine-grained and coarse-grained retrieval respectively, mainly for indirect haptic generation. Coarse-grained retrieval involves fewer available categories, while fine-grained retrieval provides more specific and detailed categories.
[0040] Step 312: After determining the category, we can search for haptic signals of the same category in the existing haptic database. Essentially, the information required for classification is a subset of the information required for complete signal reconstruction. Similarly, for retrieval tasks with different granularities, the information required for the coarse-grained task is a subset of the information required for the fine-grained task. Denote the bandwidths required to transmit these information as M 1 , M 2 and M 3 , where M 1 > M 2 > M 3 . The bandwidth is closely related to the feature dimension, the frame rate of obtaining visual signals, and the channel coding strategy.
[0041] (3-2)、Meanwhile, during the above semantic encoding process, different network conditions need to be addressed. The present invention selects to use a bandwidth prediction network that combines a heuristic-based controller and a deep reinforcement learning controller.
[0042] Step 321: Initially, when the input data is limited, it uses a heuristic-based controller.
[0043] Step 322: Once sufficient input data is collected, the bandwidth prediction network utilizes deep reinforcement learning to adapt to different network conditions. The device dynamically adjusts the scalable encoding network according to the predicted bandwidth information.
[0044] A further improvement of the present invention is that step (4) includes the following steps:
[0045] (4-1) After the edge side receives the semantic representation transmitted by the device, it performs semantic decoding and haptic generation. According to the different granularities of the semantics, the edge side adaptively executes corresponding generation strategies, namely coarse-grained cross-modal retrieval, fine-grained cross-modal retrieval, and cross-modal generation.
[0046] Step 411: First, perform semantic decoding to obtain semantic information;
[0047] Step 422: According to the granularity of the received semantic information, execute the corresponding signal recovery strategy;
[0048] (1) Retrieval model from coarse-grained to fine-grained:
[0049] Under the condition of limited bandwidth, deep feature maps containing less semantic information are used for coarse-grained retrieval, while shallow feature maps containing more semantic information are used for fine-grained retrieval. After obtaining the category information, haptic signals of the same category are retrieved from the existing haptic database and presented to the operator. The retrieval model includes a fully connected layer, and the structure of the final layer is determined by the number of categories in the database. The output of the final layer is used to calculate the scores of different categories, and then the softmax function is used for normalization. The ultimate goal is to minimize the cross-entropy between the predicted distribution and the actual label. This goal can be defined as:
[0050]
[0051] Where, is the one-hot encoding representing the true label of the fine-grained and coarse-grained retrieval models, and are the predicted multi-class probabilities.
[0052] (2) Cross-modal generation:
[0053] In the case of sufficient bandwidth, the device transmits complete fusion features, enabling cross-modal generation on the edge side. The present invention uses generative adversarial networks (GANs) to generate tactile signals. Different from standard GANs, the present invention combines two discriminator networks D 1 and D 2 , to enhance the consistency between the generated tactile signal h′ and the real tactile signal h. Specifically, D 1 is designed to distinguish two pairs of signals: h paired with itself, and h paired with the generated signal h′. This process ensures the structural similarity of the signals, effectively aligning h with h′ from a global perspective. The discriminator D 2 functions similarly to a standard GAN. D 1 and D 2 's adversarial learning loss function can be expressed as:
[0054]
[0055]
[0056] where θ g , θ d1 , θ d2 are the parameters of G, D 1 and D 2 , p(*) represents the distribution of the signal, E h′~p(h′),h~p(h) [] means sampling h′ and h from the distributions p(h′) and p(h) respectively, and calculating the expected value of the expression inside the brackets; E h~p(h) [] means sampling h from the distribution p(h) and calculating the expected value of the expression inside the brackets; E h′~p(h′) [] means sampling h′ from the distribution p(h′) and calculating the expected value of the expression inside the brackets.
[0057] In addition, the present invention also calculates the L 2 loss between the generation result and the real tactile signal, which has been proven to improve the stability of GAN training. It can be calculated as follows:
[0058] L G =||h - h′|| 2 .
[0059] Overall, the total loss of cross-modal generation can be defined as:
[0060] L gen =L D1 +L D2 +L G .
[0061] With the above objective function, the generation and discriminant models can be iteratively trained to generate tactile signals closer to the real tactile signals. Finally, the overall objective function of the cross-modal tactile generation scheme of the present invention can be written as:
[0062] L total = L contra + L fine + L coarse + L gen .
[0063] A further improvement of the present invention lies in that step (5) includes the following steps:
[0064] (5-1) Utilize visual signals, audio signals, and tactile signals to train and optimize the encoding network, fusion, and generation network. After freezing the parameters of the visual encoding network, audio encoding network, and cross-modal encoding network, then train the fine-grained retrieval network, and then freeze the parameters of the fine-grained retrieval network, and then train the fine-grained retrieval network to obtain the trained network parameters. The specific process is as follows:
[0065] Step 5-1: Utilize visual signals, audio signals, and tactile signals to train and optimize the encoding network, fusion, and generation network. After freezing the parameters of the visual encoding network, audio encoding network, and cross-modal encoding network, then train the fine-grained retrieval network, and then freeze the parameters of the fine-grained retrieval network, and then train the fine-grained retrieval network to obtain the trained network parameters. The specific process is as follows:
[0066] Step 511: Initialize the parameters of the visual encoding network, audio encoding network, and cross-modal encoding network, set the number of iterations: epoch1, set the learning rate: μ 1 , μ 2 , μ 3 , and start training the encoding, fusion, and generation network;
[0067] Step 512: Use the Adam optimizer to update θ v corresponding to the loss function L contra + L gen , update θ a corresponding to the loss function L contra + L gen , update θ cmf corresponding to the loss function L contra + L gen , update θ g corresponding to the loss function L contra + L gen , update θ d1 corresponding to the loss function L D1 , update θ d2 corresponding to the loss function L D2, the specific format is as follows:
[0068]
[0069] where θ v 、θ a 、θ cmf 、θ g 、θ d1 、θ d2 are the visual encoder, audio encoder, cross-modal feature fusion module, generation network, discriminator D1, and discriminator D2 respectively, is to take the partial derivative of each loss function;
[0070] Step 513: If epoch < epoch1, then repeat Step 512. After epoch1 rounds of iteration, a network that converges to the optimal is obtained, and at the same time, the parameters of the visual encoding network, audio encoding network, and cross-modal encoding network are updated and frozen;
[0071] Step 514: Set the number of iterations: epoch2, and start training the fine-grained retrieval network;
[0072] Step 515: Use the Adam optimizer to update the loss functions L ssc and L c-fine corresponding to θ c-coarse + L c-fine , the specific formula is as follows:
[0073]
[0074] where θ ssc 、θ c-fine are both fine-grained retrieval network parameters, is to take the partial derivative of each loss function;
[0075] Step 516: If epoch < epoch2, then repeat Step 515. After epoch2 rounds of iteration, a network that converges to the optimal is obtained, and at the same time, the parameters of the fine-grained retrieval network are updated and frozen;
[0076] Step 517: Set the number of iterations: epoch3, and start training the coarse-grained retrieval network;
[0077] Step 518: Use the Adam optimizer to update the loss function L c-coarse corresponding to θ c-coarse + L c-fine , the specific formula is as follows:
[0078]
[0079] where θ c-coarseare the parameters of the coarse-grained retrieval network, is to take the partial derivative of each loss function;
[0080] Step 519: If epoch < epoch3, repeat Step 518. After epoch3 rounds of iteration, a network that converges to the optimal is obtained, and the optimal model parameters are obtained for the subsequent test phase;
[0081] A further improvement of the present invention is that: Step (6) includes the following steps:
[0082] (6-1) Input paired audio and visual signals into the optimal network model, which will extract the features of the audio and visual signals, perform fusion, and then execute the optimal tactile signal restoration strategy according to the current network state. The specific steps are as follows:
[0083] Step 611: Collect visual and audio samples;
[0084] Step 612: Perform visual encoding and tactile encoding on the samples;
[0085] Step 613: Perform feature extraction and cross-modal feature fusion to generate features;
[0086] Step 614: Based on network state awareness, determine the current network bandwidth M;
[0087] Step 615: The device side performs different processing according to the received bandwidth information. If M > M 1 , transmit the fused feature f V-A , at the edge computing side, use the generation network for cross-modal generation to generate tactile signals; if M 2 <M<M 1 , transmit the shallow feature map f V-A-s , at the edge computing side, use the fine-grained cross-modal retrieval network for fine-grained cross-modal retrieval to search for tactile signals based on the existing database; if M 3 <M<M 2 , transmit the shallow feature map f V-A-d , at the edge computing side, use the coarse-grained cross-modal retrieval network for coarse-grained cross-modal retrieval to search for tactile signals based on the existing database;
[0088] Step 616: The edge side finally transmits the generated tactile signal to the user.
[0089] The beneficial effects of the present invention are:
[0090] 1. The present invention utilizes easily collectable audio and visual signals to perform scalable cross-modal tactile signal generation, ensuring high-quality and highly reliable real-time generation of tactile signals, and overcoming the problem of the decline in the user's immersive operation experience caused by the lack of tactile signals.
[0091] 2. The present invention designs an innovative multi-modal fusion scheme, which realizes the deep interaction between visual and auditory signals, thereby fully exploiting complementary information and improving the reliability of the model. On the one hand, it uses contrastive learning to map the features of different modalities into the same feature space, thereby aligning paired samples and reducing the heterogeneity differences between modalities. On the other hand, it realizes the exchange of information between modalities through the designed cross-modal fusion network based on Transformer and integrates the information beneficial to the tactile generation task in a complementary way, and finally obtains the multi-modal fusion semantics.
[0092] 3. The present invention constructs a collaborative framework for the edge and the device side, realizes the real-time perception of the network state in the rescue environment, and feeds back to guide the device side to execute the optimal coding strategy. Specifically, through scalable semantic coding, the hierarchical compression of multi-modal fusion semantics is realized to ensure adaptation to the fluctuating transmission environment.
[0093] 4. The present invention designs a three-level tactile signal generation strategy. Under the change of network bandwidth from low to high, the reliable and real-time generation of tactile signals is ensured respectively through cross-modal coarse-grained retrieval, cross-modal fine-grained retrieval, and cross-modal direct generation strategies, and the generality of the model is improved.
[0094] 5. The present invention evaluates the method proposed in the present invention in a standard multi-modal dataset and a simulated emergency environment. The results show that the method can adaptively generate tactile feedback, improve the user's operation experience, and outperform the existing methods in terms of the quality of tactile signals. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 is the flowchart of the method of the present invention;
[0096] Figure 2 is the schematic diagram of the complete network structure of the present invention;
[0097] Figure 3 is the comparison chart between the accuracy of the indirect tactile generation using cross-modal retrieval proposed in the present invention and other comparison methods under the condition of limited bandwidth;
[0098] Figure 4 is the result chart of the comparison between the visual qualitative results of the direct tactile generation using the cross-modal generation model proposed in the present invention and the tactile signals generated by other comparison methods under the condition of sufficient bandwidth. DETAILED DESCRIPTION OF THE INVENTION
[0099] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. It should be noted that the terms "front", "rear", "left", "right", "upper" and "lower" used in the following description refer to the directions in the accompanying drawings, and the terms "inner" and "outer" respectively refer to the directions towards or away from the geometric center of a specific component.
[0100] The present invention provides a scalable cross-modal tactile signal generation method for emergency rescue scenarios, and its flowchart is as Figure 1 shown. The method includes the following steps:
[0101] Step 1: Use a residual neural network as a feature encoder to extract low-dimensional features from visual and audio data, such as Figure 2 . However, since the features of different modalities are kept in their respective spaces, it is difficult to perform effective cross-modal interaction. Therefore, a contrastive learning method is used to align the extracted visual and audio features, alleviate the modality differences, improve the model performance, and provide support for the next step of feature fusion.
[0102] Step 1-1: In order to generate more powerful features and alleviate the problem of model overfitting caused by the increase in network depth, the present invention uses ResNet as a unified single-modal feature encoder instead of VGG. Specifically, the visual image v i and the audio spectrogram a i are input into ResNet-18 to extract the low-dimensional feature vectors f V and f A ;
[0103] Step 1-2: In order to alleviate the modality differences, contrastive learning is used to align the features of different modalities, and a similarity function is used to evaluate the similarity between the features extracted from different modalities. Contrastive learning aims to increase the similarity score between paired visual and audio signals, and vice versa. To construct paired audiovisual signals, positive samples are defined as features from different modalities but belonging to the same category, while negative samples are features from different modalities and categories.
[0104] The similarity function is defined as follows:
[0105]
[0106] where and represent the linear projections that map the features to low-dimensional representations, V represents video features, A represents audio features; the value range of S is from 0 to 1, and the higher the value, the greater the similarity between the features.
[0107] For each visual and audio feature, calculate the similarity between vision-to-audio and audio-to-vision as follows:
[0108]
[0109]
[0110] where B is the batch size, λ represents the temperature parameter, V b and A b represent the video feature and audio feature under the b-th sample respectively;
[0111] Finally, the contrastive loss is defined as cross-entropy and can be expressed as:
[0112]
[0113] where and indicate that the true similarity and the probability of the positive sample pair are 1, while the probability of the negative sample pair is 0.
[0114] Step 2: Design an effective cross-modal fusion method that can exchange information between modalities and integrate key information beneficial to haptic generation in a complementary manner. This method first normalizes the multi-modal features, calculates the intermediate representations of different modalities, then uses the multi-head attention block to extract cross-modal interaction information, and finally merges the fused feature representations through element-wise summation to obtain the fused features.
[0115] Step 2-1: Use two Transformer blocks for feature fusion. Each block consists of a multi-head attention block (MHA), a normalization layer (NL), and a multi-layer perceptron (MLP). The key factor promoting cross-modal information exchange is to use MHA. In traditional Transformer, the MHA block captures the intrinsic context relationships in images or texts. In contrast, MHA is used here to extract the semantic associations between visual and audio features. The specific steps are as follows:
[0116] Step 211: First, normalize the input multi-modal features to obtain f V-norm and f A-norm .
[0117] Step 212: Then, use linear projection to calculate the intermediate representations of different modalities: query (Q), key (K), and value (V).
[0118] Step 213: After that, obtain the cross-modal interaction information f V-iter and f A-iter through MHA. The process is described as follows:
[0119]
[0120]
[0121] where Q a , K v , V v , Q v , K a , V a are intermediate representations of visual and audio features, and d represents the dimension.
[0122] Step 214: Combine the outputs of the two Transformer blocks through element-wise summation to finally obtain f V-A . The whole process can be described as:
[0123]
[0124] Step 3: Utilize the fused features extracted in the previous step to achieve scalable semantic encoding to obtain hierarchical semantic representations. At the same time, through the network situation awareness module, gain a comprehensive understanding of the current network environment, and based on the network bandwidth, select the optimal transmission strategy such as Figure 2 .
[0125] Step 3-1: Utilize the neural network to extract a concise representation of the fused features to achieve scalable semantic encoding, that is, represent different semantic granularities through feature maps of different depths.
[0126] Step 311: The fused feature with complete semantics is represented as f V-A , while the shallow and deep feature maps extracted from the fused feature through the neural network are respectively represented as f V-A-s and f V-A-d . The feature f V-A will be used for direct haptic generation, while f V-A-s and f V-A-d will be used for fine-grained and coarse-grained retrieval respectively, mainly for indirect haptic generation. Coarse-grained retrieval involves fewer available categories, while fine-grained retrieval provides more specific and detailed categories.
[0127] Step 312: After determining the category, we can search for haptic signals of the same category in the existing haptic database. Essentially, the information required for classification is a subset of the information required for complete signal reconstruction. Similarly, for retrieval tasks of different granularities, the information required for the coarse-grained task is a subset of the information required for the fine-grained task. Denote the bandwidths required to transmit this information as M 1 , M 2 and M 3 , where M 1 > M 2 > M 3。The bandwidth is closely related to the feature dimension, the frame rate of obtaining visual signals, and the channel coding strategy.
[0128] Step 3-2. Meanwhile, during the above semantic coding process, different network conditions need to be addressed. The present invention selects to use a bandwidth prediction network that combines a heuristic-based controller and a deep reinforcement learning controller.
[0129] Step 321. Initially, when the input data is limited, it uses a heuristic-based controller.
[0130] Step 322. Once sufficient input data is collected, the bandwidth prediction network utilizes deep reinforcement learning to adapt to different network conditions. The device dynamically adjusts the scalable coding network according to the predicted bandwidth information.
[0131] Step (4) includes the following steps:
[0132] (4-1). After the edge side receives the semantic representation transmitted by the device, it performs semantic decoding and haptic generation. According to the different granularities of the semantics, the edge side adaptively executes the corresponding generation strategies, namely coarse-grained cross-modal retrieval, fine-grained cross-modal retrieval, and cross-modal generation, as Figure 2 。
[0133] Step 411. First, perform semantic decoding to obtain semantic information;
[0134] Step 422. According to the granularity of the received semantic information, execute the corresponding signal recovery strategy;
[0135] (1) From coarse-grained to fine-grained retrieval model:
[0136] Under the condition of limited bandwidth, deep feature maps containing less semantic information are used for coarse-grained retrieval, while shallow feature maps containing more semantic information are used for fine-grained retrieval. After obtaining the category information, haptic signals of the same category are retrieved from the existing haptic database and presented to the operator. The retrieval model includes a fully connected layer, and the structure of the final layer is determined by the number of categories in the database. The output of the final layer is used to calculate the scores of different categories, and then the softmax function is used for normalization. The ultimate goal is to minimize the cross-entropy between the predicted distribution and the actual label. This goal can be defined as:
[0137]
[0138] Where, is the one-hot encoding representing the true label of the fine-grained and coarse-grained retrieval models, and are the predicted multi-category probabilities.
[0139] (2) Cross-modal generation:
[0140] When the bandwidth is sufficient, the device transmits complete fusion features, enabling cross-modal generation on the edge side. The present invention uses generative adversarial networks (GANs) to generate tactile signals. Different from the standard GAN, the present invention combines two discriminator networks D 1 and D 2 , to enhance the consistency between the generated tactile signal h′ and the real tactile signal h. Specifically, D 1 is designed to distinguish two pairs of signals: h paired with itself, and h paired with the generated signal h′. This process ensures the structural similarity of the signals, effectively aligning h with h′ from a global perspective. The discriminator D 2 functions similarly to the standard GAN. The adversarial learning loss functions of D 1 and D 2 can be expressed as:
[0141]
[0142] where θ g , θ d1 , θ d2 are the parameters of G, D 1 and D 2 , p(*) represents the distribution of the signal, E h′~p(h′),h~p(h) [] means sampling h′ and h from the distributions p(h′) and p(h) respectively, and calculating the expected value of the expression inside the brackets; E h~p(h) [] means sampling h from the distribution p(h) and calculating the expected value of the expression inside the brackets; E h′~p(h′) [] means sampling h′ from the distribution p(h′) and calculating the expected value of the expression inside the brackets.
[0143] In addition, the L 2 loss between the generated result and the real tactile signal is also calculated, which has been proven to improve the stability of GAN training. It can be calculated as follows:
[0144] L G =||h - h′|| 2 .
[0145] Overall, the total loss of cross-modal generation can be defined as:
[0146] L gen =L D1 +L D2 +L G .
[0147] With the above objective function, the generation and discrimination models can be iteratively trained to generate tactile signals that are closer to real tactile signals. Finally, the overall objective function of the cross-modal tactile generation scheme of the present invention can be written as:
[0148] L total = L contra + L fine + L coarse + L gen .
[0149] A further improvement of the present invention lies in that step (5) includes the following steps:
[0150] (5-1), Utilize visual signals, audio signals, and tactile signals to train and optimize the encoding network, the fusion and generation network. After freezing the parameters of the visual encoding network, the audio encoding network, and the cross-modal encoding network, then train the fine-grained retrieval network, and then freeze the parameters of the fine-grained retrieval network, and then train the fine-grained retrieval network to obtain the trained network parameters, such as Figure 1 , The specific process is as follows:
[0151] Step 511, Initialize the parameters of the visual encoding network, the audio encoding network, and the cross-modal encoding network, set the number of iterations: epoch1, set the learning rate: μ 1 , μ 2 , μ 3 , and start training the encoding, fusion, and generation networks;
[0152] Step 512, Use the Adam optimizer to update θ v corresponding to the loss function L contra + L gen , update θ a corresponding to the loss function L contra + L gen , update θ cmf corresponding to the loss function L contra + L gen , update θ g corresponding to the loss function L contra + L gen , update θ d1 corresponding to the loss function L D1 , update θ d2 corresponding to the loss function L D2 , and the specific format is as follows:
[0153]
[0154] where θ v , θ a , θ cmf , θ g , θd1 , θ d2 are respectively a visual encoder, an audio encoder, a cross-modal feature fusion module, a generation network, discriminator D1, and discriminator D2. is to take the partial derivative of each loss function;
[0155] Step 513: If epoch < epoch1, then repeat Step 512. After epoch1 rounds of iteration, a network that converges to the optimal is obtained, and at the same time, the parameters of the visual encoding network, the audio encoding network, and the cross-modal encoding network are updated and frozen;
[0156] Step 514: Set the number of iterations: epoch2, and start training the fine-grained retrieval network;
[0157] Step 515: Use the Adam optimizer to update θ ssc and θ c-fine corresponding loss function L c-coarse +L c-fine , and the specific formula is as follows:
[0158]
[0159] where θ ssc , θ c-fine are both fine-grained retrieval network parameters, is to take the partial derivative of each loss function;
[0160] Step 516: If epoch < epoch2, then repeat Step 515. After epoch2 rounds of iteration, a network that converges to the optimal is obtained, and at the same time, the fine-grained retrieval network parameters are updated and frozen;
[0161] Step 517: Set the number of iterations: epoch3, and start training the coarse-grained retrieval network;
[0162] Step 518: Use the Adam optimizer to update θ c-coarse corresponding loss function L c-coarse +L c-fine , and the specific formula is as follows:
[0163]
[0164] where θ c-coarse is the parameter of the coarse-grained retrieval network, is to take the partial derivative of each loss function;
[0165] Step 519: If epoch < epoch3, then repeat Step 518. After epoch3 rounds of iteration, a network that converges to the optimal is obtained, and the optimal model parameters are obtained for use in the subsequent test phase;
[0166] A further improvement of the present invention lies in that step (6) includes the following steps:
[0167] (6-1) Input paired audio and visual signals into the optimal network model, which will extract the features of the audio and visual signals, perform fusion, and then execute the optimal tactile signal recovery strategy according to the current network state, such as Figure 1 , and the specific steps are as follows:
[0168] Step 611: Collect visual and audio samples;
[0169] Step 612: Perform visual coding and tactile coding on the samples;
[0170] Step 613: Perform feature extraction and cross-modal feature fusion to generate features;
[0171] Step 614: Based on network state perception, determine the current network bandwidth M;
[0172] Step 615: The device side performs different processing according to the received bandwidth information. If M > M 1 , transmit the fusion feature f V-A , at the edge computing side, use the generation network for cross-modal generation to generate tactile signals; if M 2 < M < M 1 , transmit the shallow feature map f V-A-s , at the edge computing side, use the fine-grained cross-modal retrieval network for fine-grained cross-modal retrieval to search for tactile signals based on the existing database; if M 3 < M < M 2 , transmit the shallow feature map f V-A-d , at the edge computing side, use the coarse-grained cross-modal retrieval network for coarse-grained cross-modal retrieval to search for tactile signals based on the existing database;
[0173] Step 616: The edge side finally transmits the generated tactile signals to the user.;
[0174] The following experimental results show that, compared with the existing methods, the present invention utilizes the correlation between multi-modal signals and scalable semantic encoding and decoding to achieve reliable generation of tactile signals under emergency rescue conditions and obtains better generation effects.
[0175] The present invention uses the LMT-108 multi-modal surface material dataset to evaluate the performance of the method of the present invention. This dataset includes 108 surface materials, which can be divided into nine categories according to their physical properties. Each major category contains data of 5 to 17 sub-categories. Each sample includes data of three modalities, which are paired one-to-one and uniformly labeled. And uniformly marked. In short, this dataset provides rich modal information to support cross-modal data processing.
[0176] It provides sufficient category information for both coarse-grained and fine-grained retrieval as well as cross-modal generation.
[0177] Existing Methods One, Two, and Three: The literature "Cross-modal surface material retrieval using discriminant adversarial learning" (authors W. Zheng, H. Liu, B. Wang, and F. Sun) proposed three corresponding retrieval schemes for cross-modal retrieval. That is, Method One only uses audio signals for retrieval, Method Two only uses visual signals for retrieval, and Method Three uses simple concatenation for early fusion to perform cross-modal retrieval.
[0178] Existing Method Four: The literature "Deeply supervised subspace learning for cross-modal material perception of known and unknown objects" (authors P. Xiong, J. Liao, M. Zhou, A. Song, and P. X. Liu) proposed a deeply supervised subspace learning method for cross-modal retrieval.
[0179] Existing Method Five: The literature ""touching to see"and"seeing to feel":Robotic cross-modal sensory data generation for visual-tactile perception,"" (authors J.-T. Lee, D. Bollegala, and S. Luo) proposed an innovative cross-modal generation framework that uses conditional generative networks to achieve the generation from visual signals to tactile signals.
[0180] Existing Method Six: The literature "Visual-tactile cross-modal data generation using residue-fusion gan with feature-matching and perceptual losses" (authors S. Cai, K. Zhu, Y. Ban, and T. Narumi) uses the structure of conditional generative networks and a residue fusion module to perform mutual generation from vision to touch and from touch to vision. By jointly considering feature-matching losses and perceptual losses, it significantly improves the performance of cross-modal generation.
[0181] Existing Method 7: The literature "Cross-modal generation of tactile friction coefficient from audio and visual measurements by transformer" (authors R. Song, X. Sun, and G. Liu) designed an audio-visual assisted cross-modal generation architecture, which achieved a deeper understanding of audio-visual touch by integrating traditional convolution and self-attention.
[0182] In this embodiment, the performance metrics for evaluating the super-resolution reconstruction scheme proposed by the present invention are divided into four categories: classification accuracy, root mean square error, similarity, and structural similarity.
[0183] Classification accuracy: Classification accuracy represents the accuracy when performing cross-modal retrieval.
[0184] Root mean square error: The root mean square error (RMSE) is used to measure the deviation between the real signal and the generated signal.
[0185] Similarity: Similarity (SIM) is used to quantitatively evaluate the similarity between the real tactile signal and its reconstructed signal. It uses the Chebyshev distance between two signals instead of the Euclidean distance to measure the maximum distance between corresponding points
[0186] Structure-similarity: Structural similarity (ST-SIM) is an objective quality assessment method for vibrotactile signals, which is highly consistent with human perception. It evaluates similarity in both the time domain and the frequency domain.
[0187] Table 1 Performance comparison results of cross-modal tactile signal generation schemes
[0188] Root Mean Square Error Similarity Structure - Similarity Existing Method Five 0.1322 0.6103 0.7598 Existing Method Six 0.1861 0.5622 0.7873 Existing Method Seven 0.1297 0.6272 0.8115 The present invention 0.1272 0.6488 0.8031
[0189] From Figure 3 , Figure 4 and the results in Table 1, it can be seen that compared with the above competitive schemes, the present invention has significant performance advantages. Specifically, among the seven comparison schemes, the method of cross-modal tactile signal generation based on cross-modal semantic fusion and scalable encoding and decoding shows better performance in retrieval accuracy, qualitative visualization results, and quantitative index evaluation. The results indicate that the contrastive learning constraint proposed by the present invention can promote the encoder network to capture modality-invariant class attribute information, the cross-modal fusion network is beneficial to the acquisition of deep semantic interaction and complementary features, and the designed generation and discrimination networks can enhance the global similarity of the signal.
[0190] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features.
Claims
1. A scalable cross-modal tactile signal generation method for emergency rescue scenarios, characterized by: The cross-modal tactile signal generation method comprises the following steps: Step 1: Use a residual neural network as a feature encoder to extract low-dimensional features from visual and audio data and use a contrastive learning method to align the extracted visual and audio features; Step 2: Design a cross-modal fusion method, which first normalizes the multi-modal features, calculates the intermediate representations of different modalities, then uses a multi-head attention block to extract cross-modal interaction information, and finally merges the features by element-by-element summation to obtain the fused features. Step 3: Use the fusion features extracted in the previous step to implement scalable semantic coding to obtain hierarchical semantic representation. At the same time, the network situation awareness module is used to fully understand the current network environment and select the optimal transmission strategy based on the network bandwidth. Step 4: After receiving the semantic representation transmitted by the device, the edge performs semantic decoding and cross-modal tactile signal generation, and according to the different semantic granularities, the edge adaptively adopts the corresponding generation strategy, namely, coarse-grained cross-modal retrieval, fine-grained cross-modal retrieval and cross-modal generation; Step 5: Use the optimization algorithm to train the entire model and finally obtain the optimal model parameters for the subsequent testing phase; Step 6: Input the paired audio and visual signals into the optimal network model, which extracts the features of the audio and visual signals, performs fusion, and then executes the optimal tactile signal recovery strategy according to the current network state.
2. The method for generating cross-modal tactile signals for emergency rescue scenarios according to claim 1, characterized in that: Step 1 specifically includes the following steps: Step 1-1: Transform the visual image v i And the spectrum of the audio signal a i Input into ResNet-18 to extract the low-dimensional feature vector f V and f A ; Step 1-2: To alleviate the modality difference, contrastive learning is used to align the features of different modalities, and a similarity function is used to evaluate the similarity between the features extracted from different modalities. The similarity function is defined as follows: in and represents the linear projection that maps features to low-dimensional representation; V represents video features, and A represents audio features; The value of S ranges from 0 to 1, and higher values indicate greater similarity between features; For each visual and audio feature, the similarity between visual to audio and audio to visual is calculated as follows: Where B is the batch size, λ is the temperature parameter, V b , A b Respectively represent the video features and audio features of the b-th sample; Finally, the contrastive loss is defined as the cross entropy, expressed as: in and It indicates that the probability of the true similarity and positive sample pair is 1, while the probability of the negative sample pair is 0.
3. The cross-modal tactile signal generation method for emergency rescue scenarios according to claim 1, characterized in that: Step 2 includes the following steps: Step 2-1: Use two Transformer blocks for feature fusion. Each block consists of a multi-head attention block MHA, a normalization layer NL, and a multi-layer perceptron MLP. Use MHA to extract the semantic association between visual and audio features. The specific steps are as follows: Step 211: First, normalize the input multimodal features to obtain f V-norm and f A-norm ; Step 212: Next, use linear projection to calculate intermediate representations of different modalities: query Q, key K and value V; Step 213: After that, the cross-modal interaction information f is obtained through MHA V-iter and f A-iter ; The process is described as follows: Where Q a , K v , V v , Q v , K a , V a It is the intermediate representation of visual and audio features, and d represents the dimension; Step 214: Combine the outputs of the two Transformer blocks by element-wise summation to obtain f V-A ; The whole process is described as:
4. The method for generating cross-modal tactile signals for emergency rescue scenarios according to claim 1, characterized in that: Step 3 specifically includes the following steps: Step 3-1: Use the concise representation of the fused features extracted in the previous step to achieve scalable semantic encoding, that is, use feature maps of different depths to represent different semantic granularities; Step 311: The fused feature with complete semantics is represented as f V-A , and the shallow and deep feature maps extracted from the fusion features through the neural network are represented as f V-A-s and f V-A-d ; Feature f V-A will be used for direct tactile generation, while f V-A-s and f V-A-d They will be used for fine-grained and coarse-grained retrieval, respectively, mainly for indirect contact generation; Step 312, after determining the category, search for tactile signals of the same category in the existing tactile database; the information required for classification is a subset of the information required for complete signal reconstruction; similarly, for retrieval tasks of different granularities, the information required for the coarse-grained task is a subset of the information required for the fine-grained task; the bandwidth required to transmit this information is represented as M1, M2 and M3, where M1>M2>M3; in addition, the bandwidth is closely related to the feature dimension, the frame rate of the visual signal, and the channel coding strategy; Step 3-2: In the above semantic encoding process, it is necessary to deal with different network conditions and choose to use a bandwidth prediction network that combines a heuristic-based controller and a deep reinforcement learning controller; Step 321, initially, when input data is limited, it uses a heuristic-based controller; Step 322. Once sufficient input data is collected, the bandwidth prediction network uses deep reinforcement learning to adapt to different network conditions; the device dynamically adjusts the scalable coding network based on the predicted bandwidth information.
5. The method for generating cross-modal tactile signals for emergency rescue scenarios according to claim 1, characterized in that: Step 4 specifically includes the following steps: Step 4-1: After receiving the semantic representation transmitted by the device, the edge performs semantic decoding and tactile generation; according to the different granularities of the semantics, the edge adaptively executes the corresponding generation strategy, namely, coarse-grained cross-modal retrieval, fine-grained cross-modal retrieval and cross-modal generation; Step 411: firstly perform semantic decoding to obtain semantic information; Step 422: Execute a corresponding signal recovery strategy according to the granularity of the received semantic information.
6. The method for generating cross-modal tactile signals for emergency rescue scenarios according to claim 5, characterized in that: The retrieval model from coarse-grained to fine-grained is as follows: under limited bandwidth conditions, deep feature maps containing less semantic information are used for coarse-grained retrieval, while shallow feature maps containing more semantic information are used for fine-grained retrieval; after obtaining the category information, tactile signals of the same category are retrieved from the existing tactile database and presented to the operator; the retrieval model includes a fully connected layer, and the structure of the final layer is determined by the number of categories in the database; the output of the final layer is used to calculate the scores of different categories, and then normalized using the softmax function; The ultimate goal is to minimize the cross entropy between the predicted distribution and the actual labels. This objective is defined as: in, is the one-hot encoding of the true labels for the fine-grained and coarse-grained retrieval models, and is the predicted multi-class probability.
7. The method for generating cross-modal tactile signals for emergency rescue scenarios according to claim 5, characterized in that: Cross-modal generation: When bandwidth is sufficient, the device transmits complete fusion features, enabling the edge side to perform cross-modal generation; Generative adversarial networks (GANs) are used to generate tactile signals; two discriminator networks D1 and D2 are combined to enhance the consistency between the generated tactile signal h′ and the real tactile signal h; specifically, D1 aims to distinguish between two pairs of signals: h paired with itself, and h paired with the generated signal h′; this process ensures the structural similarity of the signals and effectively aligns h with h′ from a global perspective; the role of the discriminator D2 is similar to that of the standard GAN; the adversarial learning loss function of D1 and D2 is expressed as: Among them, θ g ,θ d1 ,θ d2 are the parameters of G, D1 and D2, p(*) represents the distribution of the signal, E h′~p(h′),h~p(h) [] means sampling h′ and h from distribution p(h′) and p(h) respectively, and calculating the expected value of the expression in the brackets; E h~p(h) [] means sampling h from the distribution p(h) and calculating the expected value of the expression in the brackets; E h′~p(h′) [] means sampling h′ from the distribution p(h′) and calculating the expected value of the expression in the brackets; In addition, the L2 loss between the generated result and the real tactile signal is calculated as follows: L G =||hh′||2; Overall, the total loss for cross-modal generation is defined as: L gen =L D1 +L D2 +L G ; With the above objective function, the generation and discriminant models are iteratively trained to generate tactile signals closer to the real tactile signals; finally, the overall objective function of the cross-modal tactile generation scheme is written as: L total =L contra +L fine +L coarse +L gen .
8. The cross-modal tactile signal generation method for emergency rescue scenarios according to claim 1, characterized in that: Step 5 includes the following steps: Step 5-1: Using visual signals, audio signals, and tactile signals, train and optimize the encoding network, fusion, and generation network. After freezing the parameters of the visual encoding network, audio encoding network, and cross-modal encoding network, then train the fine-grained retrieval network, and then freeze the parameters of the fine-grained retrieval network, and then train the fine-grained retrieval network to obtain the trained network parameters. The specific process is as follows: Step 511: Initialize the parameters of the visual encoding network, audio encoding network, and cross-modal encoding network. Set the number of iterations: epoch1, and set the learning rates: μ1, μ2, μ3. Start training the encoding, fusion, and generation network; Step 512: Use the Adam optimizer to update θ v The corresponding loss function L contra +L gen , update θ a The corresponding loss function L contra +L gen , update θ cmf The corresponding loss function L contra +L gen , update θ g The corresponding loss function L contra +L gen , update θ d1 The corresponding loss function L D1 , update θ d2 The corresponding loss function L D2 , the specific format is as follows: where θ v ,θ a ,θ cmf ,θ g ,θ d1 ,θ d2 They are visual encoder, audio encoder, cross-modal feature fusion module, generative network, discriminator d1, and discriminator d2 respectively. ▽ is the partial derivative of each loss function. Step 513: If epoch < epoch1, then repeat Step 512. After epoch1 rounds of iteration, obtain the network that converges to the optimal, and at the same time update and freeze the parameters of the visual encoding network, audio encoding network, and cross-modal encoding network; Step 514: Set the number of iterations: epoch2, and start training the fine-grained retrieval network; Step 515: Use the Adam optimizer to update θ ssc and θ c-fine The corresponding loss function L c-coarse +L c-fine , the specific formula is as follows: i ssc =θ ssc -μ1▽ θssc (L c-coarse +L c-fine ), where θ ssc ,θ c-fine They are all fine-grained retrieval network parameters, and ▽ is the partial derivative of each loss function; Step 516: If epoch < epoch2, then repeat Step 515. After epoch2 rounds of iteration, obtain the network that converges to the optimal, and at the same time update and freeze the parameters of the fine-grained retrieval network; Step 517: Set the number of iterations: epoch3, and start training the coarse-grained retrieval network; Step 518: Use the Adam optimizer to update θ c-coarse The corresponding loss function L c-coarse +L c-fine , the specific formula is as follows: i c-coarse =θ c-coarse -μ1▽ θc-coarse (L c-coarse ), where θ c-coarse is the parameter of the coarse-grained retrieval network, ▽ is the partial derivative of each loss function; Step 519: If epoch < epoch3, then repeat Step 518. After epoch3 rounds of iteration, obtain the network that converges to the optimal, and obtain the optimal model parameters for the subsequent test phase.
9. The method for generating cross-modal tactile signals for emergency rescue scenarios according to claim 1, characterized in that: Step 6 includes the following steps: Step 6-1: Input the paired audio and visual signals into the optimal network model, which will extract the features of the audio and visual signals, perform fusion, and then execute the optimal tactile signal recovery strategy according to the current network state. The specific steps are as follows: Step 611: Collect visual and audio samples; Step 612: Perform visual encoding and tactile encoding on the samples; Step 613: Perform feature extraction and cross-modal feature fusion to generate features; Step 614: Based on network state perception, determine the current network bandwidth M; Step 615: The device performs different processing based on the received bandwidth information. If M>M1, the transmission fusion feature f V-A , at the edge computing end, use the generative network for cross-modal generation to generate tactile signals; if M2<M<M1, transmit the shallow feature map f V-A-s , at the edge computing end, a fine-grained cross-modal retrieval network is used to perform fine-grained cross-modal retrieval and search for tactile signals based on the existing database; if M3<M<M2, the shallow feature map f is transmitted V-A-d ,On the edge computing end, a coarse-grained cross-modal retrieval network is used to perform coarse-grained cross-modal retrieval and search tactile signals based on the existing database; Step 616: The edge side finally transmits the generated tactile signal to the user.
Citation Information
Patent Citations
Cross-modal image generation method and device based on audio-tactile signal fusion
CN113627482A
Audio and video auxiliary tactile signal reconstruction method based on cloud edge collaboration
CN113642604A
Audio-visual assisted fine-grained tactile signal reconstruction method
CN115905838A
Multi-modal feature learning efficiency optimization method based on low-rank factorization
CN118568658A
Text-conditioned image search based on transformation, aggregation, and composition of visio-linguistic features
US20220245391A1