Dance costume prop retrieval method based on causal reasoning and cross-modal matching

Through causal reasoning and cross-modal matching models, the problems of imperfect classification and mismatch of dance costume props are solved, and more accurate matching of costume props and cultural significance are achieved, and the accuracy and cultural authenticity of the search system are improved.

CN120541260APending Publication Date: 2025-08-26CHANGSHU INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510678805.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The classification system of existing dance costume props is incomplete, the cross-modal matching is difficult, and the misappropriation detection ability is insufficient, resulting in distortion and misuse of traditional culture.

Method used

The causal reasoning and cross-modal matching model are used to extract multi-scale features of images and text through Faster R-CNN and BERT models, combined with self-attention and gating fusion, the exact matching of images and text features is achieved, and the model performance is optimized using the bidirectional triple loss function.

Benefits of technology

It improves the matching accuracy of dance costume props, avoids mismatch and cultural misuse, and achieves more efficient cross-modal fusion and detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541260A_ABST
    Figure CN120541260A_ABST
Patent Text Reader

Abstract

The invention discloses a dance costume prop retrieval method based on causal reasoning and cross-modal matching, which comprises the following steps: in a matching task, calculating the similarity between costume prop images and text matching features by a model, and judging whether the costume prop images and the text matching features are matched or not according to a threshold value; and when the task is retrieved, calculating the similarity between the query text and the image feature set, and sorting to obtain a result. The process of calculating the features and the similarity comprises the following steps: firstly, extracting multi-scale visual features of an image by using Faster R-CNN, and extracting multi-granularity semantic features of a text by using BERT; then, the two types of features are subjected to self-attention and gating fusion respectively and then are input into a Transform layer, and final visual and text representation is obtained; then, multi-head attention calculation is carried out on the two representations, self-attention calculation is carried out after sequences are combined, and matching features are obtained; and finally, dividing the dot product of the two feature vectors by the modular length product to calculate the similarity. According to the invention, the matching accuracy of dance costume props can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a dance costume and props retrieval method, and in particular to a dance costume and props retrieval method based on causal reasoning and cross-modal matching. Background Art

[0002] Because ethnic costumes vary significantly across regions and ethnic groups, and because dance styles and techniques vary widely, even within ethnic groups, mixing and matching costumes is a common problem in dance performances, teaching, and creation. Dancers, choreographers, and costume designers often struggle to accurately distinguish and select costumes and props that reflect the characteristics of specific ethnic groups, leading to distorted representations of traditional culture and potentially sparking misunderstandings and controversy.

[0003] The existing clothing and props matching methods mainly have the following problems:

[0004] Imperfect classification system: The existing classification standards for ethnic minority dance costumes and props are relatively rough, often dividing them into major ethnic categories (such as Yi, Tibetan, Miao, etc.), while ignoring the significant differences between branches within the same ethnic group, resulting in inaccurate selection of costumes and props.

[0005] Difficulty in cross-modal matching: The selection of dance costumes usually relies on the combination of image and text descriptions, but most current retrieval systems are based on single-modality (image only or text only) matching and lack efficient cross-modal fusion methods, resulting in insufficient accuracy in retrieval and recommendation.

[0006] Lack of cultural context: Existing dance costume and prop management systems often only focus on external features while ignoring the cultural significance behind them, such as the symbolic meaning of the costumes and their suitability for specific dance styles. This makes it difficult for dance practitioners to balance artistry and cultural authenticity when choosing costumes.

[0007] Inadequate mismatch detection: Traditional search or recommendation systems cannot effectively identify the compatibility of clothing and props, making it easy to mismatch items across ethnic groups and ethnic lines. For example, certain accessories may only be compatible with traditional clothing from a specific ethnic lineage, but not from other ethnic lines or ethnic groups. However, existing systems cannot automatically identify and flag potential mismatches. Summary of the Invention

[0008] In response to the above-mentioned shortcomings of the existing technology, the present invention provides a dance costume and props retrieval method based on causal reasoning and cross-modal matching, which improves the matching accuracy of dance costumes and props and avoids mismatching and cultural misuse.

[0009] The technical solution of the present invention is as follows: a dance costume and prop retrieval method based on causal reasoning and cross-modal matching, comprising the following steps:

[0010] In the matching task, the causal reasoning and cross-modal matching model calculates the similarity between the image matching features and the text matching features of the clothing props, and determines whether the image and text match based on the similarity threshold. In the retrieval task, the causal reasoning and cross-modal matching model calculates the similarity between the query text features and the image feature set, and sorts them from high to low by similarity, selecting the top several images as the matching results.

[0011] The process of calculating features and similarities of the causal reasoning and cross-modal matching model includes:

[0012] S1. The Faster R-CNN model extracts multi-scale visual features of images, and the BERT model extracts multi-granular semantic features of text descriptions.

[0013] S2: Multi-scale visual features are fused through self-attention and gating and then input into the Transformer layer to obtain the final visual representation. Multi-granularity semantic features are fused through self-attention and gating and then input into the Transformer layer to obtain the final text representation.

[0014] S3, perform multi-head attention calculation on the final visual representation and the final text representation respectively, merge them into one sequence, and then perform self-attention calculation again to obtain image matching features and text matching features;

[0015] S4. The similarity is calculated as the dot product of the two feature vectors divided by the product of the moduli of the two feature vectors.

[0016] Furthermore, when calculating the matching relationship between clothing props A1 and A2, the causal reasoning and cross-modal matching model calculates the image matching features of each clothing prop based on steps S1 to S3. and and text matching features and Depend on Get the similarity to determine whether it matches, where α and β are hyperparameters used to adjust the weight of visual and text similarity in the total similarity. and represents the similarity calculated according to the method of step S4.

[0017] Furthermore, α+β=1.

[0018] Furthermore, when extracting multi-scale visual features of an image using the Faster R-CNN model, the Faster R-CNN model is used to extract K region-level original visual feature vectors for each input image, and the extracted i-th original visual feature vector is recorded as f i , Represents the original visual feature vector f iIn the space where the dimension of each original visual feature vector is Then a layer of linear transformation is used to obtain the multi-scale visual feature i i =W v ·f i +b v ,i i ∈R d ,i=1,2,…,K,W v represents a weight matrix, b v Represents a bias term.

[0019] Furthermore, the step S2 specifically includes:

[0020] Computing self-attention weights for multi-scale visual features W a Represents the weight matrix used to calculate the self-attention weights and generate a global visual representation Will With the visual representation vector i of each local region i Fusion to obtain new regional features

[0021]

[0022] in is the weight vector generated by a layer of gating network, W g represents the gate weight, σ is the sigmoid function, ⊙ represents element-by-element multiplication, [·;·] represents the vector concatenation operation, and all {i i ′ After inputting the standard Transformer layer, the final visual representation of each local area is obtained

[0023] Furthermore, the step S2 specifically includes:

[0024] Calculating sub-attention weights of multi-granularity semantic features W b Represents the weight matrix used to calculate the text self-attention weight and generate the global text representation Will With the original representation vector t of each word j Combined via a gated fusion layer:

[0025]

[0026] in W g represents the gate weight, σ is the sigmoid function, ⊙ represents element-by-element multiplication, [·;·] represents the vector concatenation operation, and the updated word vector set {t′ jAfter inputting the standard Transformer layer, the final text representation is output

[0027] Furthermore, the causal reasoning and cross-modal matching model is trained based on a bidirectional triplet loss function, and the total loss function is L total =L1+L2+L3+L4;

[0028] L1=max[0,d-S1(I,T)+S3(I,T)],

[0029] L2=max[0,d-S1(I,T)+S4(I,T)],

[0030] L3=max[0,d-S1(I,T)+S2(I,T)],

[0031] L4=max[0,d-S2(I,T)+S3(I,T)],

[0032] Where d represents the marginal value, S1(I,T)=i O ·t O ,S2(I,T)=i C ·t C , S3(I,T)=i O ·t C ,S4(I,T)=i C ·t O ,i O and t O represents the final visual representation and final text representation obtained in step S2, i C and t C Represents the final visual representation and final textual representation after the counterfactual intervention.

[0033] Furthermore, the final visual representation and final textual representation after the counterfactual intervention are artificially set using the self-attention weights and The visual representation and text representation are calculated and then obtained through gated fusion and Transformer layers.

[0034] The advantages of the technical solution provided by the present invention are:

[0035] This invention combines image features with text descriptions to achieve precise multimodal matching of ethnic minority dance costumes and props. Traditional methods often rely on manual experience to select costumes and props, which can easily lead to mismatching. However, this invention uses causal reasoning and cross-modal matching to more accurately classify and match dance costumes, avoiding mix-ups and mismatches.

[0036] Existing image-text matching techniques mostly rely on a single modality, resulting in insufficient matching accuracy. This paper uses the Faster R-CNN and BERT models to extract multi-scale visual and text features, and establishes causal relationships between image and text features through a causal graph model. This improves image-text matching and enables more efficient cross-modal fusion.

[0037] Existing technologies are weak at detecting compatibility between clothing and props, making cultural misuse more likely. This invention uses counterfactual reasoning and multimodal similarity calculation to accurately detect and correct potential mismatches. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 Schematic diagram of the flow of a dance costume and props retrieval method based on causal reasoning and cross-modal matching according to an embodiment of the present invention. DETAILED DESCRIPTION

[0039] The present invention will be further described below with reference to the following examples. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading this description, various equivalent modifications to this description by those skilled in the art fall within the scope defined by the claims appended to this application.

[0040] Please combine Figure 1 As shown, the dance costume and prop retrieval method based on causal reasoning and cross-modal matching in this embodiment includes the following steps:

[0041] The Faster R-CNN and BERT models are used to extract multi-scale features of images and texts, fuse global and local visual information, and combine with the self-attention mechanism to form detailed visual feature representations and text feature representations.

[0042] Specifically, the form of visual feature representation includes the following steps:

[0043] (1) Regional feature extraction

[0044] The Faster R-CNN model is used to extract K region-level original visual feature vectors for each input image, and the extracted i-th original visual feature vector is recorded as f i , Represents the original visual feature vector f i In the space where the dimension of each original visual feature vector is To facilitate subsequent processing, it is mapped to a unified d-dimensional space through a layer of linear transformation:

[0045] i i =W v ·f i+b v ,i i ∈R d ,i=1,2,…,K,

[0046] Thus, a set of local region visual representations is obtained where i i W represents the visual representation vector of the i-th region, which is mapped to the new space after linear transformation of the original visual feature vector. v Represents a weight matrix responsible for f i The dimension d of the target space is mapped to b. v Represents a bias term used to adjust the result after linear transformation. d Represents the space after mapping, with dimension d.

[0047] (2) Self-attention and global visual representation

[0048] First of all, Perform average pooling to obtain a global query vector q v , then q v As a query, Each visual representation vector i in i Calculate self-attention weights to obtain global visual representation Self-attention weight W a Represents the weight matrix used to calculate the self-attention weight, which is used to transform the visual representation vector i of each region i Projected to the query vector q v Same-dimensional space, generating global visual representation

[0049] (3) Fusion and Transformer Encoding

[0050] A gated fusion strategy is used to With the visual representation vector i of each local region i Fusion to obtain new regional features

[0051]

[0052] in is the weight vector generated by a layer of gating network, W g represents the gate weight, σ is the sigmoid function, ⊙ represents element-by-element multiplication, and [·;·] represents the vector concatenation operation. i ′ After inputting the standard Transformer layer, the final visual representation of each local area is obtained This process is able to capture more discriminative global context through gating and attention mechanisms while preserving local information.

[0053] The formation of text feature representation includes the following steps:

[0054] (1) Basic text encoding

[0055] On the text side, the attribute descriptions of dance costumes or props (such as "silk embroidered umbrella", "red tassel skirt", etc.) are segmented to obtain L words. The word vectors are then input into the pre-trained BERT model, and the average of the forward and reverse hidden states is taken to obtain the original representation vector t for each word. j ∈R d , forming a text representation set

[0056] (2) Text Self-Attention and Transformer Encoding

[0057] Same as the visual side, Perform self-attention aggregation and first use average pooling to obtain the global query vector q t , aggregated to the global text vector through the attention mechanism Self-attention weight W b Represents the weight matrix used to calculate the text self-attention weight, which is used to represent each word vector t j Projected to the query vector q t The space of the same dimension is used to calculate the similarity between each word and the global query vector and generate a global text representation.

[0058] (3) Fusion and Transformer Encoding

[0059] Will With the original representation vector t of each word j Combined through the gated fusion layer, the fusion formula

[0060]

[0061] in W g represents the gating weight, σ is the sigmoid function, ⊙ represents element-wise multiplication, and [·;·] represents the vector concatenation operation.

[0062] The updated word vector set {t j ′ Input the Transformer encoding with the same structure as the visual side and output the final text representation

[0063] Dance costumes and props often have multiple fine-grained attributes, such as material, color, texture, style, and accessories. These attributes may be reflected both visually (e.g., appearance characteristics) and in textual descriptions (e.g., material names, specific cultural elements). To better capture and align the semantics between visual and textual representations, this paper introduces a cross-modal fine-grained attribute modeling module based on the final visual and textual representations, enabling the model to learn more accurate attribute representations from complementary information sources.

[0064] By simultaneously considering the two directions of "visual attention to text" and "text to visual attention", we can comprehensively capture the correspondence between clothing and prop attributes in visual and textual modalities.

[0065] Text focuses on vision: Chinese text features as query (Q = W Q T), The visual features are used as key values ​​(K = W K I,V=W V I), where T and I represent text features and visual features integrated into matrix or batch form respectively. Using multi-head attention calculation,

[0066]

[0067] The text side obtains a weighted fusion representation of the visual attributes. This allows the text to understand which visual areas best support the attributes of the dance costume or props currently being described.

[0068] Visual attention text: Accordingly, the visual features are used as the query (Q = W q ′ I), text features as key values ​​(K = W K ′ T,V=W V ′ T), using multi-head attention calculation,

[0069]

[0070] A weighted fusion representation of the text description on the visual side is obtained, thereby capturing the correspondence between the visual area and each text word (especially keywords related to dance costumes and props details).

[0071] Based on bidirectional cross-modal attention, in order to fully utilize the interactive information of fine-grained attributes between text and vision to obtain more discriminative representations, self-attention re-encoding is adopted. After obtaining the output from bidirectional cross-modal attention, in order to more deeply model the correspondence between different attributes between vision and text, the fused multimodal features are input into a self-attention module again, allowing it to further learn contextual dependencies and importance weights within the same sequence.

[0072] Specifically, the text and visual features output by the bidirectional attention are first merged into a sequence in Represents the bidirectional attention output Attention(T→I) of the j-th position of the text, Represents the bidirectional attention output Attention(I→T) of the k-th visual position. Then the matching features are calculated by self-attention:

[0073]

[0074] Where SA(H) represents the output sequence obtained by the self-attention operation, which includes the text matching feature F T and image matching feature F I .

[0075] This process allows different attributes to "communicate" with each other within the same fused sequence. For example, the "silk texture" attribute recognized on the visual side can be corroborated by attention with the "gorgeous silk dress" description on the text side. Ultimately, the model more accurately captures the connections between attributes and reduces redundant or noisy information, allowing subsequent steps to make decisions based on more refined multimodal features.

[0076] Through this re-encoding step, the cross-modal alignment information captured by "bidirectional cross-modal attention" is further integrated, and the attribute features of text and vision are continuously interacted in the same multimodal space, thereby obtaining a more accurate and consistent representation.

[0077] For a clothing prop, calculate the image matching feature F I Match feature F with text T The similarity Sim(F I ,F T ), the model outputs the matching degree to achieve result retrieval. In the matching task, the threshold τ is set, when Sim(F I ,F T )>τ, the image and text are determined to match. In the retrieval task, the model calculates the query text feature F T and image feature set The similarity is calculated and sorted from high to low, and the top k images (e.g., k = 1 or k = 5) are selected as the final result. The result set is represented as

[0078] where F I is the image feature vector, F T is the text feature vector, |F I | and |F T | respectively represent their magnitudes, and · represents the dot product of vectors.

[0079] The system calculates the multimodal similarity between the visual features and text descriptions for each pair of clothing props A1 and A2. The similarity calculation method is the same as the previous formula:

[0080]

[0081] where and are the similarities between the image features and text features respectively. α and β are hyperparameters, α + β = 1, which are used to adjust the weights of visual and text similarities in the total similarity.

[0082] By calculating the similarity of each pair of clothing props, the system generates a difference heatmap, where the color depth represents the similarity between the clothing props. The color of the heatmap is calculated by the following formula:

[0083]

[0084] where represents the color value in the heatmap. The smaller the value, the higher the similarity, and the larger the value, the greater the difference. Through the heatmap, the differences between clothing props can be intuitively seen, helping to understand the similarities and differences in aesthetics and culture among different clothing props.

[0085] During the training process, a causal graph is used to model the relationship between image and text features, and cross-modal matching between vision and text is optimized through attention generation and intervention to improve the matching accuracy.

[0086] The specific steps are as follows:

[0087] First, construct a causal graph. The set of nodes in the graph is {I, T, A, M}:

[0088] where I is the image feature, including information such as the clothing of the dancer, the texture of the clothing, and the shape of the props;

[0089] T is the text feature, recording the description of ethnic dance, such as clothing features and prop uses;

[0090] A attention weight (Attention), used to highlight the key features in an image or text, is a collection of attention weights, including visual self-attention weights {α i} and text self-attention weight {β j}.

[0091] The final matching feature of M is used to measure the matching degree between I and T.

[0092] Establish directed edges I→A, T→A, (I, A)→M, (T, A)→M;

[0093] I / T→A: The model extracts attention information based on the input;

[0094] (I / T,A)→M: The final matching feature M is jointly affected by the original features and the attention-weighted features.

[0095] When evaluating the impact of attention on retrieval performance in image-text matching, we introduce the concept of counterfactual intervention. Counterfactual intervention simulates the impact of incorrect attention on the model by artificially assigning incorrect attention values ​​and severing its causal chain, thereby optimizing attention learning.

[0096] Correct attention: Correct attention is the "real" attention calculated from the image I or text T, which is used to extract key clothing or prop features. The correct attention is denoted as A = A * , which represents the influence of the original image or text on the target matching.

[0097] By using counterfactual intervention, we can force attention to have an incorrect value. (That is, artificially specifying an incorrect attention value, modifying the visual self-attention weight {α i}for and text self-attention weights {β j}for ) This intervention method can measure how the matching results of images and texts shift when attention goes wrong, so as to provide targeted constraints and corrections to the model's attention mechanism during training.

[0098] During training, the model is optimized by learning the following four pairs of matching features:

[0099] (i O ,t O ): Correct matching features obtained using the original image and original text;

[0100] (i C ,t C ): matching features obtained after counterfactual intervention (false attention);

[0101] (i O ,t C ) and (i C ,t O ): Used to measure partial matching errors (situations where one of the image or text is interfered with).

[0102] By calculating the similarity between the four sets of matching features, we can quantify the effect of each pair of image-text matching. The similarity calculation formula is as follows:

[0103] S1(I,T)=i O ·t O

[0104] S2(I,T)=i C ·t C

[0105] S3(I,T)=i O ·t C

[0106] S4(I,T)=i C ·t O

[0107] Among them, i O and t O Represent the features of the original image and text respectively, i C and t C Represents the image and text features after counterfactual intervention.

[0108] Among them, the image feature i C :

[0109] use Computing a global visual representation By gated fusion and Transformer encoding generation Take the representative features as

[0110] Text feature t C :

[0111] use Compute global text representation By gated fusion j ′ =η j ⊙t j + and Transformer encoding generation Take the representative features as

[0112] In order to ensure that the similarity of correct matches is greater than that of incorrect matches, a bidirectional triplet loss function is used to optimize performance. The loss function improves retrieval accuracy by minimizing the difference between the similarity of matches and the similarity of incorrect matches. The specific form is:

[0113] L1=max[0,d-S1(I,T)+S3(I,T)],

[0114] Ensure that the original image matches the original text better than the original image matches the intervention text.

[0115] L2=max[0,d-S1(I,T)+S4(I,T)],

[0116] Ensure that the original image matches the original text better than the intervention image matches the original text.

[0117] L3=max[0,d-S1(I,T)+S2(I,T)],

[0118] Ensure that the original match is better than the fully intervened match.

[0119] L4=max[0,d-S2(I,T)+S3(I,T)],

[0120] The relationship between constrained full intervention matching S2 and partial intervention matching S3, where d represents the marginal value.

[0121] The total loss function is L total =L1+L2+L3+L4. By optimizing this loss function, the model can learn more accurate image-text matching, thereby improving retrieval accuracy.

Claims

1. A dance costume and prop retrieval method based on causal reasoning and cross-modal matching, characterized by: The following steps are involved: In the matching task, the causal reasoning and cross-modal matching model calculates the similarity between the image matching features and the text matching features of the clothing props, and determines whether the image and text match based on the similarity threshold. In the retrieval task, the causal reasoning and cross-modal matching model calculates the similarity between the query text features and the image feature set, and sorts them from high to low by similarity, selecting the top several images as the matching results. The process of calculating features and similarities of the causal reasoning and cross-modal matching model includes: S1. The Faster R-CNN model extracts multi-scale visual features of images, and the BERT model extracts multi-granular semantic features of text descriptions. S2: Multi-scale visual features are fused through self-attention and gating and then input into the Transformer layer to obtain the final visual representation. Multi-granularity semantic features are fused through self-attention and gating and then input into the Transformer layer to obtain the final text representation. S3, perform multi-head attention calculation on the final visual representation and the final text representation respectively, merge them into one sequence, and then perform self-attention calculation again to obtain image matching features and text matching features; S4. The similarity is calculated as the dot product of the two feature vectors divided by the product of the moduli of the two feature vectors.

2. The dance costume and prop retrieval method based on causal reasoning and cross-modal matching according to claim 1 is characterized in that: Including calculating the matching relationship between clothing props A1 and A2, the causal reasoning and cross-modal matching model calculates the image matching features of each clothing prop based on steps S1 to S3 and and text matching features and Depend on Get the similarity to determine whether it matches, where α and β are hyperparameters used to adjust the weight of visual and text similarity in the total similarity. and represents the similarity calculated according to the method of step S4.

3. The dance costume and prop retrieval method based on causal reasoning and cross-modal matching according to claim 2 is characterized in that: α+β=1。 4. The dance costume and prop retrieval method based on causal reasoning and cross-modal matching according to claim 1 is characterized in that: When extracting multi-scale visual features of an image through the Faster R-CNN model, the Faster R-CNN model is used to extract K region-level original visual feature vectors for each input image, and the extracted i-th original visual feature vector is recorded as f i , Represents the original visual feature vector f i In the space where the dimension of each original visual feature vector is Then a layer of linear transformation is used to obtain the multi-scale visual feature i i =W v ·f i +b v ,i i ∈R d ,i=1,2,…,K,W v represents a weight matrix, b v Represents a bias term.

5. The dance costume and prop retrieval method based on causal reasoning and cross-modal matching according to claim 1 is characterized in that: The step S2 specifically includes: Computing self-attention weights for multi-scale visual features W a Represents the weight matrix used to calculate the self-attention weights and generate a global visual representation Will With the visual representation vector i of each local region i Fusion to obtain new regional features in is the weight vector generated by a layer of gating network, W g represents the gate weight, σ is the sigmoid function, ⊙ represents element-by-element multiplication, and [·;·] represents the vector concatenation operation. i After inputting the standard Transformer layer, the final visual representation of each local area is obtained 6. The dance costume and prop retrieval method based on causal reasoning and cross-modal matching according to claim 5 is characterized in that: The step S2 specifically includes: Calculating sub-attention weights of multi-granularity semantic features W b Represents the weight matrix used to calculate the text self-attention weight and generate the global text representation Will With the original representation vector t of each word j Combined via a gated fusion layer: in W g represents the gate weight, σ is the sigmoid function, ⊙ represents element-by-element multiplication, [·;·] represents the vector concatenation operation, and the updated word vector set {t j ′ After inputting the standard Transformer layer, the final text representation is output 7. The dance costume and prop retrieval method based on causal reasoning and cross-modal matching according to claim 6 is characterized in that: The causal reasoning and cross-modal matching model is trained based on the bidirectional triplet loss function, and the total loss function is L total =L1+L2+L3+L4; L1=max[0,d-S1(I,T)+S3(I,T)], l2=max[0,d-S1(I,T)+S4(I,T)], L3=max[0,d-S1(I,T)+S2(I,T)], L4=max[0,d-S2(I,T)+S3(I,T)], Where d represents the marginal value, S1(I,T)=i O ·t O ,S2(I,T)=i C ·t C , S3(I,T)=i O ·t C ,S4(I,T)=i C ·t O ,i O and t O represents the final visual representation and final text representation obtained in step S2, i C and t C Represents the final visual representation and final textual representation after the counterfactual intervention.

8. The dance costume and prop retrieval method based on causal reasoning and cross-modal matching according to claim 7 is characterized in that: The final visual representation and final text representation after the counterfactual intervention are artificially set using self-attention weights. and The visual representation and text representation are calculated and then obtained through gated fusion and Transformer layers.

Citation Information

Cited By

  • Distributed energy storage system state evaluation method

    CN120995280A

  • Intelligent dancing garment generation method and system based on multi-modal action analysis

    CN121302465A

  • A method and system for generating an intelligent dance costume based on multi-modal motion analysis

    CN121302465B