Multi-modal data fusion shop exploration video automatic generation method

Automatically generate store-detecting videos through multimodal data fusion, solving the problem of inefficient manual screening, achieving efficient quantification of multimodal data popularity and user-defined video content, optimizing the video structure, and improving the video promotion effect.

CN120378712APending Publication Date: 2025-07-25YANGZHOU TANBAO INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510730761.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing method of creating videos in the shop relies on manual screening of materials and editing, which is inefficient and it is difficult to quantify the popularity of multimodal data manually, resulting in underestimation of popular content. At the same time, it is difficult to adjust the video structure in real time when users customize the direction of focus.

Method used

The automatic generation method of Tandian videos using multimodal data fusion is adopted. By automatically collecting audio, text comments and video data, a popular index model is built, highly disseminated content is screened, and the score index is calculated based on user scores, and the video structure is dynamically adjusted to achieve automatic video generation and optimization.

Benefits of technology

Significantly improve data processing efficiency, avoid omissions of key materials, ensure that audio, text and visual features are expressed in a coordinated manner during video generation, support users to customize content to focus on directions, optimize video structure, and improve promotion effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378712A_ABST
    Figure CN120378712A_ABST
Patent Text Reader

Abstract

The invention, which relates to the technical field of video automatic generation, discloses a multi-modal data fusion-based shop exploration video automatic generation method comprising the following steps: acquiring the like quantity, comment number and release time of audio data, text comment data and video data in a to-be-fused data set, and performing calculation and analysis to obtain a database; obtaining a popularity index of the data in the to-be-fused data set, screening and integrating the data in the to-be-fused data set into a popularity data set, presetting an emphasis type of the exploration video, and performing calculation and analysis in combination with scores of different emphasis by a user to obtain an exploration video score index for supporting the user to customize the emphasis direction of the exploration video content. Meanwhile, a shop exploration video is automatically generated; according to the method, audio, character comments and video data are automatically collected, and the popularity index model is constructed based on the like quantity, the comment number and the time attenuation coefficient, so that quantitative screening of high-transmissibility contents is realized, the data processing efficiency is remarkably improved, and manual missing of key materials is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video automatic generation, and particularly to an automatic generation method for store exploration videos with multi-modal data fusion. Background Art

[0002] With the development and progress of technology, the promotion methods of merchants have become diverse and are no longer limited to distributing leaflets and erecting billboards at the storefront. The promotion methods have gradually shifted from offline promotion to online promotion. By spreading advertising information through the Internet, the advertising scope is no longer restricted by regions. When merchants promote online, pictures or videos are required as advertising information. Existing short video platforms are one of the important promotion platforms. Store exploration influencers release store exploration videos related to merchants and use their own traffic for dissemination and publicity to help merchants with promotion. However, in actual promotion, store exploration influencers do not necessarily go to the merchant's store to shoot in person, but generate and release store exploration videos based on the information data provided by the merchants.

[0003] However, the method for generating store exploration videos relies on manual screening of materials and editing, and its efficiency is low (for example, merchants need to manually extract effective information from a large number of comments or select video clips frame by frame, and it is very easy to miss highly transmissible content). In addition, it is difficult for humans to quantify the popularity of multi-modal data (such as unable to comprehensively calculate the popularity index through the number of likes, the number of comments, and the release time), resulting in the possible underestimation of popular audio or videos. At the same time, when users customize the focus direction, it is difficult for humans to adjust the video structure in real time (for example, when users prefer the "price" focus, traditional editing requires re-editing consumption-related clips, while the automatic generation method can dynamically allocate the duration ratio based on the score index).

[0004] In view of the above technical defects, a solution is now proposed. Summary of the Invention

[0005] The purpose of the present invention is to solve the problems that the traditional method for generating store exploration videos relies on manual screening of materials and editing, resulting in low efficiency; in addition, it is difficult for humans to quantify the popularity of multi-modal data, leading to the possible underestimation of popular audio or videos; and at the same time, it is difficult for humans to adjust the video structure in real time when users customize the focus direction.

[0006] To achieve the above purpose, the present invention adopts the following technical solution: An automatic generation method for store exploration videos with multi-modal data fusion, comprising the following steps:

[0007] Step 1: Collect information data related to the production of store exploration videos and integrate them into a dataset to be fused. The dataset to be fused includes audio data, text comment data, and video data. At the same time, preprocess the data in the dataset to be fused;

[0008] Step 2: Obtain the number of likes, the number of comments, and the release time of the audio data, text comment data, and video data in the dataset to be fused, perform calculation and analysis, obtain the popularity index of the data in the dataset to be fused, and screen and integrate the data in the dataset to be fused into a popular dataset;

[0009] Step 3: Convert the feature vectors of the audio data, text comment data, and video data in the popular dataset, perform calculation and analysis based on the converted feature vectors, and obtain the multi-modal data fusion feature vector for realizing the unified fusion of multi-modal features;

[0010] Step 4: Preset the focus types of the store exploration videos, perform calculation and analysis in combination with the scores of users on different focuses, obtain the score index of the store exploration videos, which is used to support users to customize the content focus direction of the store exploration videos and automatically generate store exploration videos;

[0011] Step 5: Automatically publish the automatically generated store exploration videos to the video platform, collect the feedback data of the published store exploration videos, and optimize the multi-modal data fusion strategy according to the feedback data.

[0012] Furthermore, the calculation and analysis process of the popularity index of the data in the dataset to be fused is as follows:

[0013] S11: Obtain the number of likes, the number of comments, and the release time of the audio data, text comment data, and video data in the dataset to be fused, and perform calculation and analysis;

[0014] S12: Calculate the popularity index P of the data in the dataset to be fused according to the following formula:

[0015]

[0016] where L is the number of likes of the data to be fused, C is the number of comments of the data to be fused, μ is the time decay coefficient, t c is the current time, t i is the release time of the i-th data to be fused, α is the preset weight coefficient of the number of likes, β is the preset weight coefficient of the number of comments, α + β = 1, and the popularity index is used to quantify the data heat through summation.

[0017] Furthermore, the process of screening the data in the dataset to be fused is as follows: Obtain the preset popularity threshold P e , when P ≥ P e , it indicates that the data heat is high, and it is retained and integrated into the popular dataset. When P < P e , it indicates that the data heat is low, and the low-heat data is screened out.

[0018] Furthermore, the calculation and analysis process of the multi-modal data fusion feature vector is as follows:

[0019] S21. Convert the audio data, text comment data, and video data in the popularity dataset into feature vectors, and perform calculation and analysis based on the converted feature vectors.

[0020] S22. Calculate the multi-modal data fusion feature vector F according to the following formula:

[0021]

[0022] where n is the number of types of data modalities involved in the generation process of the store visit video, w k is the weight coefficient of the k-th type of modality data, and f k is the feature vector of the k-th type of modality data. The multi-modal data fusion feature vector is used to achieve the unified fusion of multi-modal features through summation.

[0023] Furthermore, the analysis and calculation process of the store visit video score index is as follows:

[0024] S31. Preset the focus types of the store visit video, and perform calculation and analysis in combination with the scores of users for different focuses.

[0025] S32. Calculate the store visit video score index S according to the following formula:

[0026]

[0027] where m is the number of preset focus types of the store visit video, u j is the weight coefficient of the j-th focus of the store visit video, g j is the score of the user for the j-th focus of the store visit video. The store visit video score index is used to support users to customize the content focus direction of the store visit video.

[0028] Furthermore, the feedback data of the store visit video includes the play completion rate, interaction rate, and the analysis result of the sentiment tendency of new comments. Dynamically optimizing the multi-modal data fusion strategy according to the feedback data includes: expanding or deleting the types of modalities in the dataset to be fused, adjusting the weight coefficients of modality data, and adjusting the weight coefficients of the focuses of the store visit video.

[0029] Furthermore, the preprocessing of the dataset to be fused includes: performing noise reduction processing on the audio data, performing word segmentation and stop word filtering on the text comment data, and performing frame rate unification and key frame extraction on the video data.

[0030] Further, when automatically generating the store visit video, the duration ratio of video segments corresponding to different focuses is dynamically allocated according to the store visit video score index, and the redundant frames of the coarse-grained video script are removed and the lens transition is optimized through a hybrid converter.

[0031] Further, when optimizing the multi-modal data fusion strategy according to the feedback data, if the sentiment tendency of the newly added comment is negative, the weight coefficient of the corresponding modal data is reduced.

[0032] Further, the value range of the time decay coefficient μ is 0 < μ < 1, and it monotonically increases with the increase of the difference between the current time and the data release time, which is used to reflect the decay law of data popularity over time.

[0033] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:

[0034] The method for automatically generating store visit videos with multi-modal data fusion automatically collects audio, text comments and video data, and constructs a popularity index model based on the number of likes, the number of comments and the time decay coefficient to realize the quantitative screening of highly transmissible content, significantly improve the data processing efficiency, and avoid manual omission of key materials. Secondly, the multi-modal feature vector transformation and weighted summation strategy are adopted to unify the fusion of different modal data, reduce the interference of redundant information, and ensure that the video generation process takes into account the collaborative expression of audio, text and visual features. In addition, by presetting the types of store visit focuses and combining with user ratings to calculate the score index, it supports users to customize the focus direction of video content, dynamically allocate the duration ratio of segments, and realize the intelligent optimization of the video structure. Finally, based on the feedback data such as the play completion rate, interaction rate and comment sentiment analysis, the modal weights and focus coefficients are dynamically adjusted to form a closed-loop optimization mechanism, so that the generated videos are more in line with the platform's dissemination rules and continuously improve the promotion effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 The schematic diagram of the method flow of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] Hereinafter, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0037] Embodiment:

[0038] As Figure 1 shown, the method for automatically generating store visit videos with multi-modal data fusion includes the following steps:

[0039] Step 1: Collect information data related to the production of store visit videos and integrate them into a dataset to be fused. The dataset to be fused includes audio data (such as in-store ambient sound, customer conversation sound), text comment data (such as user evaluations of store visits), and video data (such as in-store environment videos, food production process videos, customer dining scene videos). At the same time, preprocess the data in the dataset to be fused. The preprocessing of the dataset to be fused includes: performing noise reduction on the audio data, performing word segmentation and stop word filtering on the text comment data, and unifying the frame rate and extracting key frames from the video data;

[0040] Step 2: Obtain the number of likes, the number of comments, and the release time of the audio data, text comment data, and video data in the dataset to be fused, and perform calculation and analysis to obtain the popularity index of the data in the dataset to be fused, and screen and integrate the data in the dataset to be fused into a popular dataset;

[0041] The calculation process of the popularity index analysis of the data in the dataset to be fused is as follows:

[0042] S11: Obtain the number of likes, the number of comments, and the release time of the audio data, text comment data, and video data in the dataset to be fused, and perform calculation and analysis;

[0043] S12: Calculate the popularity index P of the data in the dataset to be fused according to the following formula:

[0044]

[0045] where L is the number of likes of the data to be fused, C is the number of comments of the data to be fused, μ is the time decay coefficient, and the value range of the time decay coefficient μ is 0 < μ < 1, and it is monotonically increasing with the increase of the difference between the current time and the data release time, which is used to reflect the attenuation law of data heat over time, t c is the current time, t i is the release time of the i-th data to be fused, α is the preset weight coefficient of the number of likes, β is the preset weight coefficient of the number of comments, α + β = 1, and the popularity index is used to quantify the data heat through summation to preferentially screen high-propagation content;

[0046] The process of screening the data in the dataset to be fused is as follows: Obtain the preset popularity threshold P e , when P ≥ P e , it indicates that the data heat is high, and it is retained and integrated into the popular dataset. When P < P e , it indicates that the data heat is low, and the low-heat data is screened out to reduce the subsequent data processing pressure.

[0047] Step 3: Convert the audio data, text comment data, and video data in the popularity dataset into feature vectors, perform calculation and analysis based on the converted feature vectors, and obtain the multi-modal data fusion feature vectors for realizing the unified fusion of multi-modal features;

[0048] The analysis and calculation process of the multi-modal data fusion feature vectors is as follows:

[0049] S21: Convert the audio data, text comment data, and video data in the popularity dataset into feature vectors, and perform calculation and analysis based on the converted feature vectors;

[0050] S22: Calculate the multi-modal data fusion feature vector F according to the following formula:

[0051]

[0052] where n is the number of types of data modalities involved in the process of generating the store exploration video, and its value is 3, w k is the weight coefficient of the k-th type of modal data, and f k is the feature vector of the k-th type of modal data. The multi-modal data fusion feature vector is used to achieve the unified fusion of multi-modal features through summation and reduce the interference of redundant information.

[0053] Step 4: Preset the focus types of the store exploration video (such as in-store environment, food taste, service quality, and price), combine the scores given by users for different focuses for calculation and analysis, and obtain the score index of the store exploration video to support users to customize the content focus direction of the store exploration video and automatically generate the store exploration video;

[0054] The analysis and calculation process of the score index of the store exploration video is as follows:

[0055] S31: Preset the focus types of the store exploration video, and combine the scores given by users for different focuses for calculation and analysis;

[0056] S32: Calculate the score index S of the store exploration video according to the following formula:

[0057]

[0058] where m is the number of preset focus types of the store exploration video, u j is the weight coefficient of the j-th focus of the store exploration video, g j$r_{j}$ is the score given by the user to the focus of the $j$-th store visit video. The score index of the store visit video is used to support the user in customizing the focus direction of the store visit video content, and then adjusting the duration ratio of video segments (for example, if the user focuses on "price", then increase the consumption-related segments), generating a coarse-grained video script with a unified timeline. The redundant frames of the coarse-grained script are removed and the lens transitions are optimized through a hybrid transformer, and the final version of the store visit video that meets the preset duration is output.

[0059] Step 5: Automatically publish the automatically generated store visit video to the video platform, and collect the feedback data of the published store visit video, and optimize the multi-modal data fusion strategy according to the feedback data.

[0060] When optimizing the multi-modal data fusion strategy according to the feedback data, if the sentiment tendency of the new comment is negative, then reduce the weight coefficient of the corresponding modal data. The feedback data of the store visit video includes the play completion rate, the interaction rate, and the sentiment analysis result of the new comment. Dynamically optimizing the multi-modal data fusion strategy according to the feedback data includes: expanding or deleting the modal types of the dataset to be fused, adjusting the weight coefficients of the modal data, and adjusting the weight coefficients of the focus of the store visit video.

[0061] By automatically collecting audio, text comments, and video data, and constructing a popularity index model based on the number of likes, the number of comments, and the time decay coefficient, the quantitative screening of highly transmissible content is realized, significantly improving the data processing efficiency and avoiding the omission of key materials by manual work. Secondly, a multi-modal feature vector transformation and weighted summation strategy is adopted to unify the fusion of different modal data, reduce the interference of redundant information, and ensure that the video generation process takes into account the collaborative expression of audio, text, and visual features. In addition, by presetting the types of store visit focuses and calculating the score index in combination with user scores, the user is supported to customize the focus direction of video content, dynamically allocate the duration ratio of segments, and realize the intelligent optimization of the video structure. Finally, based on the feedback data such as the play completion rate, the interaction rate, and the comment sentiment analysis, the modal weights and focus coefficients are dynamically adjusted to form a closed-loop optimization mechanism, making the generated video more in line with the platform's dissemination rules and continuously improving the promotion effect.

[0062] The setting of the size of the interval and the threshold is for the convenience of comparison. Regarding the size of the threshold, it depends on the amount of sample data and the number of base numbers set by those skilled in the art for each group of sample data; as long as it does not affect the proportional relationship between the parameters and the quantified values.

[0063] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation;

[0064] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent replacements or changes, shall be covered by the protection scope of the present invention.

Claims

1. An automatic generation method for store exploration videos with multi-modal data fusion, characterized in that, It includes the following steps: Step 1: Collect information data related to the production of store visit videos and integrate them into a dataset to be fused. The dataset to be fused includes audio data, text comment data, and video data. At the same time, preprocess the data in the dataset to be fused; Step 2: Obtain the number of likes, the number of comments, and the release time of the audio data, text comment data, and video data in the dataset to be fused, and perform calculation and analysis to obtain the popularity index of the data in the dataset to be fused, and screen and integrate the data in the dataset to be fused into a popular dataset; Step 3: Transform the feature vectors of the audio data, text comment data, and video data in the popular dataset, and perform calculation and analysis based on the transformed feature vectors to obtain a multi-modal data fusion feature vector for realizing the unified fusion of multi-modal features; Step 4: Preset the focus types of the store visit videos, and perform calculation and analysis in combination with the scores of users on different focuses to obtain the score index of the store visit videos, which is used to support users to customize the content focus direction of the store visit videos and automatically generate store visit videos; Step 5: Automatically publish the automatically generated store visit videos to the video platform, and collect the feedback data of the published store visit videos, and optimize the multi-modal data fusion strategy according to the feedback data.

2. The automatic generation method of the store exploration video with multi-modal data fusion according to claim 1, wherein, The analysis and calculation process of the popularity index of the data in the dataset to be fused is as follows: S11: Obtain the number of likes, the number of comments, and the release time of the audio data, text comment data, and video data in the dataset to be fused, and perform calculation and analysis; S12: Calculate the popularity index P of the data in the dataset to be fused according to the following formula: Among them, L is the number of likes of the data to be fused, C is the number of commenters of the data to be fused, μ is the time decay coefficient, t c is the current time, t i is the release time of the i-th piece of data to be fused, α is the preset weight coefficient of the number of likes, β is the preset weight coefficient of the number of commenters, α + β = 1, and the popularity index is used to quantify the data heat by summation.

3. The automatic generation method of the store exploration video with multi-modal data fusion according to claim 1, characterized in that, The process of screening data in the fusion dataset is as follows: Obtain the preset popularity threshold P e , when P ≥ P e , it indicates that the data has high popularity. Retain it and integrate it into the popular dataset. When P < P e , it indicates that the data has low popularity. Screen out the low-popularity data.

4. The method for automatically generating an in-store exploration video with multi-modal data fusion according to claim 1, wherein The analysis and calculation process of the multi-modal data fusion feature vector is as follows: S21: Transform the feature vectors of the audio data, text comment data, and video data in the popular dataset, and perform calculation and analysis based on the transformed feature vectors; S22: Calculate the multi-modal data fusion feature vector F according to the following formula: Among them, n is the type of data modality involved in the process of generating the store exploration video, and w k is the weight coefficient of the k-th modality data, and f k is the feature vector of the k-th modality data. The multi-modal data fusion feature vector is used to achieve unified fusion of multi-modal features through summation.

5. The automatic generation method of the store exploration video for multimodal data fusion according to claim 1, wherein, The analysis and calculation process of the score index of the store visit videos is as follows: S31: Preset the focus types of the store visit videos, and perform calculation and analysis in combination with the scores of users on different focuses; S32: Calculate the score index S of the store visit videos according to the following formula: Among them, m is the number of preset types of focus of the store visit video, and u j is the weight coefficient of the j-th focus of the store visit video, g j is the score given by the user to the j-th focus of the store visit video. The store visit video score index is used to support the user to customize the content focus direction of the store visit video.

6. The method for automatically generating a store visit video with multi-modal data fusion according to claim 1, characterized in that The feedback data of the store visit videos includes the play completion rate, the interaction rate, and the sentiment analysis results of the new comments. Dynamically optimizing the multi-modal data fusion strategy according to the feedback data includes: expanding or deleting the modal types of the dataset to be fused, adjusting the weight coefficients of the modal data, and adjusting the weight coefficients of the focuses of the store visit videos.

7. The method for automatically generating a store visit video with multi-modal data fusion according to claim 1, characterized in that The preprocessing of the dataset to be fused includes: performing noise reduction processing on the audio data, performing word segmentation and stop word filtering on the text comment data, and performing frame rate unification and key frame extraction on the video data.

8. The method for automatically generating a store visit video with multimodal data fusion according to claim 1, characterized in that, When automatically generating a store visit video, dynamically allocate the video segment duration ratio corresponding to different focuses according to the score index of the store visit video, and perform redundant frame removal and shot transition optimization on the coarse-grained video script through a hybrid transformer.

9. The method for automatically generating an in-store exploration video with multi-modal data fusion according to claim 1, wherein When optimizing the multi-modal data fusion strategy according to the feedback data, if the sentiment of the new comment is negative, reduce the weight coefficient of the corresponding modal data.

10. The method for automatically generating a store visit video with multi-modal data fusion according to claim 2, wherein, The value range of the time decay coefficient μ is 0 < μ < 1, and it increases monotonically with the increase of the difference between the current time and the data release time, which is used to reflect the decay law of data popularity over time.