Virtual film generation method based on VR large space technology and related equipment thereof

By extracting key action frames and fusing multimodal encoding of descriptive verbs, the problem of high computer resource consumption in VR large-space technology is solved, achieving efficient virtual movie generation and improving the realism and generation efficiency of virtual videos.

CN121985200APending Publication Date: 2026-05-05METROPOLITAN AREA (SHANGHAI) INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
METROPOLITAN AREA (SHANGHAI) INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-02-03
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies consume significant computer resources and fail to effectively focus on important generated frames when generating virtual movies based on VR large-space technology.

Method used

By acquiring virtual movie generation materials, extracting the visual encoding results of key action frames, performing part-of-speech classification to extract descriptive verbs, generating paired visual encoding results and descriptive verb combination tokens, filtering out the target logarithmic number of retained combination tokens, and performing multimodal encoding fusion to finally generate the target virtual movie video stream.

Benefits of technology

It improves the efficiency of virtual video generation, ensuring a multi-dimensional and realistic experience, while reducing the processing of non-critical action frames and improving the utilization efficiency of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985200A_ABST
    Figure CN121985200A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of video processing, and relates to a virtual movie generation method based on a VR large space technology and related equipment thereof, and the method is applied to simulation generation creation of experience videos such as virtual propaganda movies and virtual driving collision. Codes of key action frames are extracted; extracting all description verbs from the video description text; generating paired visual coding results and description verb combination tokens according to the sequential correspondence between all the key action frames and all the description verbs; according to the method, the reserved combination tokens of the target pair number are screened out, only the relatively most important video is generated into the key frame, and the target virtual film is generated, so that the multi-dimensional vivid experience feeling of the virtual film is ensured, more coding and decoding processing resources are input into key action frame processing, the processing of non-key action frames is ignored and reduced, and the processing efficiency of the virtual film is improved. And the generation efficiency of the target virtual film is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and is applied to the simulation generation and creation of experiential videos such as virtual promotional films and virtual driving collisions. It involves virtual movie generation methods and related equipment based on VR large space technology. Background Technology

[0002] Virtual movie generation based on VR (Virtual Reality) technology has broad application prospects in 4D and 5D film production, as well as virtual driving and flight. For example, in virtual driving experiences, multiple live road videos can be re-edited to create a complete simulated road segment, allowing drivers to experience driving simulations. Another example is virtual flight scenarios, where turbulence can be created by incorporating aircraft acceleration, deceleration, and changes in atmospheric pressure. Live flight videos can then be re-edited to generate a complete simulated flight path, which can be used for pilot training.

[0003] Current technologies for virtual video generation based on VR large-space technology still largely rely on conventional full-movie source material feature encoding / decoding and feature fusion methods. This results in significant computational resource consumption during processing. While some improvements exist, such as simple scaling and adjustment of the source video to achieve virtual video generation, these methods cannot focus on crucial generated frames. Therefore, current VR large-space technology-based virtual movie generation still faces the challenges of excessive computational resources and the inability to focus on critical generated frames. Summary of the Invention

[0004] The purpose of this application is to propose a virtual movie generation method and related equipment based on VR large space technology, so as to solve the problems of excessive computer resources and inability to focus on important generated frames when generating virtual movies based on VR large space technology.

[0005] In a first aspect, embodiments of this application provide a virtual movie generation method based on VR large-space technology, which adopts the following technical solution:

[0006] The virtual movie generation method based on VR large-space technology includes the following steps: Acquire virtual movie generation materials, wherein the movie generation materials include source video, video description text, and speech synthesis package; The source video is processed by key action frame encoding extraction to obtain the visual encoding results corresponding to each key action frame. The video description text is subjected to part-of-speech classification and extraction to obtain all descriptive verbs contained in the video description text; Based on the temporal correspondence between all the key action frames and all the descriptive verbs, generate paired visual encoding results and descriptive verb combination tokens; The paired visual encoding results and descriptive verb combination tokens are input into a preset selection model to filter out the target number of retained combination tokens. Extract all descriptive verbs from the retained combination token, and input all descriptive verbs from the retained combination token and the speech synthesis package into a preset speech encoding component to generate a text-speech encoding result; Based on all descriptive verbs in the retained combined token, obtain the target visual encoding result; The text-speech encoding result and the target visual encoding result are fused using multimodal coding to obtain a multimodal coding fusion result; The multimodal encoding fusion result is input into a preset decoding component to decode and generate the video stream contained in the target virtual movie.

[0007] Secondly, embodiments of this application also provide a virtual movie generation device based on VR large-space technology, which adopts the following technical solution: A virtual movie generation device based on VR large-space technology includes: The movie generation material acquisition module is used to acquire virtual movie generation materials, wherein the movie generation materials include source video, video description text, and speech synthesis package; The key action frame encoding module is used to extract key action frames from the source video and obtain the visual encoding results corresponding to each key action frame. The descriptive verb extraction module is used to perform part-of-speech classification and extraction on the video description text to obtain all descriptive verbs contained in the video description text. The paired combination token generation module is used to generate paired visual encoding results and descriptive verb combination tokens based on the temporal correspondence between all key action frames and all descriptive verbs. The reserved combination token filtering module is used to input the paired visual encoding results and descriptive verb combination tokens into a preset selection model to filter out the target number of reserved combination tokens. The text-to-speech encoding module is used to extract all descriptive verbs in the retained combination token, input all descriptive verbs in the retained combination token and the speech synthesis package into a preset speech encoding component, and generate a text-to-speech encoding result; The target visual encoding acquisition module is used to acquire the target visual encoding result based on all descriptive verbs in the retained combination token; The multimodal coding fusion module is used to perform multimodal coding fusion on the text-speech coding result and the target visual coding result to obtain a multimodal coding fusion result; The encoding fusion result decoding module is used to input the multimodal encoding fusion result into a preset decoding component to decode and generate the video stream contained in the target virtual movie.

[0008] Thirdly, embodiments of this application also provide a computer device that adopts the technical solution described below: A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the virtual movie generation method based on VR large space technology described above.

[0009] Fourthly, embodiments of this application also provide a computer-readable storage medium, which adopts the technical solutions described below: A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the steps of the virtual movie generation method based on VR large-space technology as described above.

[0010] Compared with the prior art, the embodiments of this application have the following main advantages: The virtual movie generation method based on VR large-space technology described in this application involves: acquiring virtual movie generation materials; extracting key action frames from the source video to obtain visual encoding results corresponding to each key action frame; performing part-of-speech classification on the video description text to obtain all descriptive verbs contained in the video description text; generating paired visual encoding results and descriptive verb combination tokens based on the temporal correspondence between all key action frames and all descriptive verbs; selecting the target number of retained combination tokens; inputting all descriptive verbs and speech synthesis packets from the retained combination tokens into a preset speech encoding component to generate text-speech encoding results; obtaining the target visual encoding results based on all descriptive verbs; performing multimodal encoding fusion on the text-speech encoding results and the target visual encoding results; and inputting the multimodal encoding fusion results into a preset decoding component to decode and generate the video stream contained in the target virtual movie. By extracting key action frames and descriptive verbs, and then using a retention and combination filtering method, the most important video generation key frames are selected for generating the target virtual video. This method is applied to the simulation generation and creation of experiential videos such as virtual promotional movies and virtual driving collisions. It not only ensures the multi-dimensional and realistic experience of the virtual video, but also allows more encoding and decoding processing resources to be invested in the processing of key action frames, ignoring and reducing the processing of non-key action frames, thereby improving the generation efficiency of the target virtual video. Attached Figure Description

[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the virtual movie generation method based on VR large space technology according to this application; Figure 3 yes Figure 2 A flowchart of a specific embodiment of step 202 shown; Figure 4 yes Figure 3 A flowchart of a specific embodiment of step 302 shown; Figure 5 yes Figure 2 A flowchart of a specific embodiment of step 204 shown; Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 205 shown; Figure 7 This is a flowchart of a specific embodiment of the virtual movie generation method based on VR large space technology described in this application, which integrates haptic motion commands; Figure 8 This is a schematic diagram of a structure of an embodiment of the virtual movie generation device based on VR large space technology according to this application; Figure 9 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0013] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0014] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0015] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0016] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0017] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0018] It should be noted that the virtual movie generation method based on VR large space technology provided in this application embodiment is generally executed by a server, and correspondingly, the virtual movie generation device based on VR large space technology is generally set in the server.

[0019] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0020] Continue to refer to Figure 2 The diagram illustrates a flowchart of an embodiment of the virtual movie generation method based on VR large-space technology according to this application. The virtual movie generation method based on VR large-space technology includes the following steps: Step 201: Obtain virtual movie generation materials.

[0021] The film generation materials include source video, video description text, and speech synthesis package.

[0022] In this embodiment, the acquisition of virtual movie generation materials can be achieved by using a VR large-space technology virtual video generation and processing terminal to acquire the movie generation materials from a preset multi-source generation material provider.

[0023] Specifically, the VR large-space technology refers to LBE VR (Location-Based Entertainment VR), a virtual reality experience based on a specific location, typically offered in places like shopping malls, theme parks, and cinemas. "Large space" refers to a large virtual spatial scene. Users can move freely within a larger area, enhancing immersion. Through VR large-space technology, combined with VR virtual reality glasses worn by the user, and through virtual video capture and cross-regional interaction, users can freely communicate online across spaces within a target simulated venue.

[0024] Specifically, the source video includes: the video obtained by VR virtual video experience users based on virtual perception results after wearing VR virtual experience glasses; it also includes the source video to be generated provided by 5D cinemas, 4D cinemas, or 3D cinemas. Here, the source video only refers to the initial video used for virtual video generation, and its source and form are not limited, and it can be animated video, real photographic video, etc.; the video description text includes video introduction text, dialogue text, etc., given for the source video; the speech synthesis package refers to the guiding voice used by the speaking object in the video when generating the virtual video, which generally includes acoustic feature information such as voice timbre and voice pitch. For example: using the speech synthesis package of XXX celebrity.

[0025] Step 202: Extract key action frames from the source video to obtain the visual encoding results corresponding to each key action frame.

[0026] In this embodiment, since the source video is used to generate the target virtual video, the source video contains some background content and non-background content that changes with the video stream. Specifically, the background content includes, for example, the sky, distant trees, and mountains; while the non-background content that changes with the video stream includes, for example, pedestrians on the street and moving vehicles. Generally, the non-background content that changes with the video stream changes rapidly with the shooting, while the background content does not change rapidly. This will not be explained further here.

[0027] Specifically, background content can be understood as static image objects, and an image feature extraction component based on the OpenCV architecture is used to extract static features from the image. Non-background content that changes with the video stream can be understood as dynamic objects in the image, and a video feature extraction component based on the SlowFast model can be used to extract dynamic change features from the video stream, thereby determining key action frames.

[0028] In order to generate a target virtual video using the source video, it is necessary to identify key action frames in the source video to provide a basis for video stream changes in the subsequent generation of the target virtual video.

[0029] Step 203: Perform part-of-speech classification and extraction on the video description text to obtain all descriptive verbs contained in the video description text.

[0030] In this embodiment, the video description text is extracted by part-of-speech classification. For example, the video description text is the dialogue content of different dialogue objects in the source video, or the dynamic changes of different objects in the video, such as the sound description word "bang" when colliding or falling.

[0031] Specifically, a natural language processing model can be used to perform part-of-speech classification and extraction on the video description text, extracting nouns, verbs, prepositions, etc. from the video description text. Finally, the verbs that best represent the events that occur in the video can be extracted to generate the target virtual video.

[0032] Step 204: Based on the temporal correspondence between all key action frames and all descriptive verbs, generate paired visual encoding results and descriptive verb combination tokens.

[0033] In this embodiment, the step of generating paired visual encoding results and descriptive verb combination tokens based on the temporal correspondence between all key action frames and all descriptive verbs should be understood as follows: a descriptive verb, such as "a vehicle knocked down a telephone pole," may often involve multiple key action frames in the source video. Therefore, the paired visual encoding results and descriptive verb combination tokens here are not visual encoding results where one descriptive verb corresponds to one key action frame, but rather video encoding results where one descriptive verb corresponds to multiple key action frames.

[0034] By generating paired visual encoding results and descriptive verb combination tokens based on the temporal correspondence between all key action frames and all descriptive verbs, dynamic target virtual videos can be generated by fully combining the contextual information of the video and video description text.

[0035] Step 205: Input the paired visual encoding results and descriptive verb combination tokens into a preset selection model to filter out the target number of retained combination tokens.

[0036] In this embodiment, a preset selection model is used to select the reserved combination tokens of the target logarithm. That is, a secondary selection is performed on all the original key action frames, and only the key action frames contained in the reserved combination tokens of the target logarithm are selected for use in the generation of the target virtual video. This can greatly simplify the generation of the target virtual video, ignore the generation of non-key action frames, and improve the generation efficiency of the target virtual video.

[0037] Step 206: Extract all descriptive verbs from the retained combination token, and input all descriptive verbs from the retained combination token and the speech synthesis package into a preset speech encoding component to generate a text-to-speech encoding result.

[0038] In this embodiment, all descriptive verbs in the retained combination token are extracted, and all descriptive verbs in the retained combination token and the speech synthesis package are input into a preset speech encoding component. Specifically, the preset speech encoding component can be a preset TTS text-to-speech encoding component.

[0039] Specifically, all descriptive verbs in the retained combination token are used as key pronunciation characters or words, and the speech synthesis package is used as the reference acoustic feature to perform speech encoding of all descriptive verbs, thereby generating the text-speech encoding result.

[0040] Step 207: Obtain the target visual encoding result based on all descriptive verbs in the retained combination token.

[0041] Specifically, since each reserved combination token possesses a corresponding key action frame for a descriptive verb, the target visual encoding result can be obtained based on all the descriptive verbs in the reserved combination token. This facilitates the fusion of visual, auditory, and textual features during subsequent target virtual video generation.

[0042] Step 208: Perform multimodal coding fusion on the text-speech coding result and the target visual coding result to obtain the multimodal coding fusion result.

[0043] Specifically, the multimodal coding fusion to obtain the multimodal coding fusion result can be achieved using a pre-defined multimodal large language model (MLLM), such as VisualGPT, Qwen, etc. It can be a deep learning model that integrates multiple types of data such as text, images, videos, and audio for joint training. The core technologies include cross-modal encoder training, semantic alignment, and feature fusion, thereby obtaining the multimodal coding fusion result.

[0044] Step 209: Input the multimodal encoding fusion result into a preset decoding component to decode and generate the video stream contained in the target virtual movie.

[0045] In this embodiment, the preset decoding component can be a video stream generation component based on the Transformer neural network architecture. Ultimately, the video stream of the target virtual movie generated after decoding includes the decoding processing effect of the multimodal encoding fusion result. That is, when the target visual encoding result is only a visual encoding result, multimodal encoding fusion is performed on the text-speech encoding result and the target visual encoding result, achieving the fusion of visual-auditory-text features; while when the target visual encoding result is a visual encoding fusion result containing tactile action instructions, multimodal encoding fusion is performed on the text-speech encoding result and the target visual encoding result, achieving the fusion of visual-auditory-tactile-text features. This fully ensures that when the target virtual video stream is subsequently generated, combined with corresponding descriptive verbs, the viewer can simultaneously experience both visual decoding effects and tactile sensations, improving the virtual movie viewer's experience with the target virtual movie.

[0046] In this embodiment, the following steps are taken: First, virtual movie generation materials are acquired. Second, key action frames are encoded and extracted from the source video to obtain the visual encoding results corresponding to each key action frame. Third, part-of-speech classification is performed on the video description text to obtain all descriptive verbs contained in the video description text. Fourth, based on the temporal correspondence between all key action frames and all descriptive verbs, pairs of visual encoding results and descriptive verb combination tokens are generated. Fifth, a target number of reserved combination tokens are selected. Sixth, all descriptive verbs in the reserved combination tokens and a speech synthesis package are input into a preset speech encoding component to generate text-to-speech encoding results. Seventh, a target visual encoding result is obtained based on all descriptive verbs. Eighth, the text-to-speech encoding result and the target visual encoding result are fused using multimodal encoding. Finally, the multimodal encoding fusion result is input into a preset decoding component to decode and generate the video stream contained in the target virtual movie. By extracting key action frames and descriptive verbs, and then using a retention and combination filtering method, the most important video generation key frames are selected for generating the target virtual video. This method is applied to the simulation generation and creation of experiential videos such as virtual promotional movies and virtual driving collisions. It not only ensures the multi-dimensional and realistic experience of the virtual video, but also allows more encoding and decoding processing resources to be invested in the processing of key action frames, ignoring and reducing the processing of non-key action frames, thereby improving the generation efficiency of the target virtual video.

[0047] Continue to refer to Figure 3 , Figure 3 yes Figure 2 A flowchart of a specific embodiment of step 202 shown includes: Step 301: Cut the source video according to the preset cutting frame rate; Specifically, for example: if the total playback duration of the source video is 30 seconds, and the preset frame rate is 1 second, then 30 frames are obtained; if the preset frame rate is 0.02 seconds, then 1500 frames are obtained. By performing frame segmentation processing on the source video, a video frame sequence image corresponding to the source video is obtained.

[0048] Step 302: Compare all adjacent frame images obtained from the segmentation one by one to identify all key action frames; Specifically, by combining a pre-set image target detection model, dynamic changes of targets in adjacent frames can be detected, thereby identifying all key action frames.

[0049] Step 303: Input all the identified key action frames into the preset visual dynamic feature encoding component based on the SlowFast model according to the frame sequence to obtain the visual encoding results corresponding to each key action frame.

[0050] In this embodiment, when performing visual dynamic feature encoding on all identified key action frames, a preset visual dynamic feature encoding component based on the SlowFast model is used. The core of this visual dynamic feature encoding component lies in its parallel dual-channel design, which processes information at different time scales in the video. The two channels are a slow channel and a fast channel. The slow channel uses a low sampling rate and high resolution to focus on extracting the spatial semantic information of the video, such as the shape, color, background environment, and long-term change trends of objects. The fast channel uses a high sampling rate and low resolution to focus on capturing temporal dynamic information, such as fast actions, motion details, and instantaneous changes. Therefore, using the preset visual dynamic feature encoding component based on the SlowFast model to perform visual dynamic feature encoding on all key action frames can not only fully capture the rapid changes between frames, but also capture the slow changes between frames.

[0051] Continue to refer to Figure 4 , Figure 4 yes Figure 3 A flowchart of a specific embodiment of step 302 shown includes: Step 401: Input all adjacent frame images after the frame segmentation process into a preset image target detection model; Specifically, all adjacent frame images are pre-set with a distinguishable encoding sequence according to the frame segmentation sequence. Then, all adjacent frame images after the distinguishable encoding sequence is set are input into a preset image target detection model. The preset image target detection model includes a deep learning detection model based on optical flow algorithm, which can identify motion changes or displacement changes of the same image target between adjacent frames.

[0052] Step 402: Use the image target detection model to extract image features from all adjacent frame images after the frame segmentation process; Specifically, the deep learning detection model based on optical flow algorithm is used to extract image features from adjacent frames being compared, and to identify changes in motion or displacement of the same target in adjacent frames.

[0053] Step 403: Based on the image feature extraction results, obtain all static objects and dynamically changing objects in the image output by the image target detection model; Specifically, if the current target's motion or displacement remains unchanged, then the current target is a static object in the image; otherwise, the current target is a dynamically changing object in the image.

[0054] Step 404: Based on the dynamically changing objects in the image, identify all key action frames.

[0055] In this embodiment, the step of identifying all key action frames based on the dynamically changing objects in the image specifically includes: when it is detected that there are no dynamically changing objects in the current frame image compared to the previous frame image, the current frame image is a non-key action frame; when it is detected that there are dynamically changing objects in the current frame image compared to the previous frame image, the current frame image is initially identified as a key action frame.

[0056] Specifically, after initially identifying the current frame image as a key action frame, the amplitude of the preset key action frame can be used to filter out key action frames with larger amplitudes. This will not be elaborated on further here.

[0057] By identifying all key action frames, the system avoids blindly using all video frames during subsequent target virtual video generation, saving resource consumption during target virtual video generation, and to a certain extent, only generating key action frames, thereby improving the efficiency of target virtual video generation.

[0058] Continue to refer to Figure 5 , Figure 5 yes Figure 2 A flowchart of a specific embodiment of step 204 shown includes: Step 501: Obtain the playback timestamp information corresponding to each of the key action frames in the source video; Specifically, the playback timestamp information corresponding to each of the key action frames in the source video can be understood as a specific playback time point, such as 1 minute and 5 seconds after the video starts playing.

[0059] Step 502: Extract the pre-annotated playback time interval information corresponding to all the descriptive verbs in the video description text; Specifically, the pre-marked playback time interval information, for example: if the shouting starts at 1 minute and 2 seconds from the beginning of the video and ends at 1 minute and 6 seconds, then the pre-marked playback time interval information is the interval from 1 minute and 2 seconds to 1 minute and 6 seconds.

[0060] Step 503: Based on the playback timestamp information and the marked playback time interval information, determine all key action frames contained in the corresponding marked playback time interval information for all descriptive verbs. Specifically, assuming there are 10 key action frames within the marked playback time interval corresponding to the shout, then the 10 key action frames are determined.

[0061] Step 504: For each descriptive verb, organize all key action frames contained in the corresponding labeled playback time interval information according to the frame order to obtain the key action frame sequence corresponding to each descriptive verb. Specifically, for example, the key action frames corresponding to the shouts are arranged in chronological order to obtain the corresponding key action frame sequence.

[0062] Step 505: Using different descriptive verbs as token text identifiers, the visual encoding results of all key action frames in the key action frame sequence corresponding to each descriptive verb are used as encoding segments, and the encoding segments are spliced ​​sequentially according to the frame sequence relationship to obtain the splicing results of the encoding segments corresponding to different descriptive verbs. Specifically, continuing with the shouting sound as an example, according to the order of the 10 key action frames in the corresponding key action frame sequence, the visual encoding result of each key action frame is obtained as an encoding segment. These 10 encoding segments are then spliced ​​together according to the frame sequence relationship to obtain the encoding segment splicing result. Finally, the descriptive verb, i.e., the shouting sound, is used as the token text identifier.

[0063] Step 506: Based on the coded segment splicing result and the corresponding token text identifier, generate the paired visual encoding result and the descriptive verb combination token.

[0064] Specifically, pairs of visual encoding results and descriptive verb combination tokens are generated to facilitate subsequent multimodal feature fusion encoding processing.

[0065] Continue to refer to Figure 6 , Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 205 shown includes: Step 601: Identify the selection dimensions provided by the selection model; Specifically, the selection dimensions provided by the selection model can be understood as selection verbs. For example, the descriptive verbs include collision sounds, shouts, braking sounds, etc. If the selection verb is a collision sound, it means that the target virtual video generates more collision sounds between the targets of interest. In this case, the collision sound is used as the selection verb to filter out the corresponding descriptive verbs and visual encoding results.

[0066] Step 602: Based on the selection dimension, extract the descriptive verbs of the query dimension from the descriptive verbs and extract the visual encoding results of the key dimension from the visual encoding results; Specifically, in this embodiment, for example, the descriptive verbs are a set. , This indicates the number of elements in the set of descriptive verbs; the encoded segment in the visual encoding result is a set. , This indicates the number of coded segments in the visual encoding result.

[0067] Correspondingly, the descriptive verbs for query dimensions can be represented as: The encoded sub-segment of the key dimension can be represented as .

[0068] Step 603: Use the descriptive verbs of the query dimension and the visual encoding results of the key dimension as attention calculation parameters, and use a preset attention mechanism to perform attention calculation to obtain the attention weights of different descriptive verbs for different visual encoding results. Specifically, utilizing the preset attention mechanism The attention matrix is ​​obtained, where, The output attention matrix The element in the middle can be represented as This element refers to the first The descriptive verb pairs the first Attention weights for each encoded segment.

[0069] By from the attention matrix Different elements are obtained from the data to achieve the attention weights of different descriptive verbs for different encoded segments.

[0070] Step 604: Based on the attention weights of different descriptive verbs for different visual encoding results, select the target logarithmic retained combination tokens in descending order of attention weight.

[0071] Specifically, after obtaining the attention weights of different descriptive verbs for different encoded segments, attention weights are accumulated for the visual encoding results corresponding to different descriptive verbs based on the encoded segments contained in the visual encoding results, and finally the accumulated attention weight value corresponding to each selection verb is obtained. Based on the magnitude of the accumulated attention weight value corresponding to each selection verb, the target logarithmic retainable combination tokens are selected.

[0072] By combining selection dimensions and attention weighting, a selectable combination of tokens is chosen to ensure that the subsequently generated target virtual video focuses more on the specified key action frames. This avoids blindly using all video frames, saving resources during target virtual video generation. Furthermore, since key action frames are selected based on changes between adjacent frames, it's impossible to avoid the possibility of cross-frame key action frames lacking inter-frame differences, which could hinder resource reduction. This selection method, by focusing on generating only the truly relevant key action frames, improves the efficiency of target virtual video generation.

[0073] In this embodiment, the step of obtaining the target visual encoding result based on all descriptive verbs in the retained combination tokens specifically includes: parsing the retained combination tokens to obtain the descriptive verbs and visual encoding results contained in each retained combination token; and using the visual encoding result corresponding to each descriptive verb as the target visual encoding result.

[0074] Continue to refer to Figure 7 In some specific embodiments, after step 207, a step of fusing tactile motion commands is also included. Figure 7 This is a flowchart of a specific embodiment of the virtual movie generation method based on VR large space technology described in this application, which integrates haptic motion commands, including: Step 701: Generate tactile action instructions for all descriptive verbs in the reserved combination token to a preset tactile instruction generation component; In this embodiment, for example, in a virtual driving video or virtual flight video experience scenario, the user sits in a seat in the cockpit, which is equipped with a back thrust component or a swaying component. When the vehicle accelerates during virtual driving, the preset tactile command generation component sends a back tactile change command to the back thrust component. Furthermore, when the vehicle travels on a bumpy road, a tactile change command of body swaying can be sent to the swaying component.

[0075] Step 702: The tactile action instruction is fused with the target visual encoding result to obtain a visual encoding fusion result containing the tactile action instruction. Specifically, the encoding segment splicing result corresponding to each descriptive verb is obtained, and the tactile action instruction corresponding to each descriptive verb is used as a separate splicing prefix or suffix and spliced ​​with the corresponding encoding segment splicing result to generate a visual encoding fusion result containing the tactile action instruction. Step 703: Replace the visual encoding fusion result containing tactile action instructions with the target visual encoding result in an update manner.

[0076] In this embodiment, the fusion of tactile and visual encoding is realized, which facilitates the use of corresponding descriptive verbs during subsequent decoding, i.e., when the target virtual video stream is generated, so that the user can experience both the visual decoding effect and the tactile sensation effect, thereby improving the user experience of the target virtual video.

[0077] In this embodiment, the step of performing multimodal coding fusion on the text-speech coding result and the target visual coding result to obtain a multimodal coding fusion result specifically includes: identifying all descriptive verbs in the reserved combination token corresponding to the text-speech coding result as a first identification result; identifying all descriptive verbs in the reserved combination token corresponding to the target visual coding result as a second identification result; and performing multimodal coding fusion on the text-speech coding result and the target visual coding result based on the mutual correspondence between the descriptive verbs in the first identification result and the second identification result, as well as the sequential relationship of all descriptive verbs in the video description text.

[0078] Specifically, when the target visual encoding result is only a visual encoding result, multimodal encoding fusion is performed on the text-speech encoding result and the target visual encoding result to achieve the fusion of visual-auditory-text features; while when the target visual encoding result is a visual encoding fusion result containing tactile action instructions, multimodal encoding fusion is performed on the text-speech encoding result and the target visual encoding result to achieve the fusion of visual-auditory-tactile-text features. This fully ensures that when the target virtual video stream is subsequently generated, combined with the corresponding descriptive verbs, the experiencer can feel both the visual decoding effect and the tactile sensation effect, thereby improving the virtual video viewer's experience with the target virtual video.

[0079] In this embodiment, the following steps are taken: First, virtual movie generation materials are acquired. Second, key action frames are encoded and extracted from the source video to obtain the visual encoding results corresponding to each key action frame. Third, part-of-speech classification is performed on the video description text to obtain all descriptive verbs contained in the video description text. Fourth, based on the temporal correspondence between all key action frames and all descriptive verbs, pairs of visual encoding results and descriptive verb combination tokens are generated. Fifth, a target number of reserved combination tokens are selected. Sixth, all descriptive verbs in the reserved combination tokens and a speech synthesis package are input into a preset speech encoding component to generate text-to-speech encoding results. Seventh, a target visual encoding result is obtained based on all descriptive verbs. Eighth, the text-to-speech encoding result and the target visual encoding result are fused using multimodal encoding. Finally, the multimodal encoding fusion result is input into a preset decoding component to decode and generate the video stream contained in the target virtual movie. By extracting key action frames and descriptive verbs, and then using a retention and combination filtering method, the most important video generation key frames are selected for generating the target virtual video. This method is applied to the simulation generation and creation of experiential videos such as virtual promotional movies and virtual driving collisions. It not only ensures the multi-dimensional and realistic experience of the virtual video, but also allows more encoding and decoding processing resources to be invested in the processing of key action frames, ignoring and reducing the processing of non-key action frames, thereby improving the generation efficiency of the target virtual video.

[0080] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0081] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0082] Further reference Figure 8 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a virtual movie generation device based on VR large-space technology. This device embodiment is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0083] like Figure 8As shown, the virtual movie generation device 800 based on VR large-space technology described in this embodiment includes: a movie generation material acquisition module 801, a key action frame encoding module 802, a descriptive verb extraction module 803, a paired combination token generation module 804, a reserved combination token filtering module 805, a text-to-speech encoding module 806, a target visual encoding acquisition module 807, a multimodal encoding fusion module 808, and an encoding fusion result decoding module 809. Wherein: The movie generation material acquisition module 801 is used to acquire virtual movie generation materials, wherein the movie generation materials include source video, video description text, and speech synthesis package; The key action frame encoding module 802 is used to extract key action frames from the source video and obtain the visual encoding results corresponding to each key action frame. The descriptive verb extraction module 803 is used to perform part-of-speech classification and extraction on the video description text to obtain all descriptive verbs contained in the video description text. The paired combination token generation module 804 is used to generate paired visual encoding results and descriptive verb combination tokens based on the temporal correspondence between all key action frames and all descriptive verbs. The reserved combination token filtering module 805 is used to input the paired visual encoding results and descriptive verb combination tokens into a preset selection model to filter out the target number of reserved combination tokens. The text-to-speech encoding module 806 is used to extract all descriptive verbs in the retained combination token, input all descriptive verbs in the retained combination token and the speech synthesis package into a preset speech encoding component, and generate a text-to-speech encoding result. The target visual encoding acquisition module 807 is used to acquire the target visual encoding result based on all descriptive verbs in the retained combination token; The multimodal coding fusion module 808 is used to perform multimodal coding fusion on the text-speech coding result and the target visual coding result to obtain a multimodal coding fusion result; The encoding fusion result decoding module 809 is used to input the multimodal encoding fusion result into a preset decoding component to decode and generate the video stream contained in the target virtual movie.

[0084] This application obtains virtual movie generation materials; extracts key action frames from the source video to obtain visual encoding results corresponding to each key action frame; performs part-of-speech classification on the video description text to obtain all descriptive verbs contained in the video description text; generates paired visual encoding results and descriptive verb combination tokens based on the temporal correspondence between all key action frames and all descriptive verbs; filters out the target number of retained combination tokens; inputs all descriptive verbs and speech synthesis packets in the retained combination tokens into a preset speech encoding component to generate text-speech encoding results; obtains the target visual encoding results based on all descriptive verbs; performs multimodal encoding fusion on the text-speech encoding results and the target visual encoding results; inputs the multimodal encoding fusion results into a preset decoding component to decode and generate the video stream contained in the target virtual movie. By extracting key action frames and descriptive verbs, and then using a retention and combination filtering method, the most important video generation key frames are selected for generating the target virtual video. This method is applied to the simulation generation and creation of experiential videos such as virtual promotional movies and virtual driving collisions. It not only ensures the multi-dimensional and realistic experience of the virtual video, but also allows more encoding and decoding processing resources to be invested in the processing of key action frames, ignoring and reducing the processing of non-key action frames, thereby improving the generation efficiency of the target virtual video.

[0085] In this embodiment, the virtual movie generation device 800 based on VR large-space technology further includes a tactile motion command generation module, a tactile-visual encoding fusion module, and a target visual encoding result update module. Wherein: A tactile motion instruction generation module is used to generate tactile motion instructions for all descriptive verbs in the reserved combination token by feeding them into a preset tactile instruction generation component. The tactile-visual encoding fusion module is used to fuse the tactile action command with the target visual encoding result to obtain a visual encoding fusion result containing the tactile action command. Specifically, it obtains the splicing result of the encoding sub-segment corresponding to each descriptive verb, and splices the tactile action command corresponding to each descriptive verb as a separate splicing prefix or suffix with the corresponding encoding sub-segment splicing result to generate a visual encoding fusion result containing the tactile action command. The target visual encoding result update module is used to replace the visual encoding fusion result containing tactile action instructions with the target visual encoding result in an update manner.

[0086] In this embodiment, the key action frame encoding module 802 includes a frame segmentation processing unit, a key action frame recognition unit, and a visual encoding unit. Wherein: A frame segmentation processing unit is used to segment the source video according to a preset segmentation frame rate; The key action frame recognition unit is used to compare all adjacent frame images obtained from the segmentation one by one to identify all key action frames; it is also used to identify all key action frames based on the dynamically changing objects in the image.

[0087] The visual encoding unit is used to input all the identified key action frames into a preset visual dynamic feature encoding component based on the SlowFast model in the frame sequence to obtain the visual encoding results corresponding to each key action frame.

[0088] In this embodiment, the key action frame encoding module 802 further includes an image detection input unit, an image feature extraction unit, and an image static and dynamic object output unit. Wherein: An image detection input unit is used to input all adjacent frame images after the frame segmentation process into a preset image target detection model; An image feature extraction unit is used to extract image features from all adjacent frame images after the frame segmentation process using the image target detection model. The image static and dynamic object output unit is used to obtain all static and dynamically changing objects in the image output by the image target detection model based on the image feature extraction results.

[0089] In this embodiment, the paired token generation module 804 includes a playback timestamp information acquisition unit, a playback time interval information extraction unit, a key action frame determination unit, a key action frame sequence acquisition unit, an encoded segment splicing unit, and a combined token generation unit. Wherein: The playback timestamp information acquisition unit is used to acquire the playback timestamp information corresponding to each of the key action frames in the source video. The playback time interval information extraction unit is used to extract the pre-annotated playback time interval information corresponding to all the descriptive verbs in the video description text; The key action frame determination unit is used to determine, based on the playback timestamp information and the marked playback time interval information, all key action frames contained in all descriptive verbs within the corresponding marked playback time interval information. The key action frame sequence acquisition unit is used to organize all key action frames contained in each descriptive verb within the corresponding labeled playback time interval information according to the frame sequence order, so as to obtain the key action frame sequence corresponding to each descriptive verb. The encoding segment splicing unit is used to use different descriptive verbs as token text identifiers, and the visual encoding results of all key action frames in the key action frame sequence corresponding to each descriptive verb as encoding segments. The encoding segments are spliced ​​sequentially according to the frame sequence relationship to obtain the encoding segment splicing results corresponding to different descriptive verbs. The combined token generation unit is used to generate the paired visual encoding results and descriptive verb combined tokens based on the encoding segment concatenation results and the corresponding token text identifiers.

[0090] In this embodiment, the retention combined token filtering module 805 includes a selection dimension identification unit, a selection extraction unit, an attention weight calculation unit, and a retention combined token filtering unit. Wherein: The selection dimension identification unit is used to identify the selection dimensions provided by the selection model. The selection and extraction unit is used to extract the descriptive verbs of the query dimension from the descriptive verbs and the visual encoding results of the key dimension from the visual encoding results, based on the selection dimension. The attention weight calculation unit is used to take the descriptive verbs of the query dimension and the visual encoding results of the key dimension as attention calculation parameters, and use a preset attention mechanism to perform attention calculation to obtain the attention weights of different descriptive verbs for different visual encoding results. The reserved combination token filtering unit is used to filter out the target logarithmic reserved combination tokens based on the attention weights of different descriptive verbs for different visual encoding results, in descending order of attention weight.

[0091] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0092] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0093] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 9 , Figure 9 This is a basic structural block diagram of the computer device in this embodiment.

[0094] The computer device 9 includes a memory 9a, a processor 9b, and a network interface 9c that are interconnected via a system bus. It should be noted that... Figure 9 Only a computer device 9 with component memory 9a, processor 9b, and network interface 9c is shown. However, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0095] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0096] The memory 9a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 9a may be an internal storage unit of the computer device 9, such as the hard disk or memory of the computer device 9. In other embodiments, the memory 9a may also be an external storage device of the computer device 9, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 9. Of course, the memory 9a may also include both the internal storage unit and its external storage device of the computer device 9. In this embodiment, the memory 9a is typically used to store the operating system and various application software installed on the computer device 9, such as computer-readable instructions for virtual movie generation methods based on VR large space technology. In addition, the memory 9a can also be used to temporarily store various types of data that have been output or will be output.

[0097] In some embodiments, the processor 9b may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 9b is typically used to control the overall operation of the computer device 9. In this embodiment, the processor 9b is used to execute computer-readable instructions stored in the memory 9a or to process data, for example, to execute computer-readable instructions for the virtual movie generation method based on VR large-space technology.

[0098] The network interface 9c may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 9 and other electronic devices.

[0099] The computer device proposed in this embodiment belongs to the field of video processing technology. This application obtains virtual movie generation materials; extracts key action frames from the source video to obtain visual encoding results corresponding to each key action frame; performs part-of-speech classification on the video description text to obtain all descriptive verbs contained in the video description text; generates paired visual encoding results and descriptive verb combination tokens based on the temporal correspondence between all key action frames and all descriptive verbs; filters out the target number of retained combination tokens; inputs all descriptive verbs and speech synthesis packets from the retained combination tokens into a preset speech encoding component to generate text-to-speech encoding results; obtains the target visual encoding result based on all descriptive verbs; performs multimodal encoding fusion on the text-to-speech encoding result and the target visual encoding result; inputs the multimodal encoding fusion result into a preset decoding component to decode and generate the video stream contained in the target virtual movie. By extracting key action frames and descriptive verbs, and then using a retention and combination filtering method, the most important video generation key frames are selected for generating the target virtual video. This method is applied to the simulation generation and creation of experiential videos such as virtual promotional movies and virtual driving collisions. It not only ensures the multi-dimensional and realistic experience of the virtual video, but also allows more encoding and decoding processing resources to be invested in the processing of key action frames, ignoring and reducing the processing of non-key action frames, thereby improving the generation efficiency of the target virtual video.

[0100] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by a processor to cause the processor to perform the steps of the virtual movie generation method based on VR large space technology as described above.

[0101] The computer-readable storage medium proposed in this embodiment belongs to the field of video processing technology. This application obtains virtual movie generation materials; extracts key action frames from the source video to obtain visual encoding results corresponding to each key action frame; performs part-of-speech classification on the video description text to obtain all descriptive verbs contained in the video description text; generates paired visual encoding results and descriptive verb combination tokens based on the temporal correspondence between all key action frames and all descriptive verbs; filters out the target number of retained combination tokens; inputs all descriptive verbs and speech synthesis packets from the retained combination tokens into a preset speech encoding component to generate text-to-speech encoding results; obtains the target visual encoding results based on all descriptive verbs; performs multimodal encoding fusion on the text-to-speech encoding results and the target visual encoding results; inputs the multimodal encoding fusion results into a preset decoding component to decode and generate the video stream contained in the target virtual movie. By extracting key action frames and descriptive verbs, and then using a retention and combination filtering method, the most important video generation key frames are selected for generating the target virtual video. This method is applied to the simulation generation and creation of experiential videos such as virtual promotional movies and virtual driving collisions. It not only ensures the multi-dimensional and realistic experience of the virtual video, but also allows more encoding and decoding processing resources to be invested in the processing of key action frames, ignoring and reducing the processing of non-key action frames, thereby improving the generation efficiency of the target virtual video.

[0102] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0103] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to make the disclosure of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

Claims

1. A virtual movie generation method based on VR large-space technology, characterized in that, Includes the following steps: Acquire virtual movie generation materials, wherein the movie generation materials include source video, video description text, and speech synthesis package; The source video is processed by key action frame encoding extraction to obtain the visual encoding results corresponding to each key action frame. The video description text is subjected to part-of-speech classification and extraction to obtain all descriptive verbs contained in the video description text; Based on the temporal correspondence between all the key action frames and all the descriptive verbs, generate paired visual encoding results and descriptive verb combination tokens; The paired visual encoding results and descriptive verb combination tokens are input into a preset selection model to filter out the target number of retained combination tokens. Extract all descriptive verbs from the retained combination token, and input all descriptive verbs from the retained combination token and the speech synthesis package into a preset speech encoding component to generate a text-speech encoding result; Based on all descriptive verbs in the retained combined token, obtain the target visual encoding result; The text-speech encoding result and the target visual encoding result are fused using multimodal coding to obtain a multimodal coding fusion result; The multimodal encoding fusion result is input into a preset decoding component to decode and generate the video stream contained in the target virtual movie.

2. The virtual movie generation method based on VR large-space technology according to claim 1, characterized in that, The step of extracting key action frames from the source video to obtain the visual encoding results corresponding to each key action frame specifically includes: The source video is cut according to a preset cutting frame rate; By comparing each adjacent frame image obtained from the segmentation one by one, all key action frames are identified. All identified key action frames are input into a preset visual dynamic feature encoding component based on the SlowFast model according to the frame sequence to obtain the visual encoding results corresponding to each key action frame.

3. The virtual movie generation method based on VR large-space technology according to claim 2, characterized in that, The step of comparing each adjacent frame image obtained from the segmentation sequentially to identify all key action frames specifically includes: The images of all adjacent frames after the frame segmentation process are input into a preset image target detection model; The image target detection model is used to extract image features from all adjacent frames after the frame segmentation process; Based on the image feature extraction results, obtain all static objects and dynamically changing objects in the image output by the image target detection model; Based on the dynamically changing objects in the image, all key action frames were identified.

4. The virtual movie generation method based on VR large-space technology according to claim 1, characterized in that, The step of generating paired visual encoding results and descriptive verb combination tokens based on the temporal correspondence between all key action frames and all descriptive verbs specifically includes: Obtain the playback timestamp information corresponding to each of the key action frames in the source video; Extract the pre-annotated playback time interval information corresponding to all the descriptive verbs in the video description text; Based on the playback timestamp information and the marked playback time interval information, determine all key action frames contained in the corresponding marked playback time interval information for all descriptive verbs; For each descriptive verb, all key action frames contained within the corresponding labeled playback time interval are arranged in frame order to obtain the key action frame sequence corresponding to each descriptive verb. Different descriptive verbs are used as token text identifiers. The visual encoding results of all key action frames in the key action frame sequence corresponding to each descriptive verb are used as encoding segments. The encoding segments are concatenated in sequence according to the frame sequence relationship to obtain the encoding segment concatenation results corresponding to different descriptive verbs. Based on the concatenation result of the encoded segments and the corresponding token text identifier, the paired visual encoding result and the descriptive verb combination token are generated.

5. The virtual movie generation method based on VR large-space technology according to claim 4, characterized in that, The step of inputting the paired visual encoding results and descriptive verb combination tokens into a preset selection model to filter out the target number of retained combination tokens specifically includes: Identify the selection dimensions provided by the selection model; Based on the selection dimensions, extract the descriptive verbs of the query dimension from the descriptive verbs and extract the visual encoding results of the key dimension from the visual encoding results; The descriptive verbs of the query dimension and the visual encoding results of the key dimension are used as attention calculation parameters. Attention is calculated using a preset attention mechanism to obtain the attention weights of different descriptive verbs for different visual encoding results. Based on the attention weights of different descriptive verbs for different visual encoding results, the target logarithmic retained combination tokens are selected in descending order of attention weight.

6. The virtual movie generation method based on VR large-space technology according to claim 5, characterized in that, After performing the step of obtaining the target visual encoding result based on all descriptive verbs in the retained combined token, the method further includes: Generate tactile action instructions for all descriptive verbs in the reserved combination token to a preset tactile instruction generation component; The tactile action command is fused with the target visual encoding result to obtain a visual encoding fusion result containing the tactile action command. Specifically, the encoding segment splicing result corresponding to each descriptive verb is obtained, and the tactile action command corresponding to each descriptive verb is used as a separate splicing prefix or suffix and spliced ​​with the corresponding encoding segment splicing result to generate a visual encoding fusion result containing the tactile action command. The visual encoding fusion result containing tactile action instructions is replaced with the target visual encoding result in an update manner.

7. The virtual movie generation method based on VR large-space technology according to claim 1, characterized in that, The step of performing multimodal coding fusion of the text-speech coding result and the target visual coding result to obtain a multimodal coding fusion result specifically includes: Identify all descriptive verbs in the retained combination token corresponding to the text-speech encoding result, and use them as the first identification result; Identify all descriptive verbs in the retained combination token corresponding to the target visual encoding result, and use them as the second identification result; Based on the correspondence between the descriptive verbs in the first and second recognition results, and the sequential relationship of all the descriptive verbs in the video description text, multimodal coding fusion is performed on the text-speech coding result and the target visual coding result.

8. A virtual movie generation device based on VR large-space technology, characterized in that, include: The movie generation material acquisition module is used to acquire virtual movie generation materials, wherein the movie generation materials include source video, video description text, and speech synthesis package; The key action frame encoding module is used to extract key action frames from the source video and obtain the visual encoding results corresponding to each key action frame. The descriptive verb extraction module is used to perform part-of-speech classification and extraction on the video description text to obtain all descriptive verbs contained in the video description text. The paired combination token generation module is used to generate paired visual encoding results and descriptive verb combination tokens based on the temporal correspondence between all key action frames and all descriptive verbs. The reserved combination token filtering module is used to input the paired visual encoding results and descriptive verb combination tokens into a preset selection model to filter out the target number of reserved combination tokens. The text-to-speech encoding module is used to extract all descriptive verbs in the retained combination token, input all descriptive verbs in the retained combination token and the speech synthesis package into a preset speech encoding component, and generate a text-to-speech encoding result; The target visual encoding acquisition module is used to acquire the target visual encoding result based on all descriptive verbs in the retained combination token; The multimodal coding fusion module is used to perform multimodal coding fusion on the text-speech coding result and the target visual coding result to obtain a multimodal coding fusion result; The encoding fusion result decoding module is used to input the multimodal encoding fusion result into a preset decoding component to decode and generate the video stream contained in the target virtual movie.

9. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the virtual movie generation method based on VR large space technology as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the virtual movie generation method based on VR large-space technology as described in any one of claims 1 to 7.