Dynamic video display method, projection equipment and storage medium

By determining the scene image containing mosaic in the image selection page and casting dynamic video when the target mosaic is selected, the problem of difficulty in dynamic static images is solved, dynamic conversion and AR enhancement of static images are realized, and the user experience is improved.

CN119996635APending Publication Date: 2025-05-13CHENGDU XGIMI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311500590.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to convert static single images into dynamic videos, and there is a lack of effective methods to use static images to achieve dynamic effects.

Method used

By displaying the image selection page, determine the scene image containing the mosaic, and when the selection operation of the target mosaic is detected, the dynamic video corresponding to the target image is projected to the position of the target mosaic, realizing dynamic conversion of the static image.

Benefits of technology

It realizes dynamic conversion of static images, enhances the practicality and user experience of images, and supports augmented reality (AR) applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996635A_ABST
    Figure CN119996635A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic video display method, projection equipment and a storage medium, and the method comprises the steps: selecting a scene image containing at least one mosaic frame in a current scene in a display page, and when a selection operation of a target mosaic frame in the scene image is detected, displaying the target mosaic frame in the scene image; according to the technical scheme, the dynamic video corresponding to the static target image can be projected to the mosaic frame position of the target mosaic frame, so that the target object in the target image is displayed in a dynamic form, dynamic conversion of a single static image is realized, a projection picture is enhanced, and meanwhile, the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a dynamic video display method, a projection device and a storage medium. Background Art

[0002] Generally speaking, a video or animated image is composed of different frames that change in time sequence. The dynamic effect of the picture content is achieved by playing the continuous picture content at a certain frame rate on the timeline. Therefore, if you want to obtain dynamic picture content, you often need to shoot continuous pictures for a specific length of time. However, in real life where cameras are very popular, there are a huge number of static single images, which are far from being fully utilized. Therefore, how to give these static single images dynamic life and make them move is an urgent problem to be solved. Summary of the invention

[0003] The embodiments of the present application are intended to provide a dynamic video display method, a projection device, and a storage medium to solve the problem in the related art of giving a static single image a dynamic life and making the static image move.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] In a first aspect, an embodiment of the present application provides a dynamic video display method, the method comprising:

[0006] Display the image selection page;

[0007] Determine a scene image through the image selection page; wherein the scene image includes at least one mosaic frame under the current scene;

[0008] When a selection operation on a target mosaic frame in the scene image is detected, a dynamic video corresponding to the obtained target image is projected to the mosaic frame position of the target mosaic frame; wherein a target object contained in the target image is in motion in the dynamic video.

[0009] In a second aspect, an embodiment of the present application provides a dynamic video display device, the device comprising:

[0010] A display module, used for displaying an image selection page;

[0011] A determination module, used to determine a scene image through the image selection page; wherein the scene image includes at least one mosaic frame under the current scene;

[0012] The projection module is used to project the dynamic video corresponding to the obtained target image to the mosaic frame position of the target mosaic frame when a selection operation is detected for the target mosaic frame in the scene image; wherein the target object contained in the target image is in motion in the dynamic video.

[0013] In a third aspect, an embodiment of the present application provides a projection device, the projection device comprising: a processor, a memory, and a communication bus;

[0014] The communication bus is used to realize the communication connection between the processor and the memory;

[0015] The processor is used to execute the dynamic video display program stored in the memory to implement the dynamic video display method described above.

[0016] In a fourth aspect, an embodiment of the present application provides a storage medium, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the dynamic video display method described above.

[0017] Embodiments of the present application provide a dynamic video display method, a projection device, and a storage medium, by displaying an image selection page; determining a scene image through the image selection page; wherein the scene image includes at least one mosaic frame in the current scene; when a selection operation for a target mosaic frame in the scene image is detected, a dynamic video corresponding to the obtained target image is projected to the mosaic frame position of the target mosaic frame; wherein a target object included in the target image is in motion in the dynamic video; that is, the present application selects a scene image including at least one mosaic frame in the current scene on a display page, and when a selection operation for a target mosaic frame in the scene image is detected, a dynamic video corresponding to a static target image can be projected to the mosaic frame position of the target mosaic frame, so that the target object in the target image is displayed in a dynamic form, thereby realizing dynamic conversion of a single static image, thereby performing AR enhancement on the target image and improving the user experience at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 An optional schematic diagram of a network architecture for implementing a dynamic video display method provided in an embodiment of the present application;

[0019] Figure 2 An optional flowchart of a dynamic video display method provided in an embodiment of the present application;

[0020] Figure 3 A schematic diagram of an optional image selection page provided in an embodiment of the present application;

[0021] Figure 4A schematic diagram of an optional target mosaic frame determination provided in an embodiment of the present application;

[0022] Figure 5 A schematic diagram of an optional target mosaic frame determination provided in an embodiment of the present application;

[0023] Figure 6 An optional flowchart of a dynamic video display method provided in an embodiment of the present application;

[0024] Figure 7 An optional flowchart of a dynamic video display method provided in an embodiment of the present application;

[0025] Figure 8 An optional flowchart of a dynamic video display method provided in an embodiment of the present application;

[0026] Fig. 9 A schematic diagram of an optional method for determining pixel motion information provided in an embodiment of the present application;

[0027] Fig. 10A A schematic diagram of an optional fusion of multiple derivative images provided in an embodiment of the present application;

[0028] Fig. 10B A schematic diagram of an optional fusion of multiple derivative images provided in an embodiment of the present application;

[0029] Fig.11 An optional flowchart of a dynamic video display method provided in an embodiment of the present application;

[0030] Fig.12 An optional flowchart of a dynamic video display method provided in an embodiment of the present application;

[0031] Fig.13 An optional structural diagram of a dynamic video display device provided in an embodiment of the present application;

[0032] Fig.14 An optional structural schematic diagram of a projection device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0034] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0035] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0036] See also Figure 1 , Figure 1 The network architecture diagram for implementing the dynamic video display method provided in the present application at least includes a projection device 100, a terminal device 200 and a network 300; wherein the projection device 100 and the terminal device 200 are connected via the network 300. The network architecture may also include a projection device 100, a network 300 and a server 400; wherein the projection device 100 and the server 400 are connected via the network 300. Of course, the network architecture may also include a projection device 100, a terminal device 200, a network 300 and a server 400; wherein the projection device 100 and the terminal device 200 are connected to the server 400 via the network 300 respectively.

[0037] Here, the projection device 100 is a device that can project images or videos onto a projection surface such as a screen, a wall or a surface where other objects are located; the projection device 100 includes but is not limited to a projector, a projection controller, a micro-projection controller, a projection control machine, an intelligent camera, an intelligent projection control machine, etc.

[0038] Here, the terminal device 200 may be referred to as a remote device connected to the projection device 100, and the terminal device 200 includes but is not limited to a smart phone, a tablet computer, a personal digital assistant (PDA), a camera, a wearable device, a smart TV, a smart camera, a smart projector, a laptop computer, and a desktop computer, etc.; illustratively, the display interface of the terminal device 200 displays a device identifier 21 of the projection device 100 and a control 22 for controlling the projection device 100; wherein the device identifier 21 may be Figure 1The icon of the projection device 100 shown may of course also be the device number of the projection device 100 .

[0039] Here, the network 300 includes but is not limited to a local area network, a metropolitan area network, and a wide area network.

[0040] Here, the server 400 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc. This application does not make any specific restrictions on this.

[0041] See also Figure 2 , Figure 2 is a schematic diagram of an implementation flow of a dynamic video display method provided in an embodiment of the present application. The dynamic video display method can be applied to a projection device. The dynamic video display method can also be applied to Figure 1 The network architecture shown in the figure, the dynamic video display method comprises the following steps:

[0042] Step 201: Display an image selection page.

[0043] Step 202: determine a scene image through an image selection page; wherein the scene image includes at least one mosaic frame in the current scene.

[0044] In the embodiment of the present application, the image selection page is used to determine a scene image containing multiple mosaic frames in the current scene.

[0045] In the embodiment of the present application, the mosaic frame can be a structure composed of four frames, and the mosaic frame is used to contain embedded objects, and the embedded objects include but are not limited to images and mosaics. In a practical application, the mosaic frame can be a mosaic frame.

[0046] In the embodiment of the present application, the mosaic frame included in the scene image may contain an embedded object such as an image, or the mosaic frame included in the scene image may not contain an embedded object such as an image.

[0047] In the embodiments of the present application, the scene image can be acquired based on the image acquisition module of the electronic device, and the scene image can also be an image recommended by the electronic device based on the current scene. Of course, the scene image can also be an image of the current scene stored in the electronic device selected by the user. The present application does not impose any restrictions on the source of the scene image. It should be noted that if the dynamic video display method is applied to a projection device, the electronic device and the electronic device described below are projection devices; if the dynamic video display method is applied to a network architecture, the electronic device and the electronic device described below can be a projection device or a terminal device, and the present application does not impose any specific restrictions on this.

[0048] In some embodiments, the process of determining the scene image through the image selection page in step 202 can be implemented in any of the following ways:

[0049] Method 1: When the image acquisition module captures the current scene to obtain the first image, when an upload operation for uploading the image is detected in the image selection page, the first image is displayed; when a selection operation for the first image is detected, the first image is determined to be a scene image.

[0050] In the embodiment of the present application, the image acquisition module can be a camera, and the electronic device can set the configuration parameters of the built-in image acquisition module, that is, the shooting configuration parameters such as the shooting angle range and shooting distance range of the image acquisition module can be set according to actual needs. The electronic device can also analyze and process the images collected by the image acquisition module.

[0051] In the embodiment of the present application, the first image is an image obtained by an image acquisition module of the electronic device capturing the current scene at that moment.

[0052] In the embodiment of the present application, the input method of the upload operation includes but is not limited to voice input operation, click input operation, press input operation and slide input operation. It should be noted that the input methods of the operations involved in the following description of the present application, such as selection operation, touch operation, slide operation and shooting operation, include but are not limited to voice input operation, click input operation, press input operation and slide input operation.

[0053] In one possible implementation, referring to Figure 3As shown, the image selection page displays an upload image button 301. When the image acquisition module captures the current scene at the first moment to obtain the first image, the electronic device detects the user's upload operation for uploading the image on the image selection page at the second moment, that is, after detecting the touch operation of the "upload image" button 301, the first image captured by the image acquisition module at the first moment is displayed in the first area, and the proportion of the first area in the display area of ​​the electronic device is less than the proportion threshold, such as the proportion threshold is 0.2, 0.3, etc. Further, when the electronic device detects the selection operation for the first image, it determines that the selected first image is a scene image; wherein the time interval between the first moment and the second moment is within the preset interval, and the first moment is the moment closest to the second moment, so as to ensure that when the electronic device selects an image for uploading at the second moment, the image at the first moment is displayed.

[0054] Method 2: when an upload operation for an uploaded image is detected in the image selection page, one or more local images contained in the local image library are displayed; when a selection operation for a local image is detected, the selected local image is determined to be a scene image.

[0055] In one possible implementation, continue to refer to Figure 3 , the image selection page displays an upload image button 301. After the electronic device detects the user's upload operation on the image selection page, that is, detects the touch operation of the "upload image" button 301, it opens the local image library and displays one or more local images. Further, when the electronic device detects the selection operation on the local image, it determines that the selected local image is a scene image.

[0056] Method three, when a shooting operation for shooting an image in the image selection page is detected, the shooting page of the image acquisition module is displayed to shoot the current scene, and the image shot by the image acquisition module is determined to be a scene image.

[0057] In one possible implementation, continue to refer to Figure 3 The image selection page displays a shooting button 302 for shooting an image. After the electronic device detects the user's shooting operation for shooting an image on the image selection page, that is, detects the touch operation of the "shoot" button 302, it displays the shooting page of the image acquisition module of the projection device or the terminal device, thereby using the image acquisition module of the projection device or the terminal device to shoot the current scene, and determines that the image captured by the image acquisition module is the scene image.

[0058] Step 203: when a selection operation for a target mosaic frame in the scene image is detected, a dynamic video corresponding to the obtained target image is projected to the mosaic frame position of the target mosaic frame; wherein the target object contained in the target image is in motion in the dynamic video.

[0059] In the embodiment of the present application, the target mosaic frame is a mosaic frame selected from one or more mosaic frames included in the scene image.

[0060] In the embodiment of the present application, the target image is a static image containing at least one target object, and the target object can be dynamically processed. It should be noted that the target object can be an object that is static in the target image but can move in the real world, such as flowing water, waterfalls, smoke, sea water, hair, windmills, clouds, rain, snowflakes, etc.

[0061] Exemplarily, the target image can be a target person image, a target animal image, a landscape image, etc., and the target objects that can be dynamically processed can be the person's hair in the target person image, the sea water in the image background, etc., the animal hair in the target animal image, etc., and the branches, clouds, smoke and flowing water in the landscape image, etc.

[0062] In the embodiments of the present application, the target image may be an image embedded in a target mosaic frame, or an image instantly captured by a user through a camera of an electronic device such as a mobile phone, tablet computer, or projector, by photographing a landscape containing a target object, or an image pre-stored in an electronic device; of course, the target image may also be an image acquired by a terminal device from a cloud server or other terminal device through a network, and the present application does not impose any specific restrictions on this.

[0063] In the embodiment of the present application, the dynamic video includes multiple derivative images related to the target image, and the objects included in each image are consistent, but the positions of the target objects included in each image are different. Therefore, the position of the target object in the target image in the obtained dynamic video will change, and the positions of other objects except the target object in the dynamic video will not change, that is, the target object is in motion in the obtained dynamic video, and other objects are in a static state in the dynamic video. For example, taking the target image as a landscape image containing smoke and the target object as smoke as an example, the position of the smoke in the generated dynamic video will change, and when the dynamic video is displayed, the effect of smoke drifting in the wind will be presented.

[0064] In one feasible scenario, when a user has a need to generate a dynamic video corresponding to a target image, the target image can be directly input into a pre-trained model for generating dynamic video, so that the model for generating dynamic video can obtain the dynamic video of the target image. In actual application, the pre-trained model for generating dynamic video can be deployed in a projection device, or in a server, and of course, can also be deployed in various computer program products. The specific storage form of the pre-trained model for generating dynamic video is not limited in the embodiments of the present application.

[0065] In the embodiment of the present application, the mosaic frame position is the position of the mosaic frame in the projection optical machine coordinate system in the current scene. The mosaic frame position can be the center position of the target mosaic frame, the mosaic frame position can also be the corner point (or vertex) position of the target mosaic frame. Of course, the mosaic frame position can also be other positions of the target mosaic frame, and the present application does not make any specific restrictions on this.

[0066] In the embodiment of the present application, the process of determining the mosaic frame position of the target mosaic frame can be obtained by the following two methods: Method 1, after selecting the target mosaic frame from one or more mosaic frames included in the scene image, determine the mosaic frame position of the target mosaic frame in the current scene in the projection optical machine coordinate system; Method 2, determine the mosaic frame position of one or more mosaic frames included in the scene image in the current scene in the projection optical machine coordinate system, and when detecting the selection operation for the mosaic frame in the scene image, determine the selected mosaic frame as the target mosaic frame, and obtain the actual mosaic frame position corresponding to the target mosaic frame in the projection optical machine coordinate system. In this way, the mosaic frame position of the target mosaic frame can be determined in different ways.

[0067] In the case where the dynamic video display method is applied to a projection device, when a user wants to generate and project a dynamic video in real time, the projection device displays an image selection page, and determines a scene image containing at least one mosaic frame in the current scene through the image selection page. Then, the projection device detects the selection operation for the target mosaic frame in the scene image, and determines the actual mosaic frame position of the target mosaic frame in the current scene. Further, the projection device obtains the target image, and dynamically processes the target object in the target image to obtain a dynamic video of the target object in the target image in a moving state. Finally, the projection device projects the dynamic video corresponding to the target image to the mosaic frame position of the target mosaic frame, thereby performing augmented reality enhancement on the target image and improving the user experience.

[0068] In the case where the dynamic video display method is applied to a network architecture composed of a projection device, a terminal device and a network, when the user wants to generate and project a dynamic video in real time, the terminal device displays an image selection page, and determines the scene image containing at least one mosaic frame in the current scene through the image selection page, and sends the scene image to the projection device through the network. Then, the projection device detects the selection operation for the target mosaic frame in the scene image, and determines the actual mosaic frame position of the target mosaic frame in the current scene. Furthermore, the projection device obtains the target image, and dynamically processes the target object in the target image to obtain a dynamic video of the target object in the target image in motion. Finally, the projection device projects the dynamic video corresponding to the target image to the mosaic frame position of the target mosaic frame, so that the target object in the target image is displayed in a dynamic form, thereby realizing the dynamic conversion of a single static image, thereby performing AR enhancement on the target image and improving the user experience.

[0069] In the case where the dynamic video display method is applied to a network architecture composed of a projection device, a server and a network, when the user wants to generate and project a dynamic video in real time, the projection device displays an image selection page, and determines the scene image containing at least one mosaic frame in the current scene through the image selection page. Then, the projection device detects the selection operation for the target mosaic frame in the scene image, and determines the actual mosaic frame position of the target mosaic frame in the current scene. Further, the projection device obtains the target image and uploads the target image to the server through the network. The server dynamically processes the target object in the target image, obtains a dynamic video of the target object in the target image in motion, and sends the dynamic video to the projection device. Finally, the projection device projects the dynamic video corresponding to the target image to the mosaic frame position of the target mosaic frame, so that the target object in the target image is displayed in a dynamic form, thereby realizing the dynamic conversion of a single static image, thereby AR enhancing the target image and improving the user experience.

[0070] In the case where the dynamic video display method is applied to a network architecture consisting of a projection device, a terminal device, a server and a network, when the user wants to generate and project a dynamic video in real time, the terminal device displays an image selection page, and determines the scene image containing at least one mosaic frame in the current scene through the image selection page, and sends the scene image to the projection device through the network. Then, the projection device detects the selection operation for the target mosaic frame in the scene image, and determines the actual mosaic frame position of the target mosaic frame in the current scene. Further, the projection device or the terminal device obtains the target image and uploads the target image to the server through the network. The server dynamically processes the target object in the target image, obtains a dynamic video of the target object in the target image in motion, and sends the dynamic video to the projection device. Finally, the projection device projects the dynamic video corresponding to the target image to the mosaic frame position of the target mosaic frame, so that the target object in the target image is displayed in a dynamic form, thereby realizing the dynamic conversion of a single static image, thereby AR enhancing the target image and improving the user experience.

[0071] In some embodiments, after the scene image is determined through the image selection page in step 202, the following process may also be performed:

[0072] The boundaries of all mosaic frames contained in the scene image are highlighted, and when a selection operation for a target mosaic frame in the scene image is detected, the highlighting of the remaining mosaic frames except the target mosaic frame is canceled; or, when a selection operation for a target mosaic frame in the scene image is detected, the boundary of the selected target mosaic frame in the scene image is highlighted.

[0073] In a feasible scenario, refer to Figure 4 As shown, the image selection page displays a scene image, and the electronic device highlights the boundaries 402 of all the mosaic frames contained in the scene image 401; further, when the user touches the target mosaic frame in the scene image, the electronic device determines the selection of the target mosaic frame. At this time, the electronic device cancels the highlighting of the boundaries of all the remaining mosaic frames except the target mosaic frame, and only highlights the boundary of the target mosaic frame 402, thereby reminding the user of the selected mosaic frame by highlighting.

[0074] In another possible scenario, refer to Figure 5 As shown, the image selection page displays a scene image. When a selection operation is detected for a target mosaic frame 402 in the scene image 401, the boundary 402 of the selected target mosaic frame in the scene image is highlighted, so that only the boundary of the target mosaic frame 402 is highlighted, thereby reminding the user of the selected mosaic frame by highlighting.

[0075] In some embodiments, in order to make the target object move in the direction and / or speed desired by the user, when a selection operation on the target mosaic frame is detected, the following steps may also be performed:

[0076] When a selection operation on the target mosaic frame is detected, the optical flow direction prompt information is displayed, and when a sliding operation of the user is detected, the sliding direction input by the user is determined as the optical flow direction;

[0077] When a selection operation for a target mosaic frame is detected, the speed of the flow mode corresponding to the current time is determined as the optical flow speed, or the flow mode page is displayed, and when a selection operation for a flow mode in the flow mode page is detected, the speed of the selected flow mode is determined as the optical flow speed.

[0078] In the embodiment of the present application, the optical flow direction prompt information is used to prompt the user to select the movement direction of the target object.

[0079] In the embodiment of the present application, the optical flow direction is the direction in which the target object moves in the dynamic video.

[0080] In the embodiment of the present application, the optical flow speed is the distance that the target object needs to move per unit time in the dynamic video.

[0081] In the embodiment of the present application, the flow rate pattern corresponding to the current time can be understood as the flow rate pattern corresponding to the season to which the current time belongs, or the flow rate pattern corresponding to the month to which the current time belongs, or the flow rate pattern corresponding to the time period to which the current time belongs on the day. Exemplarily, the format of the current time can be year, month, day, hour, minute, and second. Of course, the format of the current time can also be other formats, and the present application does not make specific restrictions on this.

[0082] In the embodiment of the present application, the flow rate mode is a preset optical flow rate for different seasons, different months or different time periods in the same day. Exemplarily, the flow rate mode includes but is not limited to a high flow rate mode, a medium-high flow rate mode, a medium flow rate mode, a medium-low flow rate mode and a low flow rate mode, and different flow rate modes correspond to different optical flow rates.

[0083] In the embodiment of the present application, the flow rate mode page displays a variety of different flow rate modes for users to choose from.

[0084] In an achievable scenario, when a selection operation for a target mosaic frame is detected, the optical flow direction prompt information is displayed; further, the user's sliding operation is detected, and the sliding direction input by the user is determined as the optical flow direction. And / or, when a selection operation for a target mosaic frame is detected, the current time is obtained, and the flow rate mode corresponding to the current time is determined as the optical flow speed, or a flow rate mode page containing multiple flow rate modes is displayed, and a selection operation for a flow rate mode in the flow rate mode page is detected. After the selection operation is completed, the speed of the selected flow rate mode is determined as the optical flow speed. In this way, the optical flow direction and optical flow speed of the target object are determined according to the user's wishes, so that the movement direction and movement speed of the target object in the subsequently generated dynamic video meet the user's requirements.

[0085] In some embodiments, reference Figure 6 As shown, the process of determining the mosaic frame position of the target mosaic frame in step 203 is combined with Figure 6 To explain,

[0086] Step 601: Obtain the first position of each corner point of the target mosaic frame in the camera coordinate system.

[0087] In the embodiment of the present application, the first position may be the coordinates of each corner point of the target mosaic frame in the camera coordinate system, or the first position may be the coordinates that can characterize the position of the target mosaic frame in the camera coordinate system, and the present application does not make any specific limitation on this. For example, the target mosaic frame is generally a regular shape, such as a rectangle or a trapezoid, and the first position may be the coordinates of the four corner points of the target mosaic frame.

[0088] In an achievable scenario, step 601 of obtaining the first position of each corner point of the target mosaic frame in the camera coordinate system can be implemented as follows: pre-designing a mosaic frame segmentation neural network model, wherein the mosaic frame segmentation neural network model includes a feature extraction module, a backbone module and an instance segmentation head module, wherein the feature extraction module is used to extract features of the mosaic frame in the image; the backbone module is used to obtain the bounding box of the mosaic frame based on the extracted features, where the backbone module can use a standard ResNet50 network module, and the backbone module includes a number of convolution layers and a size corresponding to the commonly used size of the mosaic frame, so that it has optimal performance; the instance segmentation head module is used to segment the area within the frame through the segmentation network to obtain an instance segmentation mask, thereby obtaining the position of each corner point of the mosaic frame in the camera coordinate system. However, since there is an error in the position of each corner point obtained by the mosaic frame segmentation neural network model, that is, the position of the corner point output by the mosaic frame segmentation neural network model is different from the position of the real mosaic frame; further, the mosaic frame corner points are fine-tuned using a traditional image algorithm based on a gradient, so that the position result of each corner point of the mosaic frame in the camera coordinate system is finally obtained to be more accurate. Of course, the method of obtaining the first position of each corner point of the target mosaic frame in the camera coordinate system can also be obtained through other existing network models in the relevant technology, and this application does not make specific limitations on this.

[0089] Step 602: Based on the first position, the position of the target mosaic frame in the projection optical machine coordinate system is obtained by using the camera homography matrix between the image acquisition module and the projection optical machine in the projection device.

[0090] In the embodiment of the present application, the camera homography matrix is ​​used to describe the projection relationship between the plane where the image captured by the image acquisition module is located and the plane where the projection screen of the projection optical machine is located. It is an important concept in computer vision. Through the camera homography matrix, the points on the plane where the image captured by the image acquisition module is located can be mapped to the plane where the projection screen of the projection optical machine is located.

[0091] In an embodiment of the present application, after determining the target mosaic frame in the scene image, the projection device obtains the first position of each corner point of the target mosaic frame in the camera coordinate system; then, the camera homography matrix between the image acquisition module and the projection optical machine in the projection device is obtained; further, based on the first position, the mosaic frame position of the target mosaic frame in the projection optical machine coordinate system is obtained using the camera homography matrix; then, the homography matrix from the target image to the projection optical machine is calculated, so as to map the dynamic video corresponding to the target image to the projection optical machine and project it to the mosaic frame position of the target mosaic frame. In this way, the dynamic video corresponding to the static image is projected to the position of the target mosaic frame through the projection device, so that the target object in the target image is displayed in a dynamic form, thereby realizing the dynamic conversion of a single static image, thereby performing AR enhancement on the target image and improving the user experience.

[0092] In some embodiments, the process of generating the dynamic video corresponding to the target image in step 203 may be performed by the projection device side, or the process of generating the dynamic video corresponding to the target image may also be performed by the server side connected to the projection device, and this application does not impose any specific restrictions on this. Here, the projector device processes the target image to obtain the corresponding dynamic video as an example, wherein, with reference to Figure 7 As shown, the generation method of the dynamic video corresponding to the target image in step 203 is combined with Figure 7 To explain,

[0093] Step 701: Use the trained target model to process the target image to generate derivative images of the target object at various times.

[0094] In an embodiment of the present application, the target model is used to process the object in the image to obtain derivative images of the target object at different positions at different times.

[0095] In the embodiment of the present application, in the derived images at each moment, the target object included in each derived image is consistent, but the position of the target object included in each derived image is different.

[0096] In the embodiment of the present application, step 701 uses the trained target model to process the target image, and the process of generating the derivative image of the target object at each time is combined with Figure 8 To explain,

[0097] Step 711: using an optical flow estimation operation model, based on the target image, obtain target pixel motion information of the target object in the target image at each moment, wherein the target model includes an optical flow estimation operation model.

[0098] In the embodiment of the present application, the pixel motion information is motion information obtained by performing optical flow processing on the target object in the target image through an optical flow estimation operation model, and the pixel motion information is used to characterize the movement direction and movement speed of the pixel.

[0099] In the embodiment of the present application, the optical flow estimation operation model is a trained model, and the optical flow estimation operation model is used to perform optical flow processing on the target object in the target image.

[0100] In some embodiments, reference Fig. 9 As shown, the optical flow estimation operation model includes an encoding module (encoder) 902 and a decoding module (decoder) 903, wherein the encoding module is used to downsample the image, and the encoding module is composed of multiple convolutional layers. After the target image 901 is input into the encoding module 902 in the optical flow estimation operation model, the size of the feature map output by each convolutional layer in the encoding module decreases with the number of convolutional layers; in this way, the size of the feature map generated by the current convolutional layer in the encoding module is smaller than the size of the feature map generated by the previous convolutional layer, reducing the computing power while ensuring a high analysis accuracy, thereby extracting the key features of the object in the image. Among them, the decoding module is used to upsample the feature map output by the encoding module, and the decoding module is composed of multiple convolutional layers. After the feature map output by the encoding module 902 is input into the decoding module 903 in the optical flow estimation operation model, the pixel motion information 904 of the target object in the target image 901 is obtained. Among them, the size of the feature map output by each convolutional layer in the decoding module increases with the number of convolutional layers; in this way, the size of the feature map generated by the current convolutional layer in the encoding module is larger than the size of the feature map generated by the previous convolutional layer, which reduces the computing power while ensuring a higher analysis accuracy, thereby outputting the pixel motion information of the object in the image.

[0101] In some embodiments, step 711 uses an optical flow estimation operation model to obtain target pixel motion information of a target object in the target image at each moment based on the target image, which can be achieved by:

[0102] The first step is to use the trained optical flow estimation operation model to perform feature processing on the target image to obtain the initial pixel motion information of the target object in the target image at the first moment.

[0103] The second step is to perform pixel fusion on the first optical flow velocity and the first optical flow direction and the initial pixel motion information of the target object at the first moment to obtain the target pixel motion information of the target object at the first moment.

[0104] The third step is to perform Euler interpolation on the target pixel motion information of the target object at the first moment to obtain the target pixel motion information of the target object at each of the M+N moments, thereby obtaining M+N+1 target pixel motion information.

[0105] Among them, the first moment is the middle moment of M+N+1 moments, M is multiple moments before the first moment, N is multiple moments after the first moment, the difference between M and N is less than or equal to the threshold, M and N are positive integers, M is greater than N, and exemplarily, the threshold can be 1.

[0106] In the embodiment of the present application, the first optical flow speed may be an optical flow speed input by a user, or may be an optical flow speed pre-set for different objects, and the present application does not impose any specific limitation on this.

[0107] In the embodiment of the present application, the first optical flow direction may be an optical flow direction input by a user, or may be an optical flow direction pre-set for different objects. For example, the optical flow direction of smoke may be left-right, and the optical flow direction of waterfalls and rain may be up-down. The present application does not make any specific limitation on this.

[0108] In the embodiment of the present application, the trained optical flow estimation operation model is used to perform feature processing on the target image to obtain the initial pixel motion information of the target object in the target image at the first moment; then, the first optical flow velocity and the first optical flow direction are pixel-fused with the initial pixel motion information of the target object at the first moment to obtain the target pixel motion information of the target object at the first moment; further, taking the first moment corresponding to the target image as the intermediate moment and the target pixel motion information of the target object in the target image as the benchmark, Euler interpolation is performed on the target pixel motion information of the target object at the first moment to obtain the target pixel motion information corresponding to each of the M moments before the first moment and the target pixel motion information of each of the N moments after the first moment, thereby obtaining the target pixel motion information corresponding to M+N+1 moments. In this way, the optical flow direction and optical flow velocity given to the pixel motion information of the target object can make the target object in the generated dynamic video move according to the set optical flow direction and optical flow velocity.

[0109] Step 712: Generate a derivative image of the target object at each moment according to the target pixel motion information and the target image of the target object at each moment.

[0110] In the embodiment of the present application, the target pixel motion information corresponding to the target object at M+N+1 moments is fused with the target image respectively to obtain the derived images of the target object at each moment.

[0111] As can be seen from the above, in the embodiment of the present application, by processing the target image based on the trained optical flow estimation operation model, it is possible to accurately obtain the target pixel motion information that can reflect the target object in the target image at different time sequences. Then, the target pixel motion information of the target object at different time sequences and the target image are fused to generate derivative images of the target object at different positions at different time sequences, so as to obtain a dynamic video based on the derivative images.

[0112] Step 702: Fuse and splice the derivative images of the target object at each time in pairs to generate a dynamic video; wherein the number of intervals between the two derivative images that are fused in pairs each time is the same.

[0113] In the embodiment of the present application, the number of intervals between two derivative images to be fused may be M. The derivative images corresponding to two moments separated by M moments are fused to obtain generated images corresponding to multiple moments, and the generated images are spliced ​​in the time sequence of the corresponding moments to obtain a dynamic video.

[0114] In some embodiments, if the target pixel motion information corresponding to each of the M moments before the first moment and the target pixel motion information corresponding to each of the N moments after the first moment are obtained, the target pixel motion information corresponding to the M+N+1 moments is obtained. Further, the target pixel motion information corresponding to the target object at the M+N+1 moments is fused with the target image respectively to obtain the derivative images of the target object at each moment. Then, according to the arrangement order of the moments corresponding to the derivative images, the i-th derivative image is fused with the i+M-th derivative image to obtain the generated image at the i+M-th moment, thereby obtaining the N+1 generated images corresponding to the N+1 moments; the N+1 generated images are spliced ​​in chronological order to generate a dynamic video, where i is an integer greater than or equal to 1 and less than or equal to N+1.

[0115] In one possible scenario, refer to Fig. 10A , Fig. 10ASchematic diagram of fusing derivative images at different times. Here, taking the first time as time t0, and taking M and N as 3 as an example, the projection device adopts an optical flow estimation operation model, obtains the target pixel motion information of the target object in the target image at time t0 based on the target image, performs Euler interpolation on the target pixel motion information of the target object at t0, obtains the target pixel motion information corresponding to time t-3, time t-2 and time t-1 before time t0, and the target pixel motion information corresponding to time t1, time t2 and time t3 after time t0, thereby obtaining the target pixel motion information corresponding to 7 times. Furthermore, according to the target pixel motion information and the target image of the target object at each moment, a derivative image of the target object at each moment is generated; then, the derivative images at each interval of M=3 moments are fused, that is, the derivative image at moment t-3 is fused with the derivative image at moment t0 in pairs, the derivative image at moment t-2 is fused with the derivative image at moment t1 in pairs, the derivative image at moment t-1 is fused with the derivative image at moment t2 in pairs, and the derivative image at moment t0 is fused with the derivative image at moment t3 in pairs, thereby obtaining the fused images corresponding to moments t0, t1, t2 and t3, that is, 4 (N+1) generated images, and the 4 generated images are spliced ​​in chronological order to generate a dynamic video.

[0116] In another possible scenario, refer to Fig. 10B , Fig. 10B Schematic diagram of fusing derivative images at different times. Here, taking the first time as time t0, M as 3, and N as 2 as an example, the projection device adopts an optical flow estimation operation model, obtains the target pixel motion information of the target object in the target image at time t0 based on the target image, performs Euler interpolation on the target pixel motion information of the target object at t0, obtains the target pixel motion information corresponding to time t-3, time t-2, and time t-1 before time t0, and the target pixel motion information corresponding to time t1 and time t2 after time t0, thereby obtaining the target pixel motion information corresponding to 6 times. Furthermore, according to the target pixel motion information and the target image of the target object at each moment, a derivative image of the target object at each moment is generated; then, the derivative images at each interval of M=3 moments are fused, that is, the derivative image at moment t-3 is fused with the derivative image at moment t0 in pairs, the derivative image at moment t-2 is fused with the derivative image at moment t1 in pairs, and the derivative image at moment t-1 is fused with the derivative image at moment t2 in pairs, so as to obtain the fused images corresponding to moments t0, t1 and t2, that is, 3 (N+1) generated images, and the 3 generated images are spliced ​​in chronological order to generate a dynamic video.

[0117] As can be seen from the above, in the embodiment of the present application, the target image is processed based on the trained target model, and the derivative images that can reflect the target object in the target image at different time sequences can be accurately extracted. Then, using the derivative images at different time sequences, two derivative images with the same number of interval moments are fused in pairs, and the fused images are spliced ​​according to the time sequence, so as to automatically generate dynamic videos, which greatly improves the efficiency of dynamic video generation. In addition, the trained target model has reliable output accuracy after multiple rounds of training, which improves the accuracy of the obtained dynamic video and further improves the display effect of the dynamic video.

[0118] In some embodiments, the training process of the optical flow estimation computing model in step 711 is combined with Fig.11 To further explain,

[0119] Step 1101: Obtain a sample dynamic video, wherein the sample dynamic video includes a sample image sequence, and the sample image sequence includes a sample object.

[0120] In an embodiment of the present application, the sample dynamic video may be a short video including a sample object. The sample dynamic video includes an ordered plurality of frames of sample images. The positions of the sample objects included in each frame of the sample image are different, while the positions of other objects except the sample objects are the same.

[0121] In some embodiments, the sample dynamic video includes ordered multiple frames of images, and some image frames are extracted from the image sequence as a sample image sequence, wherein the sample image sequence is a continuous p-frame image in the image sequence, wherein p is an integer greater than 1. For example, the sample dynamic video includes 200 frames of images, and the images of the 50th to 60th frames in the video can be extracted. Since the video has continuity, extracting multiple continuous frames of images in the video can help the obtained pixel motion information to more accurately reflect the movement of the pixels after obtaining the pixel motion information.

[0122] In the embodiment of the present application, the sample dynamic video can be video information captured by any electronic device with an image camera function, and the sample dynamic video can also be video information stored in an image database. This application does not impose any specific restrictions on this.

[0123] In practical applications, the electronic device recognizes each training image in the sample dynamic video based on an image recognition algorithm to determine the target object in each sample image.

[0124] Step 1102: Use the optical flow estimation operation model to be trained to perform feature processing on the starting frame sample image in the sample image sequence to obtain training pixel motion information of the sample object in the starting frame sample image at each moment.

[0125] In the embodiment of the present application, the starting frame sample image is the first frame image in the sample image sequence.

[0126] In the embodiment of the present application, the training pixel motion information is motion information obtained by performing optical flow processing on the sample object in the starting frame sample image by the optical flow estimation operation model to be trained, and the pixel motion information is used to characterize the movement direction and movement speed of the pixel.

[0127] In an embodiment of the present application, an optical flow estimation operation model to be trained is used to perform feature processing on the starting frame sample image to obtain initial training pixel motion information of the sample object in the starting frame sample image at the second moment, and then, the second optical flow velocity and the second optical flow direction are pixel-fused with the initial pixel motion information of the sample object at the second moment to obtain training pixel motion information of the sample object at the second moment; further, taking the second moment corresponding to the starting frame sample image as the intermediate moment and the training pixel motion information of the sample object in the starting frame sample image as the reference, Euler interpolation is performed on the training pixel motion information of the sample object at the second moment to obtain training pixel motion information corresponding to each of the M moments before the second moment and training pixel motion information corresponding to each of the N moments after the second moment, thereby obtaining training pixel motion information corresponding to M+N+1 moments.

[0128] Step 1103: Generate a derived sample image of the sample object at each moment according to the training pixel motion information of the sample object at each moment and the starting frame sample image.

[0129] In the embodiment of the present application, the training pixel motion information corresponding to the sample object at M+N+1 moments is respectively fused with the starting frame sample image to obtain the derived sample images of the sample object at each moment.

[0130] Step 1104: Fuse the derived sample images of the sample object at each time in pairs to obtain sample generated images corresponding to multiple time points; wherein the number of interval time points between the two derived sample images fused in pairs each time is the same.

[0131] In the embodiment of the present application, the number of intervals between two derived sample images to be fused may be M. The derived sample images corresponding to two moments separated by M moments are fused to obtain sample generated images corresponding to multiple moments.

[0132] In some embodiments, if the training pixel motion information corresponding to each of the M moments before the first moment and the training pixel motion information of each of the N moments after the first moment of the sample object are obtained, the training pixel motion information corresponding to the M+N+1 moments is obtained. Further, the training pixel motion information corresponding to the sample object at the M+N+1 moments is fused with the starting frame sample image respectively to obtain the derived sample images of the sample object at each moment. Then, according to the arrangement order of the moments corresponding to the derived sample images, the i-th derived sample image is fused with the i+M-th derived sample image to obtain the sample generated image at the i+M-th moment, thereby obtaining the N+1 sample generated images corresponding to the N+1 moments, where i is an integer greater than or equal to 1 and less than or equal to N+1.

[0133] Step 1105: Calculate a target loss value between the sample generated image and the sample image based on the sample generated image at each of the multiple moments and the sample image at the corresponding moment.

[0134] In some embodiments, step 1105 calculates the target loss value combination between the sample generated image and the sample image based on the sample generated image at each of the multiple moments and the sample image at the corresponding moment. Fig.12 To further explain,

[0135] Step 1201: determine the similarity between the training pixel motion information of the sample object included in the sample generated image at each moment in multiple moments and the sample pixel motion information of the sample object in the sample image at the corresponding moment, as the first loss value between the training pixel motion information and the sample pixel motion information.

[0136] In an embodiment of the present application, the similarity is used to characterize the degree of matching between the training pixel motion information of the sample object included in the sample generated image and the sample pixel motion information of the sample object in the sample image at the corresponding moment. If the similarity between the training pixel motion information and the sample pixel motion information at the corresponding moment is greater, it indicates that the training pixel motion information is more matched with the sample pixel motion information at the corresponding moment, and the loss value is smaller; on the contrary, if the similarity between the training pixel motion information and the sample pixel motion information at the corresponding moment is smaller, it indicates that the training pixel motion information is more mismatched with the sample pixel motion information at the corresponding moment, and the loss value is greater.

[0137] In an embodiment of the present application, the similarity between the training pixel motion information and the sample pixel motion information can be calculated by Euclidean distance, and the similarity between the training pixel motion information and the sample pixel motion information can also be calculated by a cosine similarity formula, and the present application does not make any specific restrictions on this.

[0138] Step 1202: Determine the similarity between the sample generated image at each moment and the sample image at the corresponding moment as a second loss value between the sample generated image and the sample image.

[0139] In the embodiment of the present application, the similarity is used to characterize the degree of matching between the sample-generated image and the sample image at the corresponding moment. If the similarity between the sample-generated image and the sample image at the corresponding moment is greater, it indicates that the sample-generated image and the sample image at the corresponding moment are more matched, and the loss value is smaller; on the contrary, if the similarity between the sample-generated image and the sample image at the corresponding moment is smaller, it indicates that the sample-generated image and the sample image at the corresponding moment are more mismatched, and the loss value is greater.

[0140] In an embodiment of the present application, the similarity between the sample generated image and the sample image at the corresponding moment can be calculated by Euclidean distance, and the similarity between the sample generated image and the sample image at the corresponding moment can also be calculated by the cosine similarity formula. This application does not make any specific restrictions on this.

[0141] Step 1203: Determine a target loss value based on the first loss value and the second loss value.

[0142] In the embodiment of the present application, the target loss value can be understood as a loss value obtained by measuring the proportion of the two parameters, the first loss value and the second loss value, in the process of training the optical flow estimation operation model.

[0143] In a first achievable embodiment, the first loss value and the second loss value are added to determine the target loss value. For example, the target loss value can be expressed by the following formula (1):

[0144] L(P,P ’ ,I,I ’ )=mse(P,P ’ )+mse(I,I ’ ) (1)

[0145] Among them, L(P,P ’ ,I,I ’ ) is the target loss value, P,P ’ They are the training pixel motion information and the sample pixel motion information, mse(P,P ’ ) is the first loss value between the training pixel motion information and the corresponding sample pixel motion information, I,I ’ They are sample generated image and sample image, mse(I,I ’ ) is the second loss value between the sample generated image and the sample image.

[0146] In a second achievable embodiment, the first loss value and the second loss value are multiplied to determine the target loss value. For example, the target loss value can be expressed by the following formula (2):

[0147] L(P,P ’ ,I,I ’ )=mse(P,P ’ )×mse(I,I ’ ) (3)

[0148] Among them, L(P,P ’ ,I,I ’ ) is the target loss value, P,P ’ They are the training pixel motion information and the sample pixel motion information, mse(P,P ’ ) is the first loss value between the training pixel motion information and the corresponding sample pixel motion information, I,I ’ They are sample generated image and sample image, mse(I,I ’ ) is the second loss value between the sample generated image and the sample image.

[0149] In a third achievable embodiment, a first weight is first set for the first loss value, and a second weight is set for the second loss value; then, the first loss value and the first weight are multiplied to obtain a first target loss value, and the terminal multiplies the second loss value and the second weight to obtain a second target loss value; finally, the first target loss value and the second target loss value are added to determine the target loss value. Exemplarily, the target loss value can be expressed by the following formula (3):

[0150] L(P,P ’ ,I,I ’ )=w1×mse(P,P ’ )+w2×mse(I,I ’ ) (1)

[0151] Among them, L(P,P ’ ,I,I ’ ) is the target loss value, P,P ’ They are the training pixel motion information and the sample pixel motion information, mse(P,P ’ ) is the first loss value between the training pixel motion information and the corresponding sample pixel motion information, I,I ’ They are sample generated image and sample image, mse(I,I ’ ) is the second loss value between the sample generated image and the sample image, w1 is the first weight, and w2 is the second weight.

[0152] In the embodiment of the present application, sample pixel motion information of the sample object in the sample image is obtained, and the similarity between the training pixel motion information of the sample object included in the sample generated image at each moment in multiple moments and the sample pixel motion information of the sample object in the sample image at the corresponding moment is calculated, and the similarity is used as the first loss value between the training pixel motion information and the sample pixel motion information; further, the similarity between the sample generated image at each moment and the sample image at the corresponding moment is calculated, and the similarity is used as the second loss value between the sample generated image and the sample image. According to the first loss value and the second loss value, it is determined as the target loss value.

[0153] Step 1106: Based on the target loss value, adjust the network parameters of the optical flow estimation operation model so that the loss value between the sample generated image and the sample image meets the preset convergence condition.

[0154] In an embodiment of the present application, the loss value between the sample generated image and the sample image satisfies a preset convergence condition including but not limited to the number of training cycles reaching a preset number, and the loss value between the sample generated image and the sample image is less than or equal to a loss threshold.

[0155] In the embodiment of the present application, by designing a loss function, the similarity between the training pixel motion information of the sample object included in the sample generated image at each moment in multiple moments and the sample pixel motion information of the sample object in the sample image at the corresponding moment is determined as the first loss value between the training pixel motion information and the sample pixel motion information, thereby obtaining at least two first loss values. The similarity between the sample generated image at each moment and the sample image at the corresponding moment is determined as the second loss value between the sample generated image and the sample image, thereby obtaining at least two second loss values. Further, according to the at least two first loss values ​​and the at least two second loss values, the back propagation algorithm is used to adjust the network parameters of the encoding module and the decoding module in the optical flow estimation operation model, so that the target loss value between the sample generated image and the sample image is less than or equal to the loss threshold. In this way, the network parameters are continuously adjusted during training so that the pixel motion information of the object output by the model is most matched with its real pixel motion information, and the sample generated image is most matched with its real corresponding sample image.

[0156] The present application provides a dynamic video display device, which can be used to implement Figure 2 , Figures 6 to 8 , Figure 11 to Figure 12 A dynamic video display method provided by the corresponding embodiment refers to Fig.13 As shown, the dynamic video display device 13 includes:

[0157] Display module 1301, used for displaying an image selection page;

[0158] The determination module 1302 is used to determine the scene image through the image selection page; wherein the scene image includes at least one mosaic frame under the current scene;

[0159] The projection module 1303 is used to project the dynamic video corresponding to the obtained target image to the mosaic frame position of the target mosaic frame when a selection operation on the target mosaic frame in the scene image is detected; wherein the target object contained in the target image is in motion in the dynamic video.

[0160] The present application provides a projection device, which can be used to implement Figure 2 , Figures 6 to 8 , Figure 11 to Figure 12 A dynamic video display method provided by the corresponding embodiment refers to Fig.14 As shown, the projection device 100 ( Fig.14 The projection device 100 corresponds to Fig.13 The dynamic video display device 13) includes: a processor 1401, a memory 1402 and a communication bus 1403, wherein:

[0161] The communication bus 1403 is used to realize the communication connection between the processor 1401 and the memory 1402;

[0162] The processor 1401 is used to execute the dynamic video display program stored in the memory 1402 to implement the following steps:

[0163] Display the image selection page;

[0164] Determine a scene image through an image selection page; wherein the scene image includes at least one mosaic frame under the current scene;

[0165] When a selection operation on a target mosaic frame in a scene image is detected, a dynamic video corresponding to the obtained target image is projected to the mosaic frame position of the target mosaic frame; wherein a target object contained in the target image is in motion in the dynamic video.

[0166] An embodiment of the present application provides a computer storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement Figure 2 , Figures 6 to 8 , Figure 11 to Figure 12 The corresponding embodiment provides a dynamic video display method.

[0167] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the same method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0168] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM) and the like; it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0169] It should be understood that the "one embodiment" or "an embodiment" or "an embodiment of the present application" or "the aforementioned embodiment" or "some embodiments" or "some implementation methods" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "an embodiment of the present application" or "the aforementioned embodiment" or "some embodiments" or "some implementation methods" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0170] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0171] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0172] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0173] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0174] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0175] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0176] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.

[0177] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0178] It is worth noting that the drawings in the embodiments of the present application are only for illustrating the schematic positions of various components on the terminal device and do not represent the actual positions in the terminal device. The actual positions of various components or areas may be changed or offset accordingly according to actual conditions (for example, the structure of the terminal device), and the proportions of different parts of the terminal device in the drawings do not represent the actual proportions.

[0179] The above are only implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A dynamic video display method, characterized in that: The method comprises: Display the image selection page; Determine a scene image through the image selection page; wherein the scene image includes at least one mosaic frame under the current scene; When a selection operation on a target mosaic frame in the scene image is detected, a dynamic video corresponding to the obtained target image is projected to the mosaic frame position of the target mosaic frame; wherein a target object contained in the target image is in motion in the dynamic video.

2. The method according to claim 1, characterized in that The determining the scene image through the image selection page includes: In the case where the image acquisition module captures the current scene to obtain a first image, when an upload operation for uploading an image is detected in the image selection page, the first image is displayed; when a selection operation for the first image is detected, the first image is determined to be the scene image; or, When an upload operation for an uploaded image is detected in the image selection page, one or more local images contained in the local image library are displayed; when a selection operation for a local image is detected, the selected local image is determined to be the scene image; or, When a shooting operation for shooting an image in the image selection page is detected, the shooting page of the image acquisition module is displayed to shoot the current scene, and the image shot by the image acquisition module is determined to be the scene image.

3. The method according to claim 1, characterized in that The method further comprises: When a selection operation for the target mosaic frame is detected, the optical flow direction prompt information is displayed, and when a user sliding operation is detected, the sliding direction obtained from the user input is determined as the optical flow direction; and / or, When a selection operation for the target mosaic frame is detected, the speed of the flow rate mode corresponding to the current time is determined as the optical flow speed, or a flow rate mode page is displayed, and when a selection operation for the flow rate mode in the flow rate mode page is detected, the speed of the selected flow rate mode is determined as the optical flow speed.

4. The method according to claim 1, characterized in that: The process of determining the mosaic frame position of the target mosaic frame includes: Obtaining the first position of each corner point of the target mosaic frame in the camera coordinate system; Based on the first position, the position of the target mosaic frame in the projection optical machine coordinate system is obtained by utilizing the camera homography matrix between the image acquisition module and the projection optical machine in the projection device.

5. The method according to any one of claims 1 to 4, characterized in that: The method for generating the dynamic video includes: Processing the target image using the trained target model to generate derivative images of the target object at various times; The derivative images of the target object at each time are fused and spliced ​​in pairs to generate a dynamic video; wherein the number of intervals between the two derivative images fused in pairs each time is the same.

6. The method according to claim 5, characterized in that The step of processing the target image using the trained target model to generate derivative images of the target object at various moments includes: Using an optical flow estimation operation model, based on the target image, obtain target pixel motion information of a target object in the target image at each moment, wherein the target model includes the optical flow estimation operation model; A derivative image of the target object at each moment is generated according to the target pixel motion information of the target object at each moment and the target image.

7. The method according to claim 6, characterized in that The training process of the optical flow estimation motion model includes: Obtaining a sample dynamic video, wherein the sample dynamic video includes a sample image sequence, and the sample image sequence includes a sample object; Using the optical flow estimation operation model to be trained, feature processing is performed on the starting frame sample image in the sample image sequence to obtain training pixel motion information of the sample object in the starting frame sample image at each moment; generating a derived sample image of the sample object at each time instant according to the training pixel motion information of the sample object at each time instant and the starting frame sample image; Fusing the derived sample images of the sample object at each time in pairs to obtain sample generation images corresponding to multiple time points; wherein the number of intervals between the two derived sample images fused in pairs each time is the same; Based on the sample generated image at each of the multiple moments and the sample image at the corresponding moment, calculating a target loss value between the sample generated image and the sample image; Based on the target loss value, the network parameters of the optical flow estimation operation model are adjusted so that the loss value between the sample generated image and the sample image meets a preset convergence condition.

8. The method according to claim 6, characterized in that The calculating, based on the sample generated image at each of the multiple moments and the sample image at the corresponding moment, a target loss value between the sample generated image and the sample image, comprises: Determine a similarity between training pixel motion information of a sample object included in the sample generated image at each of the multiple moments and sample pixel motion information of the sample object in the sample image at the corresponding moment, as a first loss value between the training pixel motion information and the sample pixel motion information; Determine the similarity between the sample generated image at each moment and the sample image at the corresponding moment as a second loss value between the sample generated image and the sample image; The target loss value is determined according to the first loss value and the second loss value.

9. A projection device, characterized in that: The projection device comprises: a processor, a memory and a communication bus; The communication bus is used to realize the communication connection between the processor and the memory; The processor is used to execute the dynamic video display program stored in the memory to implement the dynamic video display method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the dynamic video display method as described in any one of claims 1 to 8.