SFX Video Determination Method, Apparatus, Electronic Device, and Storage Medium
The method and device for determining SFX videos address the limitations of current SFX video creation functions by using uploaded image content as backgrounds and generating SFX video frames, resulting in enhanced user experience and video interestingness.
Patent Information
- Application Number
- JP2024563553
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-29
- Filing Date
- 2023-03-09
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-03-09
AI Technical Summary
Current SFX video creation functions are limited, failing to enhance user experience by not allowing users to change the background screen in videos, thereby reducing the fun and personalization of SFX videos.
A method and device for determining an SFX video that uses content from an uploaded image as the background, determining a target viewpoint image from a 3D image surround scene based on the image's position information, and generating SFX video frames until the user stops shooting, thereby presenting the visual effect of the scene with the target object in the uploaded image.
This solution enhances the interestingness of SFX videos, meets personal user needs by allowing background changes, and improves user experience during SFX video creation.
Smart Images

Figure 2025516221000001_ABST
Abstract
Description
Technical Field
[0001] This application claims the priority of a Chinese patent application with the application number 202210474744.1, filed with the Chinese Patent Office on April 29, 2022, and the entire content of the application is incorporated herein by reference.
[0002] Embodiments of the present disclosure relate to the technical field of image processing, for example, a method, apparatus, electronic device, and storage medium for determining an SFX video.
Background Art
[0003] With the development of network technology, more and more application programs have entered people's lives, and in particular, a series of software that can shoot short videos is popular among users.
[0004] To improve the fun of video shooting, related application software can provide users with the function of creating multiple types of SFX videos. However, the current function of creating SFX videos provided for users is very limited, and the fun of the finally obtained SFX videos needs to be further improved. At the same time, considering the personal need of users to change the background screen in the video is not considered, which reduces the user experience.
Summary of the Invention
Means for Solving the Problems
[0005] The present disclosure provides a method, apparatus, electronic device, and storage medium for determining an SFX video, which uses some content in an image uploaded by a user as the background to present the visual effect of a scene where a target object is in the uploaded image in the SFX video, meeting the personal needs of the user.
[0006] In a first aspect, embodiments of the present disclosure provide a method for determining an SFX video. responding to an SFE trigger operation, obtaining an uploaded image; determining a target viewpoint image from a 3D image surround scene corresponding to the uploaded image based on the position information of the imaging device; generating and displaying SFE video frames until an operation to stop shooting the SFE video is received based on the target viewpoint image and the target object.
[0007] In a second aspect, an embodiment of the present disclosure further provides a determining device for SFE video, an image acquisition module configured to obtain an uploaded image in response to an SFE trigger operation; a target viewpoint image module configured to determine a target viewpoint image from a 3D image surround scene corresponding to the uploaded image based on the position information of the imaging device; an SFE video frame generation module configured to generate and display SFE video frames until an operation to stop shooting the SFE video is received based on the target viewpoint image and the target object.
[0008] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device includes one or more processors; a storage device configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining SFE video according to any one of the embodiments of the present disclosure.
[0009] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium including computer-executable instructions, which are used to execute the method for determining an SFX video according to any one of the embodiments of the present disclosure when executed by a computer processor.
Brief Description of the Drawings
[0010] Throughout the drawings, the same or similar reference numerals represent the same or similar elements. As understood, the drawings are schematic and the elements and components are not necessarily drawn to scale.
[0011]
Figure 1
Figure 2
Figure 3
Embodiments for Carrying Out the Invention
[0012] As understood, the multiple steps described in the method embodiments of the present disclosure may be executed in a different order and / or in parallel. Also, the method embodiments may include additional steps and / or the execution of the steps shown may be omitted. The scope of the present disclosure is not limited in this regard.
[0013] As used herein, the term "including" and its variations are open-ended inclusion, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" represents "at least one embodiment", the term "another embodiment" represents "at least one another embodiment", and the term "some embodiments" represents "at least some embodiments". Related definitions of other terms are given in the following description.
[0014] As noted, the concepts such as "first", "second", etc. mentioned in this disclosure are only for distinguishing different devices, modules or units, and are not for limiting the order or interdependence of the functions executed by these devices, modules or units. As noted, the modifiers "one" and "a plurality" mentioned in this disclosure are not restrictive but exemplary, and should be understood as "one or a plurality" unless specifically indicated otherwise in the context, as would be understood by those skilled in the art.
[0015] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for the purpose of explanation, and are not for limiting the scope of these messages or information.
[0016] Before introducing the technical solution, first, the application scenario of the embodiments of this disclosure can be exemplarily described. Exemplarily, when a user takes a video or makes a video call with another user using application software, it may be expected to make the captured video more interesting. At the same time, the user may have personal needs for the SFX video screen. For example, some users may expect to replace the background in the video screen with specific content. At this time, according to the technical solution of this embodiment, after obtaining the image uploaded by the user, one target viewpoint image is determined from the 3D image surround scene corresponding to the image, and further, the target viewpoint image and the target object are fused to generate an SFX video, thereby presenting the visual effect of the scene where the target object is in the uploaded image on the SFX video screen.
[0017] FIG. 1 is a flowchart of a method for determining an SFX video according to an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to situations where personal needs of users are satisfied while generating more interesting SFX videos. The method may be executed by a determination device for an SFX video, and the device may be implemented in the form of software and / or hardware. For example, it may be implemented by an electronic device, and the electronic device may be a mobile terminal, a PC terminal, a server, or the like.
[0018] As shown in FIG. 1, the method includes steps S110 to S130.
[0019] S110. In response to an SFX trigger operation, obtain an uploaded image.
[0020] The device for executing the method for determining an SFX video provided by the embodiments of the present disclosure may be integrated into application software that supports the SFX video processing function, and the software may be installed on an electronic device. For example, the electronic device may be a mobile terminal or a PC terminal. The application software may be a type of software for processing images / videos. Specific application software is not repeatedly described here as long as it can realize image / video processing. It may also be a specially developed application program that adds SFX and realizes displaying SFX in the software, or may be integrated into a corresponding page, and the user can realize processing of the SFX video through the page integrated on the PC terminal.
[0021] Note that the technical solution of this embodiment may be executed during real-time imaging by a mobile terminal, or may be executed after the system receives video data actively uploaded by a user. For example, when a user captures a video in real time using an imaging device on a terminal device, application software detects an SFX trigger operation, responds to the operation, further acquires an uploaded image, and processes the video currently being captured by the user to obtain an SFX video. Or, when a user actively uploads video data by application software and executes an SFX trigger operation, the application similarly responds to the operation, further acquires an uploaded image, and then processes the video data actively uploaded by the user to thereby obtain an SFX video.
[0022] For example, in response to an SFX trigger operation, an image upload frame is popped up, and an uploaded image is determined based on a trigger operation on the image upload frame.
[0023] The SFX trigger operation includes at least one of triggering the SFX video creation control, monitoring that the audio information includes an SFX addition command, and detecting that the display interface includes a face image. For example, application software can pre-develop a control for triggering and executing an SFX video creation program, which is the SFX video creation control. Based on this, when the application detects that the user has triggered this control, it can process the acquired upload image by executing the SFX video creation program. Further, audio information can be collected based on a microphone array arranged in the terminal device and analyzed. When the processing result includes terms for SFX video processing, it indicates that the function of performing SFX processing on the current video has been triggered. The advantage of determining whether to execute SFX video processing based on the content of the audio information is to avoid the interaction between the user and the display page and improve the intelligence of SFX video processing. As another implementation form, based on the shooting field of view range of the mobile terminal, it is determined whether the field of view range includes the user's face image. When the user's face image is detected, the application software can use the event that the face image is detected as a trigger operation for performing SFX processing on the video. As those skilled in the art will understand, specifically which events are selected as the conditions for SFX video processing can be set according to the actual situation, and the embodiments of the present disclosure do not specifically limit this here.
[0024] In this embodiment, when the application software responds to the SFX trigger operation, an uploaded image can be obtained. For example, when it is detected that an image upload frame is triggered, by calling the image library, the image selected in the image library is triggered as the uploaded image, or when it is detected that the image upload frame is triggered, by calling the imaging device, the uploaded image is captured by the imaging device.
[0025] The uploaded image is an image actively uploaded by the user, for example, a panoramic image on which a view of a scenic spot is displayed. The image upload frame is a control pre-developed and integrated into the application software, for example, a circular icon containing a plus sign. Based on this, when the user triggers the image upload control, the application software can be triggered to call the image library on the mobile terminal or the associated cloud image library, and further determine the uploaded image based on the user's selection result. The application software can be triggered to call the relevant interface of the imaging device of the mobile terminal, thereby obtaining the image captured by the imaging device and using the image as the uploaded image.
[0026] Exemplarily, when a user uses the imaging device of a mobile terminal to capture a video in real time and triggers an image upload frame displayed on a display interface, application software can automatically open the "photo album" in the mobile terminal and display the images in the "photo album" on the display interface based on the user's trigger operation on the image upload frame. When a trigger operation on a certain image of the user is detected, that is, it is shown that the user expects to use the screen of the image as the background of an SFX video. For example, the image selected by the user is uploaded to a server or a client corresponding to the application software, so that the application software uses the image as an uploaded image. Or, when a user uses the imaging device of a mobile terminal to capture a video in real time and triggers an image upload frame displayed on a display interface, application software can directly obtain the current video frame from the video captured in real time by the imaging device and use the video frame as an uploaded image. Of course, in the actual application process, when the uploaded image is a panoramic image, the application can obtain a plurality of video frames when responding to the trigger operation of the image upload frame, and perform a splicing process on the screens of the plurality of video frames, so that the finally obtained image can be used as an uploaded image. The embodiments of the present disclosure will not repeat the description thereof.
[0027] For example, after determining the uploaded image, it is also possible to determine the pixel ratio information of the uploaded image. Based on the pixel ratio information and a preset pixel ratio, the uploaded image is processed as a complementary image with a target pixel ratio, and based on the complementary image, a 3D image surround scene is determined.
[0028] The pixel ratio information of the uploaded image may be represented by the aspect ratio of the image. For example, when the width of the uploaded image is 6 unit lengths and the height is 1 unit length, the aspect ratio is 6:1, and correspondingly, the pixel ratio information is also 6:1. In this embodiment, when the application software acquires the uploaded image, the pixel ratio information of the uploaded image can be automatically determined by executing the image attribute determination program. Of course, in the actual application process, when the uploaded image carries information characterizing its aspect ratio, the application software can directly call the information and use the attribute information as the pixel ratio information of the uploaded image.
[0029] In this embodiment, the preset pixel ratio is the image aspect ratio information preset based on the application software. It can be understood that the preset pixel ratio is the basis for judging how the application software selects a method to process the uploaded image. For example, the preset pixel ratio may be set to 4:1. Of course, in the actual application process, the parameter can be adjusted according to the actual needs of the Sfx video processing, and the embodiments of the present disclosure do not specifically limit this.
[0030] In this embodiment, when the application software acquires the uploaded image and determines the pixel ratio information of the uploaded image and the preset pixel ratio, complementary processing can be performed on the uploaded image based on the above information. When the pixel ratio information of the uploaded image does not match the preset pixel ratio, the complementary image is an image obtained by embedding the content of the uploaded image and adjusting the aspect ratio of the uploaded image. For example, when the pixel ratio information of the uploaded image is larger than the preset pixel ratio, the application software can complement the upper and lower parts of the uploaded image. When the pixel ratio information of the uploaded image is smaller than the preset pixel ratio, the application software can complement the left and right sides of the uploaded image. It can be understood that the pixel ratio information of the complementary image matches the preset pixel ratio. Hereinafter, the complementary process of the uploaded image will be described.
[0031] In this embodiment, under the condition that the pixel ratio information is larger than the preset pixel ratio, the long side of the uploaded image is used as the embedding reference to perform pixel point embedding processing to obtain a complementary image with the target pixel ratio, or under the condition that the pixel ratio information is larger than the preset pixel ratio, the uploaded image is trimmed to obtain a complementary image with the target pixel ratio.
[0032] In this embodiment, when the pixel ratio of the uploaded image is larger than the preset pixel ratio, it is shown that the ratio of the long side to the short side of the uploaded image is too large. When the long side of the uploaded image corresponds to both the upper and lower sides of the image, it can be understood that the application software needs to perform complementary processing on the upper and lower parts of the uploaded image.
[0033] For example, the application software needs to determine multiple rows of pixel points in the uploaded image and select the pixel points in the top row. For example, it reads the RGB values of the multiple pixel points in this row and calculates the RGB average value of the pixel points in this row based on the pre-created average value function. It can be understood that the calculation result is the upper pixel average value of the uploaded image. Similarly, the process of determining the RGB average value of the pixel points in the bottom row of the uploaded image is the same as the above process, and the embodiments of the present disclosure will not repeat the description here. After the application determines the RGB average value of the pixel points in the top row and the RGB average value of the pixel points in the bottom row, it is necessary to determine regions at the top and bottom of the uploaded image respectively, that is, the region connected to the top of the uploaded image and the region connected to the bottom of the uploaded image. For example, fill the color of this region connected to the top based on the RGB average value of the pixel points in the top row, and at the same time, fill the color of the region connected to the bottom based on the RGB average value of the pixel points in the bottom row, that is, obtain a complementary image that meets the preset pixel ratio.
[0034] In this embodiment, when the pixel ratio information of the uploaded image is larger than the preset pixel ratio, only regions are connected to the top and bottom of the uploaded image respectively, and after filling the colors of the regions based on the two RGB average values, the display effect of the initially obtained complementary image is not good. That is, because the connection between the upper and lower two boundaries of the uploaded image and the newly added regions is too prominent, in order to improve the display effect of the obtained complementary image, it is also possible to determine transition regions with specific widths in the upper region and the bottom region of the original uploaded image respectively.
[0035] In this embodiment, when the pixel ratio information of the uploaded image is greater than a preset pixel ratio, trimming processing can also be performed on the uploaded image. For example, when the pixel ratio information of the uploaded image is 8:1 and the preset pixel ratio is 4:1, the application can directly perform trimming processing on both the left and right sides of the uploaded image respectively. That is, while trimming only the content with a length of 2 unit lengths along the long side on the left side of the uploaded image, at the same time, trimming only the content with a length of 2 unit lengths along the long side on the right side of the uploaded image. It can be understood that the complementary image obtained by the trimming processing can also meet the requirements of the preset pixel ratio.
[0036] For example, based on the target pixel ratio and pixel ratio information, determine the pixel embedding width, use one long side of the uploaded image as a reference, perform pixel point embedding based on the pixel embedding width to obtain a complementary image, or use two long sides of the uploaded image as a reference, perform embedding based on the pixel embedding width to obtain a complementary image. The pixel values of the pixel points within the pixel embedding width match the pixel values of the corresponding long side.
[0037] The application software can determine the corresponding edge width information based on a preset transition ratio. The edge width information is used to divide a certain area within the uploaded image. For example, when the preset transition ratio is 1 / 8 and the width of the short side of the uploaded image is 8 unit lengths, the application can, based on the above information, determine a first edge width of 1 unit length for the upper region of the uploaded image and, at the same time, determine a second edge width of 1 unit length for the bottom region of the uploaded image based on the above information. In the actual application process, it can be understood that when the preset transition ratios for the upper and bottom of the uploaded image are different, there are also differences in the values of the edge widths finally determined by the application for the upper and bottom of the image. The specific preset transition ratio can be adjusted according to actual needs, and the embodiments of the present disclosure do not specifically limit this.
[0038] When the applications respectively determine a total area of two unit lengths in the upper area and the bottom area of the uploaded image, the pixel values of multiple rows of pixel points within one unit length in the upper part can be read, and the pixel values of multiple rows of pixel points within one unit length in the bottom part can be read. For example, substituting the pixel values of multiple rows of pixel points in the upper part and the pixel average value in the upper part into a pre-created average value calculation function, multiple pixel average values corresponding to multiple rows of pixel points within one unit length in the upper area can be obtained. Similarly, substituting the pixel values of multiple rows of pixel points in the bottom part and the pixel average value at the bottom end into the pre-created average value calculation function, multiple pixel average values corresponding to multiple rows of pixel points within one unit length in the bottom area can be obtained. It can be understood that the pixel average values corresponding to the multiple rows of pixel points obtained by calculation are the transition pixel values of the uploaded image.
[0039] Finally, update the color attribute information of the corresponding pixel points with the transition pixel values of multiple rows of pixel points, and assign color attribute information to the corresponding pixel points based on the pixel average value in the upper part and the pixel average value in the bottom part, so as to obtain a complementary image corresponding to the uploaded image. At the same time, divide the transition area at the upper part of the uploaded image and add a complementary area, and divide the transition area at the bottom part of the uploaded image and add a complementary area, so that the obtained complementary image can meet the target pixel ratio. In the actual application process, the target pixel ratio may be 2:1. Of course, in the actual application process, the target pixel ratio can be adjusted according to the needs of real-time SFX video processing, and the embodiments of the present disclosure do not specifically limit this.
[0040] Exemplarily, when the pixel ratio information of the uploaded image is 8:1 and the preset pixel ratio is 4:1, the application software needs to increase a plurality of rows of pixel points at the top and bottom of the uploaded image respectively. In addition, in the process of adding a plurality of rows of pixel points, the number of rows of pixel points added at the top can be made to match the number of rows of pixel points added at the bottom. After the addition of a plurality of rows of pixel points is completed, the application can assign color attribute information to the plurality of rows of pixel points added at the top based on the pixel average value at the top (i.e., the RGB average value of the pixel points in the topmost row), and can assign color attribute information to the plurality of rows of pixel points added at the bottom based on the pixel average value at the bottom (i.e., the RGB average value of the pixel points in the bottommost row). For example, based on the preset transition ratio and the short side width information of the uploaded image, two transition regions can be divided in the upper region and the bottom region of the uploaded image respectively. After calculating the pixel average value of a plurality of rows of pixel points in the transition region, the original color attribute information of the pixel points in the two regions is updated based on the pixel average value, thereby obtaining a complementary image corresponding to the uploaded image with a pixel ratio of 2:1.
[0041] In this embodiment, when the pixel ratio of the uploaded image is larger than the preset pixel ratio, the advantages of increasing a plurality of rows of pixel points at the top and bottom of the uploaded image and dividing the transition region on the uploaded image based on the preset transition ratio are that the obtained complementary image meets the target pixel ratio, making it easier for the application to perform subsequent processing on the image. At the same time, the display effect of the image is also improved, and finally the content of the rendered image becomes more natural.
[0042] In this embodiment, there may also be a situation where the pixel ratio information of the uploaded image is smaller than the preset pixel ratio. For example, when the pixel ratio information is smaller than the preset pixel ratio, a complementary image with the target pixel ratio can be obtained by performing mirroring processing on the uploaded image.
[0043] As would be understood by those skilled in the art, the mirroring process of an image can be divided into three types: horizontal mirroring, vertical mirroring, and diagonal mirroring. In this embodiment, since the pixel ratio information of the uploaded image is smaller than the preset pixel ratio, it is necessary to perform horizontal mirroring processing on the uploaded image. That is, mirroring conversion is performed on the uploaded image with the left edge axis or the right edge axis of the image on the screen as the center, thereby obtaining a plurality of uploaded images arranged horizontally. It can be understood that for any two adjacent images, the visual effect of mirroring conversion is presented on the screen of the image. For example, when an image formed by joining a plurality of mirroring images meets the target pixel ratio, the joined image is the complementary image corresponding to the uploaded image.
[0044] In addition, when the pixel ratio information is smaller than the preset pixel ratio and equal to the target pixel ratio, the uploaded image is used as the complementary image. That is, before the uploaded image is processed, when the ratio of its long side to its short side becomes equal to the target pixel ratio, the application does not need to perform complementary processing on the uploaded image, and directly uses the uploaded image as the complementary image to be used in the subsequent process. The embodiments of the present disclosure will not repeat the description of this.
[0045] In this embodiment, when the application software determines the complementary image corresponding to the uploaded image, it can determine the corresponding 3D image surround scene based on the complementary image. For example, based on the complementary image, six patch maps corresponding to the boundary frame of a cuboid are determined, and based on the six patch maps, the 3D image surround scene corresponding to the uploaded image is determined.
[0046] The 3D image surround scene is composed of at least six patch maps. At the same time, the 3D image surround scene corresponds to a cuboid composed of at least six patches. As those skilled in the art will understand, a patch refers to a mesh in application software that supports image rendering processing, and can be understood as an object for carrying an image within the application software. Each patch is composed of two triangles and contains a plurality of vertices. Correspondingly, based on the information of these vertices, it is also possible to determine the patch to which these vertices belong. Based on this, in this embodiment, six patches of the 3D image surround scene each carry a part of the screen on the complementary image. Furthermore, when the virtual camera is located at the center of the cuboid, it can be understood that the screens on each patch are rendered to the display interface from different angles.
[0047] Exemplarily, when the uploaded image is an image of a scenic spot and the application software determines a corresponding complementary image for the uploaded image, six different regions can be divided on the complementary image, and a cuboid boundary frame model composed of one three-dimensional space coordinate system and six blank patch maps can be constructed in the virtual space. For example, the contents of the six parts on the complementary image can be sequentially mapped to the six patches of the cuboid boundary frame model according to the order to obtain a 3D image surround scene.
[0048] In addition, in the process of mapping the screen on the complementary image to the six patches of the cuboid boundary frame, in order to ensure the accuracy of mapping, a sphere model having the same center point as the cuboid boundary frame can be constructed within the space coordinate system of the three-dimensional space. Based on this, the attribute information (such as RGB values) of a plurality of pixel points on the complementary image is mapped to the sphere surface, and the conversion relationship between a plurality of points on the sphere surface and a plurality of points on the cuboid boundary frame is determined based on trigonometric functions. Based on the conversion relationship, the attribute information of a plurality of pixel points on the sphere surface is mapped to the six patches of the cuboid boundary frame, thereby realizing the effect of mapping the screens of the six regions on the complementary image to the cuboid boundary frame.
[0049] S120 determines a target viewpoint image from a 3D image surround scene corresponding to the uploaded image based on the position information of the imaging device.
[0050] In this embodiment, when the application software determines a 3D image surround scene corresponding to the uploaded image based on the supplementary image, the screen in the 3D image surround scene can be fused with the image acquired in real time. In the actual application process, since the size of the display interface is limited and the image rendered on the interface only includes a part of the 3D image surround scene, the application software needs to further determine the position information of the imaging device and further determine the corresponding screen in the 3D image surround scene based on the information. The screen is the content that the user can see when the imaging device is at the current position. The image including the screen is the target viewpoint image. At the same time, it can be understood that the target viewpoint image also needs to be fused with the screen acquired in real time and rendered on the display interface.
[0051] For example, the position information of the imaging device is acquired in real time or periodically. Based on the position information, the rotation angle of the imaging device is determined, and it is determined that the rotation angle corresponds to the target viewpoint image in the 3D image surround scene.
[0052] The position information is, that is, information for reflecting the current viewpoint of the user, and this information is determined based on a gyroscope or an inertial measurement unit arranged in the photographing device. As would be understood by those skilled in the art, a gyroscope is, that is, an angular motion detection device that rotates around one or two axes orthogonal to the rotation axis by using the relative inertial space of the momentum moment housing of a high-speed rotating body. Of course, an angular motion detection device manufactured by using other principles and a device having a similar function are also called gyroscopes. An inertial measurement unit is a device that measures the three-axis attitude angle (or angular velocity) and acceleration of an object. Generally, one inertial measurement unit includes three single-axis accelerometers and three single-axis gyroscopes. The accelerometer detects three-axis acceleration signals of the object that are independent of the carrier coordinate system, and the gyroscope detects angular velocity signals with respect to the navigation coordinate system of the carrier, and measures the angular velocity and acceleration of the object in three-dimensional space, and can further calculate the current attitude of the object. The embodiments of the present disclosure will not be repeatedly described here.
[0053] In this embodiment, when a user uses a shooting device on a shooting device or a mobile terminal to shoot a video, the application can determine its position information in real time by using a gyroscope or an inertial measurement unit. For example, during the shooting process by the user, the above two devices transmit the detected information to the application software in real time, thereby determining the position information in real time. When determining the position information in real time, it can be understood that in the finally obtained SFX video, the screen in the 3D image surround scene as the background may continue to change. Or, the gyroscope or the inertial measurement unit periodically transmits the detected information to the application software, thereby enabling the position information corresponding to a plurality of periods to be determined during the process of the user shooting the video. For example, the above device transmits the detected information to the application every 10 seconds to cause the application to determine the position information. In the finally obtained SFX video, the screen in the 3D image surround scene as the background may change every 10 seconds. When the application is arranged on the mobile terminal, by periodically determining the position information, the consumption of the computing resources of the terminal can be reduced, thereby improving the processing efficiency of the SFX video.
[0054] In this embodiment, after determining the position information, the rotation angle of the shooting device can be determined based on the information, and further a specific screen can be determined in the 3D image surround scene based on the angle. The image corresponding to the screen is the target viewpoint image, and it can be understood that the content of the image is the part that the user can observe in the 3D image surround scene in the current posture of the shooting device.
[0055] Exemplarily, after the origin of the virtual three-dimensional space coordinate system represents the position of the imaging device and its position information is determined, the application can determine a partial area corresponding to the current position information of the imaging device on the boundary frame of the rectangular parallelepiped surrounding the origin. It can be understood that the screen of the partial area is the screen that the user can observe at the current time. For example, after constructing a blank image and reading the information of a plurality of pixel points on the patch where the partial area is located, it can be drawn on the blank image based on the pixel point information, thereby obtaining a target viewpoint image.
[0056] S130. Based on the target viewpoint image and the target object, generate and display an SFX video frame until an operation to stop shooting the SFX video is received.
[0057] In this embodiment, when the application determines the target viewpoint image, in order to obtain the SFX video, it is necessary to determine the target object in the video screen captured in real time by the user. The target object in the video screen may be dynamic or static. At the same time, the number of target objects may be one or more. For example, a plurality of specific users can be used as target objects. Based on this, when the application identifies the face features of one or more specific users from the video screen captured in real time based on the pre-trained image recognition model, the processing process of the SFX video of the embodiments of the present disclosure can be executed.
[0058] In this embodiment, when the application determines the target perspective image and determines the target object within the video screen, the corresponding SFX video frame can be generated based on the above image data. The SFX video frame can include a background image and a foreground image. The background image is the target perspective image, and the foreground image is the screen corresponding to the target object. The foreground image is overlaid on the background image and can cover all or part of the area of the background image, thereby making the constructed SFX video frame more hierarchical. The process of generating the SFX video frame will be described below.
[0059] For example, obtain the target object in the video frame to be processed, perform a fusion process on the target object and the target perspective image, and obtain the SFX video frame corresponding to the video frame to be processed.
[0060] For example, when the application obtains the video captured by the user in real time and identifies the target object from the screen, the video can be analyzed to obtain the video frame to be processed corresponding to the current time. For example, based on a pre-created clipping program, the view corresponding to the target object is extracted from the video frame to be processed. As understood by those skilled in the art, clipping refers to the processing operation of separating an image or video from a certain part of the original image or video frame to obtain a single layer. In this embodiment, the view obtained by the clipping process is the image corresponding to the target object.
[0061] For example, after fusing a view including a target object and a target perspective image, an SFX video frame corresponding to a video frame to be processed is obtained. Exemplarily, when an application identifies a user as a target object within a video frame to be processed, a view including only the user can be obtained by a clipping process operation. At the same time, the application determines a target perspective image within a 3D image surround scene corresponding to a panoramic image of a certain scenic spot. Based on this, the application can fuse a part of the screen of the scenic spot and the screen of the user to obtain an SFX video frame. It can be understood that the screen in the SFX video frame can present the visual effect that the user is currently shooting a video within the scenic spot.
[0062] Note that in the actual application process, the application can also determine a 3D image surround scene based on a trigger operation on at least one 3D image surround scene to be selected by displaying at least one 3D image surround scene to be selected on the display interface.
[0063] Exemplarily, multiple types of 3D image surround scenes corresponding to multiple screens are pre-integrated into the application and may also be stored in a specific storage space or a cloud server associated with the application. For the user, these scenes are the 3D image surround scenes to be selected. For example, a 3D image surround scene corresponding to a specific outdoor scenic spot and a 3D image surround scene corresponding to an indoor exhibition hall. For example, multiple controls can be developed in advance. Each control is associated with a specific 3D image surround scene and carries a specific logo. For example, two controls are pre-developed in the application. The logo below the first control is "scenic spot scene", and the logo below the second control is "exhibition hall scene". Based on this, when it is detected that the user triggers the first control, the application calls the data associated with the control, that is, the 3D image surround scene corresponding to a specific outdoor scenic spot, and thereby can execute the generation process of the above-mentioned SFX video frame. Of course, in the actual application process, the content and number of the 3D image surround scenes to be selected integrated in the application can be adjusted according to the actual needs. At the same time, as those skilled in the art will understand, for the 3D image surround scene generated in real time according to the embodiments of the present disclosure, the application can also use it as the 3D image surround scene to be selected and store it, and further call the scene by the user at any time. The embodiments of the present disclosure do not specifically limit this.
[0064] In this embodiment, when an application generates an SFX video frame, information of a plurality of pixels in the SFX video frame can be written into a rendering engine, so that the rendering engine can render a corresponding screen to a display interface. The rendering engine is a program that controls the GPU, i.e., the graphics processing unit, to render related images. That is, the computer can complete the drawing task for the SFX video frame, and the embodiments of the present disclosure will not repeatedly explain this.
[0065] In this embodiment, when an application detects an operation to stop shooting an SFX video, the above processing steps of the embodiments of the present disclosure will not be executed. The operation to stop shooting an SFX video includes at least one of detecting that the shooting stop control is triggered, detecting that the shooting period of the SFX video reaches a preset shooting period, detecting that the wake-up word for shooting stop is triggered, and detecting that the body movement for shooting stop is triggered. The above conditions will be described respectively below.
[0066] For example, regarding the above-mentioned SFX video shooting stop operation, one control can be developed in advance within the application software. At the same time, a program for ending the SFX video processing can be associated with the control, and this control is the shooting stop control. Based on this, when it is detected that the user triggers the control, the application software can call the relevant program, thereby ending the processing operations for the current time and a plurality of video frames to be processed after that time. It can be understood that there are multiple ways for the user to trigger the control. Exemplarily, when the client is installed and arranged on a PC terminal, the user can trigger the shooting stop control by means of mouse click. When the client is installed and arranged on a mobile terminal, the user can trigger the shooting stop control by means of finger touch. As those skilled in the art will understand, the specific touch method can be selected according to the actual situation, and the embodiments of the present disclosure do not specifically limit this.
[0067] Regarding the above-mentioned second SFX video shooting stop operation, the application can preset one period as the preset shooting period and record the period during which the user shoots the video. For example, by comparing the recording result with the preset shooting period, when it is determined that the user's shooting period has reached the preset shooting period, the processing operations for the current time and a plurality of video frames to be processed after that time can be ended.
[0068] Regarding the above-mentioned third SFE video shooting stop operation, specific information can be preset in the application software as a wake-up word for shooting stop. For example, one or more of the terms such as "stop", "shooting stop", and "stop processing" are used as the wake-up word for shooting stop. Based on this, when the application software receives voice information sent by the user, it uses a pre-trained voice recognition model to identify the voice information, and determines whether one or more of the above-mentioned preset SFE mount wake-up words are included in the recognition result. When the determination result is yes, the application can terminate the processing operations for the current time and a plurality of video frames to be processed after that time.
[0069] Regarding the above-mentioned fourth SFE video shooting stop operation, multiple people's motion information can be input into the application software and these motion information can be used as preset motion information. For example, information reflecting the motion of a person raising both hands is used as the preset motion information. Based on this, when the application receives an image or video actively uploaded by the user or collected in real time using an imaging device, it can identify the screen in the image or a plurality of video frames based on a pre-trained limb motion information recognition algorithm. When the recognition result shows that the limb motion information of the target object in the current screen matches the preset motion information, the application can terminate the processing operations for the current time and a plurality of video frames to be processed after that time.
[0070] Note that the above SFE mount conditions may take effect simultaneously in the application software, or one or more of them may be selected to take effect in the application software. The embodiments of the present disclosure do not specifically limit this.
[0071] The technical solution of the embodiments of the present disclosure is to, in response to an SFX trigger operation, obtain an uploaded image, obtain a data basis for generating an SFX video background, determine a target viewpoint image from a 3D image surround scene corresponding to the uploaded image based on the position information of the photographing device, and generate and display SFX video frames until an operation to stop shooting the SFX video is received based on the target viewpoint image and the target object, so as to use some content of the image uploaded by the user as the background, present the visual effect of the scene where the target object is in the uploaded image in the SFX video, not only enhance the interestingness of the SFX video, but also meet the personal needs of the user and improve the usage experience of the user in the process of creating the SFX video.
[0072] FIG. 2 is a structural schematic diagram of a determination device for an SFX video according to an embodiment of the present disclosure. As shown in FIG. 2, the device includes an image acquisition module 210, a target viewpoint image module 220, and an SFX video frame generation module 230.
[0073] The image acquisition module 210 is configured to obtain an uploaded image in response to an SFX trigger operation.
[0074] The target viewpoint image module 220 is configured to determine a target viewpoint image from a 3D image surround scene corresponding to the uploaded image based on the position information of the photographing device.
[0075] The SFX video frame generation module 230 is configured to generate and display SFX video frames until an operation to stop shooting the SFX video is received based on the target viewpoint image and the target object.
[0076] Based on the above technical solution, the image acquisition module 210 includes an image upload frame generation unit and an image determination unit.
[0077] The image upload frame generation unit is configured to pop up an image upload frame in response to an SFX trigger operation.
[0078] The image determination unit is configured to determine the uploaded image based on a trigger operation on the image upload frame.
[0079] For example, when it is detected that the image upload frame is further triggered, the image determination unit calls the image library to trigger the image selected in the image library as the uploaded image, or when it is detected that the image upload frame is triggered, the imaging device is called to configure the imaging device to capture the uploaded image.
[0080] Based on the above technical solution, the SFX video determination device further includes a pixel ratio information determination module, a complementary image determination module, and a 3D image surround scene determination module.
[0081] The pixel ratio information determination module is configured to determine the pixel ratio information of the uploaded image.
[0082] The complementary image determination module is configured to process the uploaded image as a complementary image with a target pixel ratio based on the pixel ratio information and a preset pixel ratio.
[0083] The 3D image surround scene determination module is configured to determine the 3D image surround scene composed of at least six patch maps based on the complementary image.
[0084] Based on the above technical solution, the 3D image surround scene corresponds to a cuboid composed of the at least six patches.
[0085] For example, the complementary image determination module is further configured to perform pixel point embedding processing using the long side of the uploaded image as the embedding reference under the condition that the pixel ratio information is greater than the preset pixel ratio, so as to obtain a complementary image with the target pixel ratio, or, under the condition that the pixel ratio information is greater than the preset pixel ratio, trim the uploaded image to obtain a complementary image with the target pixel ratio.
[0086] For example, the complementary image determination module is further configured to determine the pixel embedding width based on the target pixel ratio and the pixel ratio information, use one long side of the uploaded image as a reference, perform pixel point embedding based on the pixel embedding width to obtain the complementary image, or use two long sides of the uploaded image as a reference, perform embedding based on the pixel embedding width to obtain the complementary image, and the pixel value of the pixel point within the pixel embedding width matches the pixel value of the corresponding long side.
[0087] For example, the complementary image determination module is further configured to obtain a complementary image with the target pixel ratio by mirroring the uploaded image when the pixel ratio information is smaller than the preset pixel ratio.
[0088] For example, the 3D image surround scene determination module is further configured to determine six patch maps corresponding to the boundary frame of the cuboid based on the complementary image, and determine the 3D image surround scene corresponding to the uploaded image based on the six patch maps.
[0089] Based on the above technical solution, the target viewpoint image module 220 includes a position information acquisition unit and a target viewpoint image determination unit.
[0090] The position information acquisition unit is configured to acquire the position information of the imaging device in real time or periodically, and the position information is determined based on a gyroscope or an inertial measurement unit arranged in the imaging device.
[0091] The target viewpoint image determination unit is used to determine the rotation angle of the imaging device based on the position information and to determine that the rotation angle corresponds to a target viewpoint image in the 3D image surround scene.
[0092] Based on the above technical solution, the SFX video frame generation module 230 includes a target object acquisition unit and an SFX video frame generation unit.
[0093] The target object acquisition unit is configured to acquire a target object in a video frame to be processed.
[0094] The SFX video frame generation unit is used to perform a fusion process on the target object and the target viewpoint image to obtain the SFX video frame corresponding to the video frame to be processed.
[0095] Based on the above technical solution, the SFX video determination device further includes a 3D image surround scene display module.
[0096] The 3D image surround scene display module is configured to determine the 3D image surround scene based on a trigger operation on at least one selectable 3D image surround scene by displaying at least one selectable 3D image surround scene on a display interface.
[0097] Based on the above technical solution, the operation to stop shooting the SFX video includes at least one of the following: being detected that the shooting stop control has been triggered; being detected that the shooting period of the SFX video has reached a preset shooting period; being detected that the wake-up word for shooting stop has been triggered; being detected that the limb movement for shooting stop has been triggered.
[0098] The technical solution according to this embodiment responds to the SFX trigger operation to obtain an uploaded image, that is, to obtain the data basis for generating the SFX video background. Based on the position information of the shooting device, a target viewpoint image is determined from the 3D image surround scene corresponding to the uploaded image. Until the operation to stop shooting the SFX video is received based on the target viewpoint image and the target object, the SFX video frame is generated and displayed, so that a part of the content of the image uploaded by the user is used as the background, presenting the visual effect of the scene where the target object is in the uploaded image in the SFX video, not only enhancing the interestingness of the SFX video, but also meeting the user's personal needs and improving the user experience in the process of creating the SFX video.
[0099] The SFX video determination device according to the embodiment of the present disclosure can execute the SFX video determination method according to any embodiment of the present disclosure, and is provided with corresponding functional modules and beneficial effects for executing the method.
[0100] It should be noted that the multiple units and modules included in the above device are only divided according to functional logic, not limited to the above division, as long as the corresponding functions can be realized. Also, the specific names of the multiple functional units are only for the convenience of distinguishing each other, not for limiting the protection scope of the embodiments of the present disclosure.
[0101] FIG. 3 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure. Hereinafter, referring to FIG. 3, a schematic structural diagram of an electronic device (for example, a terminal device or a server in FIG. 3) 300 configured to implement the embodiment of the present disclosure is shown. The terminal device in the embodiment of the present disclosure may include, for example, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (for example, in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers, but is not limited thereto. The electronic device shown in FIG. 3 is only an example and does not limit the functions and usage ranges of the embodiments of the present disclosure in any way.
[0102] As shown in FIG. 3, the electronic device 300 may include a processing device (for example, a central processor, a pattern processor, etc.) 301, and may execute a plurality of appropriate operations and processes based on a program stored in the read-only memory (ROM) 302 or a program uploaded from the storage device 306 to the random access memory (RAM) 303. A plurality of types of programs and data necessary for the operation of the electronic device 300 are further stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. The editing / output (I / O) interface 305 is also connected to the bus 304.
[0103] Typically, devices such as an editing device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc., an output device 307 including, for example, a liquid crystal display (LCD), a speaker, an oscillator, etc., a storage device 308 including, for example, a magnetic tape, a hard disk, etc., and a communication device 309 can be connected to the I / O interface 305. The communication device 309 enables the electronic device 300 to perform wireless or wired communication with other devices and exchange data. Although FIG. 3 shows an electronic device 300 having a plurality of types of devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices can be implemented or included.
[0104] In particular, according to an embodiment of the present disclosure, the above process described with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product including a computer program that runs on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network by the communication device 309, or may be installed from the storage device 306, or may be installed from the ROM 302. When the computer program is executed by the processing device 301, it executes the above functions defined by the method of the embodiment of the present disclosure.
[0105] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are merely for the purpose of description and are not intended to limit the scope of these messages or information.
[0106] The electronic device provided by the embodiment of the present disclosure belongs to the same concept as the method for determining the SFX video provided by the above embodiment. Technical details not described in detail in this embodiment may be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0107] Embodiments of the present disclosure provide a computer storage medium storing a computer program which, when executed by a processor, implements the method for determining an SFX video according to the above embodiments.
[0108] Note that the computer-readable medium of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above, but is not limited thereto. More specific examples of the computer-readable storage medium may include an electrical connection having one or more conductors, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact magnetic disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above, but is not limited thereto. In the present disclosure, the computer-readable storage medium may be any tangible medium that includes or stores a program, and the program may be used by an instruction execution system, apparatus, or device, or used in combination therewith. In the present disclosure, the computer-readable signal medium may include a data signal propagated within a baseband or as part of a carrier wave, and the computer-readable program code is carried thereon. The data signal propagated in this way may use multiple types of formats, including electromagnetic signals, optical signals, or any suitable combination of the above, but is not limited thereto. The computer-readable signal medium may be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can transmit, propagate, or transmit a program for use by an instruction execution system, apparatus, or device, or for use in combination therewith. The program code included in the computer-readable medium may be transmitted through any suitable medium, including electric wires, cables, RF (RF), etc., or any suitable combination of the above, but is not limited thereto.
[0109] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the World Wide Web (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), and any network currently known or future-developed.
[0110] The computer-readable medium may be included in the electronic device or may exist independently without being incorporated into the electronic device.
[0111] One or more programs are carried on the computer-readable medium, and when the one or more programs are executed by the electronic device, the electronic device acquires an uploaded image in response to an SFX trigger operation, determines a target viewpoint image from a 3D image surround scene corresponding to the uploaded image based on the position information of the imaging device, generates and displays SFX video frames until an operation to stop shooting the SFX video is received based on the target viewpoint image and the target object.
[0112] The computer program code for performing the operations of the present disclosure may be created in one or more programming languages or combinations thereof, and the programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and further include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computer, may be partially executed on the user computer, may be executed as one independent software package, may be partially executed on the user computer and partially executed on a remote computer, or may be executed entirely on a remote computer or server. When related to a remote computer, the remote computer may be connected to the user computer via any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, connected via the Internet using an Internet service provider).
[0113] Flowcharts and block diagrams in the drawings illustrate the possible system architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram can represent a module, a program section, or a portion of code, and the module, program section, or portion of code includes one or more executable instructions for implementing a given logic function. As noted, in some alternative implementations, the functions marked within a block may occur in an order different from the order marked in the drawings. For example, two blocks shown consecutively may actually be executed substantially in parallel, or in some cases, in the reverse order, depending on the related functions. As noted, each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented using a dedicated system based on hardware for performing a given function or operation, or may be implemented using a combination of dedicated hardware and computer instructions.
[0114] The units described in the embodiments of the present disclosure may be implemented in the form of software or in the form of hardware. In some cases, the name of the unit does not limit the unit itself. For example, the first acquisition unit may be described as "a unit for acquiring at least two Internet protocol addresses".
[0115] In this specification, the functions described above can be executed at least in part by one or more hardware logic components. For example, by way of non-limiting and exemplary hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), and complex programmable logic devices (CPLDs), etc.
[0116] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes or can store a program used in or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0117] According to one or more embodiments of the present disclosure, [Example 1] provides a method for determining an SFX video, the method comprising: responding to an SFX trigger operation to obtain an uploaded image; determining a target viewpoint image from a 3D image surround scene corresponding to the uploaded image based on the position information of the photographing device; generating and displaying SFX video frames until an operation to stop shooting the SFX video is received based on the target viewpoint image and the target object.
[0118] According to one or more embodiments of the present disclosure, [Example 2] provides a method for determining an SFX video, and the step of obtaining an uploaded image in response to an SFX trigger operation is: responding to an SFX trigger operation to pop up an image upload frame; determining the uploaded image based on a trigger operation on the image upload frame.
[0119] According to one or more embodiments of the present disclosure, [Example 3] provides a method for determining an SFX video. Based on a trigger operation on the image upload frame, the step of determining the upload image includes: When it is detected that the image upload frame is triggered, by calling an image library, triggering the image selected in the image library as the upload image, or When it is detected that the image upload frame is triggered, by calling an imaging device, including the step of photographing the upload image with the imaging device.
[0120] According to one or more embodiments of the present disclosure, [Example 4] provides a method for determining an SFX video, determining the pixel ratio information of the upload image, and processing the upload image as a complementary image with a target pixel ratio based on the pixel ratio information and a preset pixel ratio, and further including the step of determining the 3D image surround scene composed of at least six patch maps based on the complementary image.
[0121] According to one or more embodiments of the present disclosure, [Example 5] provides a method for determining an SFX video, The 3D image surround scene corresponds to a rectangular parallelepiped composed of the at least six patches.
[0122] According to one or more embodiments of the present disclosure, [Example 6] provides a method for determining an SFX video. Based on the pixel ratio information and a preset pixel ratio, the step of processing the upload image as a complementary image with a target pixel ratio includes: Under the condition that the pixel ratio information is greater than the preset pixel ratio, using the long side of the upload image as an embedding reference to perform pixel point embedding processing to obtain the complementary image with the target pixel ratio, or Including the step of obtaining a complementary image of the target pixel ratio by trimming the uploaded image under the condition that the pixel ratio information is greater than the preset pixel ratio.
[0123] According to one or more embodiments of the present disclosure, [Example 7] provides a method for determining an SFX video. Using the long side of the uploaded image as an embedding reference to perform pixel point embedding processing, the step of obtaining a complementary image of the target pixel ratio is Determining a pixel embedding width based on the target pixel ratio and the pixel ratio information; Using one long side of the uploaded image as a reference, performing pixel point embedding based on the pixel embedding width to obtain the complementary image, or Using two long sides of the uploaded image as a reference, performing embedding based on the pixel embedding width to obtain the complementary image, and The pixel values of the pixel points within the pixel embedding width match the pixel values of the corresponding long side.
[0124] According to one or more embodiments of the present disclosure, [Example 8] provides a method for determining an SFX video. Based on the pixel ratio information and the preset pixel ratio, the step of processing the uploaded image as a complementary image of the target pixel ratio is Including the step of obtaining a complementary image of the target pixel ratio by mirroring the uploaded image when the pixel ratio information is smaller than the preset pixel ratio.
[0125] According to one or more embodiments of the present disclosure, [Example 9] provides a method for determining an SFX video. Based on the complementary image, the step of determining the 3D image surround scene is Determining six patch maps corresponding to the boundary frame of a rectangular parallelepiped based on the complementary image; Determining the 3D image surround scene corresponding to the uploaded image based on the six patch maps described above.
[0126] According to one or more embodiments of the present disclosure, [Example 10] provides a method for determining an SFX video. The step of determining a target viewpoint image from the 3D image surround scene corresponding to the uploaded image based on the position information of the imaging device includes: Obtaining the position information of the imaging device in real time or periodically, where the position information is determined based on a gyroscope or an inertial measurement unit disposed on the imaging device; Determining the rotation angle of the imaging device based on the position information, and determining that the rotation angle corresponds to the target viewpoint image in the 3D image surround scene.
[0127] According to one or more embodiments of the present disclosure, [Example 11] provides a method for determining an SFX video. The step of generating an SFX video frame based on the target viewpoint image and the target object includes: Obtaining the target object in the video frame to be processed; Performing a fusion process on the target object and the target viewpoint image to obtain the SFX video frame corresponding to the video frame to be processed.
[0128] According to one or more embodiments of the present disclosure, [Example 12] provides a method for determining an SFX video. Further including, by displaying at least one 3D image surround scene to be selected on a display interface, determining the 3D image surround scene based on a trigger operation on the at least one 3D image surround scene to be selected.
[0129] According to one or more embodiments of the present disclosure, [Example 13] provides a method for determining an SFX video. The operation to stop shooting the SFX video is detected that the shooting stop control was triggered, detected that the shooting period of the SFX video has reached a preset shooting period, detected that the wake-up word for shooting stop was triggered, detected that the body movement for shooting stop was triggered, and includes at least one of them.
[0130] According to one or more embodiments of the present disclosure, [Example 14] provides a determination device for SFX video, and the device includes an image acquisition module configured to acquire an upload image in response to an SFX trigger operation; a target viewpoint image module configured to determine a target viewpoint image from a 3D image surround scene corresponding to the upload image based on the position information of the shooting device; an SFX video frame generation module configured to generate and display SFX video frames until an operation to stop shooting the SFX video is received based on the target viewpoint image and the target object.
[0131] Also, although multiple types of operations are described in a specific order, this should not be understood as requiring that these operations be executed in the specific order or sequence shown. In a specific environment, multitasking and parallel processing may be advantageous. Similarly, the above description includes multiple specific implementation details, but these should not be construed as limiting the scope of the present disclosure. Some features described in the context of a single embodiment may be implemented in combination in a single embodiment. Conversely, multiple types of features described in the context of a single embodiment may be implemented alone or in any suitable sub-combination manner in multiple embodiments.
Claims
1. A method for determining an SFX video, comprising: responding to an SFX trigger operation to obtain an uploaded image; determining a target viewpoint image from a 3D image surround scene corresponding to the uploaded image based on the position information of the imaging device; generating and displaying SFX video frames until an operation to stop shooting the SFX video is received based on the target viewpoint image and the target object.
2. The step of obtaining an uploaded image in response to an SFX trigger operation includes: responding to an SFX trigger operation to pop up an image upload frame; determining the uploaded image based on a trigger operation on the image upload frame.
3. The step of determining the uploaded image based on a trigger operation on the image upload frame includes: responding to the detection that the image upload frame is triggered, calling an image library, and triggering the selected image in the image library as the uploaded image, or responding to the detection that the image upload frame is triggered, calling an imaging device, and shooting the uploaded image with the imaging device.
4. determining pixel ratio information of the uploaded image; processing the uploaded image as a complementary image with a target pixel ratio based on the pixel ratio information and a preset pixel ratio; further comprising determining the 3D image surround scene composed of at least six patch maps based on the complementary image.
5. The method according to claim 4, wherein the 3D image surround scene corresponds to a cuboid composed of the at least six patches.
6. The step of processing the uploaded image as a complementary image with a target pixel ratio based on the pixel ratio information and a preset pixel ratio includes: In response to determining that the pixel ratio information is greater than the preset pixel ratio, performing pixel embedding processing using the long side of the uploaded image as the embedding reference to obtain a complementary image with the target pixel ratio, or, The method according to claim 4, comprising the step of obtaining a complementary image with the target pixel ratio by trimming the uploaded image in response to determining that the pixel ratio information is greater than the preset pixel ratio. **Claim 7** The step of performing pixel embedding processing using the long side of the uploaded image as the embedding reference to obtain a complementary image with the target pixel ratio is determining a pixel embedding width based on the target pixel ratio and the pixel ratio information; using one long side of the uploaded image as a reference and performing pixel embedding based on the pixel embedding width to obtain the complementary image, or, using two long sides of the uploaded image as references and performing embedding based on the pixel embedding width to obtain the complementary image, The method according to claim 6, wherein the pixel values of the pixels within the pixel embedding width match the pixel values of the corresponding long side. **Claim 8** The step of processing the uploaded image as a complementary image with the target pixel ratio based on the pixel ratio information and the preset pixel ratio is The method according to claim 4, comprising the step of obtaining a complementary image with the target pixel ratio by mirroring the uploaded image in response to determining that the pixel ratio information is less than the preset pixel ratio. **Claim 9** The step of determining the 3D image surround scene based on the complementary image is determining six patch maps corresponding to the boundary frame of a rectangular parallelepiped based on the complementary image; The method according to claim 4, comprising the step of determining the 3D image surround scene corresponding to the uploaded image based on the six patch maps. **Claim 10** The step of determining a target viewpoint image from the 3D image surround scene corresponding to the uploaded image based on the position information of the imaging device is A step of acquiring the position information of the imaging device in real time or periodically, wherein the position information is determined based on a gyroscope or an inertial measurement unit arranged in the imaging device; The method according to claim 1, further comprising: determining a rotation angle of the imaging device based on the position information, and determining that the rotation angle corresponds to a target viewpoint image in the 3D image surround scene.
11. The step of generating an SFX video frame based on the target viewpoint image and the target object includes: Acquiring a target object in a video frame to be processed; The method according to claim 1, further comprising: performing a fusion process on the target object and the target viewpoint image to obtain the SFX video frame corresponding to the video frame to be processed.
12. The method according to claim 1, further comprising: displaying at least one selectable 3D image surround scene on a display interface, and determining the 3D image surround scene based on a trigger operation on the at least one selectable 3D image surround scene.
13. The operation of stopping the shooting of the SFX video includes: Detecting that a shooting stop control is triggered; Detecting that the shooting period of the SFX video reaches a preset shooting period; Detecting that a wake-up word for shooting stop is triggered; The method according to claim 1, including at least one of detecting that a limb movement for shooting stop is triggered.
14. An SFX video determination device, comprising: An image acquisition module configured to acquire an upload image in response to an SFX trigger operation; A target viewpoint image module configured to determine a target viewpoint image from a 3D image surround scene corresponding to the upload image based on the position information of the imaging device; An SFX video frame generation module configured to generate and display an SFX video frame until an operation of stopping the shooting of the SFX video is received based on the target viewpoint image and the target object.
15. An electronic device, one or more processors, a storage device configured to store one or more programs, and when the one or more programs are executed by the one or more processors, cause the one or more processors to implement the method for determining an SFX video according to any one of claims 1 to 13. An electronic device
16. A storage medium containing computer-executable instructions, where the computer-executable instructions are used to execute the method for determining an SFX video according to any one of claims 1 to 13 when executed by a computer processor. A storage medium containing computer-executable instructions
Citation Information
Patent Citations
Real-time video data processing method and apparatus for realizing scene rendering, and computing device
CN107566853A
Panoramic image generation method and apparatus
CN108989681A
Video synthesis method and device, storage medium and electronic equipment
CN113115110A
Animation image editing method and machine-readable recording medium recorded with program for executing animation image edition
JP2001155468A
Imaging device, captured image recording method and program
JP2009253921A