Video generation method, information display method, and computing device
By constructing a 3D model and generating multiple target images using transformation parameters, combined with material information, the problem of poor video visual effects in existing technologies is solved, achieving more vivid video expression and cost reduction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2022-09-22
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, the visual effect of generating videos by stitching together multiple object images is poor, which cannot effectively express the characteristics of the target object, and the cost is high.
By constructing a 3D model of the original image and using multiple transformation parameters to convert it into multiple target images, the target video is generated by combining the source material information, thereby improving the dynamic effect of the video.
The generated videos have better visual effects, can better express the characteristics of the target object, and reduce video production costs.
Smart Images

Figure CN115908694B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer application technology, and in particular to a video generation method, an information display method, and a computing device. Background Technology
[0002] Compared to images, videos are more vivid and engaging, offering better visual effects, and have gradually become one of the main ways to promote, publicize, or beautify a subject.
[0003] Since shooting objects to generate videos is costly, existing technologies can also generate videos by stitching together multiple object images when object images exist. However, videos generated in this way have poor visual effects. Summary of the Invention
[0004] This application provides a video generation method, an information display method, and a computing device to solve the technical problem of poor visual effects in existing videos.
[0005] In a first aspect, embodiments of this application provide a video generation method, including:
[0006] Obtain at least one original image containing the target object;
[0007] Construct three-dimensional models corresponding to each of the at least one original image;
[0008] By using multiple transformation parameters, at least one 3D model can be transformed into multiple target images;
[0009] Determine the material information that matches the target object;
[0010] A target video is generated based on the multiple target images and the material information.
[0011] Secondly, this application provides a video generation method, including:
[0012] Obtain the product image of the target product;
[0013] Construct a 3D model corresponding to the product image;
[0014] The 3D model is transformed into multiple target images using multiple transformation parameters;
[0015] Determine the material information matching the target product;
[0016] Based on the material information and the multiple target images, a target video is generated.
[0017] Thirdly, this application provides an information display method, including:
[0018] Provide a display interface;
[0019] Image processing prompts are displayed on the display interface;
[0020] In response to the image processing operation triggered by the image processing prompt information, an image processing request is sent to the server; the image processing request is used by the server to determine at least one original image containing the target object, construct three-dimensional models corresponding to the at least one original image respectively; transform the at least one three-dimensional model into multiple target images using multiple transformation parameters; determine the material information matching the target object; and generate a target video based on the multiple target images and the material information.
[0021] Play the target video on the display interface.
[0022] Fourthly, this application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be invoked and executed by the processing component to implement the video generation method as described in the first aspect above, or the video generation method as described in the second aspect above, or the information display method as described in the third aspect above.
[0023] Fifthly, this application provides a computer storage medium storing a computer program, which, when executed by a computer, implements the video generation method as described in the first aspect, the video generation method as described in the second aspect, or the information display method as described in the third aspect.
[0024] In this embodiment, at least one original image containing a target object is obtained, a three-dimensional model corresponding to each of the at least one original image is constructed, and multiple transformation parameters are used to transform the at least one three-dimensional model into multiple target images. Material information corresponding to the target object is determined, and a target video is generated based on the multiple target images and the material information. This embodiment obtains multiple target images by reconstructing the original image in three dimensions and adjusting the three-dimensional model using transformation parameters, so that the target object in the video synthesized from the multiple target images has dynamic effects. Furthermore, the final target video is generated by combining the material information matching the target object, which improves the visual effect of the video and allows for a better expression of the target object.
[0025] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This application provides a schematic diagram illustrating the structure of an embodiment of an information processing system.
[0028] Figure 2 A flowchart of one embodiment of the video generation method provided in this application is shown;
[0029] Figure 3 A flowchart of yet another embodiment of the video generation method provided in this application is shown;
[0030] Figure 4 This illustration shows a schematic diagram of image synthesis display in a practical application, based on an embodiment of this application.
[0031] Figure 5 A flowchart of yet another embodiment of the video generation method provided in this application is shown;
[0032] Figure 6 A flowchart of one embodiment of an information display method provided in this application is shown;
[0033] Figure 7 This invention provides a schematic diagram of the structure of one embodiment of a video generation apparatus.
[0034] Figure 8 This invention provides a schematic diagram of the structure of an embodiment of an information display device.
[0035] Figure 9 A schematic diagram of one embodiment of a computing device provided in this application is shown. Detailed Implementation
[0036] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0037] In some of the processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.
[0038] The technical solutions of this application can be applied to scenarios where images provided by merchants, enterprise users, individual users, or design solution providers are processed to generate videos. Since videos are more vivid and engaging than images, offering better visual effects, they can achieve purposes such as advertising, promotion, or beautification of objects. The objects involved in this application can refer to people, animals, objects, etc. Of course, the object can also be a virtual product provided in an online system that users can interact with, such as purchasing or browsing. This virtual product can correspond to offline physical goods. When the online system is an online transaction system, the object can specifically refer to commodities. In online systems, images or videos are often used to describe objects to better promote them to users.
[0039] To reduce the production costs associated with shooting and generating videos, the current common practice is to stitch together multiple original images containing the target object. However, videos generated in this way lack vividness and naturalness, have poor visual effects, and fail to effectively express the characteristics of the target object. To improve visual effects, obtain high-quality videos, and reduce video production costs, the inventors, after a series of studies, proposed the technical solution of this application. In this application, the original image can be modeled, and multiple target images can be generated through multiple transformation parameters. These multiple target images can express the target object from different visual angles. Adding designed materials can significantly enhance the visual effects and fully express the characteristics of the object.
[0040] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0041] The technical solutions of the embodiments in this application can be applied to... Figure 1The information processing system shown can be a system with image processing capabilities, etc. In practical applications, the information processing system can be, for example, an online system that provides interactive processing operations, such as an online trading system that provides commodity transactions, or it can be other processing systems connected to the online trading system, thereby enabling the generation of videos to be published on the online trading system, etc. The information processing system can include a user terminal 101 and a server terminal 102.
[0042] In this system, the user terminal 101 and the server terminal 102 establish a connection via a network. The network provides the medium for the communication link between the user terminal 101 and the server terminal 102. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Optionally, the server can communicate with the user terminal via a mobile network. Accordingly, the mobile network standard can be any one of 2G (GSM), 2.5G (GPRS), 3G (WCDMA, TD-SCDMA, CDMA2000, UTMS), 4G (LTE), 4G+ (LTE+), 5G, WiMax, etc. Optionally, the user terminal can also establish a communication connection with the server via Bluetooth, WiFi, infrared, etc.
[0043] The user terminal 101 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application, etc. The user terminal 101 can be deployed on an electronic device and depends on the device to run or on certain apps within the device. The electronic device can, for example, have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. For ease of understanding... Figure 1 The user terminal is primarily represented by the device itself. Various other types of applications, such as search engines and instant messaging applications, can also be configured on electronic devices.
[0044] Server 102 may include one or more servers that provide various services. That is, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. In addition, it can be a server of a distributed system, or a server combined with blockchain, or a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology, etc.
[0045] Users can interact with the server 102 through the user terminal 101 to receive or send messages, etc. In the application scenario of this application embodiment, for example, it can obtain the original image containing the target object, and sense the user's corresponding operation to send the corresponding processing request to generate the target video from the original image.
[0046] For example, in this embodiment of the application, the server 102 can obtain the image processing request from the user terminal 101, construct a corresponding three-dimensional model for at least one original image containing the target object in the request, use multiple transformation parameters to transform the at least one three-dimensional model into multiple target images, determine the material information corresponding to the target object, generate a target video based on the multiple target images and the material information, and send the target video to the user terminal 101 so that the user terminal 101 can output the target video.
[0047] It should be noted that the video generation method provided in this application embodiment is generally executed by the server 102, and the corresponding video generation device is generally located in the server 102. Similarly, the information display method provided in this application embodiment is generally executed by the user terminal 101, and the corresponding information display device is generally located in the user terminal 101. However, in other embodiments of this application, the user terminal 101 may also have similar functions to the server 102, thereby executing the video generation method provided in this application embodiment. In other embodiments, the video generation method provided in this application embodiment may also be jointly executed by the user terminal 101 and the server 102.
[0048] It should be understood that Figure 1 The number of client and server instances shown is merely illustrative. Depending on implementation needs, there can be any number of client and server instances.
[0049] The implementation details of the technical solutions in the embodiments of this application are described in detail below.
[0050] Figure 2 A flowchart illustrating an embodiment of a video generation method provided in this application is shown. The technical solution of this embodiment can be executed by a server, and the method may include the following steps:
[0051] 201: Obtain at least one original image containing the target object.
[0052] As an alternative approach, one could receive a user's image processing request and obtain at least one original image included in the request. Alternatively, in online systems, such as online trading systems, the target object could be a product offered by the system. Each product has a corresponding product description page, typically using text and images to describe the product in detail; that is, the product description page includes product images. Therefore, as another alternative approach, one could receive a user's image processing request, determine the index information corresponding to the target object included in the request, and then, based on the index information, identify the original image containing the target object from the corresponding object description page. This index information can link to the object description page, allowing the original image containing the target object to be retrieved from the object description page.
[0053] In the two optional methods mentioned above, receiving the user's image processing request can be done by sending an image processing prompt message to the user terminal so that the user terminal can display the image processing prompt message on the display interface. The image processing request is sent in response to the image processing operation triggered by the user in response to the image processing prompt message.
[0054] The original image can be an image obtained by photographing the target object, or an image obtained by photographing the target object and then processing it accordingly. The target object is the main subject in the original image, that is, the main content of the image.
[0055] 202: Construct at least one 3D model corresponding to each of the original images.
[0056] Three-dimensional reconstruction can be used to create a three-dimensional model of at least one original image, thereby obtaining a three-dimensional model corresponding to each original image.
[0057] Optionally, the construction of a 3D model corresponding to each original image can be specifically based on the pixel depth values of the original image, using a 3D reconstruction model to construct the corresponding 3D model. This 3D reconstruction model can be obtained by training based on the pixel depth values of the sample image and the corresponding 3D model.
[0058] 203: Using multiple transformation parameters, transform at least one 3D model into multiple target images.
[0059] By performing 3D reconstruction on at least one original image, a 3D model of each original image can be obtained, thus yielding at least one 3D model.
[0060] Each 3D model can be transformed according to the multiple transformation parameters to obtain multiple target images corresponding to each 3D model; or each 3D model can be transformed according to the transformation parameters corresponding to each of the multiple transformation parameters to obtain multiple target images corresponding to each 3D model.
[0061] 204: Generate a target video based on multiple target images.
[0062] This application embodiment obtains multiple target images by reconstructing the original image in three dimensions and adjusting the three-dimensional model using transformation parameters. This results in the target object in the video synthesized from the multiple target images having dynamic effects, improving the visual effect of the video and enabling a better representation of the target object.
[0063] As an alternative approach, generating a target video based on multiple target images can be achieved by stitching together the multiple target images.
[0064] The multiple target images can be spliced together in a certain order, and the order of the multiple target images can be determined based on the order of at least one original image and the order of multiple transformation parameters.
[0065] As an alternative approach, to further enhance visual effects and video quality, generating a target video based on multiple target images may include: determining the material information matching the target object; and generating the target video based on the multiple target images and the material information. Therefore, as yet another embodiment, such as... Figure 3 The flowchart shown illustrates a video generation method, which may include the following steps:
[0066] 301: Obtain at least one original image containing the target object.
[0067] 302: Construct at least one 3D model corresponding to each of the original images.
[0068] 303: Using multiple transformation parameters, transform at least one 3D model into multiple target images.
[0069] 304: Determines the material information that matches the target object.
[0070] 305: Generate target video based on multiple target images and material information.
[0071] This application embodiment obtains multiple target images by reconstructing the original image in three dimensions and adjusting the three-dimensional model using transformation parameters. This enables the target object in the video synthesized from the multiple target images to have dynamic effects. Furthermore, by combining the material information matching the target object to generate the final target video, the visual effect of the video is improved, allowing for a better expression of the target object.
[0072] It should be noted that this embodiment is different from... Figure 2 The difference in the illustrated embodiment is that, after obtaining multiple target images, the material information matching the target object is determined, and a target video is generated based on the multiple target images and the material information.
[0073] One method for generating a target video based on multiple target images and material information is to synthesize multiple target images with the material information, and then generate a target video based on the synthesized multiple target images.
[0074] As an alternative approach, the material information may include at least one material image. Combining at least one material image with multiple target images may involve determining the correspondence between at least one material image and at least one target image, and combining any one material image into its corresponding target image.
[0075] Among them, determining the material information that matches the target object can be done by determining at least one material image based on the object category to which the target object belongs.
[0076] The category of this object can be, for example, food, clothing, personal care, etc.
[0077] Determining at least one source image based on the object category to which the target object belongs can be achieved by pre-setting a correspondence between object categories and source images. Based on this correspondence, the object category is input, and the source image corresponding to the object category is found.
[0078] In this system, one target image can correspond to at least one source image. Furthermore, for ease of processing, at least one source image and at least one target image can have a one-to-one correspondence; that is, one source image also corresponds to only one target image. By combining the object category and the number of images in the multiple target images, multiple source images corresponding to the object category and matching the number of images can be determined. These multiple source images can be in the form of a source video, corresponding to pre-defined object categories, etc., and these multiple source images are the image frames contained in the source video.
[0079] In this context, matching the number of source images with the number of target images can be achieved by having the same number of images as the target images, or by having a difference in the number of images within a specified range. When the number of images in the source images is the same as the number of images in the target images, there is a one-to-one correspondence between the source images and the target images. This one-to-one correspondence can be determined based on the arrangement order of the source images and the target images; for example, source images in the same position have a one-to-one correspondence with target images.
[0080] If the number of images in at least one source image is less than or greater than the number of images in multiple target images, then if the number is less, at least one target image can be selected from the multiple target images according to the number of images in the source images, for example, by selecting images sequentially from the first image in the order of arrangement. If the number is greater, multiple source images can be selected from at least one source image according to the number of images in the multiple target images, for example, by selecting images sequentially from the first image in the order of arrangement.
[0081] Of course, different object identifiers can be pre-defined as corresponding material images, so that at least one corresponding material image can be determined based on the object identifier; or, the object features of the target object can be identified, and at least one material image that meets the matching requirements can be determined based on the degree of matching between the image features of the material image and the object features; wherein, the degree of matching between the image features and the object features can be calculated based on a pre-trained matching model, etc.
[0082] One method for compositing any source image into its corresponding target image is as follows: based on any source image and its corresponding target image, determine the compositing region based on the object position of the target object in the target image; adjust the image size and compositing direction of the source image according to the compositing region; and composite the adjusted source image into the compositing region according to the compositing direction.
[0083] In addition, the compositing method for compositing any source image to its corresponding target image can include one or more of the following: color filtering, overlay, soft light, hard light, bright light, solid color blending, opacity, multiply, color burn, and color dodge. Different compositing methods can be preset for different source images.
[0084] As an alternative approach, the material information may include text information, and the text information may be combined with multiple target images to synthesize the text information into at least one of the multiple target images.
[0085] The text information can be generated based on object-related information about the target object. This object-related information may include one or more of the following: object description information, object reviews, object category information, and object images. The object image may refer to the original image mentioned above or other images containing the target object. The object description information may refer to relevant information on the object description page, such as the object name, price, and origin. The object reviews may include user comments about the target object.
[0086] The text information can be determined using a text generation model based on object-related information of the target object. This text generation model can be trained based on object-related information of sample objects and their corresponding sample text information.
[0087] There are several ways to synthesize the text information into at least one of multiple target images. For example, the text information can be superimposed onto the target image.
[0088] Among them, at least one of the multiple target images that are synthesized with the text information can be selected from the first image according to the order of the multiple target images, and the text information is synthesized into the selected at least one target image.
[0089] As another option, the source material information can include at least one source image and text information. A target video is generated based on multiple target images and source material information. When compositing the source material information with multiple target images, at least one source image can be prioritized for compositing, followed by the text information. Alternatively, the text information can be prioritized for compositing with multiple target images. There is no limitation on this approach.
[0090] In addition, as another option, the material information may include audio data that matches the target object. Based on multiple target images and material information, generating a target video may include: stitching multiple target images together to generate a candidate video, and fusing the audio data with the candidate video to obtain the target video.
[0091] Of course, the material information can also include at least one material image, text information, and audio data. At least one material image and text information can be combined with multiple target images, and then the combined multiple target images can be stitched together to generate a candidate video. Finally, the audio data can be fused with the candidate video to obtain the final target video.
[0092] For ease of understanding, Figure 4 This diagram illustrates how source images and text information are combined into at least one target image. Figure 4In the process, after constructing a corresponding 3D model for the original image 401, the 3D model is transformed into multiple target images using multiple transformation parameters. The figure uses two target images as examples: target image 402 and target image 403. The target object is a bowl of porridge, belonging to the object category of hot food, and the corresponding source image is a steaming image. Figure 4 Taking two source images, image 404 and image 405, and the text "steaming hot" in text 406 as an example, this will be explained. The target image 402 is combined with source image 404 and the text "steaming hot" to obtain target image 407. The target image 403 is combined with source image 405 and the text "steaming hot" to obtain target image 408. Target images 407 and 408 can be spliced together in a certain order to generate the target video. The specific method for determining the arrangement order of multiple target images can be found in the previous text and will not be repeated here.
[0093] In some embodiments, transforming at least one 3D model into multiple target images using multiple transformation parameters can be achieved by determining multiple transformation parameters corresponding to at least one camera movement effect, and then using these multiple transformation parameters to transform at least one 3D model into multiple target images.
[0094] These multiple transformation parameters can be used to adjust and project at least one 3D model to generate multiple target images. Each transformation parameter can consist of a transformation matrix and a projection matrix. The transformation matrix may include a rotation matrix, and may also include one or more of a translation matrix and a scaling matrix. The projection matrix is used to project the 3D model into a 2D image to obtain the target image.
[0095] Each camera movement effect can correspond to multiple transformation parameters. The display effect of multiple target images generated by the multiple transformation parameters corresponding to each camera movement effect can express that camera movement effect.
[0096] Among them, the at least one camera movement effect may include, for example, translation or rotation in any direction (front, back, left, right, up, down), as well as Hitchcock zoom (Dolly Zoom), etc.
[0097] Since a camera is a device that projects a three-dimensional object into a two-dimensional image, in some embodiments, transforming at least one three-dimensional model into multiple target images using multiple transformation parameters can be achieved by determining a virtual camera corresponding to each of the at least one three-dimensional model; and using the virtual camera corresponding to each of the at least one three-dimensional model, based on the multiple transformation parameters, projecting at least one three-dimensional model into multiple target images.
[0098] In other words, by setting up a virtual camera, each 3D model can be projected into multiple target images according to the transformation parameters corresponding to each 3D model.
[0099] For example, by transforming the intrinsic and extrinsic parameters of a virtual camera, the position, angle, focal length, aperture, etc., of the camera can be adjusted, and a 3D model can be projected into multiple target images.
[0100] Since the original image consists of multiple layers, which are like film strips containing elements such as text or graphics, stacked one by one in sequence to form the final image effect, the occlusion relationship between adjacent layers may change during the transformation process of the 3D model, potentially leading to color distortion. To improve video quality and further ensure visual effects, some embodiments utilize multiple transformation parameters to transform at least one 3D model into multiple target images. This may further include: determining the multiple layers contained in the original image; determining the boundary region corresponding to one layer on another layer in two adjacent layers; using the boundary region corresponding to the original image as the boundary region corresponding to each of the multiple target images; and filling any boundary region in any target image with a target color.
[0101] The boundary region can refer to the area where one layer corresponds to another layer, creating occlusion on the other layer. Due to transformation processing of the 3D model, this boundary region may be exposed instead of occluded, thus requiring color filling to ensure display quality. The boundary region corresponding to each layer serves as the boundary region corresponding to the original image. Since the layers remain unchanged when generating multiple target images based on the original image, the boundary region corresponding to the original image is also the boundary region corresponding to each target image.
[0102] One approach is to fill the entire boundary region with the target color. This target color can be a pre-defined color. Alternatively, to improve the visual effect, the target color can be determined based on the color of the layer containing the boundary region. For example, it could be the blended value of the pixel colors in the layer containing the boundary region or the color with the largest proportion among the pixel colors. Another approach is to determine the target color based on the pixel colors of the surrounding areas corresponding to the boundary region. As another optional approach, filling any boundary region in any target image with the target color can be done by determining the target region that meets the filling requirements based on the transformation parameters corresponding to the target image, determining the target color based on the pixel colors of the surrounding areas corresponding to the target region, and then filling the target region with the target color.
[0103] The filling requirement can refer to the unobstructed area, or it can be determined based on the transformation parameters corresponding to the target object.
[0104] The target color can be determined using a color filling model based on the pixel colors of the surrounding areas corresponding to the target region. The color filling model can be pre-trained based on the pixel colors of sample regions and their corresponding sample target colors.
[0105] The surrounding area can also refer to the area within a certain distance from the target area, or the area formed by several pixels around the target area, etc.
[0106] This can be done by color filling each boundary region in each target image, or by color filling only the boundary regions of the target areas that require filling, as well as the target image itself.
[0107] In some embodiments, generating a target video based on multiple target images and material information may include: performing a composite processing on the material information and multiple target images; and generating a target video based on the composited multiple target images.
[0108] Specifically, generating a target video based on the multiple target images after synthesis processing can be achieved by determining the arrangement order of the multiple target images according to the arrangement order of at least one original image and the arrangement order of multiple transformation parameters; and by stitching the multiple target images together according to the arrangement order of the multiple target images to generate a target video.
[0109] Wherein, when the multiple transformation parameters are determined based on at least one camera movement effect, the order of the multiple transformation parameters can be determined according to the order of the at least one camera movement effect. The order of the at least one original image and the order of the at least one camera movement effect can be preset or determined according to user needs, for example, the image processing request may include the order of the at least one original image and the order of the at least one camera movement effect.
[0110] For ease of understanding, assume there are two original images, arranged in the order of original image A and original image B; and two camera movement effects, arranged in the order of effect A and effect B. Original image A, according to effect A, yields two target images, arranged in the order determined by their generation time: target image A1 and target image A2. Original image A, according to effect B, yields two target images, arranged in the order determined by their generation time: target image A3 and target image A4. Original image B, according to effect A, yields target image B1, and according to effect B, yields target image B2. Therefore, the final arrangement of the multiple target images could be: target image A1, target image A2; target image A3, target image A4; target image B1, target image B2. Of course, the above is merely an example illustrating one possible implementation, and this application is not limited to this.
[0111] As an optional approach, after generating the target video, the method may further include:
[0112] The target video is sent to the user's terminal for the user to output.
[0113] In some embodiments, a download prompt message can also be sent to the user terminal, so that the user terminal outputs the download prompt message at the same time as outputting the target video. In response to the download operation triggered by the download prompt message, the user terminal can also save the target video to a corresponding local file on the user terminal.
[0114] In addition, a notification message can be sent to the user's client, so that the user can output the target video along with the update notification message. In response to the update operation triggered by the update notification message, the user's client can also send an update request to the server. The server can then update the object description page based on this update request, using the target video.
[0115] In addition, a notification message can be sent to the user's device, allowing the user to view the target video while simultaneously receiving the notification. In response to a publishing operation triggered by this notification, the user's device can also send a publishing request to the server. Based on this request, the server can publish the target video to the object's promotion page. The target video and the object's description page can be linked, so that if a trigger operation targeting the target video is detected on the object's promotion page, the user can be redirected to the object's description page, facilitating interactive operations on the target object.
[0116] As an alternative approach, after generating the target video, the method may further include:
[0117] Use the target video to update the object description page.
[0118] In other words, after the server generates the target video, it can directly use the target video to update the object description page. Alternatively, it can send an update notification to the user's client, and after the user confirms the update, the client sends an update request, and the server then uses the target video to update the object description page. Updating the object description page using the target video can involve replacing an existing video on the object description page or adding the target video to the object description page.
[0119] As another optional approach, after generating the target video, the method may further include: establishing a link between the target video and the object description page; publishing the target video to the object promotion page, so that a trigger operation targeting the target video is detected on the object promotion page, and the user is redirected to the object details page based on the link.
[0120] In addition, the server can first send a publishing prompt message to the user, and after receiving the publishing request, it can then establish a link between the target video and the object description page, and publish the target video to the object promotion page, etc.
[0121] In a practical application, this embodiment of the application can be used in an online transaction scenario. In the online transaction scenario, the target object can refer to the target product provided by the online transaction system. The technical solution of this application will be described below using the target product as an example. Figure 5 The above is a flowchart of another embodiment of a video generation method provided by this application. The technical solution of this embodiment can be executed by a server, and the method may include the following steps:
[0122] 501: Obtain the product image of the target product.
[0123] As an alternative approach, one could receive a user's image processing request and obtain the product image included in that request.
[0124] As an alternative approach, one could receive a user's image processing request, determine the index information corresponding to the target product contained in the image processing request, and then, based on the index information, identify the product image containing the target product from the product description page corresponding to the target product.
[0125] In the two optional methods mentioned above, receiving the user's image processing request can be done by sending an image processing prompt message to the user terminal so that the user terminal can display the image processing prompt message on the display interface. The image processing request is sent in response to the image processing operation triggered by the user in response to the image processing prompt message.
[0126] 502: Construct a 3D model corresponding to the product image.
[0127] Constructing a 3D model corresponding to a product image can be done by using the pixel depth values of the product image and employing a 3D reconstruction model. This 3D reconstruction model is trained based on the pixel depth values of a sample image and the corresponding 3D model.
[0128] 503: Using multiple transformation parameters, a 3D model is transformed into multiple target images.
[0129] 504: Determine the material information that matches the target product.
[0130] 505: Generate target video based on source material information and multiple target images.
[0131] This application embodiment obtains multiple target images by reconstructing the product image in three dimensions and adjusting the three-dimensional model using transformation parameters. This results in the target object in the video synthesized from the multiple target images having dynamic effects. Furthermore, by combining the material information matching the target object, the final target video is generated, which improves the visual effect of the video and allows for a better expression of the target object.
[0132] It should be noted that, Figure 5 The illustrated embodiments and Figure 3 The difference in the illustrated embodiment is that the target object is specifically the target product; other identical or corresponding steps can be found in the preceding text. Figure 3 The embodiments shown will not be repeated here.
[0133] As an optional approach, after generating the target video, the method may further include:
[0134] Send the target video to the user's device for playback.
[0135] In addition, a download notification can be sent to the user's device, so that the user can view the target video while playing the video. In response to the download operation triggered by the notification, the user can also save the target video to a corresponding local file on the device.
[0136] In addition, update notifications can be sent to the user's device, allowing the user to view the target video while simultaneously receiving the update notification. In response to the update operation triggered by this notification, the user's device can also send an update request to the server. Based on this update request, the server can update the object description page using the target video.
[0137] In addition, a notification message can be sent to the user's device, allowing the user to view the target video while simultaneously receiving the notification. In response to the publishing action triggered by this notification, the user's device can also send a publishing request to the server. Based on this request, the server can publish the target video to the product promotion page. The target video and the product details page can be linked, so that if the product promotion page detects a trigger action on the target video, the user is redirected to the product details page, facilitating product purchases. This product promotion page, also known as a product aggregation page, is used for product introduction and promotion. High-quality video can attract user clicks on the target video, thereby increasing purchase rates and conversion rates.
[0138] As an alternative approach, after generating the target video, the method may also include updating the product details page using the target video.
[0139] In other words, after the server generates the target video, it uses the target video to update the product details page. Alternatively, it can send an update notification to the user and, upon receiving the update request, use the target video to update the product details page.
[0140] As another optional approach, after generating the target video, the method may further include: establishing a link between the target video and the product details page; publishing the target video to the product promotion page, so that a trigger operation targeting the target video is detected on the product promotion page, and the user is redirected to the product details page based on the link.
[0141] In addition, the server can first send a publishing prompt message to the user, and then, after receiving the publishing request, execute the link relationship between the target video and the product details page, and publish the target video to the product promotion page.
[0142] Figure 6 A flowchart illustrating an embodiment of an information display method provided in this application is provided. The technical solution of this embodiment can be executed by a user terminal, and the method may include the following steps:
[0143] 601: Provides a display interface.
[0144] 602: Displays image processing prompts on the display interface.
[0145] 603: In response to an image processing operation triggered by an image processing prompt, an image processing request is sent to the server. The image processing request is used by the server to determine at least one original image containing the target object, construct a 3D model corresponding to each of the at least one original image, transform the at least one 3D model into multiple target images using multiple transformation parameters, determine the material information matching the target object, and generate a target video based on the multiple target images and the material information.
[0146] The video generation method, in which the server determines at least one original image containing the target object based on the image processing request and then generates the target video, can be detailed above. Figure 3 The same or corresponding steps described in the illustrated embodiments will not be repeated here.
[0147] 604: Play the target video on the display interface.
[0148] The target video generated by the server is sent to the user terminal, so that the user terminal can play the video footage of the target video in the display interface, and if the target video contains audio data, the audio data can be played in conjunction with the audio playback component.
[0149] As an alternative approach, after playing the target video on the display interface, the method may further include: saving the target video to a corresponding local file on the user's device in response to a user's download operation.
[0150] This includes the ability to retrieve download notifications sent by the server and display them on the interface; the download operation can be triggered in response to these notifications.
[0151] As an alternative approach, after playing the target video on the display interface, the method may further include:
[0152] In response to a user's update action, an update request is sent to the server. The server can then use the target video to update the object description page based on the update request.
[0153] Optionally, the update notification information sent by the server can be obtained and displayed on the display interface. The update operation can be triggered in response to the update notification information.
[0154] As another alternative, after playing the target video on the display interface, the method may further include:
[0155] In response to a user's publishing action, a publishing request is sent to the server. The server can then publish the target video to the target promotion page based on the publishing request.
[0156] Optionally, the server can obtain the publishing prompt information sent by the server and display the publishing prompt information on the display interface. The publishing operation can be triggered in response to the publishing prompt information.
[0157] The specific execution operations on the server side can be found in the corresponding embodiments described above, and will not be repeated here.
[0158] Figure 7 This application provides a schematic diagram of the structure of a video generation apparatus according to one embodiment. The apparatus may include:
[0159] First acquisition module 701: used to acquire at least one original image containing the target object;
[0160] First 3D construction module 702: used to construct at least one 3D model corresponding to each of the original images;
[0161] First projection module 703: used to transform at least one three-dimensional model into multiple target images using multiple transformation parameters;
[0162] First Material Confirmation Module 704: Used to determine the material information that matches the target object;
[0163] First video generation module 705: Used to generate target videos based on multiple target images and material information.
[0164] The first acquisition module acquires at least one original image containing the target object, which may be achieved by receiving a user's image processing request and acquiring at least one original image included in the image processing request.
[0165] As an alternative approach, one could receive a user's image processing request, determine the index information corresponding to the target object contained in the request, and then, based on the index information, identify the original image containing the target object from the object description page corresponding to the target object.
[0166] The first projection module can transform at least one 3D model into multiple target images by using multiple transformation parameters. This can be achieved by determining multiple transformation parameters corresponding to at least one camera movement effect, and then using these multiple transformation parameters to transform at least one 3D model into multiple target images.
[0167] As an alternative approach, one could determine at least one virtual camera corresponding to a 3D model, and then use that virtual camera to project the 3D model into multiple target images according to multiple transformation parameters corresponding to the virtual camera.
[0168] The first video generation module can generate a target video based on multiple target images and material information by compositing the multiple target images with the material information, and then generating the target video based on the composited target images.
[0169] The material information may include at least one material image. Combining at least one material image with multiple target images may involve determining a one-to-one correspondence between at least one material image and at least one target image, and combining any one material image into its corresponding target image.
[0170] The first material confirmation module determines the one-to-one correspondence between at least one material image and at least one target image. This can be done by determining the at least one material image based on the object category to which the target object in the at least one target image belongs.
[0171] One method for compositing any source image into its corresponding target image is to determine the compositing region based on the object position of the target object in the target image, based on any source image and its corresponding target image, adjust the image size and compositing direction of the source image according to the compositing region, and then composite the adjusted source image into the compositing region according to the compositing direction.
[0172] Material information may also include text information. Combining text information with multiple target images can be done by combining the text information into at least one of the multiple target images.
[0173] The first material confirmation module determines the text information corresponding to the target object. This text information can be generated based on the object-related information of the target object, such as object description information, evaluation information, category information, etc.
[0174] The first video generation module is used to generate a target video based on multiple target images and material information. It can determine the arrangement order of multiple target images according to the arrangement order of at least one original image and the arrangement order of the transformation parameters corresponding to at least one 3D model, and then stitch the multiple target images together to generate the target video according to the arrangement order of the multiple target images.
[0175] In some embodiments, after generating a target video based on multiple target images, the target video can also be sent to the user terminal for the user terminal to output the target video.
[0176] Figure 7 The video generation device shown can be applied to e-commerce scenarios. In e-commerce scenarios, the first acquisition module can specifically acquire the product image of the target product, and the first 3D construction module can specifically construct the 3D model corresponding to the product image.
[0177] Figure 7 The video generation device can perform Figure 3 The implementation principle and technical effects of the video generation method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit of the information processing device in the above embodiments performs its operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.
[0178] Figure 8 This application provides a schematic diagram of the structure of an information display device according to one embodiment. The device may include:
[0179] First output module 801: Provides a display interface;
[0180] First display module 802: Displays image processing prompts on the display interface;
[0181] First request processing module 803: In response to the image processing operation triggered by the image processing prompt information, it sends an image processing request to the server. The image processing request is used by the server to determine the original image containing the target object.
[0182] Second display module 804: Plays the target video on the display interface.
[0183] In one possible design, Figure 8 The information display device can be implemented as the user terminal described above. Figure 8 The information display device shown can perform Figure 6 The implementation principle and technical effects of the information display method in the illustrated embodiment will not be repeated here. The specific methods by which each module and unit of the information display device in the above embodiments performs its operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.
[0184] This application also provides a computing device, such as... Figure 9 As shown, the computing device may include a storage component 901 and a processing component 902;
[0185] The storage component 901 stores one or more computer instructions, wherein the one or more computer instructions are invoked by the processing component to achieve, for example... Figure 2 or Figure 3 or Figure 5 The video generation method described in the illustrated embodiment.
[0186] Of course, the device may also include other components, such as input / output interfaces, display components, communication components, etc.
[0187] The processing components in this computing device are used to achieve, for example Figure 6 In the case of the information display method shown, the computing device may also include a display component to perform the corresponding display operation.
[0188] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc. Communication components are configured to facilitate wired or wireless communication between computing devices and other devices.
[0189] The processing component 902 may include one or more processors to execute computer instructions to complete all or part of the steps in the above-described method. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.
[0190] Storage component 901 is configured to store various types of data to support operations at the terminal. The storage component can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The display component can be an electroluminescent (EL) element, a liquid crystal display (LCD) or a microdisplay with a similar structure, or a retina-direct-view or similar laser-scanning display.
[0191] The display component can be an electroluminescent (EL) element, a liquid crystal display or a microdisplay with a similar structure, or a retina-direct display or a similar laser scanning display.
[0192] It should be noted that the above-mentioned computing devices implement Figure 2 or Figure 3 or Figure 5 In the case of the video generation method shown, it can be a physical device or an elastic computing host provided by a cloud computing platform. It can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. The aforementioned computing devices implement... Figure 6 In the case of the information display method shown, it can be specifically implemented as an electronic device. An electronic device can refer to a device used by a user that has the computing, internet access, and communication functions required by the user, such as a mobile phone, tablet computer, personal computer, wearable device, etc.
[0193] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 2 The video generation method of the embodiment shown or Figure 3 The video generation method of the embodiment shown or Figure 5 The video generation method of the embodiment shown or Figure 6 The information display method of the illustrated embodiment. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device.
[0194] This application also provides a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program can perform the above-described functions when executed by a computer. Figure 2 The video generation method of the embodiment shown or Figure 3 The video generation method of the embodiment shown or Figure 5 The video generation method of the embodiment shown or Figure 6 The information display method of the embodiment shown.
[0195] In such embodiments, the computer program may be downloaded and installed from a network, and / or installed from a removable medium. When the computer program is executed by a processor, it performs the various functions defined in the system of this application.
[0196] The processing components involved in the corresponding embodiments described above may include, for example, one or more processors to execute computer instructions to complete all or part of the steps in the methods described above. Alternatively, the processing components may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0197] Storage components are configured to store various types of data to support operations on the terminal. Storage components can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0198] The display component can be an electroluminescent (EL) element, a liquid crystal display or a microdisplay with a similar structure, or a retina-direct display or a similar laser scanning display.
[0199] Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0200] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0201] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A video generation method, characterized in that, include: Obtain at least one original image containing the target object; Construct three-dimensional models corresponding to each of the at least one original image; By using multiple transformation parameters, at least one 3D model can be transformed into multiple target images; Determine the material information that matches the target object; Based on the multiple target images and the material information, a target video is generated; The method further includes: Determine the multiple layers contained in the original image; Determine the boundary area between two adjacent layers, where one layer corresponds to the other layer; The boundary region corresponding to the original image is used as the boundary region corresponding to the multiple target images respectively; Fill any boundary region in any target image with the target color.
2. The method according to claim 1, characterized in that, Generating a target video based on the multiple target images and the material information includes: The material information is then combined with the multiple target images; A target video is generated based on multiple target images after synthesis processing.
3. The method according to claim 2, characterized in that, The material information includes at least one material image; the process of combining the material information with the plurality of target images includes: Determine the correspondence between the at least one source image and the at least one target image; Composite any source image into its corresponding target image.
4. The method according to claim 2, characterized in that, The material information includes text information; the process of combining the material information with the multiple target images includes: The text information is synthesized into at least one of the plurality of target images.
5. The method according to claim 3, characterized in that, The material information used to determine the target object includes: Based on the object category to which the target object belongs, at least one source image is determined.
6. The method according to claim 4, characterized in that, The material information used to determine the target object includes: Based on the object-related information of the target object, textual information of the target object is generated; the object-related information includes one or more of the following: object description information, object evaluation information, object category information, and object image.
7. The method according to claim 1, characterized in that, The process of transforming at least one 3D model into multiple target images using multiple transformation parameters includes: Determine multiple transformation parameters corresponding to at least one camera movement effect; Using the multiple transformation parameters, the at least one 3D model is transformed into multiple target images.
8. A video generation method, characterized in that, include: Obtain the product image of the target product; Construct a 3D model corresponding to the product image; The 3D model is transformed into multiple target images using multiple transformation parameters; Determine the material information matching the target product; Based on the material information and the multiple target images, a target video is generated; The method further includes: Determine the multiple layers contained in the product image; Determine the boundary area between two adjacent layers, where one layer corresponds to the other layer; The boundary region corresponding to the product image is used as the boundary region corresponding to the plurality of target images respectively; Fill any boundary region in any target image with the target color.
9. The method according to claim 8, characterized in that, The acquisition of the product image of the target product includes: Receives an image processing request from a user, the image processing request including the product image; or, The system receives an image processing request from a user, determines the index information corresponding to the target product included in the image processing request, and identifies a product image with the target product as the main image from the product details page of the target product based on the index information.
10. The method according to claim 8, characterized in that, The generated target video includes: Receive a user's request and update the product details page using the target video; or, Establish a link between the target video and the product details page, publish the target video to the product promotion page, and when a trigger operation targeting the target video is detected on the product promotion page, the user is redirected to the product details page based on the link.
11. An information display method, characterized in that, include: Provide a display interface; Image processing prompts are displayed on the display interface; In response to the image processing operation triggered by the image processing prompt information, an image processing request is sent to the server; the image processing request is used by the server to determine at least one original image containing the target object, construct three-dimensional models corresponding to the at least one original image respectively; and transform at least one three-dimensional model into multiple target images using multiple transformation parameters. The process involves: determining the material information matching the target object; generating a target video based on the multiple target images and the material information; wherein the image processing request is further used by the server to determine multiple layers contained in the original image; determining the boundary region corresponding to one layer on another layer among two adjacent layers; using the boundary region corresponding to the original image as the boundary region corresponding to each of the multiple target images; and filling any boundary region in any target image with a target color. Play the target video on the display interface.
12. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement the video generation method as described in any one of claims 1 to 7, or the video generation method as described in any one of claims 8 to 10, or the information display method as described in claim 11.
13. A computer storage medium, characterized in that, The device contains a computer program that, when executed by a computer, implements the video generation method as described in any one of claims 1 to 7, or the video generation method as described in any one of claims 8 to 10, or the information display method as described in claim 11.