Image processing method, apparatus, device, and storage medium
By identifying target objects in the video stream and displaying virtual models and controlling animated expressions, the problem of low integration tightness in existing special effects applications is solved, achieving high-tightness virtual information display and personalized effects.
Patent Information
- Application Number
- CN202210079134.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-24
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-01-24
AI Technical Summary
Existing special effects applications have a low degree of integration with the real environment and cannot meet users' personalized display needs for virtual information.
By identifying target objects in real-time video streams, determining their location information, and displaying virtual models on the target objects, the system loops the target audio to control the animated expressions of the virtual models.
It enables the display of virtual models in video streams in any real-world environment, improving the integration between the real environment and virtual information, enhancing the fun of virtual information display, and meeting users' personalized display needs.
Smart Images

Figure CN114419213B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of Internet application, and particularly relates to an image processing method and device, equipment and storage medium. BACKGROUND
[0002] With the continuous development of Internet technology, various interesting special effect applications appear on the network, and users can select corresponding special effect applications to shoot videos. However, the existing special effect applications have a single form and low tightness with the real environment, and cannot meet the personalized display requirements of virtual information. SUMMARY
[0003] In view of the technical problems of the traditional method, the embodiments of the present disclosure provide an image processing method, device, equipment and storage medium.
[0004] In a first aspect, the embodiments of the present disclosure provide an image processing method, comprising:
[0005] In response to the received service execution instruction, a target object in a real-time collected video stream is identified, and position information of the target object in the video stream is determined;
[0006] According to the position information, a virtual model is displayed on the target object in the video stream;
[0007] The target audio is played in a loop, and the virtual model displays a corresponding animation expression according to the target audio.
[0008] In a second aspect, the embodiments of the present disclosure provide an image processing device, comprising:
[0009] A determination module is configured to, in response to a received service execution instruction, identify a target object in a real-time collected video stream, and determine position information of the target object in the video stream;
[0010] A first display module is configured to, according to the position information, display a virtual model on the target object in the video stream;
[0011] A control module is configured to play target audio in a loop, and control the virtual model to display a corresponding animation expression according to the target audio.
[0012] In a third aspect, the embodiments of the present disclosure provide an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the image processing method provided by the first aspect of the embodiments of the present disclosure when executing the computer program.
[0013] In a fourth aspect, the present disclosure provides a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the image processing method provided in the first aspect of the present disclosure.
[0014] The technical solution provided by the embodiments of the present disclosure can identify the target object in the real-time collected video stream and determine the position information of the target object in the video stream in response to the received service execution instruction; the virtual model is displayed on the target object in the video stream according to the position information; the target audio is played in a loop, and the virtual model display controls the corresponding animation expression according to the target audio, thereby achieving the purpose of displaying the virtual model on the object in the video stream of any real environment, and the animation expression of the displayed virtual model can be controlled based on the played audio data, which improves the combination closeness of the real environment and the virtual information, enhances the interestingness of the virtual information display, and meets the personalized display needs of the user for the virtual information. BRIEF DESCRIPTION OF DRAWINGS
[0015] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent by describing in detail the following specific embodiments with reference to the attached drawings. Throughout the drawings, the same or similar reference numerals refer to the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.
[0016] Figure 1 A flowchart of an image processing method provided by the embodiments of the present disclosure is shown in FIG. 1;
[0017] Figure 2 A schematic diagram of a virtual model provided by the embodiments of the present disclosure is shown in FIG. 3;
[0018] Figure 3 A schematic diagram of an image processing result provided by the embodiments of the present disclosure is shown in FIG. 4;
[0019] Figure 4 A flowchart of a virtual model display process provided by the embodiments of the present disclosure is shown in FIG. 5;
[0020] Figure 5 Another flowchart of a virtual model display process provided by the embodiments of the present disclosure is shown in FIG. 6;
[0021] Figure 6 A structural schematic diagram of an image processing device provided by the embodiments of the present disclosure is shown in FIG. 7;
[0022] Figure 7 A structural schematic diagram of an electronic device provided by the embodiments of the present disclosure is shown in FIG. 8. DETAILED DESCRIPTION
[0023] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein, but rather, the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure and the embodiments are only for exemplary purposes and should not be used to limit the scope of protection of the present disclosure.
[0024] It should be understood that each step recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0025] The term "comprising" and variations thereof as used herein are open-ended, that is, "comprising but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.
[0026] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0027] It should be noted that the modification of "one" or "multiple" mentioned in the present disclosure is illustrative rather than limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0028] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.
[0029] To make the purposes, technical solutions and advantages of the present disclosure clearer, the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other arbitrarily without conflict.
[0030] Figure 1 A flowchart of an image processing method provided by an embodiment of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the method can include: Figure 1
[0031] S101, in response to the received service execution instruction, identifying a target object in the real-time collected video stream and determining position information of the target object in the video stream.
[0032] The video stream can be a video stream composed of multiple frames of original images about the real world collected in real time through the rear camera of the electronic device. The target object can be understood as an object existing in the video stream, such as a person and various objects in the collected real world.
[0033] After obtaining the service execution instruction, the electronic device can identify the target object in the real-time collected video stream through a preset detection algorithm, and locate the position information of the target object in the video stream. The position information can be the coordinate information of the pixel point of the target object. As an optional implementation, a business control can be displayed in the video stream interface, which is used to trigger the service execution instruction. The user can trigger the business control through touch or voice, etc. Thus, when the triggering operation of the business control is obtained, the electronic device identifies the video stream collected by the camera in real time to determine the target object existing in the video stream and the position information of the target object.
[0034] S102, according to the position information, displaying a virtual model on the target object in the video stream.
[0035] The virtual model can be a pre-prepared three-dimensional model, such as various virtual face images as shown in FIG. 1, etc. Of course, other types of virtual models can also be prepared according to actual needs, and the specific type and style of the virtual model are not limited in the embodiments of the present disclosure. Figure 2
[0036] After obtaining the position information of the target object, the electronic device can convert the position information to a three-dimensional space to obtain the spatial position information of the target object. Further, the electronic device can display the virtual model on the target object in the video stream based on the spatial position information. Displaying the virtual model on the target object in the video stream can be understood as superimposing the virtual model on the target object in the video stream and displaying the superimposed video stream. In an exemplary implementation, the virtual model can be rendered to a second layer, and the transparency of the area in the second layer except the virtual model is set to zero, i.e. the area in the second layer except the virtual model is a transparent area. The layer where the current frame image in the video stream is located is a first layer, and the second layer and the first layer are synthesized to display the virtual model on the target object in the video stream.
[0037] S103, playing the target audio cyclically and controlling the virtual model to display a corresponding animation expression according to the target audio.
[0038] The target audio can be a pre-made audio file. Meanwhile, an initial animation expression is made for the virtual model, and the made initial animation expression is adjusted by the waveform of the target audio to obtain a target animation expression. Since the target animation expression is adjusted by the waveform of the target audio, the target animation expression of the virtual model matches the target audio. In this way, when the target audio is played in a loop, the electronic device can control the virtual model to display the corresponding animation expression based on the audio data played by the target audio.
[0039] Specifically, the process of making the initial animation expression for the virtual model can be: creating a corresponding controller through the expression code of the Max script controller, and using the controller to drive the bones in the corresponding mesh of the virtual model to make the initial animation expression. After making the controller animation, the controller animation is collapsed, thereby completing the making of the initial animation expression of the virtual model.
[0040] Optionally, when there are multiple virtual models in the video stream, different virtual models can display different animation expressions during the playing of the target audio. To this end, the target audio can include multiple audio tracks, different audio tracks are used to control different virtual models, and the animation expressions of the virtual models corresponding to different audio tracks are different.
[0041] Therefore, the process of controlling the virtual model to display the corresponding animation expression according to the target audio in S103 can be: according to the currently played audio track data, controlling the virtual model corresponding to the audio track to display the corresponding animation expression.
[0042] For example, assuming that the target audio includes 4 audio tracks, and there are also 4 virtual models displayed in the video stream, when the 4 audio tracks are played together, the 4 virtual models in the video stream display different animation expressions at the same time, when a certain audio track data is played, the virtual model corresponding to the audio track in the video stream displays the corresponding animation expression, and the other virtual models can maintain the initial state expression, and when the corresponding audio track data is played, the other virtual models display the corresponding animation expression, thereby achieving the purpose of different virtual models displaying different animation expressions with the playing of the audio data.
[0043] Optionally, the subtitle information corresponding to the target audio can also be synchronously displayed in the interface of the video stream.
[0044] The display form of the subtitle information is diversified, that is, the subtitle information can be displayed according to any display parameter associated with a font. For example, the display parameter can be a font color, a font size, a text effect, a layout, a background color, and the like.
[0045] Optionally, a shooting control can also be displayed in the video stream interface, which is used to end the above image processing process. That is, when the shooting control is triggered, the image processing process is ended and the process of shooting a video or an image is entered.
[0046] As an optional implementation, in the process of real-time collection of the video stream by the rear camera of the electronic device, a business control and a shooting control are displayed in the video stream interface. After a triggering operation on the business control is obtained, a preset first animation sequence frame is played, and in the process of playing the first animation sequence frame, a preset second animation sequence frame can be further played. The first animation sequence frame and the second animation sequence frame are different. Optionally, the first animation sequence frame can be a full-screen scanning animation sequence frame, which is used to prompt that the real world is being scanned; and the second animation sequence frame can be a color ribbon animation sequence frame, so as to enrich the effect of picture display and relieve the anxiety of the user in the image processing process. The above animation sequence frames can be set based on actual needs, and can include only the first animation sequence frame, or the second animation sequence frame, or both the first animation sequence frame and the second animation sequence frame, the purpose of which is to enrich the effect of picture display, and the present embodiment is only an example and is not specifically limited.
[0047] In the process of playing the first animation sequence frame and the second animation sequence frame, the electronic device identifies a target object in the real-time collected video stream and determines position information of the target object in the video stream, and based on the position information, a virtual model is displayed on the target object in the video stream. In this way, after the first animation sequence frame and the second animation sequence frame are played, the electronic device can display an animation expression of the virtual model in the video stream. Further, the expression of the virtual model in the video stream can also be dynamically changed based on the played audio data. When the real scene changes, that is, the business execution instruction is obtained again, the electronic device can display the virtual model on a new target object identified in the video stream, and the expression of the virtual model is dynamically changed with the played target audio. As shown in FIG. 8, after the business execution instruction is triggered by the business control 31 (that is, the “scan” button in FIG. 8), the electronic device can display the virtual model on the electric fan in the video stream, and when the real scene changes, the scanning of the real scene is triggered again, at this time, the target object scanned can change, and the electronic device can also display the corresponding virtual model on the changed target object. After a triggering operation on the shooting control 32 is obtained, the business control 31 disappears, the process of shooting a video or an image is entered, and the image processing process is ended. Figure 3 Figure 3 In the process of real-time collection of the video stream by the rear camera of the electronic device, a business control and a shooting control are displayed in the video stream interface. After a triggering operation on the business control is obtained, a preset first animation sequence frame is played, and in the process of playing the first animation sequence frame, a preset second animation sequence frame can be further played. The first animation sequence frame and the second animation sequence frame are different. Optionally, the first animation sequence frame can be a full-screen scanning animation sequence frame, which is used to prompt that the real world is being scanned; and the second animation sequence frame can be a color ribbon animation sequence frame, so as to enrich the effect of picture display and relieve the anxiety of the user in the image processing process. The above animation sequence frames can be set based on actual needs, and can include only the first animation sequence frame, or the second animation sequence frame, or both the first animation sequence frame and the second animation sequence frame, the purpose of which is to enrich the effect of picture display, and the present embodiment is only an example and is not specifically limited.
[0048] The image processing method provided by the embodiments of the present disclosure can identify a target object in a real-time collected video stream and determine position information of the target object in the video stream in response to a received service execution instruction; a virtual model is displayed on the target object in the video stream according to the position information; a target audio is played in a loop, and a corresponding animation expression of the virtual model is controlled according to the target audio, thereby achieving the purpose of displaying a virtual model on an object in a video stream in any real environment, and the animation expression of the displayed virtual model can be controlled based on the played audio data, the closeness of the real environment and the virtual information is improved, the interestingness of the virtual information display is enhanced, and the personalized display requirement of the user for the virtual information is met.
[0049] In one embodiment, when it is identified that the video stream contains multiple target objects, the display of the virtual model can be performed according to the process described in the following embodiments. On the basis of the above-mentioned embodiments, optionally, as shown in Figure 4 The process of S102 can be as follows:
[0050] S401, determining a to-be-mounted object from the multiple target objects.
[0051] The object with the virtual model mounting capability can be referred to as the to-be-mounted object. Here, the mounting can be understood as the display. The video stream is identified to obtain multiple target objects. The sizes of the target objects are different, that is, the size of some target objects is small, and the virtual model is not suitable to be displayed on the target objects. Based on this, in order to improve the display effect of the virtual information in the video stream, the electronic device can select the object with the model mounting capability from the multiple target objects as the to-be-mounted object. For example, the object with the first size in the front of the first size sorting can be determined as the to-be-mounted object, or the object with the first size greater than a preset size can be determined as the to-be-mounted object.
[0052] S402, determining a target virtual model matched with each to-be-mounted object from a preset virtual model set according to the first size of each to-be-mounted object in the video stream.
[0053] The preset virtual model set includes a plurality of virtual models with different sizes. At this time, the electronic device can determine a target virtual model matching each to-be-mounted object from the preset virtual model set based on the first size of each to-be-mounted object. For example, the target virtual model matching each to-be-mounted object can be determined from the preset virtual model set based on the sorting result of the first size of the plurality of to-be-mounted objects. For example, if the first size of the to-be-mounted object is sorted in descending order, the corresponding target virtual model can be selected from the preset virtual model set in descending order of size, so that the target virtual model matching the to-be-mounted object with a larger first size is also larger, and the target virtual model matching the to-be-mounted object with a smaller first size is also smaller.
[0054] S403, according to the position information of each to-be-mounted object, respectively display each target virtual model on each to-be-mounted object in the video stream.
[0055] After determining the target virtual model matching each to-be-mounted object, the electronic device can display each target virtual model on each to-be-mounted object in the video stream based on the position information of each to-be-mounted object. That is, a smaller target virtual model is displayed on a to-be-mounted object with a smaller first size, and a larger target virtual model is displayed on a to-be-mounted object with a larger first size, thereby realizing accurate display of the virtual model.
[0056] Further, the electronic device can also perform real-time positioning on the to-be-mounted object to determine whether the position of the to-be-mounted object in the video stream changes; if so, the target virtual model is displayed on the to-be-mounted object in the video stream according to the changed position information.
[0057] Specifically, the electronic device can use a Simultaneous Localization and Mapping (SLAM) algorithm to track the position of the to-be-mounted object and obtain the position change of the to-be-mounted object in real time. After determining that the position of the to-be-mounted object in the video stream changes, the electronic device can adjust the display position of the target virtual model based on the changed position information, so that the target virtual model is stably displayed on the to-be-mounted object.
[0058] In this embodiment, the electronic device can select a target virtual model matched with the to-be-mounted object based on the size of the to-be-mounted object in the video stream, and display the target virtual model on the to-be-mounted object in the video stream, so that the displayed virtual information can be closely combined with the object in the real environment, further improving the close combination of the real environment and the virtual information, and improving the display effect of the picture. Moreover, the display position of the target virtual model can be adjusted based on the real-time positioning result of the to-be-mounted object, realizing the stability of the target virtual model display, and improving the image processing effect.
[0059] In one embodiment, when the position information of the to-be-mounted object in the video stream changes, the size of the to-be-mounted object in the video stream also changes, that is, the farther the electronic device is from the to-be-mounted object, the smaller the size of the to-be-mounted object in the video stream, and the closer the electronic device is to the to-be-mounted object, the larger the size of the to-be-mounted object in the video stream. Based on this situation, the method can further include: the electronic device obtains a second size of the to-be-mounted object in the video stream, and scales the target virtual model according to the second size.
[0060] The second size is different from the first size. That is, after the size of the identified to-be-mounted object changes, the electronic device can scale the target virtual model displayed on the to-be-mounted object based on the changed size (i.e., the second size). When the second size is larger than the first size, the target virtual model is enlarged according to a corresponding proportion, and when the second size is smaller than the first size, the target virtual model is reduced according to a corresponding proportion, so that the scaled target virtual model is adapted to the size of the to-be-mounted object. The adaptation here can be understood as that the size ratio of the adjusted target virtual model to the to-be-mounted object is equal to the preset ratio.
[0061] Optionally, the above scaling adjustment of the target virtual model according to the second size can be obtaining a size adjustment operation for the target virtual model; and scaling the target virtual model according to the size adjustment operation.
[0062] In the embodiments of the present disclosure, automatic adjustment of the size of the target virtual model is supported, and manual adjustment of the size of the target virtual model is also supported, realizing the interactivity of the user and the virtual information. Optionally, the target virtual model can support the touch operation of the user, that is, the user adjusts the size of the target virtual model through the corresponding touch operation. In this way, when the user's size enlargement adjustment operation for the target virtual model is obtained, the electronic device enlarges the target virtual model; and when the user's size reduction adjustment operation for the target virtual model is obtained, the electronic device reduces the target virtual model. The above enlargement ratio and reduction ratio can be determined based on the obtained touch operation.
[0063] In the embodiment, the electronic device can adjust the displayed target virtual model in real time based on the size change of the object to be mounted in the video stream, so that the target virtual model is adapted to the size of the object to be mounted, thereby enriching the display effect of virtual information and meeting the personalized display needs of users for virtual information. Moreover, the target virtual model can be adjusted based on the triggering operation of the user, thereby improving the interactivity between the user and the virtual information.
[0064] In actual application, in order to make the target virtual model better integrate with the object in the real world and further improve the combination density of virtual information and the real environment, on the basis of the above embodiment, optionally, as shown in Figure 5 The process of S102 can be:
[0065] S501, acquire the background material sphere corresponding to the video stream, the object material sphere corresponding to the virtual model, and the mask image of the virtual model.
[0066] The mask image includes the facial feature information and the local skin color information of the virtual model. The facial feature information mainly includes the eye, nose, mouth, and eyebrow information of the virtual model, and the local skin color information can be the blush information in the virtual model.
[0067] Material can be understood as the combination of material and texture, and the surface has specific visual properties. In short, it is what the object looks like. These visual properties refer to the color, texture, smoothness, transparency, reflectivity, refractivity, luminosity, etc. of the surface.
[0068] Different materials often have different visual properties, so when the virtual model is displayed on the target object, the materials of the two need to be fused to display a more realistic effect.
[0069] The background material sphere can reflect the background material of the video stream. In actual application, the corresponding background material sphere can be generated based on the background of the video stream. Optionally, the edge distortion processing can also be performed on the background material sphere. The object material sphere can reflect the material of the virtual model. In actual application, the map material matching the virtual model can be selected from the preset material map library, and the corresponding object material sphere can be generated based on the map material. The mask image can be understood as an image containing only the facial feature information and the local skin color information of the virtual model, that is, the virtual model is pre-processed to obtain the mask image of the virtual model.
[0070] S502, fuse the background material sphere and the object material sphere according to the mask image to obtain a fused virtual model.
[0071] After obtaining the object material ball corresponding to the virtual model, the background material corresponding to the video stream, and the mask map of the virtual model, the background material ball and the object material ball are fused based on the mask map, so as to obtain a fused virtual model. As an optional implementation, the mask map can include three channels, namely a G (Green) channel, a B (Blue) channel, and an R (Red) channel. The electronic device can fuse the background material ball and the object material ball based on the G channel and the B channel of the mask map to obtain the fused virtual model. Specifically, the background material ball and the object material ball can be controlled by a weight between 0 and 1 based on the G channel and the B channel of the mask map, so as to fuse a new material ball. The new material ball can be understood as the fused virtual model.
[0072] S503, performing edge feathering processing on the fused virtual model to obtain a feathered virtual model.
[0073] The electronic device can add highlight processing to the fused virtual model, and perform edge feathering processing on the fused virtual model after the highlight processing by using an edge feathering algorithm, so as to obtain the feathered virtual model.
[0074] As an optional implementation, the process of S503 can be:
[0075] S5031, obtaining edge light information of the fused virtual model.
[0076] S5032, determining a feathering range of the fused virtual model based on the R channel of the mask map and the edge light information.
[0077] The R channel of the mask map and the obtained edge light information are multiplied, so as to obtain the feathering range of the fused virtual model.
[0078] S5033, performing feathering on the fused virtual model based on the feathering range to obtain the feathered virtual model.
[0079] The obtained feathering range is applied to a transparent channel A of the fused virtual model, so as to obtain an edge feathering effect, that is, the feathered virtual model.
[0080] S504, displaying the feathered virtual model on the target object in the video stream based on the position information.
[0081] In this embodiment, the background material ball corresponding to the video stream and the object material ball corresponding to the virtual model can be fused and edge feathering processed based on the mask map of the virtual model, so that the virtual model can be more naturally fused in the background of the video stream and has clear facial feature information, thereby improving the realism of virtual information display.
[0082] Figure 6 A structural schematic diagram of an image processing apparatus provided by the embodiments of the present disclosure is shown in FIG. 1. As shown in the figure, the apparatus can include a determination module 601, a first display module 602, and a control module 603. Figure 6
[0083] Specifically, the determination module 601 is configured to identify a target object in a real-time collected video stream and determine position information of the target object in the video stream in response to a received service execution instruction.
[0084] The first display module 602 is configured to display a virtual model on the target object in the video stream according to the position information.
[0085] The control module 603 is configured to cyclically play a target audio and control the virtual model to display a corresponding animation expression according to the target audio.
[0086] The image processing apparatus provided by the embodiments of the present disclosure identifies a target object in a real-time collected video stream and determines position information of the target object in the video stream in response to a received service execution instruction, displays a virtual model on the target object in the video stream according to the position information, and cyclically plays a target audio and controls the virtual model to display a corresponding animation expression according to the target audio, thereby achieving the purpose of displaying a virtual model on an object in a video stream of an arbitrary real environment and improving the closeness of the real environment and the virtual information, enhancing the interest of the virtual information display, and meeting the personalized display requirements of the user for the virtual information.
[0087] Optionally, the target audio includes multiple audio tracks, different audio tracks are used to control different virtual models, and the animation expressions of the virtual models corresponding to different audio tracks are different.
[0088] The control module 603 is specifically configured to control the virtual model corresponding to the audio track to display a corresponding animation expression according to currently played audio track data.
[0089] On the basis of the above-mentioned embodiments, optionally, when it is determined that multiple target objects are included, the first display module 602 includes a first determination unit, a second determination unit, and a first display unit.
[0090] Specifically, the first determination unit is configured to determine a to-be-mounted object from the multiple target objects.
[0091] The second determination unit is configured to determine a target virtual model matching each to-be-mounted object from a preset virtual model set according to a first size of each to-be-mounted object in the video stream, wherein the preset virtual model set includes multiple virtual models of different sizes.
[0092] The first display unit is configured to display each target virtual model on each to-be-mounted object in the video stream according to position information of each to-be-mounted object.
[0093] On the basis of the above embodiment, the first display module 602 can further include a positioning unit and a second display unit.
[0094] Specifically, the positioning unit is configured to perform real-time positioning on the to-be-mounted object to determine whether a position of the to-be-mounted object in the video stream changes.
[0095] The second display unit is configured to display the target virtual model on the to-be-mounted object in the video stream according to changed position information when it is determined that the position of the to-be-mounted object in the video stream changes.
[0096] On the basis of the above embodiment, the first display module 602 can further include a first acquisition unit and an adjustment unit.
[0097] Specifically, the first acquisition unit is configured to acquire a second size of the to-be-mounted object in the video stream when position information of the to-be-mounted object in the video stream changes, wherein the second size is different from the first size.
[0098] The adjustment unit is configured to perform scaling adjustment on the target virtual model according to the second size.
[0099] On the basis of the above embodiment, the adjustment unit is specifically configured to acquire a size adjustment operation for the target virtual model, and perform scaling adjustment on the target virtual model according to the size adjustment operation.
[0100] On the basis of the above embodiment, the first display module 602 can further include a second acquisition unit, a fusion unit, a feathering unit, and a third display unit.
[0101] Specifically, the second acquisition unit is configured to acquire a background material sphere corresponding to the video stream, an object material sphere corresponding to the virtual model, and a mask image of the virtual model, wherein the mask image includes five facial information and local skin color information of the virtual model.
[0102] The fusion unit is configured to fuse the background material sphere and the object material sphere according to the mask image to obtain a fused virtual model.
[0103] The feathering unit is configured to perform edge feathering processing on the fused virtual model to obtain a feathered virtual model.
[0104] The third display unit is configured to display the feathered virtual model on the target object in the video stream according to the position information.
[0105] On the basis of the above-mentioned embodiments, optionally, the fusing unit is specifically configured to fuse the background material ball and the object material ball according to the G channel and the B channel of the mask image, to obtain a fused virtual model.
[0106] On the basis of the above-mentioned embodiments, optionally, the feathering unit is specifically configured to acquire edge light information of the fused virtual model; determine a feathering range of the fused virtual model according to the R channel of the mask image and the edge light information; and feather the fused virtual model according to the feathering range, to obtain a feathered virtual model.
[0107] On the basis of the above-mentioned embodiments, optionally, the device further comprises a second display module.
[0108] Specifically, the second display module is configured to synchronously display subtitle information corresponding to the target audio in an interface of the video stream.
[0109] On the basis of the above-mentioned embodiments, optionally, the device further comprises a third display module.
[0110] Specifically, the third display module is configured to display a service control and a shooting control in the interface of the video stream; wherein the service control is configured to trigger the service execution instruction, and the shooting control is configured to end the image processing procedure.
[0111] Reference will now be made to the following description Figure 7 which shows a structural schematic diagram of an electronic device 700 suitable for implementing the embodiments of the present disclosure. The electronic device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 7 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0112] As Figure 7As shown, the electronic device 700 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 702 or loaded into a random access memory (RAM) 703 from a storage device 706. Various programs and data required for the operation of the electronic device 700 are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0113] Generally, the following devices can be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 706 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 709. The communication devices 709 can allow the electronic device 700 to communicate wirelessly or wired with other devices to exchange data. Although Figure 7 The electronic device 700 is shown with various devices, but it should be understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present.
[0114] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 709, or installed from the storage devices 706, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-described functions defined in the methods of embodiments of the present disclosure are performed.
[0115] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.
[0116] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0117] The computer-readable medium described above can be included in the electronic device described above; or can be separate from the electronic device and not assembled in the electronic device.
[0118] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: acquire at least two Internet protocol addresses; send a node evaluation request including the at least two Internet protocol addresses to a node evaluation device, wherein the node evaluation device selects an Internet protocol address from the at least two Internet protocol addresses and returns; receive the Internet protocol address returned by the node evaluation device; and wherein the acquired Internet protocol address indicates an edge node in a content distribution network.
[0119] Alternatively, the computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, the electronic device is caused to: receive a node evaluation request including at least two Internet protocol addresses; select an Internet protocol address from the at least two Internet protocol addresses; return the selected Internet protocol address; and wherein the received Internet protocol address indicates an edge node in a content distribution network.
[0120] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages, including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0121] The computer program product of the first aspect can include one or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform the operations of the first aspect. The one or more non-transitory computer-readable media can include, for example, magnetic media such as one or more magnetic disks, magnetic tapes or cassettes; optical media such as one or more compact discs, optical discs or Blu-ray discs; magneto-optical media such as one or more floptical discs; solid state media such as one or more solid state drives or other flash memory arrays; or any suitable combination of these. The one or more non-transitory computer-readable media can be encoded with instructions that, when executed, cause one or more processors to perform the operations of the first aspect.
[0122] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware, or by a combination of software and hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.
[0123] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0124] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0125] In one embodiment, an electronic device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program:
[0126] In response to the received service execution instruction, a target object in a real-time collected video stream is identified, and position information of the target object in the video stream is determined;
[0127] According to the position information, a virtual model is displayed on the target object in the video stream;
[0128] The target audio is played in a loop, and a corresponding animation expression of the virtual model display is controlled according to the target audio.
[0129] In one embodiment, a computer readable storage medium is also provided, which stores a computer program, and the computer program is executed by a processor to implement the following steps:
[0130] In response to the received service execution instruction, a target object in a real-time collected video stream is identified, and position information of the target object in the video stream is determined;
[0131] According to the position information, a virtual model is displayed on the target object in the video stream;
[0132] The target audio is played in a loop, and a corresponding animation expression of the virtual model display is controlled according to the target audio.
[0133] The image processing device, device and storage medium provided in the above embodiments can execute the image processing method provided by any embodiment of the present application, and have the corresponding function modules and beneficial effects of executing the method. Technical details not described in detail in the above embodiments can be referred to the image processing method provided by any embodiment of the present application.
[0134] According to one or more embodiments of the present disclosure, an image processing method is provided, comprising:
[0135] In response to the received service execution instruction, a target object in a real-time collected video stream is identified, and position information of the target object in the video stream is determined;
[0136] According to the position information, a virtual model is displayed on the target object in the video stream;
[0137] The target audio is played in a loop, and a corresponding animation expression of the virtual model display is controlled according to the target audio.
[0138] According to one or more embodiments of the present disclosure, the image processing method is provided as above, and further includes that the target audio includes a plurality of audio tracks, different audio tracks are used to control different virtual models, and different audio tracks correspond to different animation expressions of the virtual models; and according to current played audio track data, a virtual model corresponding to the audio track is controlled to display a corresponding animation expression.
[0139] According to one or more embodiments of the present disclosure, the image processing method is provided as above, and further includes that when a plurality of target objects are determined, a to-be-mounted object is determined from the plurality of target objects; according to a first size of each to-be-mounted object in the video stream, a target virtual model matching each to-be-mounted object is determined from a preset virtual model set; wherein the preset virtual model set includes a plurality of virtual models with different sizes; and according to position information of each to-be-mounted object, each target virtual model is correspondingly displayed on each to-be-mounted object in the video stream.
[0140] According to one or more embodiments of the present disclosure, the image processing method is provided as above, and further includes that the to-be-mounted object is positioned in real time to determine whether a position of the to-be-mounted object in the video stream changes; if yes, the target virtual model is correspondingly displayed on the to-be-mounted object in the video stream according to changed position information.
[0141] According to one or more embodiments of the present disclosure, the image processing method is provided as above, and further includes that a second size of the to-be-mounted object in the video stream is obtained; wherein the second size is different from the first size; and the target virtual model is scaled and adjusted according to the second size.
[0142] According to one or more embodiments of the present disclosure, the image processing method is provided as above, and further includes that a size adjustment operation for the target virtual model is obtained; and the target virtual model is scaled and adjusted according to the size adjustment operation.
[0143] According to one or more embodiments of the present disclosure, the image processing method is provided as above, and further includes that a background material sphere corresponding to the video stream, an object material sphere corresponding to the virtual model, and a mask map of the virtual model are obtained; wherein the mask map includes five facial feature information and local skin color information of the virtual model; the background material sphere and the object material sphere are fused according to the mask map to obtain a fused virtual model; the fused virtual model is subjected to edge feathering processing to obtain a feathered virtual model; and the feathered virtual model is displayed on the target object in the video stream according to the position information.
[0144] According to one or more embodiments of the present disclosure, the image processing method as above is provided, and further includes: fusing the background material ball and the object material ball according to a G channel and a B channel of the mask image, to obtain a fused virtual model.
[0145] According to one or more embodiments of the present disclosure, the image processing method as above is provided, and further includes: obtaining edge light information of the fused virtual model; determining a feathering range of the fused virtual model according to an R channel of the mask image and the edge light information; and feathering the fused virtual model according to the feathering range, to obtain a feathered virtual model.
[0146] According to one or more embodiments of the present disclosure, the image processing method as above is provided, and further includes: synchronously displaying subtitle information corresponding to the target audio in an interface of the video stream.
[0147] According to one or more embodiments of the present disclosure, the image processing method as above is provided, and further includes: displaying a business control and a shooting control in the video stream interface; wherein the business control is used to trigger the business execution instruction, and the shooting control is used to end the image processing process.
[0148] The above description is merely preferred embodiments of the present disclosure and a description of the principles of the technology applied. Those skilled in the art should understand that the disclosed range of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0149] In addition, although each operation is depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.
[0150] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. An image processing method, characterized by, The method comprises the following steps: in response to the received service execution instruction, identifying a target object in a real-time collected video stream and determining position information of the target object in the video stream; according to the position information, displaying a virtual model on the target object in the video stream; cyclically playing a target audio and controlling the virtual model to display a corresponding animation expression according to the target audio; wherein the animation expression of the virtual model dynamically changes with the played target audio; the target audio comprises multiple audio tracks, different audio tracks are used to control different virtual models, and the animation expressions of the virtual models corresponding to different audio tracks are different; controlling the virtual model corresponding to the audio track to display the corresponding animation expression according to the currently played audio track data. when it is determined that multiple target objects are included, the step of displaying a virtual model on the target object in the video stream according to the position information comprises the following steps:
2. The method of claim 1, wherein, determining a to-be-mounted object from the multiple target objects; determining a target virtual model matched with each to-be-mounted object from a preset virtual model set according to a first size of each to-be-mounted object in the video stream; wherein the preset virtual model set comprises multiple virtual models with different sizes; respectively displaying each target virtual model on each to-be-mounted object in the video stream according to the position information of each to-be-mounted object. The method further comprises the following steps:
3. The method of claim 2, wherein, real-time positioning the to-be-mounted object to determine whether the position of the to-be-mounted object in the video stream changes; if yes, displaying the target virtual model on the to-be-mounted object in the video stream according to the changed position information. when the position information of the to-be-mounted object in the video stream changes, the method further comprises the following steps:
4. The method of claim 3, wherein, obtaining a second size of the to-be-mounted object in the video stream; wherein the second size is different from the first size; scaling and adjusting the target virtual model according to the second size. the step of scaling and adjusting the target virtual model according to the second size comprises the following steps:
5. The method of claim 4, wherein, obtaining a size adjustment operation for the target virtual model; scaling and adjusting the target virtual model according to the size adjustment operation. the step of displaying a virtual model on the target object in the video stream according to the position information comprises the following steps:
6. The method according to any one of claims 1 to 5, characterized in that, obtaining a background material sphere corresponding to the video stream, an object material sphere corresponding to the virtual model, and a mask map of the virtual model; wherein the mask map comprises five facial feature information and local skin color information of the virtual model; fusing the background material sphere and the object material sphere according to the mask map to obtain a fused virtual model; performing edge feathering processing on the fused virtual model to obtain a feathered virtual model; displaying the feathered virtual model on the target object in the video stream according to the position information. the step of fusing the background material sphere and the object material sphere according to the mask map to obtain a fused virtual model comprises the following steps:
7. The method of claim 6, wherein, Fusing the background material ball and the object material ball according to the G channel and the B channel of the mask map to obtain a fused virtual model.
8. The method of claim 6, wherein, The edge feathering processing on the fused virtual model to obtain a feathered virtual model comprises: Obtaining edge light information of the fused virtual model; Determining a feathering range of the fused virtual model according to the R channel of the mask map and the edge light information; Feathering the fused virtual model according to the feathering range to obtain a feathered virtual model.
9. The method according to any one of claims 1 to 5, characterized in that, Further comprising: Synchronously displaying subtitle information corresponding to the target audio in the interface of the video stream.
10. The method according to any one of claims 1 to 5, characterized in that, Further comprising: Displaying a business control and a shooting control in the interface of the video stream; wherein the business control is used to trigger the business execution instruction, and the shooting control is used to end the process corresponding to the image processing method.
11. An image processing apparatus characterized by comprising: Comprise: A determination module, configured to identify a target object in a real-time collected video stream and determine position information of the target object in the video stream in response to a received business execution instruction; A first display module, configured to display a virtual model on the target object in the video stream according to the position information; A control module, configured to cyclically play a target audio and control the virtual model to display a corresponding animation expression according to the target audio; wherein the animation expression of the virtual model dynamically changes with the played target audio; The target audio comprises multiple audio tracks, different audio tracks are used to control different virtual models, and the animation expressions of the virtual models corresponding to different audio tracks are different; The control module is specifically configured to control the virtual model corresponding to the audio track to display a corresponding animation expression according to currently played audio track data.
12. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 10.
13. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 10.
Citation Information
Patent Citations
Man-machine interaction method and device, mobile terminal and computer readable storage medium
CN110955332A