Video image processing method and device, electronic equipment and storage medium
By determining the attribute information and baseline display information of the object to be processed in video image processing, and adjusting the target display information of the mounted material, the problem of 2D algorithms being unable to recognize 3D information and the high performance of 3D algorithms are solved. This achieves the 3D display effect of the mounted material, improves the following effect of special effects and the universality of the equipment.
Patent Information
- Application Number
- CN202111522826.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2041-12-13
AI Technical Summary
In existing technologies, although 2D algorithms are accurate in recognizing key points of limbs, they cannot obtain three-dimensional information, resulting in poor special effects. While 3D algorithms can recognize three-dimensional information, they consume a lot of resources, leading to high equipment performance requirements and poor versatility.
By determining the attribute information of the object to be processed in the video frame to be processed, the baseline display information is determined based on the attribute information, and the target display information of the mounted material is adjusted accordingly to achieve the three-dimensional display of the mounted material.
It achieves the acquisition of 3D information of key limb points without increasing equipment performance consumption, thereby achieving a 3D display effect for mounted materials, improving the following effect of special effects and the universality of equipment.
Smart Images

Figure CN114202617B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to a video image processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the increasing popularity of short videos, more and more users are shooting videos using their devices. To further enhance the entertainment value of the videos, special effects are often added to the users in the videos.
[0003] In specific scenarios, added special effects can be located using corresponding limb key points. Current methods for determining limb key points primarily employ 2D or 3D algorithms. Using 3D algorithms to determine limb key points is relatively performance-intensive, placing high demands on device performance. 2D algorithms, compared to 3D algorithms, consume less power and accurately determine the main body's key points, but they cannot obtain the three-dimensional information of the limb key points, leading to poor effect tracking. Summary of the Invention
[0004] This disclosure provides a video image processing method, apparatus, electronic device, and storage medium to achieve the effect of three-dimensional display of mounted materials.
[0005] In a first aspect, embodiments of this disclosure provide a video image processing method, the method comprising:
[0006] Determine the attribute information of the object to be processed in the video frame to be processed;
[0007] Based on the attribute information, determine the baseline display information of the object to be processed in the video frame to be processed;
[0008] Based on the reference display information, the target display information of the material mounted in the video frame to be processed is adjusted to obtain the target video frame corresponding to the video frame to be processed.
[0009] Secondly, embodiments of this disclosure also provide a video image processing apparatus, the apparatus comprising:
[0010] The attribute information determination module is used to determine the attribute information of the object to be processed in the video frame to be processed;
[0011] A reference display information determination module is used to determine the reference display information of the object to be processed in the video frame to be processed based on the attribute information.
[0012] The target video frame determination module is used to adjust the target display information of the attached material in the video frame to be processed based on the reference display information, so as to obtain the target video frame corresponding to the video frame to be processed.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the video image processing method as described in any of the embodiments of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video image processing method as described in any of the embodiments of this disclosure.
[0018] The technical solution of this disclosure, by determining the attribute information of the object to be processed in the video frame to be processed; determining the reference display information of the object to be processed in the video frame to be processed based on the attribute information; and adjusting the target display information of the mounted material in the video frame to be processed based on the reference display information, to obtain a target video frame corresponding to the video frame to be processed, solves the problem in the prior art that when using 2D algorithms to identify limb key points, although the identified limb key points are relatively accurate, the three-dimensional information of the limb key points cannot be obtained, resulting in poor tracking effect; when using 3D recognition algorithms to identify key points, although the three-dimensional information of limb key points can be identified, there is a high performance consumption, which puts high demands on the performance of terminal devices and results in poor universality. This solution realizes that the target display information of the mounted material can be determined based on the limb key information of the object to be processed in the video frame to be processed, and then based on the limb key point information, thereby obtaining the effect of three-dimensional display of the mounted material. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 This is a schematic flowchart of a video image processing method provided in Embodiment 1 of this disclosure;
[0021] Figure 2 This is a schematic diagram illustrating the determination of attribute information of an object to be processed, provided in Embodiment 1 of this disclosure;
[0022] Figure 3 This is a schematic diagram illustrating the determination of attribute information of an object to be processed, provided in Embodiment 1 of this disclosure;
[0023] Figure 4 This is a schematic diagram illustrating the determination of attribute information of an object to be processed, provided in Embodiment 1 of this disclosure;
[0024] Figure 5 This is a schematic diagram of a video image processing apparatus provided in Embodiment 2 of this disclosure;
[0025] Figure 6 This is a schematic diagram of an electronic device structure provided in Embodiment 3 of this disclosure. Detailed Implementation
[0026] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0027] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0028] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0029] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should also be noted that the modifications of "a" and "a plurality of" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0030] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0031] Before introducing this technical solution, an illustrative application scenario can be provided. This disclosed technical solution can be applied to any scene requiring special effects, such as during video recording, i.e., while recording and playing back simultaneously. The recorded video frames can be uploaded to the server, where the server can execute this technical solution to process the special effects. Alternatively, after video recording is complete, corresponding special effects can be added to each video frame. In this technical solution, the added special effects can be any kind of effect.
[0032] It should be noted that this technical solution can be implemented by the server, the client, or a combination of both. For example, the client can capture video frames and process them to add effects; or the captured video frames can be uploaded to the server, processed, and then sent back to the client for display.
[0033] Example 1
[0034] Figure 1 This is a schematic flowchart of a video image processing method provided in Embodiment 1 of this disclosure. This embodiment of the disclosure is applicable to any scenario of special effects display or special effects processing supported by the Internet, and is used to adjust the size of special effects in the video frame to be processed in order to achieve a three-dimensional display of special effects. The method can be executed by a video image processing device, which can be implemented in the form of software and / or hardware, or optionally by an electronic device, such as a mobile terminal, a PC, or a server.
[0035] like Figure 1 As shown, the method includes:
[0036] S110. Determine the attribute information of the object to be processed in the video frame to be processed.
[0037] It should be noted that this typically involves adding appropriate special effects to a target subject in the video. Consequently, each video frame may or may not include the target subject. If the target subject is included, the special effects added to it can be processed based on this technical solution.
[0038] The target subject can be the object to be processed. The object to be processed can be a person or an object, and its specific content matches the pre-set parameters. Optionally, if the pre-set parameters are for processing a person in the video frame to be processed, then the object to be processed can be a person; correspondingly, the object to be processed can also be an object. Attribute information can be the object's own characteristic information, such as the display size information of the object to be processed.
[0039] Specifically, users can capture a target video including the object to be processed and upload it to the target client. Upon receiving the target video, the target client can add corresponding effects to each frame of the video to be processed. Simultaneously, it can obtain the attribute information of the object to be processed, and adjust the display information of the object in the video frames based on this attribute information, thereby achieving a 3D display effect.
[0040] In this embodiment, determining the attribute information of the object to be processed in the video frame to be processed can be: based on a 2D point recognition algorithm, determining at least two points to be processed of the object to be processed in the video frame to be processed; determining the coordinate information to be processed of the at least one point to be processed, and using the coordinate information to be processed as the attribute information.
[0041] The 2D point recognition algorithm is used to identify the limb key points of the object to be processed. This algorithm identifies limb key points relatively accurately, and consequently, the determination of the display information for the attached material based on these accurately identified limb key points is also relatively accurate. At least two points to be processed correspond to the limb key points of the object to be processed. Limb key points can be shoulder key points, hip key points, and neck key points; correspondingly, the points to be processed can be shoulder points, hip points, and neck points. See [link to relevant documentation]. Figure 2 Each point has corresponding coordinates in the video frame to be processed. These coordinates can be used as the coordinate information to be processed. Optionally, the coordinate information to be processed can be represented by (u, v). The coordinate information of the point to be processed is used as the attribute information.
[0042] In this embodiment, determining the attribute information of the object to be processed in the video frame to be processed can also be: determining the bounding box information including the object to be processed in the video frame to be processed, and using the bounding box information as the attribute information.
[0043] The bounding box can be a rectangle whose edges are tangent to the edges of the object being processed. The bounding box can be represented by the coordinates of its four vertices, which can then be used as attribute information for the bounding box.
[0044] For example, when it is determined that the video frame to be processed includes the object to be processed, a rectangular bounding box that surrounds the object and is tangent to the edge line of the object can be determined based on the pixel coordinates of the object's edge. See [link to relevant documentation]. Figure 3 The pixel coordinates of the four vertices of the rectangular bounding box are used as the attribute information of the bounding box.
[0045] S120. Based on the attribute information, determine the reference display information of the object to be processed in the video frame to be processed.
[0046] The reference display information can be the display information of the object to be processed in the video frame to be processed. For example, the reference display information can be the display size, display ratio, or display angle of the object to be processed in the video frame to be processed.
[0047] Specifically, the attribute information of the attribute to be processed can be used as the baseline display information of the object to be processed. Alternatively, the attribute information can be further processed to determine the baseline display information. In other words, the baseline display information is the relative display information of the object to be processed within the video frame to be processed. The advantage of determining the baseline display information is that the special effects display information in the video frame to be processed can be adjusted based on this display information, thereby achieving a three-dimensional display effect for the special effects.
[0048] In this embodiment, determining the reference display information of the object to be processed in the video frame to be processed based on the attribute information includes: determining at least three width information associated with the object to be processed based on the coordinate information to be processed; and determining the reference display information of the video frame to be processed based on the at least three width information and the corresponding preset reference value.
[0049] Among these, at least three width parameters can be shoulder width, upper body length, and hip width. Shoulder width is determined based on the coordinates of the shoulder joint points to be processed. Upper body width is determined based on the ordinates of the hip and neck keypoints. Hip width is determined based on the coordinates of the hip joint points to be processed. The preset baseline values are standard proportional values for shoulder width, upper body length, and hip width. Based on these standard proportional values and at least three width parameters, the baseline display information for the video frame to be processed can be determined.
[0050] Specifically, based on the coordinate information of each joint point to be processed, the shoulder width, upper body length, and hip width can be determined. Based on these three values, a ratio value can be determined. This ratio value is compared with the standard ratio value in the preset reference value to determine the reference display information of the video frame to be processed. Based on the reference display information, the size information of the special effects in the video frame to be processed can be determined, thereby achieving the effect of three-dimensional display of special effects.
[0051] For example, based on the coordinate information of each joint point to be processed, the shoulder width X, upper body length Y, and hip width Z can be determined. A standard reference ratio for these three lengths is set (shoulder width: upper body length: hip width = x:y:z). Then, the three width values are converted into corresponding standard reference ratios; optionally, these three are scaled according to the ratios. The maximum value of the ratios is obtained, and this maximum value is compared with the set standard reference value to enlarge or reduce the special effects material in the video frame to be processed according to this ratio. In this embodiment, the purpose of determining the three length values is to reduce the problem of large changes in length information caused by the body rotation of the object to be processed, that is, to achieve the effect that best matches the actual effect.
[0052] In other words, the reference display information of the video frame to be processed is determined based on the at least three width information values and the corresponding preset reference values. This can be achieved by determining the maximum ratio based on the ratios of the at least three width information values, and then comparing the maximum ratio with the preset standard reference value to determine the reference display information. The reference display information can be the scaling ratio of the special effect.
[0053] In this embodiment, if the attribute information is bounding box information, then determining the reference display information of the object to be processed in the video frame to be processed based on the attribute information can be done by: determining the reference display information based on the bounding box information in the attribute information and the page size information of the display page to which the video frame to be processed belongs.
[0054] Specifically, the size of the bounding box can be determined based on the coordinates of its four vertices, optionally including its length and width. Simultaneously, the page size information of the displayed page when the video frame is played can be obtained. This page size information includes the page length and page width. Based on the bounding box's length and width, its area can be determined; correspondingly, based on the page length and width, the page display area can be determined. By calculating the ratio of the bounding box area to the page display area, the baseline display information can be determined.
[0055] In this embodiment, if the attribute information is bounding box information, then determining the reference display information of the object to be processed in the video frame to be processed based on the attribute information can also be: determining the proportion information of the object to be processed in the video to be processed based on a predetermined near plane and the bounding box information; wherein, the near plane is a plane determined based on when the object to be processed fills the display page to which the video frame to be processed belongs; and the reference display information is determined based on the distance information of the near plane from the virtual camera and the proportion information.
[0056] In this context, when the virtual camera captures an object, the plane corresponding to the object filling the entire screen is considered the near plane. Assuming the human body fills the entire screen, it is closest to the virtual camera. Based on the field of view (Fov) value of the virtual camera, the distance between the near plane and the camera can be calculated. As the human body shrinks, it indicates that the user's current plane is moving further away from the camera. Using the theorem of similar triangles, a ratio is calculated between the distance of the near plane to the camera and the distance of the user's current plane to the camera. (See [reference needed]). Figure 4 This ratio can be used as a benchmark for displaying information.
[0057] S130. Adjust the target display information of the material mounted in the video frame to be processed based on the reference display information to obtain the target video frame corresponding to the video frame to be processed.
[0058] The mounted material can be special effects added to the video frame to be processed; optionally, the special effects material can be something like rabbit ears. The target display information is determined based on the baseline display information. The target display information can be the display information of the mounted material after it has been enlarged or reduced. For example, the target display information can be the display size information of the mounted material after it has been enlarged or reduced.
[0059] Specifically, after determining the baseline display information, the mounted material in the video frame to be processed can be enlarged or reduced according to the baseline display information to obtain the corresponding target display information. Then, based on the target display information, the target video frame is obtained. That is, the target video frame is the video frame obtained after adjusting the mounted material of the video frame to be processed.
[0060] In this embodiment, adjusting the target display information of the mounted material in the video frame to be processed based on the reference display information to obtain the target video frame corresponding to the video frame to be processed includes:
[0061] Adjust the target display information of the target mounted material according to the reference display information; process the mounted material based on the virtual camera and the target display information to obtain a target video frame corresponding to the video frame to be processed.
[0062] The target display information can be understood as the specific display information of the mounted material within the video frame. For example, based on the baseline display information, it could be the magnification or reduction value of the mounted material. Correspondingly, the target display information could be the specific display size information of the mounted material after magnification or reduction. Alternatively, it could be the depth value information of the mounted material. Based on the virtual camera and the target display information, the mounted material can be reconstructed to obtain the corresponding target video frame. The virtual camera includes at least one of perspective cameras and orthographic cameras.
[0063] For example, after obtaining the ratio of the distance from the near plane to the camera to the distance from the current user's plane to the camera, this ratio can be used to scale the center point of the following material (mounted material). This step is mainly because the point value of this ratio is between -1 and 1, which falls within this range when near the camera. When moving away from the camera, the plane size will increase, so this ratio needs to be used to scale the position of the material. Since the z-value is changed to simulate the changes in distance of the mounted material following the user in the scene, a perspective camera is needed for rendering. Of course, if the baseline display information is determined based on 2D point-based algorithms to identify limb keypoints, or based on bounding box information to determine the area proportion, orthogonal camera rendering can be used to obtain the target video frame.
[0064] The technical solution of this disclosure, by determining the attribute information of the object to be processed in the video frame to be processed; determining the reference display information of the object to be processed in the video frame to be processed based on the attribute information; and adjusting the target display information of the mounted material in the video frame to be processed based on the reference display information, to obtain a target video frame corresponding to the video frame to be processed, solves the problem in the prior art that when using 2D algorithms to identify limb key points, although the identified limb key points are relatively accurate, the three-dimensional information of the limb key points cannot be obtained, resulting in poor tracking effect; when using 3D recognition algorithms to identify key points, although the three-dimensional information of limb key points can be identified, there is a high performance consumption, which puts high demands on the performance of terminal devices and results in poor universality. This solution realizes that the target display information of the mounted material can be determined based on the limb key information of the object to be processed in the video frame to be processed, and then based on the limb key point information, thereby obtaining the effect of three-dimensional display of the mounted material.
[0065] Example 2
[0066] Figure 5 This is a schematic diagram of a video image processing apparatus provided in Embodiment 2 of this disclosure, as shown below. Figure 5 As shown, the device includes: an attribute information determination module 210, a reference display information determination module 220, and a target video frame determination module 230.
[0067] The attribute information determination module 210 is used to determine the attribute information of the object to be processed in the video frame to be processed; the reference display information determination module 220 is used to determine the reference display information of the object to be processed in the video frame to be processed based on the attribute information; and the target video frame determination module 230 is used to adjust the target display information of the attached material in the video frame to be processed based on the reference display information to obtain a target video frame corresponding to the video frame to be processed.
[0068] Based on the above technical solution, the attribute information determination module includes:
[0069] The point identification unit is used to determine at least two points of the object to be processed in the video frame to be processed based on a 2D point identification algorithm.
[0070] An attribute information determination unit is used to determine the coordinate information to be processed of the at least one point to be processed, and to use the coordinate information to be processed as the attribute information.
[0071] Based on the above technical solution, the reference display information determination module includes:
[0072] A width information determination unit is used to determine at least three width information associated with the object to be processed based on the coordinate information to be processed;
[0073] The reference display information determination unit is used to determine the reference display information of the video frame to be processed based on the at least three types of width information and the corresponding preset reference values.
[0074] Based on the above technical solution, the attribute information determination module includes:
[0075] The bounding box information determination unit is used to determine the bounding box information including the object to be processed in the video frame to be processed, and to use the bounding box information as the attribute information.
[0076] Based on the above technical solution, the reference display information determination module is further configured to:
[0077] The baseline display information is determined based on the bounding box information in the attribute information and the page size information of the display page to which the video frame to be processed belongs.
[0078] Based on the above technical solution, the reference display information determination module is further configured to:
[0079] Based on the pre-determined near plane and the bounding box information, the proportion information of the object to be processed in the video to be processed is determined; wherein, the near plane is a plane determined when the object to be processed fills the display page to which the video frame to be processed belongs;
[0080] The baseline display information is determined based on the distance information of the near-plane distance virtual camera and the proportion information.
[0081] Based on the above technical solution, the target video frame determination module includes:
[0082] The display unit is used to adjust the target display information of the target mounted material according to the reference display information;
[0083] The target video frame determination unit is used to process the mounted material based on the virtual camera and the target display information to obtain a target video frame corresponding to the video frame to be processed.
[0084] The technical solution of this disclosure, by determining the attribute information of the object to be processed in the video frame to be processed; determining the reference display information of the object to be processed in the video frame to be processed based on the attribute information; and adjusting the target display information of the mounted material in the video frame to be processed based on the reference display information, to obtain a target video frame corresponding to the video frame to be processed, solves the problem in the prior art that when using 2D algorithms to identify limb key points, although the identified limb key points are relatively accurate, the three-dimensional information of the limb key points cannot be obtained, resulting in poor tracking effect; when using 3D recognition algorithms to identify key points, although the three-dimensional information of limb key points can be identified, there is a high performance consumption, which puts high demands on the performance of terminal devices and results in poor universality. This solution realizes that the target display information of the mounted material can be determined based on the limb key information of the object to be processed in the video frame to be processed, and then based on the limb key point information, thereby obtaining the effect of three-dimensional display of the mounted material.
[0085] The image processing apparatus provided in this disclosure can execute the image processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0086] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0087] Example 3
[0088] Figure 6 This is a schematic diagram of an electronic device structure provided in Embodiment 3 of this disclosure. Refer to the following... Figure 6 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 6 The diagram below shows the structure of the terminal device (or server) 300. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0089] like Figure 6 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An edit / output (I / O) interface 305 is also connected to the bus 304.
[0090] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0091] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of embodiments of this disclosure.
[0092] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0093] The electronic device provided in this embodiment and the video image processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0094] Example 4
[0095] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video image processing method provided in the above embodiments.
[0096] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0097] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0098] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0099] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0100] Determine the attribute information of the object to be processed in the video frame to be processed;
[0101] Based on the attribute information, determine the baseline display information of the object to be processed in the video frame to be processed;
[0102] Based on the reference display information, the target display information of the material mounted in the video frame to be processed is adjusted to obtain the target video frame corresponding to the video frame to be processed.
[0103] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0105] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0106] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0107] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0108] According to one or more embodiments of this disclosure, [Example 1] a video image processing method is provided, the method comprising:
[0109] Determine the attribute information of the object to be processed in the video frame to be processed;
[0110] Based on the attribute information, determine the baseline display information of the object to be processed in the video frame to be processed;
[0111] Based on the reference display information, the target display information of the material mounted in the video frame to be processed is adjusted to obtain the target video frame corresponding to the video frame to be processed.
[0112] According to one or more embodiments of this disclosure, [Example 2] provides a video image processing method, which further includes:
[0113] Optionally, determining the attribute information of the object to be processed in the video frame to be processed includes:
[0114] Based on a 2D point recognition algorithm, at least two points of the object to be processed in the video frame to be processed are determined.
[0115] Determine the coordinate information to be processed for the at least one point to be processed, and use the coordinate information to be processed as the attribute information.
[0116] According to one or more embodiments of this disclosure, [Example 3] provides a video image processing method, which further includes:
[0117] Optionally, determining the reference display information of the object to be processed in the video frame to be processed based on the attribute information includes:
[0118] Based on the coordinate information to be processed, determine at least three width information associated with the object to be processed;
[0119] Based on the at least three width information and the corresponding preset reference values, the reference display information of the video frame to be processed is determined.
[0120] According to one or more embodiments of this disclosure, [Example 4] provides a video image processing method, which further includes:
[0121] Optionally, determining the attribute information of the object to be processed in the video frame to be processed includes:
[0122] Determine the bounding box information of the object to be processed in the video frame to be processed, and use the bounding box information as the attribute information.
[0123] According to one or more embodiments of this disclosure, [Example 5] provides a video image processing method, which further includes:
[0124] Optionally, determining the reference display information of the object to be processed in the video frame to be processed based on the attribute information includes:
[0125] The baseline display information is determined based on the bounding box information in the attribute information and the page size information of the display page to which the video frame to be processed belongs.
[0126] According to one or more embodiments of this disclosure, [Example Six] provides a video image processing method, which further includes:
[0127] Optionally, determining the reference display information of the object to be processed in the video frame to be processed based on the attribute information includes:
[0128] Based on the pre-determined near plane and the bounding box information, the proportion information of the object to be processed in the video to be processed is determined; wherein, the near plane is a plane determined when the object to be processed fills the display page to which the video frame to be processed belongs;
[0129] The baseline display information is determined based on the distance information of the near-plane distance virtual camera and the proportion information.
[0130] According to one or more embodiments of this disclosure, [Example Seven] provides a video image processing method, which further includes:
[0131] Optionally, adjusting the target display information of the mounted material in the video frame to be processed based on the reference display information to obtain a target video frame corresponding to the video frame to be processed includes:
[0132] Adjust the target display information of the target mounted material according to the benchmark display information;
[0133] The mounted material is processed based on the virtual camera and the target display information to obtain a target video frame corresponding to the video frame to be processed.
[0134] According to one or more embodiments of this disclosure, [Example Eight] provides a video image processing apparatus, the apparatus comprising:
[0135] Determine the attribute information of the object to be processed in the video frame to be processed;
[0136] Based on the attribute information, determine the baseline display information of the object to be processed in the video frame to be processed;
[0137] Based on the reference display information, the target display information of the material mounted in the video frame to be processed is adjusted to obtain the target video frame corresponding to the video frame to be processed.
[0138] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0139] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0140] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method of processing a video image, characterized by, The method comprises the steps of: determining attribute information of a to-be-processed object in a to-be-processed video frame; determining reference display information of the to-be-processed object in the to-be-processed video frame according to the attribute information, wherein the reference display information at least comprises display size or display proportion of the to-be-processed object in the to-be-processed video frame; adjusting target display information of a mounted material in the to-be-processed video frame based on the reference display information to obtain a target video frame corresponding to the to-be-processed video frame; wherein the mounted material is a special effect material added in the to-be-processed video frame, the reference display information is relative display information of the to-be-processed object in the to-be-processed video frame, and the target display information is display information of magnification or reduction of the mounted material.
2. The method of claim 1, wherein, The method comprises the steps of: determining at least two to-be-processed point positions of the to-be-processed object in the to-be-processed video frame based on a 2D point position recognition algorithm; determining to-be-processed coordinate information of the at least one to-be-processed point position and taking the to-be-processed coordinate information as the attribute information.
3. The method of claim 2, wherein, The method comprises the steps of: determining at least three kinds of width information associated with the to-be-processed object according to the to-be-processed coordinate information; determining reference display information of the to-be-processed video frame according to the at least three kinds of width information and corresponding preset reference values.
4. The method of claim 1, wherein, The method comprises the steps of: determining bounding box information comprising the to-be-processed object in the to-be-processed video frame and taking the bounding box information as the attribute information.
5. The method of claim 4, wherein, The method comprises the steps of: determining the reference display information according to the bounding box information in the attribute information and page size information of a display page to which the to-be-processed video frame belongs.
6. The method of claim 4, wherein, The method comprises the steps of: determining proportion information of the to-be-processed object in the to-be-processed video according to a pre-determined near plane and the bounding box information, wherein the near plane is a plane determined according to the to-be-processed object filling the display page to which the to-be-processed video frame belongs; determining the reference display information according to distance information of the near plane from a virtual camera and the proportion information.
7. The method according to any one of claims 1 to 6, characterized in that, The method comprises the steps of: adjusting target display information of the mounted material according to the reference display information; processing the mounted material based on a virtual camera and the target display information to obtain a target video frame corresponding to the to-be-processed video frame.
8. A video image processing apparatus, characterized by comprising: The method comprises the steps of: an attribute information determination module configured to determine attribute information of a to-be-processed object in a to-be-processed video frame; The reference display information determination module is configured to determine reference display information of the to-be-processed object in the to-be-processed video frame according to the attribute information, wherein the reference display information at least includes a display size or a display proportion of the to-be-processed object in the to-be-processed video frame. The target video frame determination module is configured to adjust target display information of the mounted material in the to-be-processed video frame based on the reference display information, to obtain a target video frame corresponding to the to-be-processed video frame. The mounted material is special effect material added in the to-be-processed video frame, the reference display information is relative display information of the to-be-processed object in the to-be-processed video frame, and the target display information is display information of magnification or reduction of the mounted material.
9. An electronic device, comprising: The electronic device includes: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the video image processing method as claimed in any one of claims 1-7.
10. A storage medium containing computer-executable instructions for performing the video image processing method as claimed in any one of claims 1-7 when executed by a computer processor.
Citation Information
Patent Citations
Data processing method and terminal
CN110168599A
Video processing method and device, electronic equipment and computer readable storage medium
CN113518256A