View angle switching method and electronic device
By identifying the three-dimensional position and spatial rate of change of skeletal points in an image, the target viewpoint is determined, solving the problem of inaccurate viewpoint switching and improving the viewing experience.
Patent Information
- Application Number
- CN202310638659.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-31
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-05-31
AI Technical Summary
The existing technology suffers from inaccurate perspective switching, resulting in poor viewing experience.
By identifying the type and location of skeletal points in an image, determining their three-dimensional position and spatial region, calculating the volume change rate of the spatial region, identifying the most active target skeletal point, and using an innovative algorithm for the target skeletal point, the target viewpoint is determined.
It enables accurate switching of perspectives, improving the viewing experience.
Smart Images

Figure CN116843763B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer application, and in particular, to a view angle switching method and an electronic device. BACKGROUND
[0002] In an application scenario, such as a dance, fitness, a video recording or a performance scene, in order to facilitate an observer to observe, a plurality of collection devices are often arranged in the scene, and a viewing view angle is switched to a view angle that facilitates the observer to observe a target action according to a posture conversion of the human body or other targets, for example, when a right arm of a dance performer makes an action, the viewing view angle needs to be switched to a view angle that watches the right arm, and if the viewing view angle that watches the left arm is still displayed at this time, the experience is poor.
[0003] In the related art, the switching of the view angle is usually performed manually, that is, the current collection device is switched, however, the manual switching is time-consuming and laborious, and is prone to cause inaccurate switching, and thus a poor viewing effect can occur. SUMMARY
[0004] Embodiments of the present application provide a view angle switching method and an electronic device, to solve the problem of inaccurate view angle switching in the related art, and to cause a poor viewing effect.
[0005] In a first aspect, embodiments of the present application provide a view angle switching method, and the method comprises:
[0006] identifying a type and a position of a skeleton point in an image, determining a three-dimensional position of the skeleton point based on the position of the skeleton point in the image;
[0007] determining a space region in which the skeleton point is located according to the three-dimensional positions of the target and the skeleton point in the image; for any skeleton point, determining a change rate of the space region in which the skeleton point is located according to a volume of the space region in which the skeleton point is located and a volume of the space region in which the skeleton point is located determined based on a preset number of frames of images adjacent to the image in a time sequence before the image; determining a target skeleton point with the highest change rate among the skeleton points;
[0008] determining a corresponding target view angle according to the three-dimensional position of the target skeleton point, and switching a viewing view angle to the target view angle.
[0009] In a second aspect, embodiments of the present application further provide an electronic device, and the electronic device at least comprises a processor and a memory, the processor is used to implement the steps of the view angle switching method according to any one of the above when executing a computer program stored in the memory.
[0010] In the embodiment of the present application, the electronic device identifies the type and position of the skeleton point in the image, determines the three-dimensional position of the skeleton point based on the position of the skeleton point in the image, determines the space region where the skeleton point is located according to the target in the image and the three-dimensional position of the skeleton point, determines the change rate of the space region where the skeleton point is located according to the volume of the space region and the volume of the space region determined by the preset number of frame images adjacent to the image in time sequence, determines the target skeleton point with the highest change rate among the skeleton points, and switches the viewing angle to the target view angle according to the three-dimensional position of the target skeleton point. In the embodiment of the present application, the electronic device determines the space region where the skeleton point is located in the adjacent images collected, determines the target skeleton point with the highest change rate according to the change rate of the volume of the space region where the skeleton point is located, determines the target view angle corresponding to the target skeleton point according to the three-dimensional position of the target skeleton point, and switches the viewing angle to the target view angle, thereby solving the problem of inaccurate manual switching of the view angle and improving the viewing effect. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0012] Figure 1 A view angle switching process schematic diagram provided for the embodiment of the present application;
[0013] Figure 2 A collection device setting schematic diagram provided for the embodiment of the present application;
[0014] Figure 3 A planar schematic diagram of the divided region provided for the embodiment of the present application;
[0015] Figure 4 A detailed process schematic diagram of determining the target view angle provided for the embodiment of the present application;
[0016] Figure 5 A region schematic diagram where the target is located provided for the embodiment of the present application;
[0017] Figure 6 A process schematic diagram of determining the space region where the skeleton point is located provided for the embodiment of the present application;
[0018] Figure 7 A process schematic diagram of determining the target view angle provided for the embodiment of the present application;
[0019] Figure 8 A process diagram for determining the position of a preset identity part in an image is provided for an embodiment of the present application.
[0020] Figure 9 A structure diagram of a view angle switching device is provided for an embodiment of the present application.
[0021] Figure 10 A structure diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art belong to the scope of protection of the present application.
[0023] In order to accurately switch the view angle, the present application provides a view angle switching method and an electronic device.
[0024] The view angle switching method comprises: identifying the type and position of a skeletal point in an image, determining the three-dimensional position of the skeletal point based on the position of the skeletal point in the image; determining the space region where the skeletal point is located according to the target and the three-dimensional position of the skeletal point in the image; for any skeletal point, determining the change rate of the space region where the skeletal point is located according to the volume of the space region where the skeletal point is located and the volume of the space region where the skeletal point is located determined by a preset number of frames of images adjacent to the image in time sequence before the image; determining the target skeletal point with the highest change rate among the skeletal points; determining the corresponding target view angle according to the three-dimensional position of the target skeletal point, and switching the viewing view angle to the target view angle. Thus, the view angle can be accurately switched, and the viewing effect can be improved.
[0025] Figure 1 A view angle switching process diagram is provided for an embodiment of the present application, and the process comprises the following steps:
[0026] S101: identifying the type and position of a skeletal point in an image, and determining the three-dimensional position of the skeletal point based on the position of the skeletal point in the image.
[0027] The view angle switching method provided by the present application is applied to an electronic device, which can be a PC, a server or other smart devices.
[0028] In order to accurately switch the view angle, the electronic device can first receive an image, wherein the electronic device can receive images collected by multiple collection devices, and for any image collected by a collection device, identify the type of the skeletal point in the image and the position of the skeletal point in the image. Specifically, the electronic device can input the image into a pre-trained skeletal point recognition model to obtain the types of multiple skeletal points and the positions of the skeletal points in the image output by the skeletal point recognition model, wherein the types of the skeletal points include wrist, elbow, knee, neck, head, etc., and the skeletal points at least include skeletal points at limbs, head, and torso. The specific determination of which types of skeletal points and positions in the image can be set according to requirements, which is not limited here.
[0029] After obtaining the positions of the skeletal points in the image, the electronic device can determine the three-dimensional positions of the skeletal points according to the positions of the skeletal points in the image. Specifically, the electronic device can combine the positions of the skeletal points in the image with the intrinsic matrix of the collection device, and then combine the extrinsic matrix of the collection device to determine the three-dimensional positions of the skeletal points, wherein the collection device can send the number of the collection device to the electronic device when sending the image to the electronic device, and the electronic device obtains the intrinsic matrix and the extrinsic matrix saved for the number. The numbers of the multiple collection devices are all different. Specifically, how to determine the three-dimensional positions of the skeletal points based on the positions of the skeletal points in the image is a prior art, which will not be repeated here. Determining the three-dimensional positions of the skeletal points is equivalent to extracting three-dimensional skeletal points.
[0030] S102: According to the three-dimensional positions of the target and the skeletal points in the image, determine the space region where the skeletal points are located; for any skeletal point, determine the change rate of the space region where the skeletal point is located according to the volume of the space region and the volume of the space region determined based on the time sequence of the preset number of frames of images adjacent to the image before and after the image; determine the target skeletal point with the highest change rate among the skeletal points.
[0031] In order to accurately realize the view angle switching, the electronic device can determine the bone point that is active, and perform view angle switching based on the bone point, and specifically switch to a view angle that facilitates observation of the bone point. In order to determine the bone point that is active, the electronic device can determine a spatial region in which each bone point is located according to the three-dimensional positions of the target and the plurality of bone points in the image, where the target can be a human body or other targets such as animals, etc. Specifically, the electronic device can input the image into a recognition model, and the recognition model can output the region in which the bone point is located in the image, where the regions in which the plurality of bone points are located in the image can constitute the region in which the target is located in the image, and the size of the region in which the bone point is located is different when the action of the target is different. The electronic device can determine the three-dimensional position of any pixel point in the region in which any bone point is located, and the point at the three-dimensional position is a point of the spatial region in which the bone point is located. The electronic device can determine the spatial region in which the bone point is located in this way.
[0032] In order to find the target bone point with the highest change rate of the spatial region, the electronic device can calculate the change rate of the volume of the spatial region in which the bone point is located. Specifically, the electronic device can determine the volume of the spatial region in which the bone point is located based on a preset number of frames of images before and adjacent to the image in time sequence, and determine the change rate of the spatial region in which the bone point is located according to the determined volume of the spatial region in which the bone point is located. Specifically, when determining the change rate of the spatial region in which a certain bone point is located, the electronic device can calculate the ratio of the volume of the spatial region in which the bone point is located based on the image to the volume of the spatial region in which the bone point is located based on the previous frame of image for any image in the image and the preset number of frames of images, and multiply the plurality of calculated ratios. The product obtained can be determined as the change rate of the spatial region in which the bone point is located. In this way, the change rates of the spatial regions in which the plurality of bone points are located can be obtained. In order to find the target bone point with the highest change rate, the electronic device can traverse the change rates of the spatial regions in which all bone points are located, find the highest value, and the bone point corresponding to the highest value is the target bone point with the highest change rate. Wherein, the change rate of the bone point is proportional to the activity degree of the joint of the bone point. In this way, the joint of the target bone point with the highest activity degree can be determined.
[0033] S103: Determine a corresponding target view angle according to the three-dimensional position of the target bone point, and switch the viewing view angle to the target view angle.
[0034] After determining the target skeleton point with the highest change rate, the electronic device can determine a target view angle that is most suitable for collecting the activity of the target skeleton point according to the three-dimensional position of the target skeleton point. Specifically, the electronic device can determine the distances between the three-dimensional positions corresponding to a plurality of view angles and the three-dimensional position of the skeleton point, where the three-dimensional position corresponding to a view angle is the three-dimensional position of a collection device that can collect the view angle, and determine the view angle with the smallest distance as the target view angle. Since the target view angle corresponds to a three-dimensional position that is closer to the target skeleton point, the target view angle can more clearly display the changes of the target skeleton point.
[0035] Among the plurality of view angles, there are view angles of collection devices and virtual view angles virtually generated based on the plurality of collection devices. The video of a virtual view angle is fused based on videos collected by the plurality of collection devices, and the specific fusion manner is not limited herein. If a view angle is a virtual view angle, the three-dimensional position corresponding to the view angle can be determined by a business staff in advance, and the three-dimensional position corresponding to the view angle is still the three-dimensional position of a collection device that can collect the view angle. Since the number of collection devices is limited in an actual application scenario, it can be impossible to arrange a plurality of collection devices around the target. In order to more comprehensively display the target, a virtual view angle can be virtually generated for a view angle of a collection device, and subsequent view angle switching can be performed based on the plurality of view angles. For example, collection devices are arranged in front of and to the left of the target (front and back in an actual scenario, and left and right in an actual scenario). The video collected by a collection device in the front left (front, back, left, and right in an actual scenario) can be virtually generated by fusing videos collected by the two collection devices in advance, and the video is the video of the virtual view angle.
[0036] When the view angles include virtual view angles, the three-dimensional position corresponding to each virtual view angle can also be determined by a business staff in advance, so as to facilitate the determination of the target view angle according to the three-dimensional position corresponding to each view angle.
[0037] After determining the target view angle, the electronic device can switch the viewing view angle to the target view angle, thereby improving the viewing effect. The images and the preset number of images described above are images in the free view angle video, that is, the scheme provided in the embodiments of the present application can automatically switch the viewing view angle according to the change rate of the space region in which the skeleton point is located in the free view angle video, without manually switching the view angle.
[0038] The scheme provided in the present application can determine the target view angle in real time according to the images, and the determined target view angle is relatively accurate, that is, the scheme provided in the present application has real-time performance, reliability, and meets the trusted characteristics.
[0039] It should be noted that, in the continuous frames in a short time, if the volume change rate of the space region where a certain skeleton point is located is large, it indicates that the joint represented by the skeleton point changes in a short time, that is, an action is made. In the embodiment of the present application, the joint of the target making the action can be determined by judging the volume change rate of the space region where all skeleton points are located when the continuous frames of images in a short time are collected, and then the viewing angle of the collected image is adjusted.
[0040] For example, the target skeleton point is the left arm, and it can be determined that the left arm makes an action, and the viewing angle is switched to the viewing angle that is beneficial to display the left arm, specifically, the currently displayed video is switched to the video that is beneficial to display the left arm.
[0041] Figure 2 A schematic diagram of a collection device provided in the embodiment of the present application is shown.
[0042] Figure 2 The target is taken as a human body for example in the embodiment of the present application, and Figure 2 It can be known that a plurality of collection devices can be arranged around the human body in advance to obtain videos of the human body from different angles, and the video array formed has different viewing angles, and the viewing angle can be freely adjusted, that is, the free viewing angle. The image mentioned in the embodiment of the present application can be an image collected by any collection device.
[0043] It should be noted that the method provided in the embodiment of the present application can automatically switch the viewing angle according to the action of the target, and in addition, the method provided in the embodiment of the present application can also switch the viewing angle for animals. If the viewing angle is to be switched for animals, only the animal skeleton point detection algorithm needs to be used for skeleton point detection. For example, a scene in which a giant panda is to be observed can arrange a plurality of collection devices around the giant panda, and the viewing angle is switched based on the action of the giant panda, so that the action of the giant panda can be completely displayed, and the viewing effect is improved.
[0044] In the embodiment of the present application, the electronic device determines the space region where the skeleton point is located in the adjacent collected images, and determines the target skeleton point with the largest change rate according to the volume change rate of the space region where the skeleton point is located. The target viewing angle corresponding to the three-dimensional position of the target skeleton point is determined, and the viewing angle is switched to the target viewing angle, so that the problem of inaccurate manual switching of the viewing angle can be solved, and the viewing effect is improved.
[0045] In order to accurately determine the change rate of the space region where the skeleton point is located, on the basis of the above-mentioned embodiment, in the embodiment of the present application, the change rate of the space region where the skeleton point is located is determined according to the volume of the space region where the skeleton point is located, and the volume of the space region where the skeleton point is located determined based on the time sequence and the preset number of frames of images adjacent to the image before the image, comprising:
[0046] Determine the volume of the space region where the skeleton point is located based on the time sequence of the preset number of frame images adjacent to the image before the image, respectively; and determine the change rate of the space region where the skeleton point is located according to the variance of the volumes of the space regions where the plurality of skeleton points are located.
[0047] In order to determine the change rate of the space region where the skeleton point is located, the electronic device can first acquire the volume of the space region where the skeleton point is located based on the image, wherein the volume of the space region where the skeleton point is located is equivalent to the volume of the space region where the skeleton point is located in the target when the image is collected, and determine the volume of the space region where the skeleton point is located based on the time sequence of the preset number of frame images adjacent to the image before the image, wherein the acquired volume of the space region where the skeleton point is located is the volume of the space region where the skeleton point is located in the target when the corresponding image is collected. The electronic device can determine the variance of the acquired volumes, and determine the change rate of the space region where the skeleton point is located according to the variance. Specifically, the variance can be determined as the change rate of the space region where the skeleton point is located.
[0048] It should be noted that when a certain skeleton point in the target is not active for a long time, the volume of the space region where the skeleton point is located determined based on the images collected for a long time changes little, and the determined variance is small. When a certain skeleton point in the target is active, the volume of the space region where the skeleton point is located determined based on the images collected within the time of the action changes greatly, and the determined variance is large. Therefore, the method based on the variance can accurately determine the target skeleton point in the target that is active.
[0049] The electronic device can arrange the volumes of the space regions where the skeleton points are located in the order of the continuous images collected based on the same collection device for any skeleton point, for example, the volumes of the space regions where the skeleton points are located are vpi-1, vpi-2, …, vpi-n in turn. The electronic device can determine the change rate of the space region where the skeleton point is located based on the volumes, thereby improving the efficiency of acquiring the volume of the space region where the skeleton point is located.
[0050] In order to accurately determine the change rate of the space region where the skeleton point is located, on the basis of the above-mentioned embodiments, in the embodiments of the present application, the change rate of the space region where the skeleton point is located is determined according to the volume of the space region where the skeleton point is located, and the volumes of the space regions where the skeleton points are located determined based on the time sequence of the preset number of frame images adjacent to the image before the image.
[0051] determine a first change rate of the space region where the skeleton point is located between the previous frame of the image and the image according to the volume of the space region where the skeleton point is located and the volume of the space region where the skeleton point is located determined based on the previous frame of the image;
[0052] and obtain a preset number of second change rates saved for the previous frame of the image and the skeleton point;
[0053] determine the change rate of the space region where the skeleton point is located according to an average value of the first change rate and the second change rate.
[0054] To determine the change rate of the space region where the skeleton point is located, the electronic device can first obtain the volume of the space region where the skeleton point is located and the volume of the space region where the skeleton point is located in the previous frame of the image, determine the first change rate of the skeleton point according to the volume of the space region where the skeleton point is located and the volume of the space region where the skeleton point is located determined based on the previous frame of the image, specifically, can determine the absolute value of the difference between the volume of the space region where the skeleton point is located determined based on the image and the volume of the space region where the skeleton point is located determined based on the previous frame of the image, and determine the ratio of the absolute value to the volume of the space region where the skeleton point is located determined based on the image or the previous frame of the image, and determine the ratio as the first change rate of the skeleton point.
[0055] To obtain a more accurate change rate of the skeleton point, a preset number of second change rates saved for the previous frame of the image and the skeleton point can also be obtained, and the second change rates of the volume of the skeleton point in the previous preset number of frames of images can reflect the historical change trend of the skeleton point. The plurality of second change rates are determined in the manner of determining the first change rate and are determined based on the previous preset number of frames of images.
[0056] After determining the first change rate and the second change rate, the electronic device can obtain an average value of the first change rate and the second change rate, and determine the change rate of the space region where the skeleton point is located according to the average value, specifically, the average value can be determined as the change rate of the space region where the skeleton point is located. The change rate can reflect the recent change trend of the skeleton point.
[0057] It should be noted that when a certain skeleton point is inactive for a long time, even if the volume of the space region where the skeleton point is located determined is wrong, the first change rate or the second change rate is determined based on the previous several frames of images, and then the average value of the determined change rate is still small, so even if the volume of the space region where the skeleton point is located determined based on a certain image is wrong, the skeleton point will not be determined as the target skeleton point that is active, that is, the accuracy of determining the target skeleton point can be improved.
[0058] For example, if the preset number of frames of images is n frames of images, the rate of change of the space region in which a certain skeletal point is located is determined by the following formula:
[0059] t emp1 = |v pi-1 -v pi | / v pi
[0060] t emp2 = |v pi-2 -v pi-1 | / v pi-1
[0061] …
[0062] t empn = |v pi-n -v pi-(n-1) | / v pi-(n-1)
[0063] t = (t emp1 + t emp2 + … + t empn ) / n
[0064] wherein t emp1 is the first rate of change, v pi-1 is the volume of the space region in which the skeletal point is located, which is determined based on the previous frame of image of the image, v pi is the volume of the space region in which the skeletal point is located, which is determined based on the image, t emp2 , …, t empn are the second rates of change, v pi-2 is the volume of the space region in which the skeletal point is located, which is determined based on the previous two frames of image of the image, v pi-n is the volume of the space region in which the skeletal point is located, which is determined based on the previous n frames of image of the image, v pi-(n-1) is the volume of the space region in which the skeletal point is located, which is determined based on the previous n-1 frames of image of the image, and t is the rate of change of the space region in which the skeletal point is located.
[0065] In order to accurately determine the space region in which the skeletal point is located, on the basis of the above embodiments, in the embodiments of the present application, the determination of the space region in which the skeletal point is located according to the three-dimensional positions of the target and the skeletal point in the image comprises:
[0066] adopting a roll wrapping method or an incremental method to determine a three-dimensional convex hull containing the skeletal point, and determining the three-dimensional convex hull as the space region in which the target is located;
[0067] using a three-dimensional Voronoi diagram generation method to divide the space region according to the three-dimensional position of the skeletal point, and obtaining the space region in which the skeletal point is located.
[0068] To determine the spatial region where the skeleton points are located, the electronic device can determine a three-dimensional convex hull containing the plurality of skeleton points according to the three-dimensional positions of the skeleton points. Specifically, the three-dimensional convex hull containing the plurality of skeleton points can be determined by using a rolling-wrapping method or an incremental method or other methods. The three-dimensional convex hull can be regarded as the spatial region where the target is located. Specifically, how to determine the three-dimensional convex hull containing the plurality of skeleton points by using the rolling-wrapping method or the incremental method or other methods is prior art, which will not be described here.
[0069] After determining the three-dimensional convex hull containing the plurality of skeleton points, the electronic device can divide the spatial region where the target is located according to the three-dimensional positions of the plurality of skeleton points by using a three-dimensional Voronoi diagram generation method. The three-dimensional Voronoi diagram generation method is a spatial division method, which can divide the space into a plurality of non-overlapping regions. Each region after division contains a skeleton point, and the region is the spatial region where the contained skeleton point is located. It should be noted that the regions where the plurality of skeleton points are located are all polyhedrons.
[0070] The region of the skeleton point is equivalent to dividing the target into a plurality of body parts, such as left arm, right arm, left leg, right leg, torso, etc. By calculating the change rate of the spatial region where the skeleton point is located, the body part where the skeleton point with a larger motion amplitude is located can be determined.
[0071] Figure 3 A planar schematic diagram of a divided region provided by an embodiment of the present application.
[0072] Figure 3 The points in the diagram are skeleton points, and the regions are spatial regions where the skeleton points are located. Figure 3 It can be seen that any region determined contains only one skeleton point.
[0073] Figure 4 A detailed process diagram for determining a target view angle is provided by an embodiment of the present application. The process includes the following steps:
[0074] S401: Identify the type of skeleton point and the position of the skeleton point in the image.
[0075] S402: Determine the three-dimensional position of the skeleton point.
[0076] S403: Determine the three-dimensional convex hull containing the skeleton point.
[0077] S404: Divide the three-dimensional convex hull into a plurality of spatial regions where the skeleton points are located.
[0078] S405: For any skeleton point, determine the change rate of the spatial region where the skeleton point is located according to the volume of the spatial region where the skeleton point is located and the volume of the spatial region where the skeleton point is located determined in the preset number of frames of images before and adjacent to the image in time sequence.
[0079] S406: Determine a target skeletal point with the highest change rate among the skeletal points.
[0080] S407: Determine a corresponding target view angle according to the three-dimensional position of the target skeletal point.
[0081] To accurately determine the space region where the skeletal point is located, on the basis of the above embodiments, in the embodiments of the present application, the determination of the space region where the skeletal point is located according to the three-dimensional position of the target and the skeletal point in the image comprises:
[0082] Identify the region where the target is located in the image, for any pixel point in the region, determine the three-dimensional position of the pixel point according to the position of the pixel point in the image, determine the distance between the three-dimensional position and the three-dimensional position of each skeletal point, determine the skeletal point with the minimum distance from the three-dimensional position, and determine the point at the three-dimensional position as a point located in the space region where the skeletal point is located.
[0083] To determine the space region where the skeletal point is located, the electronic device can first identify the region where the target is located in the image. Specifically, the electronic device can identify the region where the target is located in the image through a Haar feature classifier, wherein the Haar feature classifier is a face detection technology based on machine learning, which can identify the region where the target is located by detecting specific Haar features in the image, such as edges, corners, etc. The electronic device can also identify the region where the target is located in the image through an edge detection method, which is a technology for converting an image into only the outline of an object. By detecting the edges in the image, the region where the target is located can be identified. Specifically, how to identify the region where the target is located in the image through the Haar feature classifier and the edge detection method is prior art, which will not be described here. The region where the target is located can be represented by the positions of the top left corner and the bottom right corner of a rectangular frame that completely contains the target (here, up, down, left and right refer to the up, down, left and right in the image), and the region where the target is located can also be represented by the height, width and position of a certain vertex of the rectangular frame that completely contains the target.
[0084] Figure 5 A schematic diagram of the region where the target is located is provided in the embodiments of the present application.
[0085] Figure 5 The rectangular frame that completely contains the target is the region where the target is located, and the position related to the rectangular frame can be determined in the embodiments of the present application, for example, the positions of the top left corner and the bottom right corner of the rectangular frame (here, up, down, left and right refer to the up, down, left and right in the image).
[0086] After determining the region where the target is located, the electronic device can determine the three-dimensional position of any pixel point in the region according to the position of the pixel point in the image. The point at the three-dimensional position is a point in the spatial region where the target is located. The electronic device can determine the distance between the pixel point in the spatial region and the plurality of skeletal points according to the three-dimensional position of the pixel point and the three-dimensional positions of the plurality of skeletal points, respectively, and determine the skeletal point with the smallest distance. The point in the spatial region of the pixel point is determined to be a point in the spatial region of the skeletal point. In this way, the electronic device can determine each point in the spatial region of the skeletal point, and further determine the spatial region of the skeletal point.
[0087] Figure 6 A process diagram for determining the spatial region of the skeletal point is provided in the embodiments of the present application. The process includes the following steps:
[0088] S601: Identify the region where the target is located in the image.
[0089] S602: For any pixel point in the region where the target is located, determine the three-dimensional position of the pixel point according to the position of the pixel point in the image.
[0090] S603: For any pixel point in the region where the target is located, determine the distance between the three-dimensional positions of the plurality of skeletal points and the three-dimensional position of the pixel point.
[0091] S604: For any pixel point in the region where the target is located, determine the skeletal point with the smallest distance from the three-dimensional position of the pixel point. The point at the three-dimensional position is determined to be a point in the spatial region of the skeletal point.
[0092] To accurately determine the spatial region of the skeletal point, on the basis of the above embodiments, in the embodiments of the present application, identifying the region where the target is located in the image includes:
[0093] Input the image into a pre-trained target recognition model to obtain the region where the target is located in the image output by the target recognition model.
[0094] To determine the region where the target is located in the image, the electronic device can locally save a pre-trained target recognition model. The electronic device can input the image into the pre-trained target recognition model, and obtain the output of the target recognition model. The output is the region where the target is located in the image. The region where the target is located can be represented by the positions of the top left corner and the bottom right corner of a rectangular frame that completely contains the target. The region where the target is located can also be represented by the height, width and position of a certain vertex of the rectangular frame in the image.
[0095] To accurately switch the view angle, on the basis of the above embodiments, in the embodiments of the present application, the determining of the corresponding target view angle according to the three-dimensional position of the target skeletal point comprises:
[0096] identifying the position of the preset body part in the image, and determining the three-dimensional position of the preset body part according to the position;
[0097] determining a ray with the three-dimensional position of the preset body part as an end point and pointing to the direction of the three-dimensional position of the target skeletal point, determining the distance between the three-dimensional position corresponding to the view angle and the ray according to the three-dimensional position corresponding to the view angle and the ray, and determining the view angle with the minimum distance as the target view angle.
[0098] To accurately switch the view angle, the electronic device can identify the position of the preset body part in the image, and specifically, the preset body part can be the belly. The electronic device can also determine the position of the target center according to the position of the skeletal point in the image, and determine the target center as the preset body part, and specifically, the center point of the skeletal point can be determined as the target center. The electronic device can determine the three-dimensional position of the preset body part according to the position.
[0099] After determining the three-dimensional position of the preset body part, the electronic device can determine a ray with the three-dimensional position of the preset body part as an end point and pointing to the direction of the three-dimensional position of the target skeletal point, and specifically, the electronic device can calculate the vector of the point at the three-dimensional position of the preset body part pointing to the point at the three-dimensional position of the target skeletal point, and determine the unit vector of the vector, with the point at the three-dimensional position of the preset body part as an end point and the unit vector as a direction vector, to obtain the sought ray. After determining the ray, the electronic device can determine the distance between the three-dimensional position corresponding to each view angle and the ray, and the electronic device determines the view angle with the minimum distance, wherein the view angle with the minimum distance can more accurately and more closely observe the target skeletal point, and the electronic device can determine the view angle with the minimum distance as the target view angle, and based on the target view angle, the changes of the target skeletal point can be accurately captured, thereby improving the viewing effect.
[0100] Figure 7 A process diagram for determining a target view angle is provided for the embodiments of the present application, and the process comprises the following steps:
[0101] S701: Identify the position of the preset body part in the image.
[0102] S702: Determine the three-dimensional position of the preset body part according to the position of the preset body part in the image.
[0103] S703: Determine a ray with the three-dimensional position of the preset body part as an end point and pointing to the direction of the three-dimensional position of the target skeletal point.
[0104] S704: Determine the distance between each view and the ray according to the three-dimensional position corresponding to each view and the ray.
[0105] S705: Determine the view with the minimum distance as the target view.
[0106] To determine the position of the preset body part in the image, on the basis of the above embodiments, in the embodiments of the present application, the identification of the position of the preset body part in the image comprises:
[0107] inputting the image into the pre-trained recognition model to obtain the position of the preset body part in the image output by the recognition model.
[0108] To determine the position of the preset body part in the image, the electronic device locally saves a pre-trained recognition model, and the electronic device can input the image into the pre-trained recognition model to obtain the output of the recognition model, which is the position of the preset body part in the image.
[0109] Figure 8 A process diagram for determining the position of a preset identity part in an image is provided in the embodiments of the present application, which includes the following steps:
[0110] S801: Obtain an image.
[0111] S802: Input the image into a pre-trained recognition model.
[0112] S803: Obtain the position of the preset body part in the image output by the recognition model.
[0113] To accurately switch the view, on the basis of the above embodiments, in the embodiments of the present application, the identification of the type and position of the skeletal point in the image comprises:
[0114] The type and position of the skeletal point in the image are identified by a skeletal point detection algorithm.
[0115] To determine the type and position of the skeletal point in the region, the electronic device can identify the type and position of the skeletal point in the image by a skeletal point detection algorithm. Specifically, how to identify the type and position of a skeletal point in a certain image by a skeletal point detection algorithm is known in the art and will not be described here. For example, a 21-point skeletal point detection algorithm can be used.
[0116] Any skeletal point detection algorithm can be used in the embodiments of the present application, and the skeletal point detection algorithm used is not limited, nor is the type of skeletal point identified by the skeletal point detection algorithm.
[0117] Figure 9 A perspective switching device structure schematic diagram is provided for an embodiment of the present application, the device comprising:
[0118] A recognition determination module 901 is configured to recognize the type and position of the skeleton point in the image, determine the three-dimensional position of the skeleton point based on the position of the skeleton point in the image;
[0119] A processing module 902 is configured to determine the spatial region where the skeleton point is located according to the three-dimensional position of the target and the skeleton point in the image, determine the change rate of the spatial region where the skeleton point is located according to the volume of the spatial region where the skeleton point is located and the volume of the spatial region where the skeleton point is located determined based on a preset number of frame images adjacent to the image in time sequence before the image, for any skeleton point, determine the target skeleton point with the highest change rate among the skeleton points;
[0120] A determination collection module 903 is configured to determine the corresponding target perspective according to the three-dimensional position of the target skeleton point, and switch the viewing perspective to the target perspective.
[0121] Further, the processing module 902 is specifically configured to determine the volume of the spatial region where the skeleton point is located based on a preset number of frame images in time sequence before the image and adjacent to the image, respectively; and determine the change rate of the spatial region where the skeleton point is located according to the variance of the volumes of the spatial regions where the skeleton point is located determined.
[0122] Further, the processing module 902 is specifically configured to determine a first change rate of the volume of the spatial region where the skeleton point is located between the last frame image and the image based on the volume of the spatial region where the skeleton point is located and the volume of the spatial region where the skeleton point is located determined based on the last frame image of the image; and obtain a preset number of second change rates saved for the last frame image and the skeleton point; and determine the change rate of the spatial region where the skeleton point is located according to the average value of the first change rate and the second change rate.
[0123] Further, the processing module 902 is specifically configured to determine a three-dimensional convex hull containing the skeleton point by using a roll wrapping method or an incremental method, determine the target spatial region as the three-dimensional convex hull, and divide the spatial region according to the three-dimensional position of the skeleton point by using a Voronoi diagram generation method to obtain the spatial region where the skeleton point is located.
[0124] Further, the processing module 902 is specifically configured to identify a region where the target is located in the image, for any pixel point in the region, determine a three-dimensional position of the pixel point according to a position of the pixel point in the image, determine a distance between the three-dimensional position and a three-dimensional position of each skeleton point, determine a skeleton point with a minimum distance from the three-dimensional position, and determine a point at the three-dimensional position as a point located in a space region where the skeleton point is located.
[0125] Further, the processing module 902 is specifically configured to input the image into a pre-trained target recognition model, and acquire a region where the target is located in the image output by the target recognition model.
[0126] Further, the determination acquisition module 903 is specifically configured to identify a position of a preset body part in the image, and determine a three-dimensional position of the preset body part according to the position; determine a ray in a direction from the three-dimensional position of the preset body part to the three-dimensional position of the target skeleton point, and determine a target view angle according to a three-dimensional position corresponding to the view angle and the ray, a distance between the three-dimensional position corresponding to the view angle and the ray, and a minimum distance.
[0127] Further, the determination acquisition module 903 is specifically configured to input the image into a pre-trained recognition model, and acquire a position of a preset body part in the image output by the recognition model.
[0128] Further, the identification determination module 901 is specifically configured to identify a type and a position in an image of a skeleton point in the image by using a skeleton point detection algorithm.
[0129] Figure 10 An electronic device structure schematic diagram provided by an embodiment of the present application is provided on the basis of the above embodiments, and the present application further provides an electronic device, as shown in the figure, which includes a processor 1001, a communication interface 1002, a memory 1003 and a communication bus 1004, wherein the processor 1001, the communication interface 1002 and the memory 1003 complete mutual communication through the communication bus 1004. Figure 10
[0130] The memory 1003 stores a computer program, and when the program is executed by the processor 1001, the processor 1001 performs the following steps:
[0131] identify a type and a position in an image of a skeleton point in the image, and determine a three-dimensional position of the skeleton point based on a position of the skeleton point in the image;
[0132] determine a change rate of the space region where the target skeleton point is located according to the volume of the space region where the target skeleton point is located and the volume of the space region where the target skeleton point is located determined based on a preset number of frame images before and adjacent to the image in time sequence; determine a target skeleton point with the highest change rate among the skeleton points;
[0133] determine a corresponding target view angle according to the three-dimensional position of the target skeleton point, and switch the viewing view angle to the target view angle.
[0134] Further, the processor 1001 is specifically configured to determine the volume of the space region where the target skeleton point is located based on a preset number of frame images before and adjacent to the image in time sequence; and determine the change rate of the space region where the target skeleton point is located according to the variance of the volumes of the space regions where the target skeleton point is located determined.
[0135] Further, the processor 1001 is specifically configured to determine a first change rate of the volume of the space region where the target skeleton point is located between the last frame image and the image according to the volume of the space region where the target skeleton point is located and the volume of the space region where the target skeleton point is located determined based on the last frame image of the image.
[0136] and obtain a preset number of second change rates saved for the last frame image and the target skeleton point;
[0137] determine the change rate of the space region where the target skeleton point is located according to the average of the first change rate and the second change rate.
[0138] Further, the processor 1001 is specifically configured to determine a three-dimensional convex hull containing the skeleton points by using a roll wrapping method or an incremental method, and determine the three-dimensional convex hull as the target space region.
[0139] divide the space region according to the three-dimensional positions of the skeleton points by using a Voronoi diagram generation method, and obtain the space region where the skeleton points are located.
[0140] Further, the processor 1001 is specifically configured to identify a region where the target is located in the image, determine the three-dimensional position of each pixel point in the region according to the position of the pixel point in the image, determine the distance between the three-dimensional position and each skeleton point, determine the skeleton point with the smallest distance to the three-dimensional position, and determine the point at the three-dimensional position as a point located in the space region where the skeleton point is located.
[0141] Further, the processor 1001 is specifically configured to input the image into a pre-trained target recognition model, and obtain a region where a target in the image is located, which is output by the target recognition model.
[0142] Further, the processor 1001 is specifically configured to identify a position of a preset body part in the image, and determine a three-dimensional position of the preset body part according to the position.
[0143] A ray is determined, which has the three-dimensional position of the preset body part as an end point and points to a direction of the three-dimensional position of the target skeletal point, a distance between a three-dimensional position corresponding to the view angle and the ray is determined according to the three-dimensional position corresponding to the view angle and the ray, and a view angle with a minimum distance is determined as a target view angle.
[0144] Further, the processor 1001 is specifically configured to input the image into a pre-trained recognition model, and obtain a position of a preset body part in the image, which is output by the recognition model.
[0145] Further, the processor 1001 is specifically configured to identify a type and a position of a skeletal point in an image by using a skeletal point detection algorithm.
[0146] The communication bus mentioned in the foregoing server can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0147] The communication interface is used for communication between the electronic device and other devices.
[0148] The memory can include a Random Access Memory (RAM) and can also include a Non-Volatile Memory (NVM), for example, at least one disk memory. Optionally, the memory can also be at least one storage device located away from the foregoing processor.
[0149] The processor can be a general processor, including a central processing unit, a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic, a discrete hardware component, etc.
[0150] On the basis of the above embodiments, the embodiments of the present application further provide a computer readable storage medium, wherein a computer program executable by an electronic device is stored in the computer readable storage medium, and when the program runs on the electronic device, the electronic device is caused to perform the following steps:
[0151] The memory stores a computer program, and when the program is executed by the processor, the processor performs the following steps:
[0152] Identify the type and position of the skeleton point in the image, determine the three-dimensional position of the skeleton point based on the position of the skeleton point in the image;
[0153] According to the three-dimensional positions of the target and the skeleton point in the image, determine the space region where the skeleton point is located; for any skeleton point, determine the change rate of the space region where the skeleton point is located according to the volume of the space region where the skeleton point is located and the volume of the space region where the skeleton point is located determined based on a preset number of frame images adjacent to the image in time sequence before the image; determine the target skeleton point with the highest change rate among the skeleton points;
[0154] According to the three-dimensional position of the target skeleton point, determine the corresponding target view angle, and switch the viewing view angle to the target view angle.
[0155] In a possible implementation, determining the change rate of the space region where the skeleton point is located according to the volume of the space region where the skeleton point is located and the volume of the space region where the skeleton point is located determined based on a preset number of frame images adjacent to the image in time sequence before the image includes:
[0156] Determine the volume of the space region where the skeleton point is located based on a preset number of frame images adjacent to the image in time sequence before the image; and determine the change rate of the space region where the skeleton point is located according to the variance of the determined volumes of the space region where the skeleton point is located.
[0157] In a possible implementation, the determining the change rate of the space region where the skeleton point is located according to the volume of the space region where the skeleton point is located and the volume of the space region where the skeleton point is located determined based on a preset number of frame images adjacent to the image in time sequence before the image includes:
[0158] determining a first change rate of the volume of the space region where the skeleton point is located between the previous frame of images and the image according to the volume of the space region where the skeleton point is located and the volume of the space region where the skeleton point is located determined based on the previous frame of images of the image;
[0159] and obtaining a preset number of second change rates saved for the previous frame of images and the skeleton point;
[0160] determining the change rate of the space region where the skeleton point is located according to the average of the first change rate and the second change rate.
[0161] In a possible implementation, the determining of the space region where the skeleton point is located according to the three-dimensional positions of the target and the skeleton point in the image comprises:
[0162] determining a three-dimensional convex hull containing the skeleton point by using a roll wrapping method or an incremental method, and determining the three-dimensional convex hull as the space region where the target is located;
[0163] dividing the space region according to the three-dimensional position of the skeleton point by using a Voronoi diagram generation method to obtain the space region where the skeleton point is located.
[0164] In a possible implementation, the determining of the space region where the skeleton point is located according to the three-dimensional positions of the target and the skeleton point in the image comprises:
[0165] identifying the region where the target is located in the image, determining the three-dimensional position of any pixel point in the region according to the position of the pixel point in the image, determining the distance between the three-dimensional position and the three-dimensional position of each skeleton point, determining the skeleton point with the minimum distance to the three-dimensional position, and determining the point at the three-dimensional position as a point located in the space region where the skeleton point is located.
[0166] In a possible implementation, the identifying of the region where the target is located in the image comprises:
[0167] inputting the image into a pre-trained target recognition model to obtain the region where the target is located in the image output by the target recognition model.
[0168] In a possible implementation, the determining of the corresponding target view angle according to the three-dimensional position of the target skeleton point comprises:
[0169] identifying the position of a preset body part in the image, and determining the three-dimensional position of the preset body part according to the position;
[0170] A ray is determined, with an endpoint being a three-dimensional position of the preset body part, and pointing to a direction of a three-dimensional position of the target skeletal point. A distance between the three-dimensional position corresponding to the view angle and the ray is determined according to the three-dimensional position corresponding to the view angle and the ray, and a view angle with a minimum distance is determined as a target view angle.
[0171] In a possible implementation, the determining the position of the preset body part in the image comprises:
[0172] The image is input into a pre-trained recognition model, and a position of the preset body part in the image output by the recognition model is obtained.
[0173] In a possible implementation, the recognizing the type and the position of the skeletal point in the image comprises:
[0174] The type and the position of the skeletal point in the image are recognized by a skeletal point detection algorithm.
[0175] Those skilled in the art should understand that embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0176] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and a combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device implemented in accordance with the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The functions specified in a flow or multiple flows and / or blocks
[0177] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction devices, which implement the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The functions specified in a flow or multiple flows and / or blocks
[0178] These computer program instructions can also be loaded into computer or other programmable data processing devices, so that a series of operation steps are performed on the computer or other programmable data processing devices to generate computer-implemented processes, thus the instructions executed on the computer or other programmable data processing devices provide a process for implementing the functions specified in the flowchart Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0179] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A view switching method, characterized by, The method comprises: identifying the type and location of the skeleton points in the image, determining the three-dimensional position of the skeleton points based on the location of the skeleton points in the image; determining the spatial region where the skeleton points are located according to the three-dimensional positions of the target and the skeleton points in the image; for any skeleton point, determining the change rate of the spatial region where the skeleton point is located according to the volume of the spatial region where the skeleton point is located and the volume of the spatial region where the skeleton point is located determined based on the preset number of frame images adjacent to the image in time sequence before the image; determining the target skeleton point with the highest change rate among the skeleton points; identifying the location of the preset body part in the image and determining the three-dimensional position of the preset body part according to the location; determining a ray with the three-dimensional position of the preset body part as an end point and pointing to the direction of the three-dimensional position of the target skeleton point; determining the distance between the three-dimensional position corresponding to the viewing angle and the ray according to the three-dimensional position corresponding to the viewing angle and the ray, and determining the target viewing angle as the viewing angle with the minimum distance; switching the viewing angle to the target viewing angle; The method comprises: determining the volume of the spatial region where the skeleton point is located based on the volume of the spatial region where the skeleton point is located and the volume of the spatial region where the skeleton point is located determined based on the preset number of frame images adjacent to the image in time sequence before the image; determining the change rate of the spatial region where the skeleton point is located according to the volume of the spatial region where the skeleton point is located and the volume of the spatial region where the skeleton point is located determined based on the preset number of frame images adjacent to the image in time sequence before the image; determining the target skeleton point with the highest change rate among the skeleton points; The method comprises: determining the volume of the spatial region where the skeleton point is located based on the volume of the spatial region where the skeleton point is located and the volume of the spatial region where the skeleton point is located determined based on the preset number of frame images adjacent to the image in time sequence before the image; determining the change rate of the spatial region where the skeleton point is located according to the volume of the spatial region where the skeleton point is located and the volume of the spatial region where the skeleton point is located determined based on the preset number of frame images adjacent to the image in time sequence before the image; determining the target skeleton point with the highest change rate among the skeleton points; The method comprises: using the roll wrapping method or the incremental method to determine the three-dimensional convex hull containing the skeleton points, and determining the three-dimensional convex hull as the spatial region where the target is located; using the generation method of the three-dimensional Voronoi diagram to divide the spatial region according to the three-dimensional positions of the skeleton points, and obtaining the spatial region where the skeleton points are located; or 2. The method of claim 1, wherein, identifying the region where the target is located in the image, determining the three-dimensional position of any pixel point in the region according to the location of the pixel point in the image, determining the distance between the three-dimensional position of each skeleton point and the three-dimensional position, determining the skeleton point with the minimum distance from the three-dimensional position, and determining the point at the three-dimensional position as a point located in the spatial region where the skeleton point is located. The method comprises: inputting the image into a pre-trained target recognition model to obtain the region where the target is located in the image output by the target recognition model.
3. The method of claim 1, wherein, The identifying the position of the preset body part in the image comprises: inputting the image into a pre-trained recognition model to obtain the position of the preset body part in the image output by the recognition model.
4. The method of claim 1, wherein, The identifying the type and position of the skeleton point in the image comprises: identifying the type and position of the skeleton point in the image by a skeleton point detection algorithm.
5. An electronic device, comprising: The electronic device at least comprises a processor and a memory, and the processor is configured to execute a computer program stored in the memory to implement the steps of the view angle switching method according to any one of claims 1-4.
Citation Information
Patent Citations
Interactive free three-dimensional display method for real three-dimensional scene
CN110087059A
Pedestrian attribute information extraction and re-identification method combined with graph convolutional neural network
CN114511870A