Image processing device, image processing method and program
By determining the correspondence relationship between frames of 3D shape data and using this information for shape fitting, the image processing device achieves highly accurate shape alignment in sports scenes captured by multiple cameras.
Patent Information
- Application Number
- JP2023075944
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-05-02
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2039-02-21
AI Technical Summary
In sports scenes captured by multiple cameras, the correspondence relationship between frames of three-dimensional shape data is often undetermined, making it difficult to achieve highly accurate shape fitting during the shape fitting process.
An image processing device that acquires and analyzes 3D shape data from multiple frames, determines the correspondence relationship between frames based on object movement analysis, and performs shape fitting using this correspondence information to align three-dimensional models across frames.
This approach enables highly accurate shape fitting by establishing clear correspondence relationships between frames, improving the precision of shape alignment and reducing data redundancy.
Smart Images

Figure 0007676463000001 
Figure 0007676463000002 
Figure 0007676463000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a shape fitting technique for three-dimensional shape data. [Background technology]
[0002] Conventionally, with regard to a technology for generating a virtual viewpoint image from images captured by multiple cameras, a technology has been proposed that realizes a compact data stream by performing time differential encoding when transmitting 3D shape data of people and objects in the captured scene (Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Alvaro Collet, and 8 others, "High-Quality Streamable Free-Viewpoint Video", [online], [Retrieved May 30, 2018], Internet<URL:http: / / hhoppe.com / fvv.pdf> Summary of the Invention [Problem to be solved by the invention]
[0004] When a sports game such as soccer or basketball is shot, a situation occurs in which multiple players and a ball move freely within the shooting space. In the 3D shape data across multiple frames generated from the shot images (moving images) of such a sports scene, the correspondence between the frames of the individual 3D shape data included in each frame is undetermined. When performing shape fitting processing on 3D shape data in which the correspondence between frames is undetermined, there is a possibility that shape fitting with high accuracy cannot be performed.
[0005] Therefore, an object of the present invention is to perform shape fitting processing with high accuracy. [Means for solving the problem]
[0006] The image processing device according to the present invention includes an acquisition unit for acquiring first three-dimensional shape data indicating a three-dimensional shape of an object in a first frame and second three-dimensional shape data indicating a three-dimensional shape of the object in a second frame, and a calculation unit for calculating the three-dimensional shape of elements constituting the first three-dimensional shape data. In 3D space The position of the element and the corresponding element constituting the second three-dimensional shape data In 3D space The direction that an element moves between frames based on its position and the amount of movement A means for identifying the moving direction and the movement amount and analyzing means for analyzing a movement of the object in a period including the first frame and the second frame based on the motion of the object. Effect of the Invention
[0007] According to the present invention, it is possible to perform highly accurate shape fitting processing. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram showing an example of a hardware configuration of an image processing device that performs shape fitting of three-dimensional shape data. [Diagram 2] FIG. 1 is a functional block diagram showing a software configuration related to shape fitting processing of an image processing device according to a first embodiment; [Diagram 3] FIG. 13 is a diagram showing an example of vertex coordinate information and mesh connection information. [Figure 4] 1 is a flowchart showing a process flow in an image processing device according to a first embodiment. [Diagram 5] (a) to (c) are diagrams explaining how to derive the correspondence between frames. [Figure 6] An example of the result of shape fitting processing [Figure 7] FIG. 11 is a functional block diagram showing a software configuration related to shape fitting processing of an image processing device according to a second embodiment. [Figure 8]Flowchart showing details of object tracking process [Figure 9] FIG. 1 is a diagram for explaining the effects of the second embodiment. [Figure 10] Diagram explaining grouping processing [Figure 11] FIG. 13 is a diagram showing an example of a result of grouping processing. [Figure 12] FIG. 11 is a functional block diagram showing a software configuration related to shape fitting processing of an image processing device according to a third embodiment. [Figure 13] A schematic diagram showing an example of the shape fitting result. [Figure 14] (a) and (b) are conceptual diagrams explaining the difference values of vertex coordinates. [Figure 15] A diagram explaining the data structure of a code string DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Note that the following embodiment does not limit the present invention, and not all of the combinations of features described in the present embodiment are necessarily essential to the solution of the present invention. Note that the same components will be described with the same reference numerals.
[0010] [Embodiment 1] In this embodiment, a correspondence relationship between frames is obtained for input 3D models for multiple frames, and shape fitting between 3D models is performed based on the obtained correspondence. In this specification, the term "object" refers to a dynamic object (foreground object) that moves within a target 3D space, such as a person or a ball, among various objects present in a shooting scene. In this specification, the term "3D model" refers to data (3D shape data) that represents the 3D shape of such a dynamic object, and may be referred to as a 3D model below.
[0011] (Hardware configuration of image processing device) FIG. 1 is a diagram showing an example of a hardware configuration of an image processing device that performs shape fitting of a three-dimensional model according to this embodiment. The image processing device 100 includes a CPU 101, a RAM 102, a ROM 103, a HDD 104, a communication I / F 105, an input device I / F 106, and an output device I / F 107. The CPU 101 is a processor that uses the RAM 102 as a work memory, executes various programs stored in the ROM 103, and generally controls each unit of the image processing device 100. The CPU 101 executes various programs to realize the functions of each unit shown in FIG. 2 described later. Note that the image processing device 100 may include one or more dedicated hardware or a GPU different from the CPU 101, and at least a part of the processing by the CPU 101 may be performed by the GPU or the dedicated hardware. Examples of the dedicated hardware include an ASIC (application specific integrated circuit) and a DSP (digital signal processor). The RAM 102 temporarily stores data supplied from the outside via the communication I / F 105, such as programs read from the ROM 103 and calculation results. The ROM 103 holds programs and data such as an OS that do not require modification. The HDD 104 is a large-capacity storage device that stores various data such as a three-dimensional model input from an external information processing device, and may be, for example, an SSD. The communication I / F 105 is an interface for receiving various data such as a three-dimensional model. The input device I / F 106 is an interface for connecting a keyboard 104 and a mouse 105 for a user to perform an input operation. The output device I / F 107 is an interface for connecting to a display device such as a liquid crystal display that displays information required by the user.
[0012] (Software configuration of image processing device) 2 is a functional block diagram showing a software configuration related to shape fitting processing of the image processing device 100 according to this embodiment. The image processing device 100 of this embodiment has a 3D model acquisition unit 201, a correspondence relationship derivation unit 202, and a registration unit 203. An overview of each unit will be described below.
[0013] The 3D model acquisition unit 201 acquires data of a three-dimensional model generated by an external information processing device (not shown) for multiple frames on a frame-by-frame basis. For example, a scene of a game such as soccer being played in a stadium is synchronously shot using multiple cameras in a video mode, and the external information processing device estimates the shapes of multiple objects such as players and a ball for each frame. It is assumed that the data of the three-dimensional model for multiple frames thus obtained is input. The 3D model acquisition unit 201 may acquire image data acquired based on shooting by multiple cameras, and generate a three-dimensional model based on the image data. This embodiment can also be applied to scenes other than sports scenes, such as concerts and plays.
[0014] A known technique such as a visual volume intersection method may be applied to estimate the three-dimensional shape of an object from its contour, which is included in a plurality of images (a group of images of the same frame) obtained by synchronous shooting with a plurality of cameras. In addition, there are a point cloud format, a voxel format, a mesh format, and the like as a representation format of a three-dimensional model. Although any representation format may be used for the three-dimensional model handled in this embodiment, the following description will be given taking the mesh format as an example. In addition, it is assumed that a plurality of objects are captured in each frame, and the number of objects is constant between frames without changing. However, one three-dimensional model is not necessarily generated for one object. There are also cases where one three-dimensional model is generated for a plurality of objects that exist in close positions. An example of this is when players come into contact with each other.
[0015] When expressing the three-dimensional shape of an object in a mesh format, it is necessary to determine a reference point (origin) in the target three-dimensional space in order to define the object shape by the three-dimensional coordinates (x, y, z) of each vertex and the connection information connecting the vertices. In this embodiment, the center position of the field of the stadium is defined as the origin (0,0,0). Then, the vertex coordinates (x, y, z) of each mesh expressing the three-dimensional shape of each object and the connection information between the vertices are input as a three-dimensional model. FIG. 3 shows an example of vertex coordinate information and mesh connection information when the surface shape of an object is expressed by a triangular mesh. In FIG. 3(a), T0 to T2 represent triangular meshes, and V0 to V4 represent vertices, respectively. The vertices of mesh T0 are V0, V1, and V2, the vertices of mesh T1 are V1, V2, and V3, and the vertices of mesh T2 are V0, V2, and V4. In this case, since the vertices V1 and V2 are common to mesh T0 and mesh T1, it can be seen that both meshes are adjacent to each other. Similarly, since vertices V0 and V2 are common to mesh T2 and mesh T0, it can be seen that the two meshes are adjacent to each other. The information defining the connection relationships between such vertices makes it possible to understand how each mesh is constructed, and to identify the surface shape of the object. As a specific data structure, for example, as shown in the table in FIG. 3(b), the vertices V constituting each triangular mesh T and the three-dimensional coordinates (x, y, z) of each vertex V may be held in a list format. Of course, data structures other than a list format may also be used.
[0016] The correspondence derivation unit 202 performs object tracking processing on the input 3D models for multiple frames to derive correspondence between frames. Then, based on the derivation result, information (hereinafter referred to as "correspondence information") is generated that allows the user to understand that the 3D models express the same object between a target frame and a next frame among the input multiple frames. The correspondence information may be, for example, information that allows the user to know from where to where a player moving in the field has moved between previous and next frames. For object tracking, a known method such as template matching or feature point matching may be used. By object tracking, it is determined which of the 3D models of multiple objects present in the previous frame corresponds to which of the 3D models of multiple objects present in the subsequent frame, and the correspondence information is generated based on the determination result. The generated correspondence information between frames of each 3D model is used for shape fitting processing in the positioning unit 203.
[0017] The positioning unit 203 performs shape fitting based on the correspondence relationship information of the three-dimensional models between the frame of interest and the next frame that is time-advanced therefrom. Here, shape fitting means matching the corresponding positions of the three-dimensional models between frames, and is synonymous with "nonlinear positioning" or "shape positioning". In the present embodiment, the specific processing content is a process of calculating the amount of movement for moving the coordinate positions of the vertices of the mesh that represents the shape of the three-dimensional model of interest so as to be close to the shape of the three-dimensional model that corresponds to the three-dimensional model of interest. Note that this amount of movement is a vector amount consisting of two components, a direction and a magnitude. By moving the vertex coordinate positions of the mesh that represents the shape of the three-dimensional model of interest according to the amount of movement obtained by this shape fitting, it is possible to generate three-dimensional shape data equivalent to the three-dimensional model that corresponds to the three-dimensional model of interest. In other words, if the three-dimensional model in the reference frame has vertex coordinate information and mesh connection information, it is possible to reconstruct the three-dimensional model in the next frame that corresponds to it by using the amount of movement of each vertex. This means that it is no longer necessary to have vertex coordinate information and mesh connection information required to identify the 3D shape of an object in every frame, and as a result, the amount of data can be reduced. For shape fitting, a known technique such as the ICP (Iterative Closest Point) algorithm can be used. The ICP algorithm is a method that defines a cost function by the sum of squares of the amount of movement between each point representing a 3D shape and its corresponding point, and performs shape fitting so as to minimize the cost function.
[0018] In this embodiment, the correspondence derivation unit 202 provided in the image processing device 100 executes a correspondence derivation process between frames of a three-dimensional model, but is not limited to this. For example, the correspondence derivation process may be executed in an external device, and the correspondence information resulting from the process may be input to the image processing device 100 and used in the registration unit 203. Furthermore, although data of a three-dimensional model spanning multiple frames is acquired from an external device and the correspondence therebetween is derived, image data from multiple viewpoints may be acquired, a three-dimensional model may be generated within the device itself, and the correspondence therebetween may be derived.
[0019] (Image processing flow) 4 is a flowchart showing the flow of processing in the image processing device 100 according to this embodiment. This flowchart is implemented by reading a control program stored in the ROM 103 into the RAM 102 and executing it by the CPU 101. In the following description, "S" stands for step.
[0020] In S401, the 3D model acquisition unit 201 acquires data of N frames (N is an integer equal to or greater than 2) of 3D models for which a virtual viewpoint video is to be generated from an external device or the HDD 104. For example, when generating a virtual viewpoint video from a 10-second video shot at 60 fps, 600 frames of 3D models are input. As described above, each frame includes multiple 3D models corresponding to two or more objects. The input data of the N frames of 3D models is stored in the RAM 202.
[0021] In S402, the correspondence derivation unit 202 determines two frames to be processed (a frame of interest and a next frame that is temporally advanced from the frame of interest) and derives a correspondence between the three-dimensional models between the two frames. Specifically, out of the input N frames of three-dimensional model data, the two frames of three-dimensional model data determined to be processed are read from the RAM 202, object tracking processing is performed, and correspondence information between the two frames is generated. For example, when 600 frames of three-dimensional models are input in S401, to derive correspondence between all frames, (N-1) pairs=599 pairs of frames are processed. However, the frame of interest and the next frame do not necessarily need to be consecutive, and the next frame may be determined by thinning out frames, for example, when the frame of interest is the first frame, the next frame is the third or fourth frame.
[0022] FIG. 5 (a) to (c) are diagrams for explaining the derivation of the correspondence between frames. Here, the explanation will be given assuming that, among the input N frames, the frame at a certain time t is the frame of interest, and the frame at time (t+1) is the next frame. Also, the three-dimensional models present in each frame are represented as M(1), M(2), ..., M(I). Here, "I" represents the total number of three-dimensional models contained in each frame. Now, FIG. 5 (a) is the frame of interest, and FIG. 5 (b) is the next frame, and in each frame, there are three-dimensional models corresponding to each of the three objects (player A, player B, and the ball). In other words, I=3. M(t,1) and M(t+1,1) represent the three-dimensional model of player A, M(t,2) and M(t+1,2) represent the three-dimensional model of the ball, and M(t,3) and M(t+1,3) represent the three-dimensional model of player B, and the correspondence is shown in the table of FIG. 5(c). That is, M(t,1) of the frame of interest corresponds to M(t+1,1) of the next frame, M(t,2) of the frame of interest corresponds to M(t+1,2) of the next frame, and M(t,3) of the frame of interest corresponds to M(t+1,3) of the next frame. In this way, information indicating the correspondence between the three-dimensional models between the frame of interest and the next frame is obtained. In this embodiment, each of the N frames includes the same number (I) of three-dimensional models, and the number is assumed to be constant without change.
[0023] In S403, the registration unit 203 initializes (sets i=1) a variable i for identifying one of the three-dimensional models M(t,1) to M(t,I) present in the frame of interest. The one three-dimensional model M(t,i) to be focused on in the subsequent processing is determined by this variable i.
[0024] In S404, the positioning unit 203 performs a shape fitting process between the three-dimensional model of interest M(t,i) in the frame of interest and the three-dimensional model of interest M(t+1,i) in the next frame based on the correspondence information obtained in S402. In the shape fitting process of this embodiment, a movement amount (hereinafter, referred to as a "movement vector") is obtained, which indicates how much and in which direction the vertices of the mesh representing the three-dimensional model of interest move while the frame of interest progresses in time from the frame of interest to the next frame. FIG. 6 shows the result of performing the shape fitting process in the specific example of FIG. 5(a) to (c) described above. In FIG. 6, the shaded portion represents the three-dimensional model of the frame of interest, and the non-shaded portion represents the three-dimensional model of the next frame. Then, a plurality of arrows pointing from the shaded three-dimensional model to the non-shaded three-dimensional model represent the movement vectors for each vertex of the three-dimensional model in the frame of interest. Note that there are more vertices than those shown in the figure, but here, representative vertices are illustrated.
[0025] In this way, the movement vector for each vertex held by the 3D model of the target frame is obtained by shape fitting. Even if the object shape is almost the same between frames, the 3D model is generated each time, so the position and number of vertices of the mesh will differ for each 3D model. For example, in the example of FIG. 6, even if only the feet of player A move during the transition from the target frame to the next frame, the position and number of vertices of the non-moving parts such as the head and torso are highly unlikely to match between frames. In other words, it is important to note that the movement vector for each vertex obtained by shape fitting is not a difference vector between each vertex of the 3D model in the target frame and each vertex of the corresponding 3D model in the next frame. In this way, the movement vector indicating how the vertex coordinates of the target 3D model in the target frame move in the next frame is output as the result of the shape fitting process.
[0026] In S405, it is determined whether or not there is an unprocessed 3D model among the 3D models present in the frame of interest. If the shape fitting process has been completed for all 3D models, the process proceeds to S407. On the other hand, if there is a 3D model for which the shape fitting process has not been completed, the process proceeds to S406, where the variable i is incremented (+1) and the next 3D model of interest is determined. After the variable i is incremented, the process returns to S404, and the same process is continued for the next 3D model of interest.
[0027] In S407, it is determined whether or not there are any unprocessed combinations for the input N frames of 3D models. If the derivation of correspondence relationships and shape fitting processing have been completed for all combinations, this processing ends.
[0028] The above is the processing content of the image processing device 100 according to this embodiment. The output result of the positioning unit 203 is not limited to a movement vector, and may be, for example, coordinate information of the vertices after movement. Furthermore, a motion analysis process may be performed using the output movement vector and the vertex coordinates after movement, and the results may be output. For example, the output movement vector may be used to perform an analysis of the movement (direction and speed of movement) of the ball or a player in a specific section (between specific frames), or the movement of the hands and feet of a specific player, and further, a prediction of the future movement of the player or the ball, and the results may be output.
[0029] <Modification> Furthermore, for the purpose of improving shape accuracy, the result of shape fitting based on the above-mentioned correspondence information can be used to compensate for defects in a specific 3D model. For example, due to the posture and position of the object, the shooting conditions, the generation method, etc., defects may occur in the shape of the 3D model to be generated, and the object shape may not be correctly estimated in a single frame. Even in such a case, the shape of the object may be correctly reproduced in the corresponding 3D models in the previous and next frames. Therefore, shape fitting is performed between the 3D model of interest and each 3D model in the corresponding other frame. For example, shape fitting is performed with the 3D model of interest using two 3D models that correspond in the previous and next frames of the frame containing the 3D model of interest. As a result, even if there is, for example, a defect in the shape of the 3D model of interest, the defect can be compensated for by referring to the shapes of the corresponding 3D models in the previous and next frames. By performing shape fitting with the 3D models of other frames in this way, it is possible to obtain a 3D model with higher accuracy than one generated independently in each frame.
[0030] According to this embodiment, when performing shape fitting, information that defines the correspondence relationship of three-dimensional models between frames is used, thereby making it possible to realize shape fitting processing with high accuracy.
[0031] [Embodiment 2] In the first embodiment, an example was described in which shape fitting is performed by deriving correspondence between three-dimensional models between two target frames, on the premise that the same number of three-dimensional models are included in all of the input N frames. Next, an aspect in which shape fitting is performed by deriving correspondence between three-dimensional models between consecutive frames, on the premise that the number of three-dimensional models included in each frame may change, will be described as the second embodiment. Note that the description of the contents common to the first embodiment will be omitted or simplified, and the following description will focus on the differences.
[0032] 7 is a functional block diagram showing a software configuration related to shape fitting processing of the image processing device 100 according to this embodiment. A major difference from the first embodiment is that a grouping unit 701 is provided. By performing grouping processing of 3D models in the grouping unit 701, accurate shape fitting is possible even when merging / separation of 3D models occurs and the number of 3D models changes between frames.
[0033] (3D model combination / separation) For example, in a game such as soccer, when objects such as players and a ball exist at separate positions on the field (within the target three-dimensional space), each object is represented as a separate three-dimensional model. However, individual objects do not always exist at separate positions during a game. For example, in a scene where players are fighting over the ball or dribbling, players or a player and a ball may come close to each other or may even come into contact with each other. In such a case, two or more players or a player and a ball are combined and represented as one three-dimensional model. The representation of two or more objects as one three-dimensional model in this way is called "combination of three-dimensional models." In addition, the case where objects that were close to each other or in contact in a certain frame move away from each other, and a single combined three-dimensional model is represented as separate three-dimensional models in the next frame, is called "separation of three-dimensional models."
[0034] (Details of the correspondence derivation part) Before describing the grouping unit 701, the correspondence derivation unit 202' of this embodiment will be described first. FIG. 8 is a flowchart showing details of the object tracking process performed by the correspondence derivation unit 202' of this embodiment. In the following, a frame of interest among the input N frames to be processed will be expressed as the "nth frame". In this case, "n" is an integer equal to or greater than 1, and "N" is an integer equal to or greater than 2, where N≧n. As described above, the number of 3D models included in each frame may vary in this embodiment. Therefore, the number of 3D models included in the nth frame is defined as I(n). Then, the 3D model with number i included in the nth frame is defined as M(n, i). In this case, "i" is an integer equal to or greater than 1, where I≧i. In the following description, "S" means a step. It is assumed that the numbers of the 3D models are assigned without duplication.
[0035] First, in S801, the input data of the three-dimensional model for N frames is read from the RAM 202, and the position P(n,i) of the three-dimensional model M(n,i) in each frame is obtained. An example of this position P(n,i) is the center of gravity of each three-dimensional model. The center of gravity can be obtained by calculating the average value of all vertex coordinates of the mesh that constitutes the three-dimensional model. Note that the position P(n,i) is not limited to the center of gravity as long as the position of each three-dimensional model can be specified. Once the position P(n,i) of each three-dimensional model M(n,i) for N frames has been obtained, the process proceeds to S802.
[0036] In S802, a variable n indicating the number of a frame of interest is initialized (set to n=1). Then, in the following S803, the similarity between each three-dimensional model present in the nth frame, which is the frame of interest, and each three-dimensional model present in the (n+1)th frame is calculated. In this embodiment, the distance D between the position P(n,1)-P(n,I) of each three-dimensional model in the nth frame and the position P(n+1,1)-P(n+1,I) of each three-dimensional model in the (n+1)th frame is calculated as an index indicating the similarity. For example, when the number of three-dimensional models included in both frames is three (I=3), the distance between the center of gravity positions (the distance between two points) is calculated for each of the following combinations:
[0037] Position P(n,1) and position P(n+1,1) Position P(n,1) and position P(n+1,2) Position P(n,1) and position P(n+1,3) Position P(n,2) and position P(n+1,1) Position P(n,2) and position P(n+1,2) Position P(n,2) and position P(n+1,3) Position P(n,3) and position P(n+1,1) Position P(n,3) and position P(n+1,2) Position P(n,3) and position P(n+1,3)
[0038] In this embodiment, the positional relationship between the three-dimensional models in the target three-dimensional space is used as a criterion, and the closer the distance (the closer the distance) between the three-dimensional model of interest is, the higher the similarity is evaluated. The evaluation index of the similarity is not limited to the distance between the three-dimensional models. For example, by further considering the shape and size of the three-dimensional model, texture data, and the moving direction of the object as elements other than the distance, a more accurate similarity can be obtained.
[0039] In S804, based on the similarity calculated in S803, each three-dimensional model included in the nth frame is associated with a three-dimensional model included in the (n+1)th frame. Specifically, a process is performed in which each three-dimensional model included in one frame is associated with a three-dimensional model with the highest similarity among the three-dimensional models included in the other frame. As described above, in the case of this embodiment in which the closeness of the distance between the center positions is used as the similarity, the three-dimensional models with the smallest value of the distance D calculated in S803 are associated with each other. At that time, the same identifier (ID) is given to both of the associated three-dimensional models. Any ID may be used as long as it can identify the object, but an integer value is used in this embodiment. For example, it is assumed that "2" is given as the ID of the three-dimensional model M(n,3). Then, if the three-dimensional model M(n+1,5) in the (n+1)th frame exists at the closest distance, the same ID "2" is given to the three-dimensional model M(n+1,5).
[0040] In S805, it is determined whether or not there is a 3D model that is not associated with any 3D model in either the nth frame or the (n+1)th frame. This is because in the present embodiment, in which the number of 3D models included in the nth frame and the (n+1)th frame may differ, as a result of associating 3D models in descending order of similarity, there may be cases in which there is no 3D model remaining in the other frame to which it can be associated. If there is a 3D model remaining in either frame that is not associated with any 3D model, proceed to S806, and if there is no 3D model remaining, proceed to S807.
[0041] In S806, a process is performed in which a 3D model that is not associated with any other 3D model is associated with a 3D model that is located closest to the 3D model among the 3D models included in the other frame.
[0042] In S807, it is determined whether the current (n+1)th frame is the final frame of the input N frames. If the result of the determination is that the current (n+1)th frame is not the final frame, the process proceeds to S808, where the variable n is incremented (+1), and the next frame of interest is determined. After the variable n is incremented, the process returns to S803, and the same process is continued for the next nth frame and the (n+1)th frame of the 3D model. On the other hand, if the current (n+1)th frame is the final frame, this process ends.
[0043] The above is the content of the object tracking process performed by the correspondence deriving unit 202' according to this embodiment.
[0044] In the object tracking results of this embodiment obtained as described above, each three-dimensional model is assigned an ID, and three-dimensional models that are associated with each other have the same ID throughout N frames of captured scenes. The effect obtained by this embodiment will be described with reference to FIG. 9. FIG. 9 is a diagram that shows a schematic diagram of a result of performing inter-frame association of each three-dimensional model included in each frame with three frames (N=3) as processing targets. Now, three three-dimensional models M(n,1), M(n,2), and M(n,3) exist in each frame. In FIG. 10, the lines connecting the three-dimensional models represent their respective correspondences. For example, the three-dimensional model M(1,1) in the first frame corresponds to the three-dimensional model M(2,2) in the second frame, and further, the M(2,2) corresponds to the three-dimensional model M(3,3) in the third frame. In this case, the three-dimensional models M(1,1), M(2,2), and M(3,3) have the same ID. In this case, M(1,1) to M(1,I) existing in the first frame are assigned ID values of "1 to I(1)", and the IDs of the corresponding three-dimensional models are assigned to the three-dimensional models in the next frame and thereafter. In the example of FIG. 9, M(1,1), M(2,2), and M(3,3) are assigned ID=1, M(1,2), M(2,1), and M(3,2) are assigned ID=2, and M(1,3), M(2,3), and M(3,1) are assigned ID=3. In the example shown in FIG. 9, one three-dimensional model is always associated with one three-dimensional model. However, as described above, two or more three-dimensional models may be associated with one three-dimensional model. For example, when multiple three-dimensional models in the (n+1)th frame are associated with one three-dimensional model in the nth frame, the IDs of the three-dimensional models in the nth frame are assigned to the multiple three-dimensional models. Similarly, when multiple 3D models in the nth frame correspond to one 3D model in the (n+1)th frame, all of the IDs that the multiple 3D models in the nth frame each have are assigned to that one 3D model.Furthermore, when a three-dimensional model having multiple IDs in the nth frame corresponds to multiple three-dimensional models in the (n+1)th frame, all of the multiple IDs held by the three-dimensional model in the nth frame are assigned to each of the multiple three-dimensional models. In this manner, in this embodiment, at least one ID (hereinafter, referred to as a "tracking result ID") assigned as a result of object tracking is assigned to each three-dimensional model. Note that the tracking result ID assigned to each three-dimensional model constitutes a part of the data of the three-dimensional model together with vertex coordinate information. The way in which the tracking result ID is held is not limited to this, and a list describing the IDs assigned to each three-dimensional model may be generated separately from the three-dimensional model.
[0045] (Grouping details) Next, the grouping unit 401 will be described. For example, when two independent three-dimensional models in a certain frame are combined into one three-dimensional model in the next frame, even if one of the three-dimensional models is fitted to the shape of the combined three-dimensional model, the shape fitting cannot be performed correctly. In other words, when a combination or separation of three-dimensional models occurs, the process may fail unless the shape fitting is performed taking this into consideration. Therefore, the grouping unit 701 judges the combination or separation of the above-mentioned three-dimensional models based on the correspondence information obtained by the correspondence derivation unit 202', and performs a process (grouping process) of grouping the corresponding three-dimensional models into one. For the sake of simplicity, it is assumed that no new objects appear or no objects disappear during the shooting scene (no increase or decrease of objects between N frames).
[0046] FIG. 10 is a diagram showing grouping processing in a certain shooting scene. In FIG. 10, ellipses indicate 3D models present in each frame, and the numbers inside the ellipses indicate the tracking result IDs of the 3D models. The grouping processing is performed based on the tracking result ID given to the 3D model present in the final frame. In the example of FIG. 10, the tracking result IDs of 3D models 3_1 to 3_7 in the third frame, which is the final frame of the input N frames (N=3), are the basis for grouping. Each group generated by the grouping processing is given an ID (here, an alphabet) to distinguish it from other groups.
[0047] For example, if we focus on 3D model 3_1 (tracking result ID = 1) in the third frame, there is no other 3D model with the same ID in the third frame. Therefore, it can be determined that no separation or combination of 3D models occurred throughout the shooting scene. In this case, three 3D models (1_1, 2_1, 3_1) with the same tracking result ID of "1" form one group (group A).
[0048] Next, when we look at 3D model 3_2 in the third frame, it has tracking result IDs of "2" and "3". In this case, it can be determined that the 3D model with ID=2 and the 3D model with ID=3 were combined in the middle of the shooting scene. In fact, 3D models 1_2 and 1_3, which existed independently in the first frame, are combined in the second frame to become one 3D model 2_2. And, in the third frame, there are no other 3D models with ID=2 or ID=3. Therefore, the four 3D models (1_2, 1_3, 2_2, 2_3) with ID=2 or ID=3 are grouped together (group B).
[0049] Next, looking at 3D models 3_3 and 3_4 in the third frame, both have a tracking result ID of "4". In this case, it can be determined that the 3D model with ID=4 separated during the shooting scene. In fact, 3D model 1_4, which was one in the first frame, separates into separate 3D models 2_3 and 2_4 in the second frame. Therefore, five 3D models with ID=4 (1_4, 2_3, 2_4, 3_3, 3_4) form one group (group C).
[0050] Next, when focusing on the 3D model 3_5 in the third frame, it has the tracking result IDs "5" and "6". In this case, it can be determined that the 3D model having ID=5 and the 3D model having ID=6 were combined in the middle of the shooting scene. In fact, the 3D model 1_5 and the 3D model 1_6, which existed independently in the first frame, were combined in the second frame to become one 3D model 2_5. And, since there are other 3D models having ID=5 or ID=6 in the third frame, it can be determined that separation also occurred in the middle of the shooting scene. In fact, the 3D model 2_5 in the second frame is separated into the 3D model 3_5 and the 3D model 3_6 in the third frame. Also, the 3D model 3_6 has the tracking result IDs "5" and "6" as well as "7". Therefore, it can be determined that the 3D model 3_6 was generated by combining with the 3D model having ID=7. Therefore, eight 3D models having IDs of 5 to 7 (1_5, 1_6, 1_7, 2_5, 2_6, 3_5, 3_6, 3_7) are configured into one group (Group D).
[0051] As described above, in the example of FIG. 10, the images are grouped into four groups A to D. The above-mentioned grouping method is merely an example, and any method may be used as long as it is possible to group together the mutually related 3D models that have been separated and combined between frames into one group. Furthermore, the results of the grouping process may be output in any data format as long as it is known which 3D models each group is made up of. For example, a list showing the IDs of the 3D models that belong to a group may be generated for each group, or each 3D model may be assigned the ID (alphabet or number) of the group to which it belongs.
[0052] (Details of the alignment part) Next, the alignment unit 203' of this embodiment will be described. The alignment unit 203' performs shape fitting processing on a group basis (a unit of a set of 3D models) that is the processing result of the grouping unit 701. Specifically, first, shape fitting is performed from the target 3D model M(n,i) in the nth frame to the 3D model M(n+1,i) in the (n+1)th frame that corresponds thereto. Here, the approximate shape data of the corresponding 3D model obtained by moving the vertex coordinates of the target 3D model is called an "estimated 3D model". Then, shape fitting is performed from the obtained estimated 3D model to the 3D model M(n+2,i) in the (n+2)th frame that corresponds to the 3D model M(n+1,i) in the (n+1)th frame, and similar processing is repeated within the group. Here, "i" is a value that can change for each frame and is not necessarily the same value for all frames. As a result of the processing, the positioning unit 203' outputs the vertex coordinate information of the reference 3D model in the group (target 3D model) and each estimated 3D model obtained by repeating the shape fitting. In the case of this embodiment, it is sufficient to hold the mesh connection information of the reference 3D model in each group (it is not necessary to hold the mesh connection information for all 3D models belonging to the group). Therefore, it is possible to reduce the amount of data by sharing the mesh connection information.
[0053] 11 is a diagram showing an example of a result of grouping processing. Shape fitting in units of groups according to this embodiment will be described with reference to FIG.
[0054] The group indicated by the thick frame 1201 is a group composed of three-dimensional models in which no joining or separation has occurred. In the case of this group 1201, first, shape fitting is performed from the three-dimensional model M(1,1) of the first frame to the three-dimensional model M(2,2) of the second frame. The estimated three-dimensional model generated by this shape fitting is represented as "M'(n,i)". That is, the estimated three-dimensional model M'(2,2) is generated in the first shape fitting. Then, shape fitting is performed between the generated estimated three-dimensional model M'(2,2) and the three-dimensional model M(3,3) of the third frame, and the estimated three-dimensional model M'(3,3) is generated. In this case, it should be noted that the estimated three-dimensional model M'(3,3) is generated from M'(2,2) and not from M(2,2). In this way, the 3D models M(2,2) and M(3,3) of the second and third frames belonging to group 1201 can be approximated by deformation caused by moving the vertices of the 3D model M(1,1) of the first frame, if the data (here, vertex information and mesh connection information) of the 3D model M(1,1) is available.
[0055] The group indicated by the thick frame 1202 is a group in which a combination of three-dimensional models occurs. In the case of this group 1202, first, shape fitting is performed from two three-dimensional models M(1,2) and M(1,3) in the first frame to one three-dimensional model M(2,1) in the second frame. Then, shape fitting is performed between the estimated three-dimensional model M'(2,1) obtained by the shape fitting and M(3,2) in the third frame to generate an estimated three-dimensional model M'(3,2). In this way, the three-dimensional models M(2,1) and M(3,2) in the second and third frames belonging to group 1202 can also be approximated by deformation by moving the vertices if data of the three-dimensional models M(1,2) and M(1,3) in the first frame is available.
[0056] The group indicated by the thick frame 1203 is a group in which a combination and separation of three-dimensional models occurs. In the case of group 1203, similar to the case of group 1202, shape fitting is performed from two three-dimensional models M(1,4) and M(1,5) in the first frame to one three-dimensional model M(2,3) in the second frame. Then, shape fitting is performed between the estimated three-dimensional model M'(2,3) obtained by the shape fitting and two three-dimensional models M(3,1) and M(3,4) in the third frame to generate estimated three-dimensional models M'(3,1) and M'(3,4). In this way, the three-dimensional models M(2,3), M(3,1), and M(3,4) in the second and third frames belonging to group 1203 can be approximated by deformation by moving the vertices of the three-dimensional models M(1,4) and M(1,5) in the first frame, if there is data of those models.
[0057] As described above, according to this embodiment, by performing shape fitting based on the grouping process result, even when three-dimensional models are combined or separated, it is possible to perform correct shape fitting.
[0058] [Embodiment 3] Next, an aspect of compressing the amount of data of a time-series three-dimensional model by calculating the time difference of the output result from the position matching unit and encoding it will be described as embodiment 3. Note that the description of the contents common to embodiments 1 and 2 will be omitted or simplified, and the following description will focus on the differences.
[0059] FIG. 12 is a functional block diagram showing a software configuration related to shape fitting processing of the image processing device 100 according to this embodiment. A major difference from the second embodiment is that a time difference encoding unit 1301 is provided. The time difference encoding unit 1301 performs time difference encoding using as input the result of the shape fitting processing in the position adjustment unit 203". This will be described in detail below.
[0060] (Details of the alignment part) First, as a preprocessing step for performing time differential encoding, the alignment unit 203" of this embodiment selects a reference frame for performing shape fitting for each group. In the second embodiment, shape fitting was performed using the first frame of the input N frames as the reference frame, but in this embodiment, the reference frame is first selected in accordance with the following conditions. If the 3D model in the reference frame does not accurately reproduce the shape of the object, or if two or more objects are in contact, the accuracy of shape fitting will decrease. Therefore, the decrease in accuracy is suppressed by selecting a frame that is more suitable as the reference frame.
[0061] <Reference frame selection criteria> The surface area of the 3D model in the frame is larger than that of other frames. The shape of the 3D model in the frame does not contain any loops (holes) (or has fewer loops than other frames) - The number of 3D models in one frame is greater than in other frames
[0062] The first condition is related to the surface area of the 3D model. If the surface area is smaller than that of other 3D models expressing the same object, there is a possibility that the object shape is not reproduced accurately due to defects in the 3D shape. Therefore, priority is given to frames that contain 3D models with a larger surface area. The second condition is related to the shape represented by the 3D model. If a 3D model contains a ring or hole, it is highly likely that the shape of the part that should be the tip of the object is hidden. For example, in a 3D model in which a person has his or her hand on his or her waist, a "ring" is formed by the body and arm. In other words, the "hand", which is one part of the object, is in contact with the "waist", which is another part, and the shape of the "hand" cannot be obtained. If a frame containing such a 3D model is used as the reference frame, shape fitting will be performed from a 3D model in which parts that would not normally be in contact are in contact, and shape accuracy will decrease. Therefore, priority is given to frames that contain 3D models that do not contain rings (holes) in their shape. The third condition is related to the number of 3D models. When the number of 3D models present in a frame is small, there is a possibility that two or more objects are in contact and are represented by a single 3D model. Therefore, frames with a large number of 3D models are prioritized. By selecting the frame that is most suitable as the reference frame in light of these three conditions and performing shape fitting using the 3D models present in that frame as the reference, a more accurate estimated 3D model can be obtained.
[0063] Fig. 13 is a diagram showing a schematic example of a result of shape fitting in this embodiment. In the example of Fig. 13, the scene to be processed is composed of five frames (N=5). Each frame includes two (first and fifth frames) or three (second to fourth frames) three-dimensional models shown in thick frames, which are grouped into two groups 1401 and 1402. In each of the groups 1401 and 1402, the third and second frames are selected as reference frames, respectively, based on the above three conditions. In this case, for group 1401, the alignment unit 203" performs shape fitting using the 3D model existing in the third frame as a reference. As a result, as shown in the upper part of Figure 13, estimated 3D models that do not have mesh connection information are generated for 3D models other than the 3D model in the third frame that serves as reference. Similarly, for group 1402, as a result of shape fitting, estimated 3D models that have vertex coordinate information but do not have mesh connection information are generated for 3D models other than the two 3D models in the second frame that serve as reference. Then, the alignment unit 203" outputs both vertex coordinate information and mesh connection information for 3D models that exist in the reference frame, and outputs only vertex coordinate information for estimated 3D models that exist in non-reference frames.
[0064] (Details of the time differential encoding section) The time difference encoding unit 1301 performs a time difference encoding process using the result of shape fitting for each group obtained by the position adjustment unit 203". Specifically, for each group, a difference value of vertex coordinates between corresponding 3D models is calculated in sequence starting from the first frame, and the obtained difference value is quantized and encoded. (a) and (b) of FIG. 14 are conceptual diagrams for explaining the difference value of vertex coordinates. In FIG. 14(a), 3D models of a person for four frames from t=0 to 3 are present. The vertex coordinates near the right elbow in each 3D model are indicated by black circles, and the passage of time is indicated by arrows. As shown in FIG. 14(a), the difference value p between the vertex coordinate V0 of the black circle in the frame at t=0 and the vertex coordinate V1 of the black circle in the frame at t=1 is p=V0-V1. FIG. 14(b) is a diagram showing a three-dimensional representation of this difference value p. The difference value p may be expressed in any data format, such as a vector having two components, direction and distance.
[0065] FIG. 15 is a diagram for explaining the data structure of a code string output as a result of a time difference encoding process. Header information necessary for decoding encoded data, such as information on the number of frames and the number of groups, is added to the beginning of the code string. The header information is followed by encoded data for each group. The encoded data for each group is subdivided as follows. First, at the beginning, there is group header information such as information on a reference frame necessary for decoding a three-dimensional model constituting a group. Then, mesh connection information common to the group is followed, followed by vertex coordinate information of a three-dimensional model present in the first frame, and time difference encoded data from the second frame to the final frame, in that order. Note that the data structure of the code string shown in FIG. 15 is an example, and any format may be used as long as the data structure allows the three-dimensional model of each frame constituting a shooting scene to be decoded. In addition, any method may be used for the encoding method as long as it is an entropy encoding method such as Golomb encoding. In addition, in the present embodiment, an example has been described in which, for each group, difference values are calculated in order using the vertex coordinates of a three-dimensional model included in the first frame as a reference, and time difference encoding is performed, but the encoding method is not limited to this. For example, the start position of the encoding may be changed, such as using the vertex coordinates of the central frame or the final frame as a reference, or interframe prediction may be performed using a frame set called a GOP (Group of Pictures). In this embodiment, mesh connection information and vertex coordinate information of a three-dimensional model existing in the first frame are not subject to encoding, but these may also be subject to encoding.
[0066] As described above, according to this embodiment, by performing a time difference encoding process using the results of shape fitting for each group, it is possible to achieve both a reduction in the amount of data and accurate shape reproduction of each three-dimensional model.
[0067] (Other Examples) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions. [Explanation of symbols]
[0068] 201 3D Model Acquisition Department 202 Correspondence Derivation Part 203 Alignment section
Claims
1. an acquiring means for acquiring first three-dimensional shape data indicating a three-dimensional shape of an object in a first frame and second three-dimensional shape data indicating a three-dimensional shape of the object in a second frame; a specifying means for specifying a moving direction and a moving amount of an element between frames based on a position in a three-dimensional space of an element constituting the first three-dimensional shape data and a position in a three-dimensional space of an element constituting the second three-dimensional shape data corresponding to the said element; an analysis means for analyzing a movement of the object during a period including the first frame and the second frame based on the movement direction and the movement amount; An image processing device comprising:
2. The image processing device according to claim 1 , wherein the analysis comprises predicting a movement of the object.
3. The image processing apparatus according to claim 1 , wherein the analyzing means analyzes, when the object is a person, a movement of a hand or a foot of the person.
4. the acquiring means acquires information indicating a correspondence relationship between the first three-dimensional shape data and the second three-dimensional shape data; the specifying means associates positions in a three-dimensional space of elements constituting the first three-dimensional shape data with positions in a three-dimensional space of elements constituting the second three-dimensional shape data based on the information; The image processing device according to claim 1 .
5. 5. The image processing device of claim 4, wherein the information indicating the correspondence is information indicating that, when two or more objects in a first frame among a plurality of frames are represented by independent three-dimensional shape data, and the two or more objects are represented by a single three-dimensional shape data in the second frame, the two or more independent three-dimensional shape data in the first frame correspond to a single three-dimensional shape data in the second frame.
6. 6. The image processing device according to claim 5, wherein the information indicating the correspondence is information indicating that, when two or more objects are represented by one three-dimensional shape data in the first frame among a plurality of frames, and the two or more objects are each represented by independent three-dimensional shape data in the second frame, the one three-dimensional shape data in the first frame corresponds to two or more independent three-dimensional shape data in the second frame.
7. A program for causing a computer to function as each of the means of the image processing apparatus according to any one of claims 1 to 6.
8. an acquiring step of acquiring first three-dimensional shape data indicating a three-dimensional shape of an object in a first frame and second three-dimensional shape data indicating a three-dimensional shape of the object in a second frame; a specifying step of specifying a movement direction and a movement amount of an element between frames based on a position in a three-dimensional space of an element constituting the first three-dimensional shape data and a position in a three-dimensional space of an element constituting the second three-dimensional shape data corresponding to the element; an analysis step of analyzing a movement of the object during a period including the first frame and the second frame based on the movement direction and the movement amount; An image processing method comprising the steps of:
Citation Information
Patent Citations
Method and device for detecting motion of moving image
JP1990239376A
Moving image processor
JP1996149458A
Moving image coder
JP1998013840A
Image processing apparatus and image processing method
JP2018055643A
Image processing apparatus and image processing method
WO2019003953A1