Image processing apparatus, image processing method, and program

By estimating the three-dimensional shape of the moving object and generating predicted images, the problem of identifying the same moving object under different camera perspectives is solved, and the accurate identification of the moving object identity and the improvement of robustness are achieved.

JP2025076600APending Publication Date: 2025-05-16CANON KK
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
JP2023188252
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately identify the same moving objects at different cameras' installation positions and perspective angles, because the feature vectors will have a large difference at different perspectives.

Method used

By extracting the moving objects on the first and second cameras, the three-dimensional shape of the moving object is estimated, and a predicted image is generated that displays the shape and state of the first moving object at the perspective of the second camera, thereby identifying the identity of the moving object by matching the predicted image and the image of the second moving object.

Benefits of technology

It realizes accurate identification of the same moving objects under different camera perspectives, improving the robustness of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025076600000001_ABST
    Figure 2025076600000001_ABST
Patent Text Reader

Abstract

To accurately identify the same movable body whose images are picked up by different cameras.SOLUTION: An image processing method includes: extracting a first movable body from a first image picked up by a first imaging unit; extracting a second movable body from a second image picked up by a second imaging unit; estimating a three-dimensional shape of the first movable body on the basis of a result of extraction of the first movable body; estimating the state of the second movable body on the basis of a result of extraction of the second movable body; creating a predicted image when an image of the first movable body is picked up by the second imaging unit, on the basis of a result of estimation of the three-dimensional shape of the first movable body and a result of estimation of the state of the second movable body; and collating the predicted image with an image of the second movable body extracted from the second image to identify the sameness of the first movable body and the second movable body.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image processing device, an image processing method, and a program. [Background technology]

[0002] In recent years, there has been an increasing demand for improving business activities in large commercial facilities and maintaining security in smart cities. In such use cases, it is expected that multiple cameras will be installed in a large area. In order to record a person's behavior log from camera footage and use it for analysis, it is useful to track the behavior of a specific person, and for this purpose, a system that can identify the same person across cameras is required. Non-Patent Document 1 discloses a method for identifying a person using a feature vector based on pixel information of each area obtained by dividing a person's image in the height direction. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Y. Sun, “Beyond Part Models: Person Retrieval with Refined Part Pooling”, ECCV2018. Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the method disclosed in Non-Patent Document 1, it may be difficult to identify the same person between cameras because the installation positions and angles of view of the cameras are different. For example, in the upper body area of ​​a person wearing a white shirt and carrying a large black backpack, white pixel information is prominent when viewed from the front, whereas black pixel information is prominent when viewed from the back, resulting in a large difference in the feature vector.

[0005] Therefore, an object of the present invention is to accurately identify the same moving object captured by different cameras. [Means for solving the problem]

[0006] The image processing device of the present invention is characterized by having a first extraction means for extracting a first moving body from a first image captured by a first imaging unit, a second extraction means for extracting a second moving body from a second image captured by a second imaging unit, a first estimation means for estimating a three-dimensional shape of the first moving body based on a first extraction result which is the extraction result by the first extraction means, a second estimation means for estimating a state of the second moving body based on a second extraction result which is the extraction result by the second extraction means, a generation means for generating a predicted image when the first moving body is captured by the second imaging unit based on the estimation result of the three-dimensional shape of the first moving body and the estimation result of the state of the second moving body, and an identification means for identifying the identity of the first moving body and the second moving body by comparing the predicted image with the image of the second moving body extracted by the second extraction means. Effect of the Invention

[0007] According to the present invention, the same moving object captured by different cameras can be identified with high accuracy. [Brief description of the drawings]

[0008] [Figure 1] FIG. 2 is a diagram illustrating an example of a hardware configuration of an image processing system. [Diagram 2] 1 is a diagram illustrating an example of a functional configuration of an image processing device according to a first embodiment. [Diagram 3] 4 is a flowchart showing a process performed by the image processing device according to the first embodiment. [Figure 4] FIG. 2 is a diagram showing a specific example of an image to be processed in the first embodiment. [Diagram 5] FIG. 2 is a diagram showing a data flow in the image processing device according to the first embodiment. [Figure 6] 5 is a diagram for explaining the processing of a determination unit according to the first embodiment. FIG. [Figure 7]FIG. 11 is a diagram illustrating an example of a functional configuration of an image processing device according to a second embodiment. [Figure 8] FIG. 11 is a diagram showing a specific example of an image to be processed in the second embodiment. [Figure 9] FIG. 11 is a diagram showing a data flow in an image processing device according to a second embodiment. [Figure 10] FIG. 11 is a diagram illustrating an example of a functional configuration of an image processing device according to a third embodiment. [Figure 11] FIG. 13 is a diagram showing a specific example of an image to be processed in the third embodiment. [Figure 12] FIG. 11 is a diagram showing a data flow in an image processing device according to a third embodiment. [Figure 13] FIG. 13 is a diagram showing a specific example of an image to be processed in the fourth embodiment. [Figure 14] 13 is a flowchart showing the processing of an image processing device according to a fourth embodiment. [Figure 15] FIG. 11 is a diagram showing a data flow in an image processing device according to a fourth embodiment. [Figure 16] FIG. 13 is a diagram illustrating an example of a functional configuration of an image processing device according to a fifth embodiment. [Figure 17] FIG. 13 is a diagram showing a specific example of an image to be processed in the fifth embodiment. [Figure 18] FIG. 23 is a diagram showing a specific example of an image to be processed in the sixth embodiment. [Figure 19] FIG. 13 is a diagram showing a data flow in an image processing device according to a sixth embodiment. [Figure 20] FIG. 13 is a diagram illustrating an example of a functional configuration of an image processing device according to a seventh embodiment. [Figure 21] FIG. 13 is a diagram showing a data flow in an image processing device according to a seventh embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Hereinafter, an embodiment of the present invention will be described in detail with reference to the accompanying drawings. Note that the configurations shown in the following embodiments are merely examples, and the present invention is not limited to the configurations shown in the drawings.

[0010] <Embodiment 1> 1 shows an example of the hardware configuration of an image processing system. The image processing system includes multiple cameras 11A, 11B, . . . , 11C arranged in a predetermined space, and an image processing device 10. The image processing device 10 is connected to each of the multiple cameras 11A to 11C, and transmits and receives information between the multiple cameras 11A to 11C. The image processing device 10 is a computer device such as a PC (personal computer), a smartphone, or a tablet terminal device.

[0011] First, the image processing device 10 will be described. As shown in Fig. 1, the image processing device 10 is configured to include a CPU 101, a RAM 102, a ROM 103, a large-capacity storage device 104, an operation unit 105, and a display unit 106. The CPU 101, the RAM 102, the ROM 103, the large-capacity storage device 104, the operation unit 105, and the display unit 106 are connected to each other via a bus 107.

[0012] The CPU 101 executes various processes using computer programs and data stored in the ROM 103 and the large-capacity storage device 104. In this way, the CPU 101 controls the entire image processing device 10.

[0013] The RAM 102 has an area for storing computer programs and data loaded from the ROM 103 or the mass storage device 104, and an area for storing images received from the cameras 11A to 11C. The RAM 102 also has a work area used when the CPU 101 executes various processes. In this way, the RAM 102 can provide various areas as needed.

[0014] The ROM 103 stores setting data for the image processing device 10, computer programs and data relating to startup, computer programs and data relating to basic operations, and the like. The mass storage device 104 is a hard disk drive device or the like. The mass storage device 104 stores an OS (operating system), as well as computer programs and data for causing the CPU 101 to execute or control various processes of the image processing device 10. The data stored in the mass storage device 104 also includes data related to a trained model (such as model parameters). The computer programs and data stored in the mass storage device 104 are loaded into the RAM 102 as appropriate under the control of the CPU 101, and become targets for processing by the CPU 101.

[0015] The operation unit 105 is an input device such as a keyboard, a mouse, a touch panel, etc. The CPU 101 accepts various instructions input by a user operating the operation unit 105. The display unit 106 is a display device such as a liquid crystal panel or a touch panel display. The CPU 101 displays images, characters, and the like representing the processing results on the display unit 106. Note that the display unit 106 may be a projection device such as a projector that projects images, characters, and the like.

[0016] The image processing device 10 receives images captured from each of the cameras 11A to 11C. The images captured by each camera are given the time of capture, and the image processing device 10 can recognize the chronological order of each image after receiving the images by referring to the time.

[0017] 2 is a block diagram showing the functional configuration of the image processing device 10 according to this embodiment. The CPU 101 executes a computer program stored in the ROM 103 or the mass storage device 104 to realize the functions of an image acquisition unit 201, an extraction unit 202, a form estimation unit 203, a state estimation unit 204, an image generation unit 205, and an identification unit 206.

[0018] Image acquisition unit 201 sequentially acquires images from each of cameras 11A to 11C at a predetermined time interval, and provides the images to extraction unit 202. Note that image acquisition unit 201 is not limited to receiving captured images from cameras 11A to 11C. For example, images may be input by streaming input via a network and by reading video data (recorded video) from mass storage device 104.

[0019] The extraction unit 202 extracts moving objects in images acquired by the image acquisition unit 201, and generates a moving object extraction result list for each camera acquired from the image acquisition unit 201. The moving object extraction result list is time-series data arranged by the time at which the moving object extraction results are added to the images. The moving object extraction result includes visible detection frame information (center position, size) indicating a visible detection frame 224 (described later in FIG. 4B) that includes the visible area of ​​the moving object, and a moving object image 222 (described later in FIG. 5) that indicates an image of the visible area of ​​the moving object. The visible detection frame information of the moving object indicates the position of the moving object in the image.

[0020] The shape estimation unit 203 estimates the three-dimensional shape of the moving body based on the moving body extraction result acquired by the extraction unit 202. The three-dimensional shape represents a shape in a three-dimensional Euclidean space having height, width, and depth. The viewpoint in the three-dimensional Euclidean space for the three-dimensional shape of the moving body estimated by the shape estimation unit 203 can be freely changed. In addition, the estimation result of the three-dimensional shape of the moving body includes a color distribution based on the color information of the moving body image included in the moving body extraction result.

[0021] The state estimation unit 204 estimates the moving body state of the moving body based on the moving body extraction result acquired by the extraction unit 202. The estimation result of the moving body state includes posture information of the moving body and image quality parameters. The posture information of the moving body indicates the tilt angle of the moving body around three axes in three-dimensional space. The image quality parameters indicate parameters related to the image quality and texture of the moving body image, as well as hue, saturation, and brightness.

[0022] The image generating unit 205 generates a predicted image of the moving object based on the moving object extraction result, the estimation result of the three-dimensional shape of the moving object, and the estimation result of the moving object state of the moving object. The image generating unit 205 first determines a background image of the moving object from the visible detection frame information of the moving object included in the moving object extraction result. Next, the image generating unit 205 changes the viewpoint in three-dimensional space with respect to the three-dimensional shape of the moving object based on the posture information of the moving object included in the estimation result of the moving object state of the moving object, and generates a predicted image of the moving object by image conversion processing. Furthermore, the image generating unit 205 converts the image quality parameters of the generated predicted image of the moving object based on the image quality parameters included in the estimation result of the moving object state of the moving object.

[0023] The identification unit 206 identifies the identity of the moving object in the two images by comparing the moving object image included in the moving object extraction result with the predicted image of the moving object generated by the image generation unit 205. Specifically, the identification unit 206 performs a process of outputting an identity score using the feature vectors extracted from the moving object image and the predicted image as input, and determines whether the moving objects in the two images are the same based on the identity score.

[0024] According to the functional configuration of the image processing device 10 shown in the present embodiment, for any moving object, it is possible to generate an image of a new moving object having a posture and image quality similar to that of the extracted moving object at a position in an image where the moving object is extracted. Note that any moving object may be a pair of moving objects, that is, a first moving object and a second moving object.

[0025] FIG. 3 is a flowchart showing the processing of the image processing device 10 according to this embodiment. In the following description of the flowchart, each process (step) is represented by adding S to the beginning, and the description of the process (step) is omitted. The processing of this flowchart is executed by the image acquisition unit 201 sequentially acquiring images from each of the cameras 11A to 11C at a predetermined time interval and providing them to the extraction unit 202. In the following, for simplicity, the case of two cameras is taken as an example, and the image acquisition unit 201 acquires images from each of the cameras 11A and 11B. Note that the camera 11A may be referred to as the first imaging unit, and the camera 11B may be referred to as the second imaging unit. Also, the moving object imaged by the first imaging unit may be referred to as the first moving object, and the moving object imaged by the second imaging unit may be referred to as the second moving object.

[0026] First, in S901, the image acquisition unit 201 acquires an image captured by the first imaging unit. Then, the extraction unit 202 extracts a first moving object from the acquired image and acquires a first moving object extraction result that is the extraction result. The first moving object extraction result includes visible detection frame information of the first moving object and a moving object image of the first moving object (hereinafter referred to as the first moving object image). In S902, the shape estimation unit 203 estimates the three-dimensional shape of the first moving object based on the first moving object extraction result in S901.

[0027] After an image is acquired in S901, in S903, the image acquisition unit 201 acquires an image captured by the second imaging unit. Then, the extraction unit 202 extracts a second moving object from the acquired image and acquires a second moving object extraction result that is the extraction result. The second moving object extraction result includes visible detection frame information of the second moving object and a moving object image of the second moving object (hereinafter referred to as the second moving object image). In S904, the state estimation unit 204 estimates a moving body state of the second moving body (hereinafter referred to as a second moving body state) based on the second moving body extraction result in S903. The estimation result of the second moving body state includes posture information and image quality parameters of the second moving body.

[0028] In S905, the image generating unit 205 generates a predicted image of the first moving object in the second imaging unit based on the estimation result of the three-dimensional shape of the first moving object in S902, the second moving object extraction result in S903, and the estimation result of the second moving object state in S904. This predicted image is generated from a background image obtained from visible detection frame information of the second moving object included in the second moving object extraction result, and a foreground image obtained by converting the three-dimensional shape of the first moving object into a state close to the posture information and image quality parameters of the second moving object.

[0029] In S906, the identification unit 206 compares the second moving object image included in the second moving object extraction result in S903 with the predicted image generated in S905 to determine whether the moving objects in the two images are the same. If the identification unit 206 determines that the first moving object and the second moving object are the same, it adds information to the second moving object extraction result indicating that it is the same moving object as the first moving object extraction result. Then, the process of this flowchart ends.

[0030] According to the process of this flowchart, a predicted image is generated by converting a moving object captured first by one camera according to the orientation, image quality, background, etc. of the moving object captured later by another camera, and the predicted image can be used for matching with the moving object image. This makes it possible to reduce the possibility of losing sight of the moving object due to changes in the way the moving object is captured when tracking a moving object that moves from one camera to another.

[0031] Next, an example of a captured image to be the subject of this embodiment will be described with reference to FIG. 4. FIG. 4(a) shows an example in which two cameras 11A and 11B are arranged in a predetermined space where the traffic of moving objects is predicted. In the example of FIG. 4(a), the orientations of the cameras 11A and 11B are significantly different, and there is no common range in the imaging ranges R1 and R2 of each other. Such an arrangement of the imaging units generally makes it difficult to identify a moving object. This is because the imaging units are spatially discontinuous, making it impossible to apply a method that focuses on spatial continuity, and because the orientation of the moving object relative to each imaging unit is different, causing the visible information of the moving object to vary among the imaging units.

[0032] FIG. 4(b) shows an example of an image captured by the first imaging unit and the second imaging unit in this embodiment. Here, the first imaging unit is assumed to be the camera 11A, and the second imaging unit is assumed to be the camera 11B. Although the moving body is assumed to be a human body, the moving body is not limited to a human body, and may be an animal such as a dog or an object such as a car. The left diagram of FIG. 4(b) shows an image 211A captured by the camera 11A at time T1. The right diagram of FIG. 4(b) shows an image 211B captured by the camera 11B at time T2 after time T1. A first moving body 301A is extracted from the captured image 211A. A visible detection frame 224A including a visible area of ​​the first moving body 301A is displayed in the captured image 211A. A second moving body 301B is extracted from the captured image 211B. Also, the captured image 211B displays a visible detection frame 224B including the visible area of ​​the second moving body 301B. The coordinate system in the captured image 211A is expressed by X1 extending rightward and Y1 extending downward, with the upper left corner of the image as the origin (0,0). Similarly, the coordinate system in the captured image 211B is expressed by X2, Y2, with the upper left corner of the image as the origin. Note that the captured image 211B displays a center O of three-dimensional coordinate axes related to the attitude of the second moving body 301B. These three-dimensional coordinate axes are used when the state estimation unit 204 represents the attitude of the second moving body 301B.

[0033] FIG. 4(c) is a diagram for explaining three-dimensional coordinate axes related to the attitude of the moving body. The center O of the three-dimensional coordinate axes is located at the center of the part of the moving body that is in contact with the ground. There are also three coordinate axes, x, y, and z, that extend in three directions from the coordinate axis center O. The y axis is the axis that defines the front of the moving body. The x axis is another axis that is perpendicular to the y axis and defines the ground. The z axis is the axis that is perpendicular to the ground. The rotation angles around the x, y, and z axes are represented by α, β, and γ, respectively.

[0034] FIG. 5 is a data flow diagram of the image processing device 10 according to this embodiment. First, data acquired from the captured image 211A by the first imaging unit will be described. The first moving object extraction result 221A extracted from the captured image 211A includes a first moving object image 222A and visible detection frame information 223A of the first moving object. The cropping range of the first moving object image may be an arbitrary magnification with respect to the size of the moving object. In addition, whether the outside of the captured image 211A is included or the properties when including it may also be arbitrary. Hereinafter, the visible detection frame information 223A of the first moving object is referred to as first position information 223A.

[0035] The shape estimation unit 203 estimates the three-dimensional shape 231A of the first moving body based on the first moving body image 222A. When estimating the three-dimensional shape of the first moving body, color distribution and position information are generated for the visible surface of the first moving body image 222A based on its color information and geometric information, and processing of the color information of the invisible surface may be arbitrary. The three-dimensional shape 231A of the first moving body is data that expresses the surface of the first moving body as a position in a three-dimensional Euclidean space, and the data type may be point cloud data or mesh data. The reference point of the three-dimensional shape 231A of the first moving body is preferably the center of the part of the moving body that contacts the ground, but may be another position. The three-dimensional shape 231A of the first moving body may include information of an object (such as a bag or an umbrella) associated with the first moving body reflected in the first moving body image 222A.

[0036] Next, data acquired from the captured image 211B by the second imaging unit will be described. The second moving object extraction result 221B extracted from the captured image 211B includes a second moving object image 222B and visible detection frame information 223B of the second moving object. The cropping range of the second moving object image may be an arbitrary magnification with respect to the size of the moving object. In addition, whether the outside of the captured image 211B is included or the properties when including it may also be arbitrary. Hereinafter, the visible detection frame information 223B of the second moving object is referred to as second position information 223B.

[0037] The state estimation unit 204 estimates the posture information 241B and the image quality parameter 242B of the second moving object based on the second moving object image 222B. When obtaining the image quality parameter, a process such as changing the cutout range of the cut out second moving object image 222B may be included. The image quality parameter 242B is an arbitrary parameter related to hue, saturation, brightness, image quality, texture, etc. The image quality parameter 242B is an example of an image quality parameter. The process of estimating the image quality parameter may be a process of dividing the area of ​​the second moving object image 222B to obtain a set of image quality parameters, and is not particularly limited as long as it is a process of inputting an image and outputting image quality parameters. The state estimation unit 204 may estimate the posture information 241B and the image quality parameter 242B of the second moving object using information such as the camera setting, installation position, and lighting conditions of the second imaging unit.

[0038] The image generating unit 205 generates a predicted image 251 of the first moving object in the second imaging unit from the three-dimensional shape 231A of the first moving object, the second position information 223B, the posture information 241B of the second moving object, and the image quality parameter 242B. As a method for the image generating unit 205 to acquire a background image in the second imaging unit using the second position information 223B, for example, a method using a background image of the second imaging unit stored in advance in the ROM 103 or the like is available, but other methods may also be used. In addition, when the image generating unit 205 determines a viewpoint for image conversion of the three-dimensional shape 231A of the first moving object based on the posture information 241B of the second moving object, the scale of the three-dimensional shape 231A of the first moving object is determined. As a method for this, for example, a method of referring to the area size of the second position information 223B is available, but other methods may also be used. Further, the method by which the image generating unit 205 converts the three-dimensional form 231A of the first moving object into an image is ray tracing, but other methods may also be used.

[0039] 5, a predicted image 251 reflecting the position, background, moving body posture, and image quality parameters of the second moving body image 222B is generated from the first moving body image 222A. Therefore, if the first moving body and the second moving body are the same moving body, the second moving body image 222B and the predicted image 251 will be substantially the same image. The identification unit 206 executes a process of outputting a similarity score 261 between the second moving object image 222B and the predicted image 251. The identification unit 206 determines whether the first moving object and the second moving object are the same moving object based on the similarity score 261.

[0040] FIG. 6 is a diagram for explaining the processing by the identification unit 206 in this embodiment. FIG. 6(a) shows an example of processing for outputting the identity score 261 between the second moving object image 222B and the predicted image 251. In the processing shown in FIG. 6(a), the identification unit 206 first inputs each image to a CNN (Convolutional Neural Network). This CNN is an example of a network model for extracting a feature vector of an image. The CNN 410 to which the second moving object image 222B is input and the CNN 410 to which the predicted image 251 is input are configured with the same network model. The input image is converted into a feature vector by the CNN 410. Next, the vector distance calculation unit 411 calculates the distance between the feature vectors based on each image, converts the calculated vector distance into the identity score 261, and outputs it.

[0041] FIG. 6(b) shows another example of a process for outputting an identity score 261 between the second moving object image 222B and the predicted image 251. In the process shown in FIG. 6(b), the identification unit 206 first subtracts the predicted image 251 from the second moving object image 222B. The subtraction process between images is performed based on luminance information or RGB information. For example, since the background images of the predicted image 251 and the second moving object image 222B are substantially identical, the RGB information of the background portion is approximated to 0 by the subtraction process, and only the result of the subtraction process of the foreground portion (moving object portion) remains. If the predicted image 251 is approximated to be substantially identical to the second moving object image 222B, the RGB information of the subtraction result of the moving object portion can be approximated to 0. The identification unit 206 inputs the image after subtraction to the CNN 410, converts the output result into an identity score 261, and outputs it. In this embodiment, CNN is used for the process of vectorizing images and the process of calculating a similarity score between images. However, similarity between images may be evaluated without using CNN.

[0042] According to the present embodiment as described above, when identifying the same moving object captured by different cameras, it is possible to improve robustness against variations in how the moving object is captured between the cameras.

[0043] <Embodiment 2> In this embodiment, a method of estimating the joint positions of a moving object and generating a predicted image of the moving object by further using the estimation result of the joint positions of the moving object will be described. In the following, the description of the same parts as in the first embodiment will be omitted, and the description will focus on the differences.

[0044] 7 is a block diagram showing the functional configuration of an image processing device 10 according to this embodiment. The image processing device 10 differs from the image processing device 10 in FIG. The joint position estimation unit 207 estimates joint position information of the moving object based on the moving object extraction result acquired by the extraction unit 202 . The image generating unit 205 converts the joint positions of the three-dimensional form of the moving object based on the joint position information of the moving object estimated by the joint position estimating unit 207, and generates a predicted image of the moving object by image conversion processing.

[0045] Next, an example of a captured image targeted in this embodiment will be described with reference to FIG. 8. The spatial arrangement of each imaging unit is the same as that in FIG. 4(a). The left diagram in FIG. 8 shows a captured image 211A captured by camera 11A at time T1. The right diagram in FIG. 8 shows a captured image 211B captured by camera 11B at time T2 after time T1. FIG. 8 shows a case in which the joint positions of a first moving object 301A extracted from captured image 211A are different from the joint positions of a second moving object 301B extracted from captured image 211B.

[0046] 8, when a change occurs in the joint position of a moving object, the accuracy of identifying the moving object may decrease due to a difference in color distribution of the edge of the moving object. Therefore, in this embodiment, a predicted image that reflects the change in the joint position is generated.

[0047] 9 is a data flow diagram of the image processing device 10 according to this embodiment. Here, the differences from FIG. 5 will be mainly explained. The joint position estimation unit 207 estimates the joint position information 271B of the second moving object based on the second moving object image 222B. The joint position information includes, for example, the positions in a three-dimensional space of 17 joints of the human body and their estimated likelihoods. However, the definition of the joint points may be anywhere based on the parts of the moving object, and does not have to be 17 points, and is not limited to this.

[0048] The three-dimensional form 231A of the first moving body is converted into a three-dimensional form 232A of the first moving body in which the joint positions have been changed by a process of converting the positions of each joint part in three-dimensional space based on the joint position information 271B of the second moving body.

[0049] The image generation unit 205 generates a predicted image 251 of the first moving body in the second imaging unit from the three-dimensional shape 232A of the first moving body whose joint position has been changed, the second position information 223B, the posture information 241B of the second moving body, and the image quality parameters 242B.

[0050] 9, a predicted image reflecting the same position, background, joint positions and posture of the moving body, and image quality parameters as those of the second moving body image 222B is generated from the first moving body image 222A. Therefore, if the first moving body and the second moving body are the same moving body, the second moving body image 222B and the predicted image 251 will be substantially the same image.

[0051] According to the present embodiment as described above, when identifying the same moving object captured by different cameras, it is possible to improve robustness against fluctuations in the joint positions of the moving object that may occur between the cameras.

[0052] <Embodiment 3> In this embodiment, a method for detecting an occluded area of ​​a moving object and generating a predicted image of the moving object by using the detection result of the occluded area will be described. In the following, the description of the same parts as in the second embodiment will be omitted, and the description will focus on the differences.

[0053] 10 is a block diagram showing the functional configuration of an image processing device 10 according to this embodiment. The image processing device 10 differs from the image processing device 10 in FIG. The occlusion detection unit 208 detects an occluded area in the moving object image based on the joint position information of the moving object estimated by the joint position estimation unit 207 . The image generation unit 205 converts the joint positions of the three-dimensional shape of the moving body based on the joint position information of the moving body estimated by the joint position estimation unit 207, and generates a predicted image of the moving body by image conversion processing. In this embodiment, the image generation unit 205 performs processing to exclude a part of the predicted image from the target of matching by the identification unit 206 based on the detection result of the occluded area.

[0054] Next, an example of a captured image targeted in this embodiment will be described with reference to FIG. 11. The spatial arrangement of each imaging unit is the same as that in FIG. 4(a). The left diagram of FIG. 11 shows a captured image 211A captured by the camera 11A at time T1. The right diagram of FIG. 11 shows a captured image 211B captured by the camera 11B at time T2 after time T1. The first moving object 301A extracted from the captured image 211A is not blocked, but the lower half of the second moving object 301B extracted from the captured image 211B is blocked by a blocking object S. FIG. 11 shows a case in which the blocking state of the first moving object image in the captured image 211A is different from the blocking state of the second moving object image in the captured image 211B.

[0055] 11, in the case where the occlusion state of the moving object has changed, the accuracy of identifying the moving object may decrease due to differences in the geometric features of the moving object in the visible detection frames 224A and 224B of the moving object. Therefore, in this embodiment, a predicted image that reflects the occlusion state is generated.

[0056] FIG. 12 is a data flow diagram of the image processing device 10 according to this embodiment. The occlusion detection unit 208 detects an occluded area of ​​the second moving object based on the joint position information 271B of the second moving object. In this embodiment, an area other than the occluded area indicated by the occlusion detection result 281B is used for matching with the second moving object image. A specific example of the processing of the occlusion detection unit 208 is a rule-based processing in which, when the estimated likelihood of a joint belonging to the lower half of the joint position information 271B of the second moving object is low, the lower half is determined to be occluded. Another specific example is a rule-based processing in which, when the position of a joint belonging to the lower half of the joint position information 271B of the second moving object is approximately the same as or exists below the lower side of the detection frame in the second position information 223B, the lower half is determined to be occluded. The processing of the occlusion detection unit 208 is not limited to the above specific example, and may be any processing that detects an occluded area based on the joint position information of the moving object and the moving object extraction result.

[0057] The image generating unit 205 generates a predicted image 251 of the first moving object in the second imaging unit from the three-dimensional shape 232A of the first moving object, the second position information 223B, the posture information 241B of the second moving object, and the image quality parameters 242B shown in FIG. 9. In this embodiment, the image generating unit 205 converts the generated predicted image 251 into a predicted image 252 by excluding the occlusion area indicated by the occlusion detection result 281B of the second moving object. The process of converting the predicted image by the image generating unit 205 based on the occlusion detection result 281B is not particularly limited as long as it is a process of enlarging a part of the predicted image so as not to include the occluding object S, or moving the predicted image area.

[0058] 12, a predicted image reflecting the same position, background, joint positions and posture of the moving body, image quality parameters, and occlusion area as the second moving body image 222B is generated from the first moving body image 222A. Therefore, if the first moving body and the second moving body are the same moving body, the second moving body image 222B and the predicted image 252 will be substantially the same image.

[0059] According to the present embodiment as described above, when identifying the same moving object captured by different cameras, it is possible to improve robustness against variations in occluded areas that may occur between the cameras.

[0060] <Embodiment 4> In this embodiment, a method is described in which one or more predicted images are generated assuming that a first moving object is imaged by a first imaging unit and then the first moving object is imaged by a second imaging unit, and the generated predicted images are compared with an image of a second moving object imaged by the second imaging unit. In the following, the same parts as in the first embodiment are not described, and differences are mainly described.

[0061] First, an example of a captured image that is a target of this embodiment will be described with reference to FIG. 13. The spatial arrangement of each imaging unit is the same as that in FIG. 4(a). The left diagram of FIG. 13 shows a captured image 211A captured by camera 11A at time T1. A first moving object 301A is extracted from the captured image 211A. Moreover, the right diagram of FIG. 13 shows predicted images 3011B, 3012B, and 3013B of the first moving object by dashed lines in a captured image 2111B captured by camera 11B at time T1+t after time T1.

[0062] Unlike the first to third embodiments, the image generating unit 205 according to this embodiment generates a predicted image assuming that the first moving object will be captured by the second imaging unit as soon as the first moving object is extracted from the captured image by the first imaging unit, without waiting for the second moving object to be captured by the second imaging unit. In this embodiment, the process of generating a predicted image is executed before the second moving object is extracted. Therefore, the state estimating unit 204 according to this embodiment estimates the state of the first moving object in the second imaging unit without using the second moving object extraction result.

[0063] FIG. 14A is a flowchart showing the process of the image processing device shown in this embodiment. First, in S911, the state estimation unit 204 uses learning data stored in a storage unit such as the mass storage device 104 to learn a model for estimating the state of a moving object. The learning data here is, for example, data that is a set of the state of the moving object in the first imaging unit and the state of the moving object in the second imaging unit when the moving object captured by the first imaging unit is captured by the second imaging unit. Note that a DNN (Deep Neural Network) model or the like is used. The state estimation unit 204 saves the learned model learned from the learning data in the storage unit. Note that the processing of this step may be executed in advance, independently of this flowchart.

[0064] In S912, the image acquisition unit 201 acquires an image captured by the first imaging unit. Then, the extraction unit 202 extracts a first moving object from the acquired image and acquires a first moving object extraction result that is the extraction result. In S913, the form estimation unit 203 estimates the three-dimensional form of the first moving object based on the first moving object extraction result in S912. In S914, the state estimation unit 204 estimates the state of the first moving object in the second imaging unit, such as the position and the attitude, based on the first moving object extraction result in S912. Details of the estimation process executed in this step will be described later with reference to the flowchart in FIG.

[0065] In S915, the image generating unit 205 generates a predicted image of the first moving object in the second imaging unit based on the estimation result of the three-dimensional shape of the first moving object in S913 and the estimation result of the state in S914. In this embodiment, the estimation result of the state in S914 includes a plurality of candidates. Therefore, a plurality of predicted images of the first moving object in the second imaging unit are generated according to the number of candidates (e.g., n). After the image is acquired in S912, in S916, the image acquisition unit 201 acquires the image captured by the second imaging unit. Then, the extraction unit 202 extracts the second moving object from the acquired image and acquires the second moving object extraction result, which is the extraction result. In S917, the identification unit 206 compares the multiple predicted images generated in S915 with the second moving object image included in the second moving object extraction result in S916, and determines whether or not there is a predicted image including the same moving object as the second moving object image. Then, the process of this flowchart ends.

[0066] According to the process of this flowchart, the state of a moving object that is imaged later by a different camera, such as its position and attitude, can be estimated from the extraction result of the moving object that was imaged earlier by a different camera, and a predicted image that reflects the estimation result can be generated and used for matching with the moving object image. This makes it possible to reduce the possibility of losing sight of the moving object due to changes in how the moving object is captured when tracking a moving object that moves from one camera to another.

[0067] FIG. 14B is a flowchart showing details of the estimation process executed in S914 of FIG. 14A. First, in S921, the state estimation unit 204 acquires the first moving object extraction result in S912. In S922, the state estimation unit 204 acquires the first position information included in the first moving object extraction result acquired in S921. In S923, the state estimation unit 204 acquires a first moving object image included in the first moving object extraction result acquired in S921.

[0068] In S924, the state estimation unit 204 acquires posture information and image quality parameters of the first moving object from the first moving object image acquired in S923. In S925, the state estimation unit 204 sets the number of state candidates to be generated for the first moving object in the second imaging unit to n. In S926, the state estimation unit 204 generates n sets of position information, attitude information, and image quality parameters as state candidates of the first moving object in the second imaging unit. In this embodiment, the state estimation unit 204 estimates the state of the first moving object in the second imaging unit by inputting the information representing the state of the first moving object acquired in S922 and S924 into the model trained in S911. Then, the process proceeds to S915.

[0069] Here, the number n of state candidates generated for the first moving object in the second imaging unit may be 1 or a positive integer greater than 1. In this embodiment, a predicted image of the first moving object captured by the second imaging unit is generated before the second moving object is captured by the second imaging unit, and an identity score between the second moving object image captured by the second imaging unit and the generated predicted image is evaluated. Therefore, it is necessary to allow a certain degree of uncertainty. Therefore, the image generating unit 205 may generate multiple predicted images of the first moving object in the second imaging unit, taking into account robustness against uncertainty. In other words, the number n of state candidates generated for the first moving object in the second imaging unit is set to an integer greater than 1.

[0070] 15 is a data flow diagram of the image processing device 10 according to this embodiment. Here, the differences from FIG. 5 will be mainly explained. The upper state estimation section 204 outputs state information 240A (position information 243A, attitude information 241A, image quality parameters 242A) of the first moving object in the second imaging section from the first moving object extraction result 221A.

[0071] The image generating section 205 generates a predicted image 251 of the first moving object in the second imaging section from the three-dimensional form 231A of the first moving object and the state information 240A of the first moving object. The identification unit 206 executes a process of outputting an identity score 261 between the second moving object image 222B and the predicted image 251. The first moving object extraction result 221A and the second moving object extraction result 221B are labeled with an identity label 229 indicating whether they are caused by the same moving object or not.

[0072] Similarly to the state estimation unit 204 in the upper row, the state estimation unit 204 in the lower row executes S921 to S924 and outputs state information 240B (position information 243B, posture information 241B, image quality parameters 242B) of the second moving object from the second moving object extraction result 221B. Note that since the second moving object image 222B and the second position information 223B are used as the correct answer data, the position information 243B inherits the second position information 223B. The loss 263 is calculated by inputting the difference between the identity label 229 and the identity score 261 and the difference between the state information 240A of the first moving object and the state information 240B of the second moving object to a loss function. The state estimation unit 204 can improve the estimation accuracy by proceeding with the learning of a model that estimates the state of the moving object using the loss 263.

[0073] According to the present embodiment as described above, a predicted image of a moving object that will be captured later by another camera is generated from an image of the moving object captured by a previous camera, and the predicted image is compared with an image of the moving object that was actually captured, thereby making it possible to identify the same moving object.

[0074] <Embodiment 5> In this embodiment, it is assumed that the first moving object is imaged by the first imaging unit, and then the first moving object is imaged by the second imaging unit, and a method for generating one or more predicted images of the moving object that take into account changes in the joint positions of the moving object will be described. Descriptions of the same parts as in the second and fourth embodiments will be omitted, and differences will be mainly described.

[0075] 16 is a block diagram showing the functional configuration of an image processing device 10 according to this embodiment. The image processing device 10 differs from the image processing device 10 in FIG. The data storage unit 400 stores data such as the moving object extraction result of the extraction unit 202, the moving object state estimated by the state estimation unit 204, the joint position information estimated by the joint position estimation unit 207, and the identity score output by the identification unit 206, in association with the moving object extraction result. The data storage unit 400 uses the large-capacity storage device 104 or the like. The joint position generating unit 209 generates predicted joint position information based on the learning result obtained by using the joint position information stored in the data storage unit 400. The joint position generating unit 209 is an example of a joint position prediction means. The data used for learning is various data based on the moving object extraction result in a predetermined imaging unit. The joint position generating unit 209 performs extensive learning on the joint position information that can be taken by a moving object captured by a predetermined imaging unit.

[0076] Next, an example of a captured image to be captured in this embodiment will be described with reference to Fig. 17. The spatial arrangement of each imaging unit is the same as that in Fig. 4(a). 17(a) shows a captured image 2112B captured by the second imaging unit. A plurality of different moving objects are shown in the captured image 2112B. The moving objects in the captured image 2112B each have different joint position information. The data storage unit 400 acquires and stores all of the joint position information of these moving objects. This makes it possible to generate joint position information with a wide variety.

[0077] The left diagram of Fig. 17(b) shows an image 211A captured by camera 11A at time T1. A first moving object 301A is extracted from the captured image 211A. The right diagram of Fig. 17(b) shows predicted images 3011B, 3012B, and 3013B of the first moving object by dashed lines in an image 2111B captured by camera 11B at time T1+t after time T1.

[0078] The joint position generating unit 209 generates one or more candidates of predicted joint position information using the joint position information stored in the data storage unit 400. The joint position generating unit 209 changes the position in the three-dimensional space for each joint part of the three-dimensional form of the first moving body based on the generated predicted joint position information, and generates a three-dimensional form of the first moving body in which each joint position has been changed. The image generating unit 205 generates predicted images 3011B, 3012B, 3013B of the first moving body in one or more second imaging units based on the one or more three-dimensional forms of the first moving body generated in this way and one or more state candidates of the first moving body.

[0079] According to the present embodiment as described above, when identifying the same moving object captured by different cameras, it is possible to improve robustness against fluctuations in the joint positions of the moving object that may occur between the cameras.

[0080] <Embodiment 6> In this embodiment, a method is described in which a likelihood distribution is also estimated when estimating the three-dimensional shape of a moving object from a moving object image, and the likelihood distribution is used when determining the identity of the moving object. In the following, the same parts as in the first embodiment are omitted, and the differences are mainly described.

[0081] FIG. 18 shows an example of a captured image according to this embodiment. FIG. 18(a) shows an image 1801 of a moving body when viewed from approximately the right rear. It can be seen that the moving body has a large baggage attached to its back, and the sleeves of the moving body's jacket are relatively bright. FIG. 18(b) shows a predicted image 1802 of a moving body when the three-dimensional shape is estimated based on FIG. 18(a) and the three-dimensional shape is viewed from the front. If it has been learned that the front side of the jacket is the same color based on the sleeve color of the moving body's jacket in the image 1801 of the moving body, an image like the predicted image 1802 can be acquired. FIG. 18(c) shows a different viewpoint image 1803 of the moving body. As shown in FIG. 18(c), the bottom of the moving body's jacket is the bottom of a front-opening object such as a cardigan, and in the upper half of the moving body when viewed from the front, the color information of an object worn inside the front-opening object may be predominant.

[0082] In the above case, a discrepancy may occur in the image information between a predicted image 1802 obtained from a three-dimensional shape estimated based on an image 1801 of a moving body and another viewpoint image 1803 of the moving body, which may result in a decrease in the accuracy of identification between the predicted image and the moving body image. Therefore, in this embodiment, a predicted image is generated by excluding invisible parts (for example, the front and left sides of the moving body) in the image 1801 of the moving body. This is expected to improve the accuracy of identification between the predicted image and the moving body image.

[0083] 19 is a data flow diagram of the image processing device 10 according to this embodiment. Here, the differences from FIG. 5 will be mainly explained. When estimating a three-dimensional shape 231A of a moving object based on a first moving object extraction result 221A extracted by the extraction unit 202, the shape estimation unit 203 also estimates a three-dimensional likelihood distribution 233A at the same time. In the three-dimensional likelihood distribution 233A in Fig. 19, surfaces with high likelihood are shown bright and surfaces with low likelihood are shown dark.

[0084] When generating a predicted image 251 of the first moving body based on the three-dimensional shape 231A of the first moving body and the moving body state of the second moving body, the image generating unit 205 simultaneously generates a moving body likelihood distribution 253. Specifically, when performing image conversion by changing the viewpoint with respect to the three-dimensional shape 231A of the moving body based on the posture information 241B of the second moving body, the viewpoint with respect to the three-dimensional likelihood distribution 233A is similarly changed and projected as an in-plane likelihood distribution.

[0085] The identification unit 206 executes a process of outputting an identity score 261 between the second moving object image 222B and the predicted image 251, using the moving object likelihood distribution 253. The identification unit 206 may use the moving object likelihood distribution 253 for pre-processing of the predicted image 251, or may combine it with the predicted image 251 to generate a vector, and the method of using the moving object likelihood distribution is arbitrary and is not limited to this. Furthermore, the identification unit 206 may further execute a process of adding a dummy likelihood distribution to the second moving object image 222B, etc.

[0086] According to the present embodiment described above, by applying the estimated likelihood of a three-dimensional shape to a predicted image to deal with the uncertainty that arises due to differences in the orientation of cameras relative to a moving object, it is possible to ensure robustness by adding information on the image's focus point at a particular time.

[0087] <Embodiment 7> In this embodiment, a method for generating a predicted image based on a first moving body image and a second moving body state without estimating the three-dimensional shape of the moving body will be described. In the following, the same parts as in the first embodiment will be omitted, and the differences will be mainly described.

[0088] 20 is a block diagram showing the functional configuration of an image processing device 10 according to this embodiment. This differs from the image processing device 10 in FIG.

[0089] 21 is a data flow diagram of the image processing device 10 according to this embodiment. Here, the differences from FIG. 5 will be mainly explained. The image generating unit 205 generates a predicted image 251 from the first moving body image 222A, the posture information 241B of the second moving body, and the image quality parameters 242B, as if the first moving body were viewed under the conditions of the posture and image quality parameters of the second moving body. In order to generate a predicted image having a posture different from the first moving body image 222A, the image generating unit 205 uses a trained network model having a structure of, for example, GAN (Generative Adversarial Networks). By using GAN, it is possible to generate a predicted image in which only the posture and image quality parameters are changed while preserving the appearance features of the first moving body image 222A. In addition, various variations are possible for the combination of tilt angles around three axes in a three-dimensional space, which is the posture of the second moving body, and it is difficult to prepare learning data that covers all of these variations in the real space. For this reason, learning of GAN using CG data is considered. It is possible to prepare a CG moving body model that is the same as or different from the first moving body image and the second moving body image, and generate a learning dataset of a moving body having a different moving body posture (tilt angle around three axes in CG space) by changing the camera parameters in the CG space.

[0090] According to this embodiment, similarly to the first embodiment, when identifying the same moving object captured by different cameras, it is possible to improve robustness against variations in how the moving object is captured between the cameras.

[0091] <Other embodiments> Although the embodiment has been described in detail above, the present invention can be embodied as, for example, a system, an apparatus, a method, a program, or a recording medium (storage medium), etc. Specifically, the present invention may be applied to a system composed of multiple devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or may be applied to an apparatus composed of a single device.

[0092] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0093] The disclosure of each of the above-described embodiments includes the following configurations, methods, and programs. (Configuration 1) a first extraction means for extracting a first moving object from a first image captured by the first imaging unit; a second extraction means for extracting a second moving object from a second image captured by the second imaging unit; a first estimation means for estimating a three-dimensional shape of the first moving object based on a first extraction result which is an extraction result by the first extraction means; a second estimation means for estimating a state of the second moving object based on a second extraction result which is an extraction result by the second extraction means; a generating means for generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on the estimation result of the three-dimensional shape of the first moving object and the estimation result of the state of the second moving object; an identification means for identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted by the second extraction means; 13. An image processing device comprising: (Configuration 2) 2. The image processing device according to configuration 1, wherein the estimation result of the state of the second moving object includes information on the posture of the second moving object and image quality parameters of the second moving object. (Configuration 3) the estimation result of the state of the second moving body includes information on the attitude of the second moving body; The image processing device described in configuration 1 or 2, characterized in that the generation means generates a background image of the predicted image based on the position of the second moving body in the second image, and generates a foreground image of the predicted image by converting a viewpoint with respect to a three-dimensional shape of the first moving body based on the attitude of the second moving body. (Configuration 4) the estimation result of the state of the second moving object includes information of an image quality parameter of the second moving object; 4. The image processing apparatus according to any one of configurations 1 to 3, wherein the generating means converts the predicted image based on an image quality parameter of the second moving object. (Configuration 5) 5. The image processing device according to any one of configurations 1 to 4, wherein the estimation result of the three-dimensional shape of the first moving object includes information on a position and color of the first moving object in a three-dimensional space. (Configuration 6) The image processing device described in any one of configurations 1 to 5, characterized in that the identification means determines the identity of the first moving body and the second moving body based on the distance between feature vectors obtained by inputting the predicted image and the image of the second moving body into a network model. (Configuration 7) The image processing device described in any one of configurations 1 to 6, characterized in that the identification means determines the identity of the first moving body and the second moving body based on the output result obtained by inputting the result obtained by subtracting the predicted image and the image of the second moving body into a network model. (Configuration 8) 8. The image processing device according to any one of configurations 1 to 7, wherein there is no common range between the imaging range of the first imaging unit and the imaging range of the second imaging unit. (Configuration 9) a joint position estimation means for estimating a joint position of the second moving object based on the second extraction result, The image processing device according to any one of configurations 1 to 8, wherein the generating means generates the predicted image by changing joint positions of a three-dimensional form of the first moving body based on an estimation result of a joint position of the second moving body. (Configuration 10) The apparatus further includes a detection means for detecting a blocking area for the second moving object, 10. The image processing device according to any one of configurations 1 to 9, wherein the specifying means excludes an area in the predicted image that corresponds to the occluded area from objects to be compared. (Configuration 11) a joint position estimation means for estimating a joint position of the second moving object and a likelihood of the joint position based on the second extraction result, 11. The image processing device according to configuration 10, wherein the detection means determines an occluded area for the second moving object based on an estimation result of a likelihood of a joint position of the second moving object. (Configuration 12) the estimation result of the first three-dimensional shape includes information of a likelihood distribution of the three-dimensional shape of the first moving object; The image processing device according to any one of configurations 1 to 11, characterized in that the identification means utilizes a likelihood distribution of a three-dimensional shape of the first moving body when comparing the predicted image with the image of the second moving body. (Configuration 13) a first extraction means for extracting a first moving object from a first image captured by the first imaging unit; a second extraction means for extracting a second moving object from a second image captured by the second imaging unit; a first estimation means for estimating a three-dimensional shape of the first moving object based on a first extraction result which is an extraction result by the first extraction means; a second estimation means for estimating a state of the first moving object when the first moving object is imaged by the second imaging unit based on the first extraction result; a generating means for generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on the estimation result of the three-dimensional shape of the first moving object and the estimation result of the state of the first moving object; an identification means for identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted by the second extraction means; 13. An image processing device comprising: (Configuration 14) the second estimation means estimates a plurality of candidates for a state of the first moving object when the first moving object is imaged by the second imaging unit; 14. The image processing device according to configuration 13, wherein the generating means generates a plurality of the predicted images in accordance with the plurality of candidates. (Configuration 15) a joint position prediction unit that predicts a joint position of the first moving object imaged by the second imaging unit, The image processing device according to configuration 13 or 14, wherein the generation means generates the predicted image by changing joint positions of a three-dimensional form of the first moving body based on the joint positions predicted by the joint position prediction means. (Configuration 16) The apparatus further includes a storage unit that stores a joint position of the second moving object imaged by the second imaging unit, The image processing device according to configuration 15, characterized in that the joint position prediction means predicts the joint position of the first moving body imaged by the second imaging unit based on the joint position of the second moving body stored in the memory means. (Configuration 17) a first extraction means for extracting a first moving object from a first image captured by the first imaging unit; a second extraction means for extracting a second moving object from a second image captured by the second imaging unit; a second estimation means for estimating a state of the second moving object based on a second extraction result which is an extraction result by the second extraction means; a generating means for generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on a first extraction result which is an extraction result by the first extraction means and an estimation result of the state of the second moving object; an identification means for identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted by the second extraction means; 13. An image processing device comprising: (Method 1) a first extraction step of extracting a first moving object from a first image captured by a first imaging unit; a second extraction step of extracting a second moving object from a second image captured by the second imaging unit; a first estimation step of estimating a three-dimensional shape of the first moving object based on a first extraction result that is an extraction result obtained by the first extraction step; a second estimation step of estimating a state of the second moving object based on a second extraction result that is an extraction result obtained by the second extraction step; a generation step of generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on the estimation result of the three-dimensional shape of the first moving object and the estimation result of the state of the second moving object; a step of identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted in the second extraction step; 13. An image processing method comprising: (Method 2) a first extraction step of extracting a first moving object from a first image captured by a first imaging unit; a second extraction step of extracting a second moving object from a second image captured by the second imaging unit; a first estimation step of estimating a three-dimensional shape of the first moving object based on a first extraction result that is an extraction result obtained by the first extraction step; a second estimation step of estimating a state of the first moving object when the first moving object is imaged by the second imaging unit based on the first extraction result; a generation step of generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on the estimation result of the three-dimensional shape of the first moving object and the estimation result of the state of the first moving object; a step of identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted in the second extraction step; 13. An image processing method comprising: (Method 3) a first extraction step of extracting a first moving object from a first image captured by a first imaging unit; a second extraction step of extracting a second moving object from a second image captured by the second imaging unit; a second estimation step of estimating a state of the second moving object based on a second extraction result that is an extraction result obtained by the second extraction step; a generation step of generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on a first extraction result, which is an extraction result obtained by the first extraction step, and an estimation result of a state of the second moving object; a step of identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted in the second extraction step; 13. An image processing method comprising: (Program 1) The computer of the image processing device a first extraction means for extracting a first moving object from a first image captured by the first imaging unit; a second extraction means for extracting a second moving object from a second image captured by the second imaging unit; a first estimation means for estimating a three-dimensional shape of the first moving object based on a first extraction result which is an extraction result by the first extraction means; a second estimation means for estimating a state of the second moving object based on a second extraction result which is an extraction result by the second extraction means; a generating means for generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on the estimation result of the three-dimensional shape of the first moving object and the estimation result of the state of the second moving object; an identification means for identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted by the second extraction means; A program that functions as a (Program 2) The computer of the image processing device a first extraction means for extracting a first moving object from a first image captured by the first imaging unit; a second extraction means for extracting a second moving object from a second image captured by the second imaging unit; a first estimation means for estimating a three-dimensional shape of the first moving object based on a first extraction result which is an extraction result by the first extraction means; a second estimation means for estimating a state of the first moving object when the first moving object is imaged by the second imaging unit based on the first extraction result; a generating means for generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on the estimation result of the three-dimensional shape of the first moving object and the estimation result of the state of the first moving object; an identification means for identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted by the second extraction means; A program that functions as a (Program 3) The computer of the image processing device a first extraction means for extracting a first moving object from a first image captured by the first imaging unit; a second extraction means for extracting a second moving object from a second image captured by the second imaging unit; a second estimation means for estimating a state of the second moving object based on a second extraction result which is an extraction result by the second extraction means; a generating means for generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on a first extraction result which is an extraction result by the first extraction means and an estimation result of the state of the second moving object; an identification means for identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted by the second extraction means; A program that functions as a [Explanation of symbols]

[0094] 10: image processing device, 11A to 11C: cameras

Claims

1. a first extraction means for extracting a first moving object from a first image captured by the first imaging unit; a second extraction means for extracting a second moving object from a second image captured by the second imaging unit; a first estimation means for estimating a three-dimensional shape of the first moving object based on a first extraction result which is an extraction result by the first extraction means; a second estimation means for estimating a state of the second moving object based on a second extraction result which is an extraction result by the second extraction means; a generating means for generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on the estimation result of the three-dimensional shape of the first moving object and the estimation result of the state of the second moving object; a determination means for determining whether the first moving object and the second moving object are identical by comparing the predicted image with an image of the second moving object extracted by the second extraction means; 13. An image processing device comprising:

2. The image processing apparatus according to claim 1 , wherein the estimation result of the state of the second moving object includes information on the attitude of the second moving object and image quality parameters of the second moving object.

3. the estimation result of the state of the second moving body includes information on the attitude of the second moving body; The image processing device described in claim 1, characterized in that the generation means generates a background image of the predicted image based on the position of the second moving body in the second image, and generates a foreground image of the predicted image by converting a viewpoint relative to the three-dimensional shape of the first moving body based on the posture of the second moving body.

4. the estimation result of the state of the second moving object includes information of an image quality parameter of the second moving object; 2. The image processing apparatus according to claim 1, wherein said generating means converts said predicted image based on an image quality parameter of said second moving object.

5. 2 . The image processing apparatus according to claim 1 , wherein the estimation result of the three-dimensional shape of the first moving object includes information on a position and a color of the first moving object in a three-dimensional space.

6. The image processing device according to claim 1, characterized in that the identification means determines the identity of the first moving body and the second moving body based on the distance between feature vectors obtained by inputting the predicted image and the image of the second moving body into a network model.

7. The image processing device described in claim 1, characterized in that the identification means determines the identity of the first moving body and the second moving body based on the output result obtained by inputting the result obtained by subtracting the predicted image from the image of the second moving body into a network model.

8. 2 . The image processing device according to claim 1 , wherein the imaging range of the first imaging unit and the imaging range of the second imaging unit do not have a common range.

9. a joint position estimating means for estimating a joint position of the second moving object based on the second extraction result, The image processing device according to claim 1 , characterized in that the generating means generates the predicted image by changing joint positions of a three-dimensional form of the first moving body based on an estimation result of the joint positions of the second moving body.

10. The apparatus further includes a detection unit for detecting a blocking area for the second moving object, The image processing apparatus according to claim 1 , wherein the specifying means excludes an area in the predicted image that corresponds to the occluded area from objects to be compared.

11. a joint position estimation unit that estimates a joint position of the second moving object and a likelihood of the joint position based on the second extraction result, 11. The image processing apparatus according to claim 10, wherein the detection means determines an occluded area for the second moving object based on an estimation result of a likelihood of a joint position of the second moving object.

12. the estimation result of the first three-dimensional shape includes information of a likelihood distribution of the three-dimensional shape of the first moving object; 2. The image processing apparatus according to claim 1, wherein the specifying means uses a likelihood distribution of a three-dimensional shape of the first moving object when comparing the predicted image with the image of the second moving object.

13. a first extraction means for extracting a first moving object from a first image captured by the first imaging unit; a second extraction means for extracting a second moving object from a second image captured by the second imaging unit; a first estimation means for estimating a three-dimensional shape of the first moving object based on a first extraction result which is an extraction result by the first extraction means; a second estimation means for estimating a state of the first moving object when the first moving object is imaged by the second imaging unit based on the first extraction result; a generating means for generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on the estimation result of the three-dimensional shape of the first moving object and the estimation result of the state of the first moving object; a determination means for determining whether the first moving object and the second moving object are identical by comparing the predicted image with an image of the second moving object extracted by the second extraction means; 13. An image processing device comprising:

14. The second estimation means estimates a plurality of candidates for a state of the first moving object when the first moving object is imaged by the second imaging unit, The image processing apparatus according to claim 13 , wherein the generating means generates a plurality of the predicted images in accordance with the plurality of candidates.

15. a joint position prediction unit that predicts a joint position of the first moving object imaged by the second imaging unit, The image processing device according to claim 13 , wherein the generating means generates the predicted image by changing joint positions of a three-dimensional form of the first moving body based on the joint positions predicted by the joint position predicting means.

16. a storage unit that stores a joint position of the second moving object imaged by the second imaging unit, 16. The image processing device according to claim 15, wherein the joint position prediction means predicts the joint position of the first moving object imaged by the second imaging unit based on the joint position of the second moving object stored in the memory means.

17. a first extraction means for extracting a first moving object from a first image captured by the first imaging unit; a second extraction means for extracting a second moving object from a second image captured by the second imaging unit; a second estimation means for estimating a state of the second moving object based on a second extraction result which is an extraction result by the second extraction means; a generating means for generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on a first extraction result which is an extraction result by the first extraction means and an estimation result of the state of the second moving object; a determination means for determining whether the first moving object and the second moving object are identical by comparing the predicted image with an image of the second moving object extracted by the second extraction means; 13. An image processing device comprising:

18. a first extraction step of extracting a first moving object from a first image captured by a first imaging unit; a second extraction step of extracting a second moving object from a second image captured by the second imaging unit; a first estimation step of estimating a three-dimensional shape of the first moving object based on a first extraction result that is an extraction result obtained by the first extraction step; a second estimation step of estimating a state of the second moving object based on a second extraction result that is an extraction result obtained by the second extraction step; a generation step of generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on the estimation result of the three-dimensional shape of the first moving object and the estimation result of the state of the second moving object; a step of identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted in the second extraction step; 13. An image processing method comprising:

19. a first extraction step of extracting a first moving object from a first image captured by a first imaging unit; a second extraction step of extracting a second moving object from a second image captured by the second imaging unit; a first estimation step of estimating a three-dimensional shape of the first moving object based on a first extraction result that is an extraction result obtained by the first extraction step; a second estimation step of estimating a state of the first moving object when the first moving object is imaged by the second imaging unit based on the first extraction result; a generation step of generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on the estimation result of the three-dimensional shape of the first moving object and the estimation result of the state of the first moving object; a step of identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted in the second extraction step; 13. An image processing method comprising:

20. a first extraction step of extracting a first moving object from a first image captured by a first imaging unit; a second extraction step of extracting a second moving object from a second image captured by the second imaging unit; a second estimation step of estimating a state of the second moving object based on a second extraction result that is an extraction result obtained by the second extraction step; a generation step of generating a predicted image of the first moving object when the first moving object is imaged by the second imaging unit, based on a first extraction result that is an extraction result obtained by the first extraction step and an estimation result of a state of the second moving object; a step of identifying the identity of the first moving object and the second moving object by comparing the predicted image with an image of the second moving object extracted in the second extraction step; 13. An image processing method comprising:

21. A program for causing a computer to execute the image processing method according to any one of claims 18 to 20.

Citation Information

Cited By

  • Multi-camera object tracking system and multi-camera object tracking program

    JP7752458B1