Image processing device, method and program

By applying VM processing to the entire frame or local subject area and synthesizing with input data, the method addresses noise issues in VM-processed video, achieving high-quality, noise-free visualization of minute changes.

JP7754423B2Active Publication Date: 2025-10-15NIPPON TELEGRAPH & TELEPHONE CORP +2
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022023162
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-17
Publication Date
2025-10-15
Estimated Expiration
2042-02-17

AI Technical Summary

Technical Problem

Existing VM processing methods leave noise from subject tracking in video data, degrading the quality of processed video.

Method used

Apply VM processing to the entire frame image or local subject area, separately detect the subject area, and synthesize VM-processed subject area data with input data to remove tracking noise.

Benefits of technology

Obtain high-quality, noise-free VM-processed video data by aligning time and coordinate positions, ensuring clear visualization of minute changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007754423000005
    Figure 0007754423000005
  • Figure 0007754423000006
    Figure 0007754423000006
  • Figure 0007754423000007
    Figure 0007754423000007
Patent Text Reader

Abstract

To obtain high-quality video data in which a minute change for a subject region is emphasized while limiting a total amount of data.SOLUTION: An image processing apparatus is configured to: acquire first video data that includes an image of a subject in an image frame; detect a minute change of pixels from the acquired first video data; perform minute-change enhancement processing to generate second video data by including image components obtained by emphasizing the detected minute change of the pixels in the first video data; perform processing to detect, in parallel with the minute change enhancement processing, as a subject region, a limited region including the image of the subject from the first video data; extract third video data corresponding to the detected subject region from the second video data; and synthesizing the extracted third video data with the first video data in accordance with time positions and coordinate positions to generate and output fourth video data.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] One aspect of the present invention relates to an image processing device, method, and program having a video magnification (VM) processing function. [Background technology]

[0002] Minute changes in images of real-world subjects can sometimes contain important messages. For example, differences in the way an athlete uses their muscles or joints can be recorded as minute changes in the images of the subject, and these minute changes can be used as one factor in judging the athlete's athletic performance. Furthermore, minute changes such as heartbeat, chest movement due to breathing, changes in the brightness of the color of the body surface, and vibrations from cranes or engines can be used as a basis for detecting abnormal conditions based on images. However, human vision has difficulty detecting minute changes in images.

[0003] Therefore, VM processing technology has attracted attention. VM is a technology that visualizes minute changes in an image by detecting and highlighting the minute changes in the image. For example, Non-Patent Documents 1, 2, and 3 describe a technology that detects minute changes in the subject in the image by using optical flow, Euler's algorithm, phase change, or the like to detect changes in color or movement of the subject in the image and applying a time-series filter. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Ce Liu, Antonio Torralba, William T. Freeman, Fredo Durand, Edward H. Adelson. “Motion Magnification”. ACM Transactions on Graphics (2005). [Non-patent document 2] Yichao Zhang, Silvia L. Pintea, and Jan C. van Gemert. “Video Acceleration Magnification”. IEEE International Conference on Computer Vision and Pattern Recognition (2017). [Non-patent document 3] Shoichiro Takeda, Kazuki Okami, Dan Mikami, Megumi Isogai, Hideaki Kimata. “Jerk-Aware Video Acceleration Magnification”. IEEE International Conference on Computer Vision and Pattern Recognition (2018). Summary of the Invention [Problem to be solved by the invention]

[0005] Incidentally, when analyzing the minute state of a subject based on VM-processed video, it is usually sufficient to analyze only the subject region. One possible solution is to detect the area of ​​the subject by subject tracking and then apply VM processing only to this area. However, this method leaves subject tracking noise in the video data after VM processing, degrading the quality of the video data after VM processing.

[0006] This invention was made in light of the above circumstances, Noise-free High quality VM processed The present invention aims to provide a technology that makes it possible to obtain video data. [Means for solving the problem]

[0007] To solve the above problem, one aspect of an image processing device or method according to the present invention acquires first video data including an image of an object in an image frame, detects minute changes in pixels from the acquired first video data, and generates second video data including image components in the first video data that have been subjected to an emphasis process on the detected minute changes in pixels. In parallel with the minute change emphasis process, the image processing device or method performs a process of detecting a limited area including the image of the object from the first video data as an object area. Then, third video data corresponding to the detected object area is extracted from the second video data, and the extracted third video data is combined with the first video data by aligning the time position and coordinate position to generate fourth video data, and outputs the generated fourth video data.

[0009] In one aspect of this invention, VM processing is applied to the entire area of ​​a frame image of input video data or to a local area including a subject area, and the subject area is detected from the input video data separately from this VM processing, and VM-processed video data of the subject area is extracted from the video data after the VM processing based on information about the detected subject area, and the VM-processed video data of the extracted subject area is synthesized with the input video data. Therefore, for example, if the subject area is detected by subject tracking processing and VM processing is applied only to this subject area, noise due to the subject tracking processing may remain in the video data after VM processing, One aspect of the present invention According to this method, it is possible to obtain composite image data that does not contain noise caused by the tracking process of the subject area. [Effects of the Invention]

[0010] That is, according to one aspect of the present invention, Noise-free High quality VM processed It is possible to provide a technique that makes it possible to obtain video data. [Brief explanation of the drawings]

[0011] [Figure 1]FIG. 1 is a block diagram showing an example of the hardware configuration of an image processing apparatus according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing an example of the software configuration of the image processing apparatus according to the embodiment of the present invention. [Figure 3] FIG. 3 is a flowchart showing an example of a processing procedure and processing contents of a series of image processing executed by the image processing device shown in FIG. [Figure 4] FIG. 4 is a flowchart showing an example of the processing procedure and processing contents of the VM processing of the image processing procedure shown in FIG. [Figure 5] FIG. 5 is a flowchart showing an example of the processing procedure and processing contents of the subject tracking processing from the image processing procedure shown in FIG. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0013] [One embodiment] (Configuration example) FIG. 1 is a block diagram showing an example of the hardware configuration of an image processing apparatus according to an embodiment of the present invention, and FIG. 2 is a block diagram showing an example of the software configuration of the image processing apparatus.

[0014] The image processing device VD is provided as one of the processing functions in an information processing device such as a server computer or a personal computer, etc. A camera and an input / output device (not shown) are connected to the image processing device VD via a signal cable or a network, for example.

[0015] The camera may be, for example, a high-resolution camera, and captures subjects such as a person's body during exercise, the movement of a living body such as the heart or chest, or machinery such as a crane or engine in operation, and outputs video data including the image of the subject. In addition to the camera, an external storage device storing video data of the subject may be connected to the image processing device VD. The subject to be captured may also be various natural objects, animals, artwork, scenery, etc.

[0016] The input / output device is, for example, a monitor device, and is used to display VM-processed video data, etc., output from the image processing device VD. In addition to the input / output device, an external storage device, an information processing device for management or analysis, etc. may be connected to the image processing device VD.

[0017] The image processing device VD has a control unit 1 that uses a hardware processor such as a central processing unit (CPU), and this control unit 1 is connected via a bus 5 to a memory unit having a program memory unit 2 and a data memory unit 3, and an input / output interface (hereinafter, interface will be abbreviated as I / F) unit 4.

[0018] The input / output I / F unit 4 has a communication interface function, and transmits and receives input video data and VM-processed video data to and from the above-mentioned cameras and input / output devices via a signal cable or a network.

[0019] The program storage unit 2 is configured by combining, for example, a non-volatile memory such as a solid-state drive (SSD) as a storage medium that can be written to and read from at any time, and a non-volatile memory such as a read-only memory (ROM), and stores middleware such as an operating system (OS), as well as application programs required to execute various control processes according to one embodiment. Hereinafter, the OS and each application program will be collectively referred to as the program.

[0020] The data storage unit 3 is, for example, a combination of a non-volatile memory such as an SSD that can be written to and read from at any time as a storage medium, and a volatile memory such as a RAM (Random Access Memory), and is equipped with an input image storage unit 31, a VM processing image storage unit 32, a subject area detection information storage unit 33, and a composite image storage unit 34 as the main storage units required to implement one embodiment.

[0021] The input video storage unit 31 is used to store high-resolution input video data acquired from a camera or an external storage device.

[0022] The VM processed video storage unit 32 is used to store input video data that has been subjected to VM processing.

[0023] The object region detection information storage unit 33 is used to store information representing an object region detected from the input video data before VM processing.

[0024] The composite image storage unit 34 is used to store composite image data in which only the subject area of ​​the input image data is replaced with image data that has been subjected to VM processing.

[0025] The control unit 1 includes, as processing functions necessary to implement one embodiment, an input video acquisition processing unit 11, a VM processing unit 12, an object tracking processing unit 13, a video synthesis processing unit 14, and a synthesized video output processing unit 15. These processing units 11 to 15 are all realized by causing a hardware processor of the control unit 1 to execute an application program stored in the program storage unit 2.

[0026] Note that part or all of the processing units 11 to 15 may be realized using hardware such as an LSI (Large Scale Integration) or an ASIC (Application Specific Integrated Circuit).

[0027] The input video acquisition processing unit 11 acquires input video data from a video source such as a camera or an external storage device via the input / output I / F unit 4, and performs processing to store the acquired input video data in the input video storage unit 31.

[0028] The VM processing unit 12 reads input video data frame by frame from the input video storage unit 31, and performs VM processing on the image of each frame that has been read. Then, the VM processing unit 12 stores the VM-processed input video data frame by frame in the VM processed video storage unit 32. An example of the VM processing will be described in the operation example.

[0029] The subject tracking processing unit 13 performs subject tracking processing in parallel with the VM processing by the VM processing unit 12, reading input video data frame by frame from the input video storage unit 31, and detecting the subject area from each of the read frame images. The subject tracking processing unit 13 then stores information representing the subject area detected from each of the frame images in the subject area detection information storage unit 33 in frame order. An example of the subject area detection processing will be described in the operation example.

[0030] The video synthesis processing unit 14 extracts VM processed video data of the subject region from the VM processed video data stored in the VM processed video storage unit 32 based on the information representing the subject region stored in the subject region detection information storage unit 33, and synthesizes the VM processed video data of the extracted subject region with the input video data stored in the input video storage unit 31 by matching the time position and coordinate position. The video synthesis processing unit 14 then stores the synthesized video data generated by the synthesis process in the synthesized video storage unit 34.

[0031] The composite video output processing unit 15 reads the composite video data from the composite video storage unit 34 and outputs the read composite video data from the input / output I / F unit 4 to an input / output device or information processing device (not shown).

[0032] (Example of operation) Next, an example of the operation of the image processing device VD configured as above will be described. FIG. 3 is a flowchart showing an example of the processing procedure and processing contents of the image processing operation executed by the control unit 1 of the image processing device VD.

[0033] (1) Acquiring input video data For example, when a target user starts exercising, a camera captures the target user's entire body or specific parts while exercising, and the resulting video data is output. The video data may consist of, for example, high-resolution color moving images, with each frame of the image containing brightness information and color information for each pixel. Note that the video data does not necessarily have to be high-resolution color images, and may also be monochrome images with normal resolution. Furthermore, the video data may be video data stored in advance in an external storage device other than a camera.

[0034] In response to this, in step S1, the control unit 1 of the image processing device VD, under the control of the input video acquisition processing unit 11, receives the video data output from the camera or external storage device as input video data via the input / output I / F unit 4. Then, the received input video data is stored in the input video storage unit 31 in chronological order for each frame.

[0035] (2) VM processing Next, in step S2, the control unit 1 of the image processing device VD, under the control of the VM processing unit 12, reads the input video data from the input video storage unit 31 frame by frame, and performs VM processing on the image of each frame that has been read.

[0036] As an example of VM processing, an outline of the operation will be described below using the VM processing relating to minute changes in movement described in Non-Patent Document 2. Fig. 4 is a flowchart showing an example of the processing procedure and processing content.

[0037] For example, if each frame image of the input video data has pixel coordinates (x, y), frame time positions t=1,...,T, and color space c∈C, then I c In this case, the VM processing unit 12 performs VM processing on the entire image area of ​​each frame image or on a local area (x, y) ∈ (X, Y) including the subject image as follows, and outputs the VM-processed video data I^ c We get (x,y,t).

[0038] That is, first, in step S21, the VM processing unit 12 converts the input video data I c For each frame (x,y,t), the Y color signal component I y (x,y,t) is extracted. Note that this Y color signal component I y As a result of the extraction process of (x,y,t), I i (x,y,t) ,I q (x,y,t) remains.

[0039] Next, in step S32, the VM processing unit 12 uses a filter called a Complex Steerable Filter (hereinafter abbreviated as CFS) with a filter characteristic of ψ ω,θ (x,y) to extract the Y color signal component I y K (x,y,t) is transformed into an analytic signal representation divided into multiple band frequencies ω∈Ω and multiple directions θ∈Θ as follows:

number

[0040] where R y ω,θ (x,y,t) represents the analytic signal representation, and A y ω,θ (x,y,t) is the amplitude signal representation, φ y ω,θ(x,y,t) represent the phase signal representation. The phase signal representation is known to represent local motion changes centered on the coordinate (x,y) in the video data. This is described in detail, for example, in Neal Wadhwa, Michael Rubinstein, Fredo Durand, William T. Freeman. "Phase-based Video Motion Processing." ACM Transactions on Graphics (2013).

[0041] Subsequently, in step S23, the VM processing unit 12 calculates the phase signal representation φ y ω,θ (x,y,t) with a pre-specified arbitrary time frequency f t Using time series filtering h(t;f t ) which results in a time frequency f t A small phase signal representation C^ in y ω,θ (x,y,t) is generated as follows, where * indicates a convolution operation.

number

[0042] Subsequently, in step S24, the VM processing unit 12 calculates the generated minute phase signal representation C^ y ω,θ (x, y, t) is amplified using an arbitrary enhancement factor α, and then the phase signal representation φ y ω,θ (x,y,t)

number

[0043] Next, in step S25, the VM processing unit 12 generates a phase signal representation φ̂ in which only the minute phase signal representation is emphasized. y ω,θ Using (x, y, t), an analytic signal representation to which VM processing has been applied is constructed as follows:

number

[0044] Subsequently, in step S26, the VM processing unit 12 calculates the VM-processed Y color signal image I^ using the CSF in equation (5). y Then, the VM processing unit 12 converts the inversely converted Y color signal image I^ into (x, y, t). y (x, y, t) is the I remaining by the extraction process in step S21. i (x,y,t) ,I q (x,y,t) and the VM processed video data I^ C Generate (x,y,t).

[0045] Finally, the VM processing unit 12 processes the generated VM-processed video data I^ C (x, y, t) is stored in the VM processed video storage unit 32. Note that the processing procedure and processing contents of the VM processing are not limited to those described above, and any method can be applied.

[0046] (3) Subject tracking processing In parallel with the VM processing by the VM processing unit 12, the control unit 1 of the image processing device VD, under the control of the subject tracking processing unit 13, executes a process of tracking the image of the subject contained in the input video data in a time series in step S3 as follows.

[0047] That is, the object tracking processing unit 13 reads input video data from the input video storage unit 31. Then, at each preset timing t, the object tracking processing unit 13 extracts an object region (x, y, t) ∈(X obj ,Y obj ,t) is detected.

[0048] FIG. 5 is a flowchart showing an example of the processing procedure and processing content of the subject detection processing by the subject tracking processing unit 13.

[0049] As shown in FIG. 5, first, in step S31, the object tracking processing unit 13 extracts a plurality (N) of joint points {(x joint ,y joint ,t)} N i=1 Detect.

[0050] Next, in step S32, the subject tracking processing unit 13 calculates the number of detected N joint points {(x joint ,y joint ,t)} N i=1 From the frame image, an image region having a predetermined size and shape is extracted, centered on each of the joints. For example, a range of radius M pixels (x, y, t) ∈ {N(x joint ,y joint ,t)} N i=1 is extracted from the frame image. The extracted image area is then called the subject area (X obj ,Y obj ,t) ∈{N(x joint ,y joint ,t)} N i=1 Let's say.

[0051] The above-mentioned subject region detection process is described in detail in, for example, the following document. Z. Cao, G. Hidalgo, T. Simon, S. Wei and Y. Sheikh, "OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields". IEEE Transactions on Pattern Analysis & Machine Intelligence,2021.

[0052] In the above description, the range for extracting the subject region is set in advance to a range with a radius of M pixels. However, this is not limiting. For example, the joint length of the subject may be calculated in pixel units from the input video data, and the value of M may be set based on the calculated joint length. In this way, it is possible to set an optimal range of radius M that takes into account individual differences between subjects.

[0053] As a method for setting the subject region from the input video data, an image segmentation method or the like can be applied. The image segmentation method is described in detail in, for example, the following document. Chuong Huynh, Anh Tuan Tran, Khoa Luu, and Minh Hoai. “Progressive Semantic Segmentation”, IEEE Transactions on Pattern Analysis & Machine Intelligence,2021.

[0054] Furthermore, the subject region tracking method is not limited to the above-described method, and any method can be used depending on the type and shape of the subject. For example, if the subject is an object other than a human body, the length between the movable base points of this object is calculated. Furthermore, the range from which the subject region is extracted may be other shapes such as an ellipse or a rectangle in addition to a circle.

[0055] Finally, the subject tracking processing unit 13 stores information representing the subject region detected as described above in the subject region detection information storage unit 33 in step S33.

[0056] (4) Video composition processing Next, in step S4, the control unit 1 of the image processing device VD, under the control of the video synthesis processing unit 14, executes processing to generate synthesized video data as follows.

[0057] That is, video synthesis processing unit 14 first reads information representing the detected subject area from subject area detection information storage unit 33, and based on this information representing the subject area, extracts VM processed video data corresponding to the subject area from the VM processed video data stored in VM processed video storage unit 32. Then, video synthesis processing unit 14 synthesizes the VM processed video data of the extracted subject area with the input video data stored in input video storage unit 31 by matching the time position and coordinate position.

[0058] The video synthesis processing unit 14 stores the video data obtained by the synthesis process in the synthesized video storage unit 34. Thus, the video synthesis process obtains synthesized video data in which only the subject area of ​​the input video data is replaced with VM processed video data.

[0059] When combining the VM-processed video data of the subject region with the input video data, pixel smoothing processing may be performed on the boundary between the VM-processed video data of the subject region and the input video data. For example, image processing such as alpha blending may be used to perform image processing so that the boundary between the VM-processed video data of the subject region and the input video data is smoothly combined. In this way, the contours of the subject region that have been emphasized by VM processing are not lost, making it possible to generate more natural VM-combined video data.

[0060] (5) Output of composite video data Finally, in step S5, under the control of the composite video output processing unit 15, the control unit 1 of the image processing device VD reads out the composite video data in which only the subject area of ​​the input video data has been replaced with VM processed video data from the composite video memory unit 34, and outputs the read-out composite video data from the input / output I / F unit 4 to an input / output device not shown.

[0061] Therefore, the input / output device displays the composite video data on a monitor, for example, and an administrator can clearly see even minute changes in the movements of the entire body or specific parts of the subject, for example, a target user exercising, by viewing the displayed VM-processed composite video.

[0062] The composite video data may be transmitted to an external information processing device, which may then perform an analysis process, such as extracting features of minute changes in the subject by performing image processing on the composite video data, inputting the extracted features of minute changes into a learning model, and obtaining information representing an analysis result of the movement characteristics of the subject.

[0063] (effect) As described above, in one embodiment, the image processing device VD performs VM processing on video data acquired from, for example, a camera or an external storage device, and simultaneously performs processing to detect a subject area from the input video data. Then, video data corresponding to the subject area is extracted from the VM-processed video data, and the VM-processed video data of the extracted subject area is combined with the input video data by aligning the time position and coordinate position.

[0065] Therefore, according to one embodiment, the following effects can be achieved.If a subject area is detected by subject tracking processing and VM processing is applied only to the detected subject area, noise due to the subject tracking processing may remain in the video data after VM processing. In response to this, in one embodiment, VM processing is applied to the entire area of ​​a frame image of input video data or a local area including the subject area, and the subject area is detected from the input video data separately from this VM processing. VM-processed video data of the subject area is extracted from the video data after VM processing based on information about the detected subject area, and the VM-processed video data of the extracted subject area is synthesized with the input video data. Therefore, high-quality synthesized video data that does not include noise due to the subject area tracking processing can be obtained.

[0066] Furthermore, in one embodiment, the subject area is tracked by detecting multiple joint positions of a person from input video data, setting ranges each having a predetermined radius M centered on each of the detected joint positions, and determining the area surrounded by these ranges as the subject area. As a result, it is possible to detect the subject area with simpler processing than when detecting the contour of the subject.

[0067] Furthermore, when setting the subject area, the joint lengths of the person are detected from the input video data, and the radius M is set based on these joint lengths, making it possible to set an optimal radius M that takes into account the individual differences between subjects.

[0068] Furthermore, when combining VM processed video data of the subject area with input video data, image processing techniques such as alpha blending can be applied to smoothly combine the boundary areas between the subject area and the input video data, thereby preventing the contour areas of the subject area that have been emphasized by VM processing from being lost, making it possible to generate natural VM combined video data.

[0069] [Other embodiments] (1) In the embodiment, the input video data acquisition process, VM process, subject area tracking process, video synthesis process, and synthesized video output process are performed in one image processing device VD. However, the present invention is not limited to this. Part of the series of processes from the VM process, subject area tracking process, video synthesis process, super-resolution process, and video output process may be distributed among multiple information processing devices.

[0070] Furthermore, the image processing device does not need to be configured to exclusively perform a series of processes from the acquisition of input video data to the output of composite video data, but may also be configured to perform other processing functions, such as the analysis processing function of the subject area.

[0071] (2) In the embodiment, a series of processes from the acquisition of input video data to the output of composite video data is performed in real time. However, the present invention is not limited to this. For example, video data for the entire period or a part of a person's exercise may be temporarily stored in input video storage unit 31, and VM processing, subject area tracking processing, and video composition processing may be performed collectively on this stored video data.

[0072] In addition, all or part of the composite video data may be temporarily stored in the composite video storage unit 34, and in this state, when a request to acquire the composite video data is sent, for example, from an administrator's terminal device, the composite video data may be read from the composite video storage unit 34 and transmitted in bulk to an input / output device or an external information processing device, etc.

[0073] (3) In one embodiment, an example was given of generating video data consisting of moving images with VM processing applied only to the subject area, but it is also possible to generate video data consisting of still images with VM processing applied only to the subject area.

[0074] (4) In addition, the functional configuration, processing procedures and contents, VM processing, subject area tracking processing, and video synthesis processing procedures and contents, etc. of the image processing device can be modified and implemented in various ways without departing from the spirit of this invention.

[0075] Although the embodiments of the present invention have been described in detail above, the above description is merely an example of the present invention in every respect. It goes without saying that various improvements and modifications can be made without departing from the scope of the present invention. In other words, when implementing the present invention, specific configurations according to the embodiments may be appropriately adopted.

[0076] In short, this invention is not limited to the above-described embodiments, and in the implementation stage, the components can be modified and embodied without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]

[0077] VD: Image processing device 1...Control unit 2...Program memory section 3...Data storage unit 4...Input / output interface 5...Bus 11...Input image acquisition processing unit 12...VM processing section 13...Subject tracking processing unit 14...Video synthesis processing unit 15...Synthetic image output processing section 31...Input image storage unit 32...VM processing video storage unit 33…Subject area detection information storage unit 34...Synthetic image memory unit

Claims

1. a first processing unit that acquires first video data including an image of a subject in an image frame; a second processing unit that performs minute change enhancement processing to detect minute changes in pixels from the acquired first video data, and generate second video data that includes image components that have been subjected to enhancement processing on the detected minute changes in pixels in the first video data; and a third processing unit that performs processing to detect a limited area including an image of the subject from the first video data as a subject area in parallel with the small change emphasis processing; a fourth processing unit that extracts third video data corresponding to the detected object region from the second video data, and synthesizes the extracted third video data with the first video data by aligning the time position and coordinate position, thereby generating fourth video data; a fifth processing unit that outputs the generated fourth video data; An image processing device comprising:

2. The image processing device according to claim 1 , wherein the third processing unit detects, from the first video data, an image area that is set in advance to include a range of motion of the subject, as the subject area.

3. 2. The image processing device according to claim 1, wherein the third processing unit calculates a length between a plurality of movable base points of the subject in pixel units from the first video data, sets an image area including a movable range of the subject based on the calculated length between the movable base points, and detects the set image area from the first video data as the subject area.

4. 2. The image processing device according to claim 1, wherein the fourth processing unit further performs a process of smoothing a boundary portion between the third video data and the first video data when synthesizing the third video data with the first video data.

5. An image processing method executed by an information processing device, a first process for acquiring first video data including an image of a subject in an image frame; a second processing step of performing minute change enhancement processing to detect minute changes in pixels from the acquired first video data, and to generate second video data including image components that have been subjected to enhancement processing on the detected minute changes in pixels in the first video data; a third processing step of detecting a limited area including an image of the subject from the first video data as a subject area in parallel with the small change emphasis processing; a fourth processing step of extracting third image data corresponding to the detected object region from the second image data, and synthesizing the extracted third image data with the first image data by aligning the time position and coordinate position, thereby generating fourth image data; a fifth processing step of outputting the generated fourth video data; An image processing method comprising:

6. 5. A program for causing a processor included in the image processing device to execute processing performed by at least one of the first to fifth processing units included in the image processing device according to claim 1.

Citation Information

Patent Citations

  • Non-contact mental stress assessment system

    CN113229790A

  • Merchandise proposition system, merchandise sales system, and merchandise design support system

    JP2006323804A

  • Image processing apparatus, and image processing method

    JP2010041682A

  • Image processing device, image processing method and image processing program

    JP2019087126A

  • Change timing detector, change timing detection method and change timing detection program

    JP2021013608A