Image recognition device, image recognition method, and program
The image recognition system enhances accuracy in time-series images by converting visual components and calculating reliability, addressing the challenges of varying colors and brightness in existing technologies.
Patent Information
- Application Number
- JP2024054303
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2025-10-09
AI Technical Summary
Existing image recognition technologies, particularly those using machine learning, struggle with accuracy when recognizing subjects in images with varying colors, vividness, and brightness, especially in time-series images such as videos, as they do not account for sensitivity adjustments.
An image recognition system that converts the values of visual components like hue, saturation, and luminance in time-series images to generate multiple converted images, recognizes subjects in each, calculates reliability based on the variation of these components, and selects the most reliable recognition result.
Improves the accuracy of subject recognition in time-series images by accounting for variations in color, brightness, and luminance, ensuring consistent and reliable identification regardless of these factors.
Smart Images

Figure 2025152419000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image recognition device, an image recognition method, and a program for recognizing a subject from an image including the subject. [Background technology]
[0002] Conventionally, techniques for performing image recognition using machine learning, such as deep learning, have been known. In order to improve the accuracy of image recognition, it is necessary to create a learning model from training data of various images of the subject to be recognized, in particular from training data of various images in which the subject appears differently in images with similar configurations, i.e., different colors, vividness, brightness, etc.
[0003] There are various other technologies for improving recognition accuracy. For example, a technology has been disclosed that improves subject recognition accuracy by capturing multiple images with different sensitivities in one frame period to generate image data, recognizing the subject from each piece of image data, and recognizing the subject that appears in one frame of image based on the recognition results (see Patent Document 1). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2020-187409 Summary of the Invention [Problem to be solved by the invention]
[0005] Incidentally, when recognizing a subject contained in an image, the accuracy of recognizing the subject can decrease depending on the appearance of the subject, i.e., the color / vividness / brightness of the subject and any unevenness in these, so it is desirable to prepare a variety of images in which the subject appears differently as training data. The technology in Patent Document 1 mentioned above uses multiple images captured at the same time with different sensitivities (of the camera's imaging sensor) to recognize the subject. However, there is no mention of changing the sensitivity for video to improve the accuracy of subject recognition.
[0006] The present invention has been made based on this background, and aims to provide an image recognition device, an image recognition method, and a program that can improve recognition accuracy when recognizing a subject from time-series images, which are a group of images arranged in chronological order, such as a video. [Means for solving the problem]
[0007] One aspect of the present invention for achieving the above-mentioned goal is to provide an image recognition device that recognizes a subject from time-series images, comprising: an image acquisition unit that acquires the time-series images including the subject; an image conversion unit that converts the values of visual components to predetermined values for at least a portion of the image area of the subject included in the time-series images, and generates each converted time-series image in which the values of the visual components have been converted; a subject recognition unit that extracts the subject from each converted time-series image and recognizes the subject in time series; a reliability calculation unit that calculates the reliability of the recognition result of each converted time-series image based on the amount of variation of the subject recognized in time series from each converted time-series image; and a recognition result selection unit that selects the recognition result based on the reliability.
[0008] Although the present invention is categorized as an image recognition device, the same effects and advantages can be achieved with an image recognition method and program. [Effects of the Invention]
[0009] According to the present invention, it is possible to improve the recognition accuracy when recognizing a subject from time-series images. [Brief explanation of the drawings]
[0010] [Figure 1]FIG. 1 is a schematic diagram illustrating skeleton recognition by an image recognition system according to an embodiment of the present invention. [Figure 2] 1 is an explanatory diagram illustrating a schematic configuration example of an image recognition system according to an embodiment of the present invention. [Figure 3] FIG. 2 is an explanatory diagram illustrating an example of the hardware configuration of the image recognition device according to the present embodiment. [Figure 4] 1 is a block diagram showing a functional configuration of an image recognition device according to an embodiment of the present invention. [Figure 5] 10 is a flowchart illustrating an example of processing performed by the image recognition device according to the present embodiment. [Figure 6] FIG. 10 is an explanatory diagram showing an example of a conversion parameter table according to the embodiment; [Figure 7] FIG. 10 is a diagram showing skeletons extracted by general skeleton extraction according to the present embodiment. [Figure 8] 10A and 10B are diagrams for explaining calculation of reliability from skeletal information along a time series by the image recognition device according to the present embodiment. [Figure 9] FIG. 2 is an explanatory diagram showing an example of an operation screen of the image recognition device according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments will be described with reference to the drawings. In the following description, identical or similar components will be designated by the same reference numerals, and duplicate descriptions will be omitted. In the following description, when it is necessary to distinguish between components of the same type, an identifier (numbers, letters, etc.) will be written in parentheses after the reference numeral that collectively refers to the components.
[0012] An image recognition system is a device that recognizes subjects included in time-series images. Here, a time-series image is a group of images, such as a video, that are composed of multiple frame-by-frame images (hereinafter also referred to as frame images) arranged in chronological order. Specifically, the image recognition system is a device that acquires time-series images including a subject, generates each time-series image (hereinafter also referred to as a converted time-series image) in which the values of visual components are converted for at least some of the image areas of the subject included in the time-series image, recognizes the subject from each converted time-series image, and selects the optimal recognition result from among the recognized subjects, thereby improving the accuracy of subject recognition.
[0013] Here, the visual components are elements that visually reproduce the color and / or light expressed in an image using pixels, and are elements that quantitatively represent the color and / or light that constitute the image. Specifically, these are hue, saturation, lightness, and luminance. Converting the values of the visual components means converting the value of at least one of the hue, saturation, lightness, and luminance to adjust the color and light expressed in the image.
[0014] In this specification, the subject may be any object, either a moving object or a still object. The object may be any object that can be configured with two parts, in other words, two points that are separated from each other, and is preferably a human or robot having a skeleton. The robot may be either an industrial robot or a service robot, and may be not only humanoid but also arm-type or hand-type. The subject may be a part of any object, such as a part of a human hand or face.
[0015] To specifically explain a method for generating transformed time-series images in which the values of visual components are transformed for time-series images including a subject in an image recognition system, the values of the visual components for at least some of the subjects (or non-subjects) in each frame image included in the time-series images are transformed to different values, thereby generating transformed time-series images in which the values of the visual components are transformed. Note that the values of the visual components for some of the subjects may be transformed to different values, or the values of the visual components for the entire time-series images may be transformed to different values. Here, a visual component is an element that visually reproduces the color and / or light expressed in an image when it is displayed by pixels on an image display means.
[0016] In addition, to specifically explain the selection of recognition results in the image recognition system, the subject is recognized in time series from each of the converted time series images in which the values of the visual components have been converted, the reliability is calculated from the amount of variation of the subject recognized in time series, and the selection is made based on the calculated reliability. The image recognition system recognizes the subject based on at least two body parts that have been set using positional information of the subject's body parts that have been registered in advance as extraction targets, and if the subject is a person or a robot, it recognizes it based on its skeleton.
[0017] The amount of movement of the subject is the amount of movement between any parts of the subject, for example, the amount of movement in the distance between one part and another part of the subject. The any part is selected from parts of the subject that have been set as extraction targets. If the subject is a person, skeletal parts (nodes) such as the shoulders and waist are selected as the any part, and if the subject is a robot, the joints connecting the links are selected as the any part.
[0018] The image recognition system described above will be described in further detail below, taking as an example a case where the subject is a person.
[0019] <Summary> FIG. 1 is a schematic diagram of an image recognition system 100 according to this embodiment when the subject is a person. The image recognition system 100 is a device that recognizes people included in time-series images. It generates transformed time-series images by converting the values of visual components for the time-series images that include the person, recognizes the person's skeleton from each transformed time-series image, and selects the optimal recognition result from the recognized skeletons, thereby improving the accuracy of skeleton recognition.
[0020] First, the image processing system 100 acquires time-series images TI including a person for which the image processing system 100 performs skeletal recognition. In FIG. 1, the time-series images TI include frame-by-frame images, frame image I(t-2), frame image I(t-1), and frame image I(t). The number of frame images included in the time-series images TI is not limited to three and may be any number. The frame images included in the time-series images TI may be composed of still images extracted from a video, or may be composed of still images captured continuously by one or more imaging devices 30. It is desirable that the frame images included in the time-series images TI be images captured at short time intervals.
[0021] The image processing system 100 converts the values of the visual components of the acquired time-series images TI into time-series images (TI A ,TI B, TI C Specifically, the values of the visual components of the frame images I(t-2), I(t-1), and I(t) included in the time-series image TI are converted to generate a converted frame image I A (t-2),I A (t-1),I A (t) is generated, and the converted frame images I A (t-2),I A (t-1),I A (t) is the transformed time series image TI A Generate.
[0022] Similarly, the image processing system 100 converts the values of the visual components of each of the frame images I(t-2), I(t-1), and I(t) included in the time-series image TI into the converted time-series image TI A The converted image I is then converted to a different visual component value. B (t-2),I B (t-1),I B (t) is the transformed time series image TI B Similarly, the image processing system 100 converts the values of the visual components of the frame images I(t-2), I(t-1), and I(t) included in the time-series image TI into the converted time-series image TIA ,TI B The values of the visual components are changed to different values, and the converted frame images I are displayed in a time series. C (t-2),I C (t-1),I C (t) is the transformed time series image TI C The image processing system 100 converts the values of the visual components of a person or a part of a person in the time-series image TI to generate a converted time-series image TI A ,TI B, TI C may be generated.
[0023] For example, in an image of a person wearing gloves or a mask, by converting the hue and changing the color of the gloves or mask to a color close to the skin color, it is possible to convert the image into an image that looks like a person not wearing gloves or a mask. Also, if the image is black and white (monochrome), by converting the face and hands to a skin color and also converting the mask and gloves to an equivalent skin color, it is possible to convert the image into an image that looks like a person not wearing gloves or a mask, as described above. Of course, even a monochrome image can be converted into an image that looks like a person not wearing gloves or a mask by changing the brightness of the mask and gloves to the same brightness as the pixels of the skin area without converting the skin, mask, gloves, etc. to a skin color.
[0024] In Fig. 1, the time-series image TI is converted into the converted time-series image TI A ,TI B, TI C Although three converted time-series images are generated, it is sufficient to generate at least two or more converted time-series images. Also, the values of the visual components may not be converted, and the original time-series image TI may be considered as the converted time-series image. Note that the recognition accuracy of skeleton recognition improves when there are more converted time-series images than fewer, but recognition accuracy saturates when the number of converted time-series images exceeds a certain number.
[0025] The image processing system 100 generates the transformed time series image TI A ,TI B, TI C In each converted frame image, the skeleton of the person in the image is recognized and the skeleton information TFA ,TF B, TF C Here, the skeleton information is information indicating the skeleton of a person, and includes information such as position information of the human body parts that make up the skeleton. Specifically, the image processing system 100 converts the converted time-series image TI A Converted frame image I included in A (t-2),I A (t-1),I A (t) In each step, the skeleton of the person is recognized and the skeleton information F A (t-2),F A (t-1),F A (t) and obtain the skeletal information F A (t-2),F A (t-1),F A Time series skeleton information TF consisting of (t) A Obtain the transformed time series image TI B, TI C Similarly, the time series skeleton information TF B ,TF C Get.
[0026] In particular, the image processing system 100 generates a transformed time series image TI B Converted frame image I included in B (t-2),I B (t-1),I B (t) In each step, the skeleton of the person is recognized and the skeleton information F B (t-2),F B (t-1),F B (t) and obtain the skeletal information F B (t-2),F B (t-1),F B Time series skeleton information TF consisting of (t) B Also, the transformed time series image TI C Converted frame image I included in C (t-2),I C (t-1),I C (t) In each case, the skeleton of the person is recognized and the skeleton information F C (t-2),F C (t-1),FC (t) and obtain the skeletal information F C (t-2),F C (t-1),F C Time series skeleton information TF consisting of (t) C Get.
[0027] The image processing system 100 generates time-series skeletal information TF A ,TF B ,TF C Based on the amount of variation between any human body parts calculated from the time series skeletal information TF A The skeleton information F contained in A (t-2),F A (t-1),F A (t) The amount of variation between any human body parts is calculated for each, and a reliability S is calculated from each calculated amount of variation. A (t-2),S A (t-1),S A (t) and calculate the reliability S A (t-2),S A (t-1),S A (t) is the time series reliability TS A Here, the arbitrary body part is, for example, between both shoulders, and the amount of variation is the distance between both shoulders.
[0028] Time-series skeletal information TF B ,TF C Similarly, the image processing system 100 calculates the reliability S along the time series. B (t-2),S B (t-1),S B (t) is the time series reliability TS B and the reliability S along the time series C (t-2),S C (t-1),S C (t) is the time series reliability TS C In Figure 1, the time series reliability TS A, TS B, TS C is shown in the graph.
[0029] The image processing system 100 then calculates the time series reliability TS A, TS B, TS C For example, the image processing system 100 selects the most reliable skeletal information at each time point as the skeletal information of the person in the time-series images T1. In the case of FIG. 1, the image processing system 100 selects the most reliable skeletal information at time point t-2 as the skeletal information of the person in the time-series images T1. A (t-2), TS at time t-1 B (t-1), TS at time t c (t) is selected as the skeleton information of the person in the time-series image TI.
[0030] <Image recognition system configuration example> FIG. 2 is a schematic configuration example of an image recognition system 100. As shown in the figure, the image recognition system 100 is configured by an imaging device 30 and an image recognition device 10, which is an information processing device, connected to each other so as to be able to communicate with each other via wired or wireless communication means 20. Although FIG. 2 illustrates one imaging device 30, it is preferable to install multiple imaging devices 30. One or more imaging devices 30 are installed to cover a range in which a subject moves and output an image of the subject. The communication means 20 is, for example, a communication means conforming to various communication standards such as USB (Universal Serial Bus) and RS-232C, a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, a dedicated line, etc. Other devices such as a mobile terminal may be connected to the communication means 20. The image recognition device 10 recognizes a subject based on an image of the subject captured by the imaging device 30.
[0031] <Example of hardware configuration for image recognition device> 3 shows an example of the hardware configuration of the image recognition device 10. As shown in the figure, the image recognition device 10 is an information processing device (computer) and includes a processor 11, a main memory device 12, an auxiliary memory device 13, an input device 14, an output device 15, and a communication device 16.
[0032] The processor 11 is, for example, a device that performs arithmetic processing, and is a central processing unit (CPU), a micro processing unit (MPU), a graphics processing unit (GPU), an artificial intelligence (AI) chip, or the like.
[0033] The main memory device 12 is a device that stores programs and data, and is, for example, a ROM (Read Only Memory), SRAM (Static Random Access Memory), NVRAM (Non Volatile RAM), Mask ROM (Mask Read Only Memory), PROM (Programmable ROM), RAM (Random Access Memory), DRAM (Dynamic Random Access Memory), etc.
[0034] The auxiliary storage device 13 is a hard disk drive, a flash memory, a solid state drive (SSD), an optical storage device (CD (Compact Disc) or DVD (Digital Versatile Disc)), etc. The programs and data stored in the auxiliary storage device 13 are read into the main storage device 12 as needed.
[0035] The input device 14 is a user interface that accepts information from a user, and is, for example, a keyboard, a mouse, a card reader, a touch panel, or the like.
[0036] The output device 15 is a user interface that outputs various types of information (display output, audio output, print output, etc.), and is, for example, a display device (LCD (Liquid Crystal Display), graphics card, etc.) that visualizes various types of information, an audio output device (speaker), a printing device, etc.
[0037] The communication device 16 is a communication interface that communicates with other devices via communication means 20. The configuration of the communication means is not necessarily limited, and examples thereof include communication means conforming to various communication standards such as USB (Universal Serial Bus) and RS-232C, LAN (Local Area Network), WAN (Wide Area Network), the Internet, and dedicated lines. Examples of the communication interface include NIC (Network Interface Card), wireless communication module, USB (Universal Serial Interface) module, and serial communication module. The communication device 16 can also function as an input device that receives information from other devices that are communicatively connected. The communication device 16 can also function as an output device that transmits information to other devices that are communicatively connected.
[0038] The various functions of the image recognition device 10 are realized by the processor 11 reading and executing a program stored in the main memory device 12, or by the hardware (FPGA, ASIC, AI chip, etc.) that constitutes the image recognition device 10.
[0039] <Functional configuration of image recognition device> 4 is a block diagram showing the functional configuration of an image recognition device 10 according to this embodiment. The image recognition device 10 includes an image acquisition unit 1, an image conversion unit 2, a skeleton recognition unit (object recognition unit) 3, a reliability calculation unit 4, a recognition result selection unit 5, and a conversion parameter table storage unit 6.
[0040] The image acquisition unit 1 acquires time-series images including a person from an image capture device 30 or the like.
[0041] The image conversion unit 2 converts the values of the visual components of each frame image included in the time-series images acquired by the image acquisition unit 1 based on a conversion parameter table 61 stored in the conversion parameter table storage unit 6 described later, and generates converted time-series images (images in frame units) made up of converted frame images in which the values of the visual components have been converted. Note that the image conversion unit 2 generates changed time-series images in which the values of the visual components have been converted for each conversion parameter table 61.
[0042] The skeleton recognition unit (subject recognition unit) 3 recognizes the skeleton of the person in each image contained in each changed frame image that constitutes the converted time series image for each converted time series image in which the values of the visual components have been converted, and obtains time series skeleton information arranged in chronological order.
[0043] The reliability calculation unit 4 calculates the time-series reliability from the time-series skeletal information for each converted time-series image. Note that the reliability here means how reliable (uncertain) the image recognition result is, that is, the accuracy of the acquired skeletal information of the person in the image and the actual skeleton.
[0044] The recognition result selection unit 5 selects skeletal information for each frame image of a person in the time series image acquired by the image acquisition unit 1 from the time series skeletal information for each converted time series image based on the time series reliability calculated for each converted time series image.
[0045] <Operation flow of image recognition device 10> The image recognition process of the image recognition device 10 in this embodiment will be described in detail with reference to Fig. 5. Fig. 5 is a flow chart showing the operation of the image recognition device 10.
[0046] It is assumed that a person is imaged by the imaging device 30. On this assumption, first, the image acquisition unit 1 acquires time-series images from the imaging device 30 in step S1.
[0047] Next, steps S3 to S5 are repeated n times (S2). Here, the time-series images acquired by the image acquisition unit 1 are converted into n converted time-series images in which the values of the visual components are converted based on n different conversion parameter tables 61 in the conversion parameter table storage unit 6. Then, time-series skeleton information is acquired for each of the n converted time-series images, and time-series reliability is calculated. Note that in this embodiment, the following procedure is described as a loop, but these may also be processed in parallel.
[0048] Next, in step S3, the image conversion unit 2 converts the values of the visual components of the time-series images acquired by the image acquisition unit 1 based on the conversion parameter table 61 in the conversion parameter table storage unit 6, to generate converted time-series images. Note that the image conversion may use a classical rule-based conversion method, or may use a conversion using a learning model such as machine learning.
[0049] FIG. 6 shows an example of the conversion parameter table 61. One or more conversion parameter tables 61 are stored in the conversion parameter table storage unit 6. In the conversion parameter table 61, conversion targets that determine the content of image conversion are associated with conversion parameters. A visual component to be converted is set as the conversion target, and the conversion content of the visual component is set by the conversion parameters. For example, if the conversion target is "hue," the conversion content of the hue is set by the conversion parameters associated with the hue.
[0050] A conversion parameter consists of whether to perform an action and a setting value. If the whether to perform an action is set to "1," the conversion target of the image is converted according to the content of the setting value, and if the whether to perform an action is set to "0," the content of the setting value is ignored and the conversion target is not converted. The user can change whether to convert the conversion target by changing the value of whether to perform an action in the conversion parameter table 61. Also, all conversion targets with an whether to perform an action set to "1" in the conversion parameter table 61 are converted. For example, the conversion parameter table 61 indicates that hue and luminance with an whether to perform an action set to "1" are to be converted.
[0051] The setting values are set to numerical values that specify the visual component elements after conversion. In this embodiment, the setting values are set to values that are assumed to be in the HSV color space and the HSL color space and are normalized to values between 0 and 1. The user can change the conversion details of the visual component elements that are the conversion target by changing the setting values in the conversion parameter table 61. Since the HSV color space and the HSL color space are assumed, not only color images but also black and white images are the conversion target.
[0052] As shown in conversion parameter table 61, the setting value "0.1" for the conversion target "hue" is a normalized value (0 to 360°) that specifies the hue, and indicates orange, and indicates that the hue of the image is to be converted to orange. Also, the setting value "0.5" for the conversion target "brightness" is a normalized value (0 to 100%) that specifies the brightness, and indicates 50% brightness, and indicates that the image is to be converted to 50% brightness. Meanwhile, in conversion parameter table 61, although the implementation status is "0," and therefore the image is not a conversion target, the setting value "0.3" for "saturation" indicates that the image is to be converted to 30% saturation, and the setting value "0.6" for "lightness" indicates that the image is to be converted to 60% brightness.
[0053] If the execution / non-execution columns for all conversion targets are "0", the image conversion unit 2 does not perform image conversion, and the time-series images acquired by the image acquisition unit 1 are obtained as converted time-series images.
[0054] Next, in step S4, the skeleton recognition unit 3 recognizes the skeleton of the person from each converted frame image of the converted time-series image, extracts skeleton information indicating the skeleton of the person, and acquires the time-series skeleton information.
[0055] Specifically, the skeleton recognition unit 3 first recognizes the skeleton of a person from each converted frame image of the converted time-series image using skeleton definition information that defines in advance the position information of the human body parts to be recognized and the state of connection with other body parts.The skeleton recognition unit 3 then extracts the position information of the skeleton recognized in each converted frame image as skeleton information, and arranges the extracted skeleton information in chronological order to obtain time-series skeleton information.To obtain the skeleton information, an existing skeleton recognition technology using a learning model such as machine learning can be used.
[0056] FIG. 7 is a diagram showing an example of skeletal information recognized by the skeleton recognition unit 3. The skeletal information is recognized by identifying and sequentially connecting the following human body parts from an image: head P0, neck P1, left and right shoulders P2, P3, left and right elbows P4, P5, left and right wrists P6, P7, left and right hips P8, P9, left and right knees P10, P11, and left and right feet P12, P13. The skeletal information includes, for example, image coordinates as position information for each of the above-mentioned human body parts. The skeletal information may also include a score indicating the reliability of the image coordinates for each human body part. The score may take a value between 0 and 1, for example. The skeleton recognition unit 3 may also extract position information for small human body parts such as fingers and toes, and their connection status with other human body parts, as the skeleton.
[0057] Returning to FIG. 5, next, in step S5, the reliability calculation unit 4 calculates the reliability of each piece of skeletal information based on the amount of variation in the time-series skeletal information. For example, the reliability calculation unit 4 may use the difference in distance between different body parts at different times as the amount of variation based on the skeletal information, and calculate the reliability of each piece of skeletal information based on the amount of variation. More specifically, the reliability calculation unit 4 calculates the reliability of each piece of skeletal information from the time-series skeletal information in the following procedure.
[0058] First, the reliability calculation unit 4 calculates the distance between different human body parts from the skeleton information. For example, the X and Y coordinates of the left and right shoulders P2 and P3 shown in Fig. 7 are (x2, y2) and (x3, y3), respectively. In this case, the distance d between the human body parts calculated from the left shoulder P2 and the right shoulder P3 is calculated using Equation 1.
[0059]
number
[0060] The reliability calculation unit 4 calculates the distance between the human body parts at each time point in the time-series images, i.e., for each frame image constituting the time-series images, to obtain the distance between the human body parts along the time series. Fig. 8(A) shows an example of the distance between the human body parts along the time series calculated by the reliability calculation unit 4 in a graph with the distance on the vertical axis and time on the horizontal axis.
[0061] Next, the reliability calculation unit 4 calculates the amount of variation along the time series from the distance d between the human body parts at different times. Here, the distances between the human body parts at time t and time t-1 are assumed to be d(t) and d(t-1), respectively. In this case, for example, the amount of variation along the time series at time t, h(t), is calculated using, for example, Equation 2.
[0062]
number
[0063] The reliability calculation unit 4 calculates the amount of fluctuation along the time series for each time point to obtain the amount of fluctuation along the time series. Fig. 8(B) shows an example of the amount of fluctuation along the time series obtained by the reliability calculation unit 4 in a graph with the amount of fluctuation on the vertical axis and time on the horizontal axis.
[0064] Next, the reliability calculation unit 4 calculates the reliability S(t) of the skeletal information at time t by finding a negative exponential function of the amount of fluctuation h(t) along the time series using the following Equation 3.
number
[0065] The reliability S(t) takes a value between 0 and 1. The reliability calculation unit 4 calculates the reliability at each time point to obtain the time-series reliability. Figure 8(C) shows an example of the time-series reliability obtained by the reliability calculation unit 4 in a graph with the reliability on the vertical axis and the time on the horizontal axis.
[0066] As described above, in the present invention, the reliability calculation unit 4 calculates the absolute value of the difference between skeletal information at different points in time as the amount of variation along a time series using the distance between human body parts calculated from the image coordinates of the left and right shoulders based on the skeletal information, and calculates the reliability along a time series using a negative exponential function. However, this is not limited to this. For example, the reliability calculation unit 4 may use the image coordinates and score of each human body part as the skeletal information, or may use statistical information such as the average and variance of the image coordinates and scores of multiple human body parts. Furthermore, the reliability calculation unit 4 may use the variance of the skeletal information over a certain time range (predetermined period) as the amount of variation. Furthermore, the reliability calculation unit 4 may use a step function or a sigmoid function instead of a negative exponential function to calculate the reliability S(t) of the skeletal information.
[0067] Although the above description has been given of calculating the reliability of each piece of skeletal information from the amount of variation for each piece of time-series skeletal information, the present invention is not limited to this. For example, the reliability calculation unit 4 may calculate a reliability for each amount of variation along multiple time series, and calculate a weighted average of the calculated multiple reliabilities as the final reliability. For example, the reliability may be calculated for each amount of variation calculated from the distances between multiple different human body parts, and the weighted average of the reliabilities may be calculated as the final reliability.
[0068] Although the reliability is calculated from the amount of variation over time in the above description, the present invention is not limited to this. For example, the reliability calculation unit 4 may calculate the reliability from a score included in the skeletal information in addition to the reliability calculated from the amount of variation in the distance between human body parts using image coordinates. For example, when the score is a value between 0 and 1, the score value itself may be used as the reliability.
[0069] Returning to FIG. 5, when the above-described process from step S3 to step S5 has been repeated n times, the loop is terminated (step S6).
[0070] Next, in step S7, the recognition result selection unit 5 selects skeletal information of the person in the time-series image acquired by the image acquisition unit 1 based on the reliability calculated by the reliability calculation unit 4 from the n time-series skeletal information recognized in each transformed time-series image (i.e., each of the n transformed time-series images) in which the values of the visual components have been transformed.
[0071] One of the selection methods is that the recognition result selection unit 5 selects the skeleton information with the highest reliability calculated by the reliability calculation unit 4 at each time point from the n pieces of time-series skeleton information recognized in each converted time-series image as the skeleton information of the person in the image acquired by the image acquisition unit 1. In this selection method, as explained in FIG. A (t-2),F B (t-1),F C (t) is selected. Another selection method is that the recognition result selection unit 5 selects, from the n pieces of time-series skeletal information recognized in each converted time-series image, the skeletal information with the highest average reliability over a certain time range (predetermined period) including that time point as the skeletal information of the person in the image acquired by the image acquisition unit 1.
[0072] Furthermore, as another method, there is a method in which the recognition result selection unit 5 selects, as the skeletal information of a person in the time-series image acquired by the image acquisition unit 1, time-series skeletal information having skeletal information with the highest reliability calculated by the reliability calculation unit 4 or time-series skeletal information having skeletal information with the highest overall average value of the reliability calculated by the reliability calculation unit 4. Furthermore, the recognition result selection unit 5 selects, from the n pieces of skeleton information along the time series recognized in each converted time series image, skeleton information arranged based on the reliability at each time point, for example, F A (t-2),F B (t-2),F C Among the skeleton information arranged in the order of (t-2), any number of top-ranked pieces of skeleton information or pieces of skeleton information above any threshold may be selected. In this case, multiple pieces of skeleton information may be selected at each time point. For example, at time point t-2, F A (t-2) and F BAlternatively, two pieces of skeleton information may be selected at time t-1 and time t-2. Alternatively, different numbers of pieces of skeleton information may be selected at each time point, for example, one piece of skeleton information may be selected at time t-1 and two pieces of skeleton information may be selected at time t-2.
[0073] Furthermore, the arbitrary threshold value may be the reliability of skeletal information extracted from a transformed time-series image on which no image transformation has been performed in the image transformation unit 2. In this case, since the transformed time-series image on which no image transformation has been performed is a time-series image acquired by the image acquisition unit 1, the recognition result selection unit 5 can select skeletal information having a reliability equal to or higher than the reliability of the skeletal information extracted from the time-series image acquired by the image acquisition unit 1.
[0074] As described above, by selecting skeletal information using the reliability along the time series based on the amount of variation for each piece of time-series skeletal information, it is possible to eliminate skeletal information with a large amount of variation along the time series. For example, when the accuracy of skeletal recognition of a transformed time-series image is low, the variation in the time-series skeletal information extracted from the transformed time-series image is larger than the variation in the time-series skeletal information that actually occurs in the transformed time-series image, and the amount of variation along the time series increases. By eliminating such skeletal information of a transformed time-series image with low accuracy of skeletal recognition using the reliability for each time series based on the amount of variation for each piece of time-series skeletal information, it is possible to select more accurate skeletal information that is closer to the variation that actually occurs in the time-series image.
[0075] 9 shows an example of an operation screen in the recognition result selection unit 5. The operation screen 1000 shown in the figure is displayed on a display of a terminal (not shown) that is capable of communicating with the output device 15 of the image recognition device 10 via the communication device 16 in step S7 of FIG. The operation screen 1000 displays the skeleton information, reliability, and selection result selected by the recognition result selection unit 5 in step S7. The operation screen 1000 displays the converted frame image at time t and the skeleton information extracted from that converted frame image in the skeleton information column, and displays the reliability value for time t calculated in step S5 of FIG. 5 in the reliability column. The selection result for time t selected by the recognition result selection unit 5 in step S7 is displayed in the selection result column. If the recognition result selection unit 5 selected the result, "1" is displayed, and if the result was not selected, "0" is displayed.
[0076] In addition, the column of time-series reliability displays reliability along a time series for a certain time range (predetermined period) including time point t. On the operation screen 1000 of FIG. 9, the column of reliability along a time series displays "1" for the selection result of skeleton information at time point t having reliability equal to or greater than the threshold indicated by a dotted line. The user can change the result selected by the recognition result selection unit 5 in step S7 by changing the value in the selection result column on the operation screen 1000. For example, by changing "0" to "1," the user can select skeleton information that was not selected by the recognition result selection unit 5.
[0077] As described above in detail, the image recognition system 100 of this embodiment acquires time-series images including a person, generates converted time-series images of the person image in units of multiple frame images by converting the values of the visual components of the acquired time-series images, acquires skeletal information for each generated converted time-series image by skeletal recognition, calculates reliability along the time series from the skeletal information along the time series, and selects skeletal information of the person in the acquired time-series images from the skeletal information recognized from the multiple converted time-series images based on the calculated reliability.
[0078] When the accuracy of skeleton extraction decreases due to the color, brightness, and luminance of the clothing of a person in an image, i.e., the visual components, the image recognition device 10 selects skeletal information based on the reliability of time-series skeletal information recognized from converted time-series images (images in frame units) in which the values of the visual components have been converted by image transformation, thereby selecting skeletal information obtained from converted time-series images that are suitable for skeleton recognition, thereby improving the accuracy of skeleton recognition. Therefore, it becomes possible to recognize skeletons regardless of the color, brightness, luminance, etc. of the clothing of a person in an image.
[0079] It goes without saying that the present invention is not limited to the above-described embodiments and can be modified in various ways without departing from the spirit of the present invention. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those having all of the described configurations. Furthermore, it is possible to add, delete, or replace part of the configuration of the above-described embodiments with other configurations.
[0080] Furthermore, the above-described configurations, functional units, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a memory, a hard disk, a recording device such as an SSD (Solid State Drive), an IC card, an SD card, a DVD, or other recording media.
[0081] In addition, in the above figures, the control lines and information lines shown are those that are considered necessary for explanation, and do not necessarily show all the control lines and information lines that are actually implemented. For example, it can be considered that almost all components are actually connected to each other.
[0082] Furthermore, the above-described layout of the various functional units, processing units, and databases of the image recognition device 10 is merely an example. The layout of the various functional units, processing units, and databases can be changed to an optimal layout in terms of the performance, processing efficiency, communication efficiency, etc. of the hardware and software provided in these devices.
[0083] Furthermore, the configuration (schema, etc.) of the database that stores the various types of data described above can be flexibly changed from the viewpoint of efficient use of resources, improved processing efficiency, improved access efficiency, improved search efficiency, and the like. [Explanation of symbols]
[0084] 1 Image acquisition unit 2 Image conversion section 3 Skeleton Recognition Unit 4. Reliability calculation section 5 Recognition result selection section 6. Conversion parameter table storage section 10 Image recognition device 11 processors 12 Main storage 13 Auxiliary storage device 14 Input Devices 15 Output Devices 16. Communications equipment 20. Means of communication 30 Imaging device 100 Image Recognition System 1000 operation screens
Claims
1. An image recognition device that recognizes a subject from time-series images, an image acquisition unit that acquires the time-series images including the subject; an image conversion unit that converts values of visual components into predetermined values for at least a part of an image region of the subject included in the time-series images, and generates converted time-series images in which the values of the visual components have been converted; an object recognition unit that extracts the object from each of the converted time-series images and recognizes the object in time series; a reliability calculation unit that calculates the reliability of the recognition result of each of the converted time-series images based on the amount of change of the subject recognized in time series from each of the converted time-series images; a recognition result selection unit that selects the recognition result based on the reliability; An image recognition device comprising:
2. At least two parts are set for the subject, The image recognition device according to claim 1, characterized in that the reliability calculation unit calculates the reliability of the recognition result of each of the transformed time-series images based on the amount of variation between parts of the subject recognized in time series from each of the transformed time-series images.
3. The image recognition device according to claim 1 , wherein the visual component is at least one of hue, lightness, saturation, and luminance.
4. 2. The image recognition device according to claim 1, wherein each of the transformed time-series images includes an image in which the values of the visual components are not transformed.
5. The image recognition device according to claim 2 , wherein the amount of variation between the parts is an amount of variation in distance between the parts.
6. The image recognition device according to claim 2, characterized in that the reliability calculation unit calculates the reliability of the recognition result of each of the transformed time-series images based on the amount of variation between the parts of the subject recognized in time series from each of the transformed time-series images.
7. 2. The image recognition device according to claim 1, wherein the subject is a whole or a part of a person, or a whole or a part of a robot.
8. 2. The image recognition device according to claim 1, wherein the result selected by the recognition result selection unit can be changed by a user.
9. An image recognition method for an image recognition device that recognizes a subject from time-series images, comprising: acquiring the time-series images including the subject; converting values of visual components into predetermined values for at least a part of image regions of the subject included in the time-series images, and generating converted time-series images in which the values of the visual components have been converted; extracting the object from each of the transformed time-series images and recognizing the object in time series; calculating a reliability of the recognition result of each of the converted time-series images based on a variation amount of the object recognized in time series from each of the converted time-series images; selecting the recognition result based on the confidence; An image recognition method comprising:
10. An image recognition device that recognizes subjects from time-series images, an image acquisition unit that acquires the time-series images including the subject; an image conversion unit that converts values of visual components into predetermined values for at least a part of an image region of the subject included in the time-series images, and generates converted time-series images in which the values of the visual components have been converted; an object recognition unit that extracts the object from each of the converted time-series images and recognizes the object in time series; a reliability calculation unit that calculates the reliability of the recognition result of each of the converted time-series images based on the amount of change of the subject recognized in time series from each of the converted time-series images; a recognition result selection unit that selects the recognition result based on the reliability; A program that functions as a
Citation Information
Patent Citations
Image recognition device, solid-state imaging device, and image recognition method
JP2020187409A