Information processing device, information processing method, and program
The information processing device addresses the limitations of existing behavior recognition by integrating personal and environmental data, enhancing accuracy and robustness in human behavior estimation.
Patent Information
- Application Number
- JP2023567318
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2041-12-14
AI Technical Summary
Existing behavior recognition technologies fail to accurately estimate human behavior due to reliance on limited personal information without considering environmental context and lack of detailed positional information.
An information processing device that extracts, aggregates, and integrates instance information for each instance in a video, using a combination of extraction, aggregation, and integration units to generate a recognition result, considering both personal and environmental data.
Enhances recognition accuracy by reducing information loss and improving the robustness of behavior recognition, enabling more precise identification of human actions and their context.
Smart Images

Figure 0007722469000001 
Figure 0007722469000002 
Figure 0007722469000003
Abstract
Description
[Technical Field]
[0001] The present invention , love The present invention relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] In recent years, behavior recognition technology that recognizes human behavior has been put into practical use and is being applied in various fields. For example, behavior recognition technology is being used in various fields to reduce human workload.
[0003] For example, in the field of nursing care, an image processing device has been proposed that estimates the behavior of a moving object in a detection area based on the posture of the moving object (see, for example, Patent Document 1).
[0004] Furthermore, a technique has been proposed in which the relationship between a person detected by a rectangle and an object is expressed by a gaze mechanism, and the features required for activity label prediction are extracted (see, for example, Non-Patent Document 1). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 2021-65617 [Non-patent literature]
[0006] [Non-Patent Document 1] Attend and Interact: Higher-Order Object Interactions for Video Understanding, Ma et. al., CVPR, 2018 Summary of the Invention [Problem to be solved by the invention]
[0007] However, for example, the image processing device described in Patent Document 1 has the problem that it cannot accurately estimate human behavior because it estimates human behavior based only on information about the person without considering the environment other than the person and the small amount of information.The technology described in Non-Patent Document 1 does not identify the object, and recognizes people and objects as rectangular information using only image features without considering detailed position information, so it has the problem that it cannot accurately recognize human behavior.
[0008] One aspect of the present invention has been made in consideration of the above-mentioned problems, and aims to provide a recognition processing technique that is robust against information loss. [Means for solving the problem]
[0009] An information processing device according to one aspect of the present invention includes an extraction means for extracting a plurality of pieces of instance information for each of one or more instances included in an input video; an aggregation means for aggregating the plurality of pieces of instance information for each instance; an integration means for generating instance integrated information by integrating the plurality of pieces of instance information aggregated by the aggregation means for each instance; and a recognition means for generating a recognition result for at least one of the one or more instances by referring to the instance integrated information generated by the integration means.
[0010] An information processing method according to one aspect of the present invention includes extracting a plurality of pieces of instance information for each of one or more instances included in an input video; aggregating the plurality of pieces of instance information for each instance; generating instance integrated information by integrating the instance information aggregated by the aggregating means for each instance; and generating a recognition result for at least one of the one or more instances by referring to the instance integrated information generated by the integrating means. Ruko and,
[0011] A program according to one aspect of the present invention causes a computer to execute the following: extraction means for extracting multiple pieces of instance information for each of one or more instances included in an input video; aggregation means for aggregating the multiple pieces of instance information for each instance; integration means for generating instance integration information by integrating the aggregated instance information for each instance; and recognition means for generating a recognition result for at least one of the one or more instances by referring to the generated instance integration information. [Effects of the Invention]
[0012] According to one aspect of the present invention, it is possible to provide a recognition processing technique that is robust against information loss. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a block diagram showing an example of the configuration of an information processing device according to a first exemplary embodiment of the present invention. [Figure 2] 1 is a flowchart showing the flow of an information processing method according to the first exemplary embodiment of the present invention. [Figure 3] FIG. 10 is a block diagram showing an example of the configuration of an information processing device according to a second exemplary embodiment of the present invention. [Figure 4] 10A and 10B are diagrams illustrating an example of extraction processing executed by an extraction unit according to the second exemplary embodiment of the present invention. [Figure 5] 10A and 10B are diagrams illustrating an example of aggregation processing executed by an aggregation unit according to the second exemplary embodiment of the present invention. [Figure 6] 10A and 10B are diagrams illustrating an example of aggregation processing executed by an aggregation unit according to the second exemplary embodiment of the present invention. [Figure 7] 10A to 10C are diagrams illustrating an example of integration processing executed by an integration unit according to the second exemplary embodiment of the present invention. [Figure 8] 10A to 10C are diagrams illustrating an example of integration processing executed by an integration unit according to the second exemplary embodiment of the present invention. [Figure 9] 10A to 10C are diagrams illustrating an example of integration processing executed by an integration unit according to the second exemplary embodiment of the present invention. [Figure 10] FIG. 10 is a diagram showing an example of a recognition result output by an output unit according to the second exemplary embodiment of the present invention. [Figure 11] FIG. 10 is a block diagram showing an example of the configuration of an information processing device according to a third exemplary embodiment of the present invention. [Figure 12] 10 is a flowchart showing the flow of an information processing method according to a third exemplary embodiment of the present invention. [Figure 13] FIG. 2 is a block diagram showing an example of a hardware configuration of an apparatus in each exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0014] Exemplary Embodiment 1 A first exemplary embodiment of the present invention will be described in detail with reference to the drawings. This exemplary embodiment is a basic form of the exemplary embodiments described below.
[0015] <Configuration of information processing device 1> The configuration of an information processing device 1 according to this exemplary embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the information processing device 1.
[0016] As shown in FIG. 1, the information processing device 1 includes an extraction unit 11, an aggregation unit 12, an integration unit 13, and a recognition unit 14. The extraction unit 11 is configured to realize the extraction means in this exemplary embodiment. The aggregation unit 12 is configured to realize the aggregation means in this exemplary embodiment. The integration unit 13 is configured to realize the integration means in this exemplary embodiment. The recognition unit 14 is configured to realize the recognition means in this exemplary embodiment.
[0017] The extraction unit 11 extracts a plurality of pieces of instance information for each of one or a plurality of instances included in the input video.
[0018] Here, one or more instances refer to objects included in the video, such as a person or an object other than a person.
[0019] The instance information is, for example, information expressed as a character string or a number string. The instance-related information is, for example, information necessary to identify an instance, and is information that characterizes the instance.
[0020] The extraction unit 11 may extract a plurality of pieces of instance information for each of one or a plurality of instances included in each frame of a plurality of image frames included in the input video.
[0021] The extraction unit 11 may have a tracking function or may use an existing tracking engine. In this case, the extraction unit 11 may extract multiple pieces of instance information in an integrated manner from two or more frames among multiple image frames included in the input video.
[0022] The aggregation unit 12 aggregates a plurality of pieces of instance information for each instance.
[0023] Aggregation for each instance means, for example, associating an instance with instance information based on that instance. Here, aggregation refers to associating multiple pieces of instance information with a given instance when multiple pieces of instance information exist for that instance. In other words, aggregation means generating data in which instance information is associated for each instance.
[0024] The integration unit 13 generates instance integrated information by integrating the plurality of pieces of instance information aggregated by the aggregation unit 12 for each instance.
[0025] The instance integrated information is generated, for example, by at least one of concatenation and summation of the instance information aggregated by the aggregation means for each instance. Note that concatenation means, for example, arranging two or more pieces of data having the same or different dimensions to create one piece of data having a dimension greater than the data before concatenation. Summation means, for example, adding two or more pieces of data having the same dimension without changing the dimension to create one piece of data.
[0026] The recognition unit 14 refers to the instance integration information generated by the integration unit 13 and generates a recognition result for at least one of the one or more instances.
[0027] The recognition result is generated for each instance by, for example, referring to the instance integration information of each instance. The recognition result may be, for example, text data consisting of words and sentences, graph data, or image data.
[0028] <Effects of information processing device 1> As described above, the information processing device 1 according to this exemplary embodiment employs a configuration in which, for each of one or more instances, a recognition result for the instance is generated using a plurality of pieces of instance information. Therefore, the information processing device 1 according to this exemplary embodiment can provide a recognition processing technology that is robust against information loss in recognition processing for recognizing information about targets such as people and objects, and events relating to people and objects. This has the effect of enabling more accurate recognition of instance behavior.
[0029] <Flow of information processing method by information processing device 1> The flow of an information processing method executed by the information processing device 1 according to this exemplary embodiment will be described with reference to Fig. 2. Fig. 2 is a flowchart showing the flow of the information processing method. As shown in the figure, the information processing includes steps S11 to S14.
[0030] (Step S11) In step S11, the extraction unit 11 extracts a plurality of pieces of instance information for each of one or a plurality of instances included in the input video.
[0031] (Step S12) In step 12, the aggregation unit 12 aggregates the multiple pieces of instance information for each instance.
[0032] (Step S13) In step 13, the integration unit 13 generates integrated instance information by integrating the aggregated pieces of instance information for each instance.
[0033] (Step S14) In step 14, the recognition unit 14 references the generated instance integration information and generates a recognition result for at least one of the one or more instances.
[0034] <Effects of information processing methods> The information processing method according to this exemplary embodiment employs a configuration in which, for each of one or more instances, a recognition result for the instance is generated using a plurality of pieces of instance information. Therefore, the information processing method according to this exemplary embodiment can provide a recognition processing technique that is robust against information loss in recognition processing for recognizing information about targets such as people and objects.
[0035] Exemplary Embodiment 2 A second exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the first exemplary embodiment are given the same reference numerals, and their description will be omitted as appropriate.
[0036] <Configuration of information processing device 1A> The configuration of an information processing device 1A according to this exemplary embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram showing an example configuration of the information processing device 1A. As shown in Fig. 3, the information processing device 1A includes a storage unit 20A, a communication unit 21, an input unit 22, a display unit 23, and a control unit 10A.
[0037] The storage unit 20A is configured, for example, with a semiconductor memory device or the like, and stores data. In this example, the storage unit 20A stores inference video data VDP, model parameters MP, and recognition results RR. Here, the model parameters are weighting coefficients obtained by machine learning, which will be described later. The model parameters MP include model parameters used in the integration process of the integration unit 13 and model parameters used in the recognition process of the recognition unit 14.
[0038] The communication unit 21 is an interface for connecting the information processing device 1A to a network. The specific configuration of the network does not limit the present exemplary embodiment, but as an example, a wireless LAN (Local Area Network), a wired LAN, a WAN (Wide Area Network), a public line network, a mobile data communication network, or a combination of these networks can be used.
[0039] The input unit 22 accepts various inputs to the information processing device 1A. The specific configuration of the input unit 22 is not limited to this exemplary embodiment, but as an example, the input unit 22 may be configured to include input devices such as a keyboard and a touchpad. The input unit 22 may also be configured to include a data scanner that reads data via electromagnetic waves such as infrared rays or radio waves, and a sensor that senses the environmental state.
[0040] The display unit 23 displays the recognition result output from the control unit 10 A. The display unit 23 may be realized by a display device such as a liquid crystal display device capable of black and white or color display, or an organic EL (Electroluminescence) display device.
[0041] The control unit 10A has the same functions as those of the information processing device 1 described in exemplary embodiment 1. The control unit 10A includes an extraction unit 11, an aggregation unit 12, an integration unit 13, a recognition unit 14, and an output unit 15.
[0042] The extraction unit 11 extracts multiple pieces of instance information for each of one or multiple instances included in the input video. The extraction unit 11 may include a person instance information extraction unit that extracts instance information related to a person. While FIG. 3 exemplarily shows a configuration including two person instance information extraction units (person instance information extraction unit 11-1 and person instance information extraction unit 11-2), the configuration is not limited to this. The extraction unit 11 may include three or more person instance information extraction units. Each of the person instance information extraction units may extract one type of instance information.
[0043] Examples of instance information related to a person include rectangle information that is a rectangle surrounding the person, pose information that indicates the posture of the person, segmentation information that indicates the surrounding environment of the person, etc. Furthermore, when multiple pieces of rectangle information are extracted from one image frame of the target video data, identification information for identifying each piece of rectangle information may be assigned to each person instance as instance information.
[0044] Specifically, the rectangle information may include the position and size of a rectangular area in the image. The position and size of a rectangular area in the image may be represented by the x-coordinate value, the y-coordinate value, or values obtained by normalizing the x-coordinate and the y-coordinate by the image size of an image element (pixel) in the image.
[0045] The pose information may specifically include information about the skeleton and joints of a person. For example, the pose information may be information in which characteristic points of the skeleton and joints of a person are represented by x-coordinate values and y-coordinate values of image elements in an image. The pose information may also include a circumscribing rectangle that surrounds the characteristic points of the skeleton and joints.
[0046] The segmentation information may be, for example, the area of the person included in the rectangle information, information on the part other than the person included in the rectangle information, or information on the part other than the person included in the circumscribing rectangle that is pose information.
[0047] The plurality of pieces of instance information may be extracted using different engines depending on the type of the instance information, or may be extracted using a single engine.
[0048] If the extraction unit 11 has a tracking function, the results of tracking at least one of rectangles, poses, and segmentations between multiple image frames included in the video may be extracted as rectangle information, pose information, and segmentation information, respectively. Also, movement information indicating a person's movement detected based on at least one of rectangle information and pose information in multiple image frames may be extracted as person instance information. The movement information may be extracted by further referring to segmentation information.
[0049] The extraction unit 11 may also include a general instance information extraction unit that extracts general instance information related to instances other than people. The instances other than people may be objects. While FIG. 3 illustrates a configuration including two general instance information extraction units (general instance information extraction unit 11-3 and general instance information extraction unit 11-4) as an example, the configuration is not limited to this. The extraction unit 11 may also include three or more general instance information extraction units. Each of the general instance information extraction units may extract one type of instance information.
[0050] Examples of general instance information include rectangle information that is a rectangle surrounding an object, feature information that constitutes the object, and segmentation information that indicates the surrounding environment of the object. The feature information that constitutes the object may be, for example, points or lines that indicate the edges of the object. Furthermore, when multiple pieces of rectangle information are extracted from one image frame of the target video data, identification information for identifying each piece of rectangle information may be assigned to each general instance as general instance information.
[0051] FIG. 4 is a schematic diagram illustrating an example of extraction processing. As an example, FIG. 4 shows an image captured of a construction site, which includes a person and a rolling machine operated by the person. The image also includes a building and the ground around the person and the rolling machine. Here, rectangle information r1 and pose information p1 are extracted as person instance information, and rectangle information r2 and pose information p2 are extracted as general instance information. The building and the ground are also extracted as segmentation information s1 and s2, respectively.
[0052] The aggregating unit 12 aggregates multiple pieces of instance information for each instance. Here, aggregation refers to associating instance information with an instance. Specifically, the aggregating unit 12 associates different types of instance information, such as the above-mentioned rectangle information, pose information, motion information, and segmentation information, with one instance. The aggregating unit 12 may aggregate, for one instance, each piece of instance information extracted from multiple image frames captured at different times.
[0053] The aggregation unit 12 may aggregate, for example, a plurality of rectangle information, a plurality of pose information, a plurality of segmentation information, etc. extracted from a plurality of image frames captured at different times as instance information for each instance.
[0054] The aggregation unit 12 may aggregate multiple rectangle information (instance information) extracted from multiple image frames captured at different times into the same instance, for example, by referring to the size and position of the rectangle contained in the rectangle information.
[0055] In addition, the aggregation unit 12 may aggregate multiple pieces of pose information (instance information) extracted from multiple image frames captured at different times into the same instance, for example, by referring to the positions of the skeleton and joints contained in the pose information.
[0056] Furthermore, segmentation information indicating the surrounding environment may not change significantly even between multiple image frames captured at different times. Therefore, the aggregation unit 12 may aggregate multiple pieces of instance information into an instance by referring to the relationship between the position of the segmentation included in the segmentation information in the image, the position of the rectangle included in the rectangle information, and the position of the skeleton included in the pose information. As an example, the aggregation unit 12 may detect the distance between the segmentation and the rectangle and skeleton, and aggregate, into the same instance, pieces of instance information including a rectangle and skeleton whose distance is within a predetermined range between multiple image frames captured at different times.
[0057] A specific example of the aggregation process performed by the aggregation unit 12 will be described. Fig. 5 is a diagram illustrating an example of the aggregation process performed by the aggregation unit 12. Fig. 5 shows frames f(t) and f(t1), which are image frames captured at the same shooting time t. As an example, the extraction unit 11, specifically the person instance information extraction unit, extracts person W as rectangle information 1101 and person X as rectangle information 1102. As an example, the extraction unit 11, specifically the person instance information extraction unit 11-2, extracts person W1 as pose information 1111, person X1 as pose information 1112, and person Y1 as pose information 1113.
[0058] The aggregating unit 12 may also output data indicating the result of associating the instances with the instance information. Data D1 in FIG. 5 illustrates an example of a data structure indicating the result of the aggregation process executed by the aggregating unit 12. In the case of FIG. 5, the aggregating unit 12 may, for example, identify that person W and person W1 are the same instance based on the positions of the rectangle information and the pose information, and aggregate the instance information. That is, the aggregating unit 12 may associate rectangle information 1101 and pose information 1111 with the same instance. Specifically, the aggregating unit 12 may identify as the same instance an instance in which the rectangle of the rectangle information and the circumscribing rectangle of the pose information overlap greatly. Furthermore, although rectangle information for person Y was not extracted, the aggregating unit 12 may, for example, identify person Y and person Y1 as the same instance by a process of elimination, and aggregate the instance information. That is, the aggregating unit 12 may associate pose information 1113 with person Y and person Y1, who are the same instance.
[0059] Between multiple image frames captured at different times, the aggregation unit 12 may aggregate the instance information by referring to the trajectory of each piece of instance information between each frame. Specifically, the aggregation unit 12 may compare the trajectories of different instances between each frame and associate trajectories with a large degree of overlap with the same instance.
[0060] Fig. 6 is a diagram illustrating an example of aggregation processing executed by the aggregation unit 12. Fig. 6 shows frames f(t) and f(t1), which are image frames at photographing time t, and frames f(t+1) and f(t1+1), which are image frames at photographing time t+1. Frames f(t) and f(t1) in Fig. 6 are the same as frames f(t) and f(t1) described in Fig. 5.
[0061] As an example, the extraction unit 11 extracts person P as rectangle information 1104, person Q as rectangle information 1105, and person R as rectangle information 1106 in frame f(t1+1).
[0062] Moreover, as an example, the extraction unit 11 extracts person P1 as pose information 1114, person Q1 as pose information 1115, and person R1 as pose information 1116 in frame f(t1+1).
[0063] In FIG. 6, the aggregation unit 12 may calculate the trajectory of the rectangle for each instance using the rectangle information included in frame f(t) and the rectangle information included in frame f(t+1), for example, by using values calculated from the x-coordinate values and y-coordinate values of the pixels of each rectangle.
[0064] Furthermore, the aggregation unit 12 may calculate the pose trajectory for each instance by using the pose information included in frame f(t) and the pose information included in frame f(t+1), for example, by using values calculated from the x-coordinate values and y-coordinate values of the pixels of each joint point or the circumscribed rectangle of the joint point.
[0065] 6, graph G1 is a graph showing trajectories of rectangles and poses. For example, trajectory L4 is a trajectory obtained from the positions of rectangles included in rectangle information 1101, rectangle information 1104, and rectangle information of a frame at time t+2 (not shown). Also, trajectory L1 is a trajectory obtained from the positions of poses included in pose information 1111, pose information 1114, and pose information of a frame at time t+2 (not shown). Aggregation unit 12 may treat trajectories obtained between multiple frames captured at different times as one piece of instance information.
[0066] The aggregation unit 12 may associate rectangle information and pose information with one instance based on the degree of similarity between the shape of the rectangle trajectory and the shape of the pose trajectory. For example, trajectory L1 and trajectory L4 may be aggregated as instance information belonging to the same instance. In this way, by using a trajectory determined from the instance information of multiple frames in the aggregation process, the aggregation unit 12 can aggregate instance information even if there is missing information, such as rectangle information not being extracted in frame f(t).
[0067] The aggregating unit 12 may assign attribute information to each of the multiple pieces of instance information. Attribute information is information that represents the attributes of an instance, and examples include a person's name, an object's name, and a model number. The attribute information may be any information that can identify the instance, and may be a predetermined management number or the like. Furthermore, when there are multiple instances of the same type, a number may be added after the object's name, and different attribute information may be added so that the instances of the same type can be distinguished from each other.
[0068] The integrating unit 13 generates integrated instance information by integrating, for each instance, the multiple pieces of instance information aggregated by the aggregating unit 12. The integrating unit 13 also includes one or more transformation layers 130 that apply transformation processing to each piece of instance information, and one or more integration layers 131 that integrate the instance information after the transformation processing. The transformation layer 130 may include, for example, a multilayer perceptron, or may include two or more types of multilayer perceptrons. For example, different types of multilayer perceptrons may be applied depending on the type of input instance information.
[0069] FIG. 7 is a diagram modeling the integration process executed by the integration unit 13. The model shown in FIG. 7 includes a transformation layer 130 and an integration layer 131. In FIG. 7, as an example, the transformation layer 130 includes a first transformation layer 1301 to which instance information E1 is input and a second transformation layer 1302 to which instance information F1 is input. Here, the first transformation layer 1301 and the second transformation layer 1302 may each be a different multilayer perceptron. Furthermore, the instance information to which the transformation process has been applied in the transformation layer 130 is integrated in the integration layer 131 and output as a single piece of instance information G1 (instance integrated information, described later). Specifically, each piece of instance information expanded into a one-dimensional tensor may be input to the transformation layer 130, and the tensor may be converted to have the same dimension between each piece of information in the transformation layer.
[0070] As described in the first exemplary embodiment, the integration layer 131 may integrate two pieces of instance information by concatenating two pieces of instance information or by adding up two pieces of instance information. The concatenated instance information is a single piece of data with a larger dimension than the data before concatenation, in which two or more pieces of data with the same dimension are arranged, as in the instance information G1 shown in Fig. 7. The added-up instance information is a single piece of data obtained by adding two or more pieces of data with the same dimension without changing the dimension, as in the instance information G2 shown in Fig. 7.
[0071] The integration unit 13 assigns importance to the instance information after conversion processing by one or more conversion layers 130, and the integration layer 131 integrates the instance information using the importance.
[0072] The importance may be a weight to be multiplied by the instance information. That is, the integrating unit 13 may weight the instance information after the conversion process and integrate the weighted instance information.
[0073] Fig. 8 is a diagram modeling the integration process executed by the integration unit 13. The integration unit 13 shown in Fig. 8 includes a conversion layer 130 and an integration layer 131, similar to the integration unit 13 shown in Fig. 7, but differs in that importance is assigned to instance information after the conversion process, and the instance information is integrated using the importance. The integration unit 13 may also include a pooling layer 132.
[0074] In FIG. 8 , for example, the instance information E2 and F2 after the conversion process are input to the pooling layer 132, where global average pooling is applied. Then, the instance information is input to the conversion layer 130, where the importance w1 and w2 of the instance information E2 and F2, respectively, are output as numerical values. The importance may be output by applying a sigmoid function to the information after the conversion process by the conversion layer 130. The importance may be a numerical value between 0 and 1. As an example, in FIG. 8 , the importance w1 is output as 0.4, and the importance w2 is output as 0.6. Furthermore, the output importance w1 and w2 are multiplied by the instance information E2 and F2 after the conversion process by the conversion layer 130, respectively, to assign importance to each piece of instance information. The instance information to which importance has been assigned is integrated by the integration layer 131 and output as instance information G1.
[0075] The integration unit 13 may include one or more conversion layers, each of which applies a conversion process to each instance information in series, and the integration layer may include an integration layer that assigns importance to the instance information after the conversion process in the conversion layer and integrates the instance information using the importance.
[0076] Also, while FIG. 8 shows an example in which the integration unit 13 includes one conversion layer 130, the integration unit 13 may include two or more conversion layers. The more times multiple pieces of instance information are converted, the smaller the gap between the pieces of information becomes. In other words, the more conversion layers are input, the higher the similarity between the pieces of information becomes. For instance information with small gaps between pieces of information, it may be preferable to add them up in the integration layer 131. Conversely, for instance information with large gaps between pieces of information, i.e., instance information with low similarity between pieces of information, it may be preferable to connect them in the integration layer 131. For this reason, the integration layer 131 may determine whether to connect or add up the instance information after the conversion process depending on the number of conversion layers.
[0077] Furthermore, when the integrating unit 13 includes multiple transformation layers, the integrating unit 13 may further include a gaze block. As an example, the gaze block calculates a weighting coefficient from input instance information as an index indicating whether or not the instance information should be gazed upon.
[0078] The weighting coefficient may represent, for example, the mutual similarity between the input multiple pieces of instance information. The weighting coefficient may be set to a real value between 0 and 1. The weighting coefficient may be set, for example, according to the level of recognition accuracy when the input multiple pieces of instance information are integrated. Specifically, the weighting coefficient may be set to a value close to 1 if integrating the input multiple pieces of instance information will increase the recognition accuracy, or may be set to a value close to 0 if integrating the input multiple pieces of instance information will decrease the recognition accuracy. In other words, the weighting coefficient may be set to a value closer to 1 as the recognition accuracy increases, and a value closer to 0 as the recognition accuracy decreases.
[0079] Depending on the similarity between the pieces of information, it may be better to integrate them at a shallower layer or at a deeper layer. Therefore, by using attention blocks, it is possible to perform appropriate integration processing according to the similarity between multiple pieces of instance information.
[0080] 9 is a diagram modeling the integration process executed by the integration unit 13. The integration unit 13 includes, as an example, a plurality of conversion layers 130, 130A and attention blocks 133, 134. Instance information E1 and instance information F1 input to the integration unit 13 are converted by a first conversion layer 1301 and a second conversion layer 1302 of the conversion layer 130, respectively, and post-conversion instance information E2 and post-conversion instance information F2 are output. Here, the post-conversion instance information E2 and F2 are input to the attention block 133, and weighting coefficients are assigned based on the similarity of the information between them.
[0081] Note that the instance information to which a weighting factor has been assigned in the attention block 133 may be input to an integration layer (not shown) in accordance with the weighting factor, rather than being input to a subsequent conversion layer (e.g., conversion layer 130A). Also, the instance information E3 and F3 after conversion processing performed in conversion layer 130A may be input to attention block 134, and a weighting factor based on the similarity of the information may be assigned to each piece, similar to attention block 133. That is, by including attention block 133 in the integration unit 13, it may be possible to automatically select, in a plurality of conversion layers, which conversion layer's conversion processing to integrate the instance information after.
[0082] There is no limitation on the number of gaze blocks provided in the integration unit 13. The integration unit 13 may have the same number of gaze blocks as the number of transformation layers in the thickness direction.
[0083] The recognition unit 14 generates a recognition result related to the human behavior from one or more instances. The recognition unit 14 generates the recognition result related to the human behavior by referring to the integrated information generated by the integration unit 13. The recognition unit 14 executes the recognition process using the model parameters MP stored in the memory unit 20A. An existing behavior recognition engine may be used for the recognition unit 14. The recognition unit 14 may also generate the recognition result related to the human behavior by using both the instance integrated information related to the person and the instance integrated information related to the object.
[0084] The recognition unit 14 may, for example, refer to the integrated information and generate, as the recognition result, information in which scores are assigned to multiple actions that are estimated to be performed by each instance (person). As an example, the recognition unit 14 may generate, as the recognition result, information in which probabilities are assigned to predetermined actions, such as "(1) 70% probability that the worker A is compacting the ground with a roller, (2) 20% probability that the worker A is repairing the roller, (3) 10% probability that the worker A is carrying the roller."
[0085] The recognition unit 14 applies different identification processes to the instance integration information related to a person and the instance integration information related to an object among one or more instances. The recognition unit 14 may use different model parameters or different behavior recognition engines for the instance integration information related to a person and the instance integration information related to an object.
[0086] The output unit 15 outputs the recognition result generated by the recognition unit 14. The output unit 15 may output the recognition result generated by the recognition unit 14 as is, or may output a part of the recognition result. For example, when the recognition unit 14 generates information in which scores are assigned to a plurality of estimated behaviors as the recognition result, the output unit 15 may output only the behavior with the highest score.
[0087] For example, if the recognition unit 14 generates the recognition result as described above regarding the behavior of worker A as "(1) 70% probability that he is working to compact the ground with a roller, (2) 20% probability that he is repairing the roller, (3) 10% probability that he is carrying the roller," the output unit 15 may output the recognition result as "Worker A is working to compact the ground with a roller."
[0088] FIG. 10 is a diagram showing an example of the recognition result output by the output unit 15. In FIG. 10, the recognition result is shown as a table, for example. In FIG. 10, the actions of each of the three people, which are instances, are shown in chronological order. In addition, in the recognition result in FIG. 10, the actions of each of the three people indicate a relationship with an object. According to the recognition result shown in FIG. 10, for example, a manager who manages workers can accurately know the work status of each worker.
[0089] <Effects of information processing device 1A> As described above, the information processing apparatus 1A according to this exemplary embodiment employs a configuration in which conversion processing is applied to each piece of instance information and the instance information after conversion processing is integrated.
[0090] According to this configuration, it is possible to apply a conversion process to each piece of instance information and integrate the instance information after the conversion process. Therefore, in addition to the effects achieved by the information processing device 1 according to the exemplary embodiment 1, it is possible to reduce information loss in the conversion process and integration process. Furthermore, because multiple pieces of instance information are integrated, it is possible to improve the recognition accuracy of the recognition process even if information is lost.
[0091] Furthermore, the information processing apparatus 1A according to this exemplary embodiment employs a configuration in which importance is assigned to instance information after conversion processing, and the instance information is integrated using the importance.
[0092] According to this configuration, importance according to the instance information is assigned to each piece of instance information, and the instance information to which importance has been assigned can be integrated into one piece of information. Therefore, in addition to the effects achieved by the information processing device 1 according to the exemplary embodiment 1, loss of information in the integration process can be reduced. Furthermore, since multiple pieces of instance information are integrated, the recognition accuracy of the recognition process can be improved even if information is lost.
[0093] In addition, the information processing device 1A according to this exemplary embodiment is equipped with multiple conversion layers that apply conversion processing to each piece of instance information, and is configured to assign importance to the instance information after the conversion processing and integrate the instance information using the importance.
[0094] According to this configuration, the conversion process is applied multiple times serially to each piece of instance information. Furthermore, according to this configuration, importance can be assigned according to the instance information after the conversion process, and the instance information to which importance has been assigned can be integrated into a single piece of information. Therefore, in addition to the effects achieved by the information processing device 1 according to the first exemplary embodiment, it is possible to appropriately convert the instance information and reduce information loss in the conversion process and integration process. Furthermore, because multiple pieces of instance information are integrated, it is possible to improve the recognition accuracy of the recognition process even if information is lost.
[0095] Moreover, the information processing device 1A according to this exemplary embodiment employs a configuration for performing a recognition process for generating a recognition result relating to a person's behavior among one or more instances.
[0096] This configuration makes it possible to generate a recognition result related to human behavior, thereby enabling human-based behavior recognition processing in addition to the effects achieved by the information processing device 1 according to the first exemplary embodiment.
[0097] Furthermore, the information processing apparatus 1A according to this exemplary embodiment employs a configuration in which different identification processes are applied to the instance integrated information related to a person and the instance integrated information related to an object among one or more instances.
[0098] According to this configuration, it is possible to apply an identification process according to the instance integrated information related to a person and the instance integrated information related to an object, respectively. Therefore, in addition to the effects achieved by the information processing device 1 according to the exemplary embodiment 1, it is possible to reduce the cost of the identification process and to reduce information loss in the identification process.
[0099] Furthermore, the information processing device 1A according to this exemplary embodiment employs a configuration in which attribute information is assigned to each of one or more instances.
[0100] According to this configuration, attribute information is assigned to each of one or more instances, and therefore, in addition to the effects achieved by the information processing device 1 according to the exemplary embodiment 1, each instance can be identified even when there are multiple similar instances, thereby improving the recognition accuracy of the recognition process.
[0101] Exemplary Embodiment 3 A third exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in the first exemplary embodiment are denoted by the same reference numerals, and their description will not be repeated.
[0102] <Configuration of information processing device 1B> The configuration of an information processing device 1B according to this exemplary embodiment will be described with reference to Fig. 11. The information processing device 1B is a device that further has a function of learning model parameters of the storage unit 20A in the information processing device 1A.
[0103] Fig. 11 is a block diagram showing an example of the configuration of an information processing device 1B. The information processing device 1B shown in Fig. 11 differs from the information processing device 1A shown in Fig. 3 in that a learning unit 16 is provided in a control unit 10B.
[0104] The learning unit 16 trains at least one of the integration unit 13 and the recognition unit 14 by referring to training data TD which includes multiple pairs of video and recognition information RI relating to at least one or more instances contained in the video.
[0105] The training data TD includes learning video data VDL, which may be, for example, video captured by a surveillance camera.
[0106] The training data TD also includes recognition information RI. This recognition information RI may be text, a graph, a table, or an image. The recognition information RI may be, for example, a behavior label of a person appearing in a video image assigned by an operator of the information processing device 1B.
[0107] The learning unit 16 may have the functions of the extraction unit 11, aggregation unit 12, integration unit 13, and recognition unit 14, similarly to the information processing device 1A of the second exemplary embodiment.
[0108] The training data TD is generated, for example, as follows: A video from a surveillance camera is acquired by the learning unit 16, and multiple instances related to one or more instances included in the video are extracted. Recognition information RI corresponding to this video is also acquired by the learning unit 16.
[0109] For example, the operator of the information processing device 1B determines the behavior of each person appearing in the acquired video, such as what behavior the person is performing and what task the person is performing, and assigns an behavior label to the person. The operator of the information processing device 1B may select a corresponding behavior label from a plurality of behavior labels prepared in advance for the behavior of the person. Alternatively, the operator of the information processing device 1B may further input the name of an object the person is handling. The operator of the information processing device 1B assigns a behavior label to each person appearing in the acquired video via the input unit 22.
[0110] Thereafter, another video is acquired by the learning unit 16, and the same process is performed. By repeating this process, training data TD including multiple pairs of video and recognition information RI related to the instances included in the video is generated.
[0111] Note that the work for generating the training data TD described above is an example and does not limit this exemplary embodiment. Furthermore, the expression "training data" in this specification is not limited beyond the fact that it is data that is referenced for updating (learning) model parameters. Instead of the expression "training data" in this specification, expressions such as "training data" or "reference data" may be used.
[0112] After the training data having a sufficient number of pairs is generated, machine learning is performed by the learning unit 16. That is, the learning unit 16 refers to the training data and learns a prediction model that represents the correlation between the video and the recognition information RI regarding the instances included in the video.
[0113] The learning unit 16 inputs the image contained in the training data TD to the extraction unit 11, and updates at least one of the parameters of the integrated model used by the integration unit 13 and the parameters of the recognition model used by the recognition unit 14 so as to reduce the difference between the recognition result generated by the recognition unit 14 and the recognition information contained in the training data.
[0114] The learning unit 16 may update the parameters of the integrated model and the parameters of the recognition model simultaneously.
[0115] <Flow of learning process by information processing device 1B> The flow of the learning process executed by the information processing device 1B configured as above will be described with reference to Fig. 12. Fig. 12 is a flowchart showing the flow of the learning process.
[0116] (Step S21) In step S21, the learning unit 16 inputs the learning video data VDL included in the teacher data TD to the extraction unit 11.
[0117] (Step S22) In step S22, the extraction unit 11 extracts a plurality of pieces of instance information for each of one or a plurality of instances included in the learning video data VDL input in step S21.
[0118] (Step S23) In step S23, the aggregation unit 12 aggregates the multiple pieces of instance information for each instance.
[0119] (Step S24) In step S24, the integration unit 13 generates integrated instance information by integrating the plurality of pieces of instance information aggregated in step S23 for each instance.
[0120] (Step S25) In step S25, the recognition unit 14 references the instance integration information generated in step S24 and generates a recognition result for at least one of one or more instances.
[0121] (Step S26) In step S26, the learning unit 16 updates the model parameters MP so as to reduce the difference between the recognition result generated in step S25 and the recognition information RI included in the training data TD. In updating the model parameters MP, at least one of the parameters of the integrated model used by the integration unit 13 and the parameters of the recognition model used by the recognition unit 14 is updated.
[0122] In this way, the learning process shown in FIG. 12 is completed.
[0123] In the above-described learning process, learning may be performed by adjusting hyperparameters as appropriate.
[0124] <Effects of information processing device 1B> As described above, the information processing device 1B according to this exemplary embodiment is configured to train at least one of the integration means and the recognition means by referring to training data that includes multiple pairs of video and recognition information relating to at least one or more instances contained in the video.
[0125] According to this configuration, at least one of the integration means and the recognition means can be trained by referring to the training data, which not only provides the effects of the information processing device 1 according to the first exemplary embodiment but also improves the recognition accuracy of the recognition process.
[0126] The information processing device 1B according to this exemplary embodiment is configured to input an image contained in training data and update at least one of the parameters of the integrated model and the parameters of the recognition model so as to reduce the difference between the generated recognition result and the recognition information contained in the training data.
[0127] According to this configuration, at least one of the parameters of the integrated model and the parameters of the recognition model is updated so as to output a recognition result that matches the recognition information. Therefore, in addition to the effects achieved by the information processing device 1 according to the exemplary embodiment 1, the recognition accuracy of the recognition process can be improved by using the updated model parameters.
[0128] [Software implementation example] Some or all of the functions of the information processing devices 1, 1A, and 1B may be realized by hardware such as an integrated circuit (IC chip), or by software.
[0129] In the latter case, the information processing devices 1, 1A, and 1B are realized, for example, by a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in FIG. 13. The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for operating the computer C as the information processing devices 1, 1A, and 1B. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing each function of the information processing devices 1, 1A, and 1B.
[0130] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0131] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, mouse, display, and printer.
[0132] Furthermore, the program P can be recorded on a non-transitory tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.
[0133] [Appendix 1] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means disclosed in the above-described embodiments are also included in the technical scope of the present invention.
[0134] [Appendix 2] Some or all of the above-described embodiments can also be described as follows: However, the present invention is not limited to the following described aspects.
[0135] (Appendix 1) an extraction means for extracting a plurality of pieces of instance information for each of one or a plurality of instances included in the input video; aggregating means for aggregating the plurality of pieces of instance information for each instance; an integration unit that generates integrated instance information by integrating, for each instance, the plurality of pieces of instance information aggregated by the aggregation unit; a recognition means for generating a recognition result for at least one of the one or more instances by referring to the instance integration information generated by the integration means; An information processing device comprising:
[0136] (Appendix 2) The integration means It has one or more conversion layers that apply conversion processing to each instance information, and one or more integration layers that integrate the instance information after the conversion processing. 10. The information processing device according to claim 1.
[0137] (Appendix 3) The integration means assigning importance to the instance information after the conversion process by the one or more conversion layers; 3. The information processing device according to claim 2, wherein the integration layer integrates instance information using the importance.
[0138] (Appendix 4) The integration means The one or more transformation layers include a plurality of transformation layers that serially apply transformation processing to each piece of instance information; the integration layer assigns importance to the instance information after the conversion process in the conversion layer and integrates the instance information using the importance; 3. The information processing device according to claim 2, comprising:
[0139] (Appendix 5) 5. The information processing device according to claim 1, wherein the recognition means generates a recognition result relating to a human behavior from among the one or more instances.
[0140] (Appendix 6) The recognition means applies different identification processes to the instance integrated information related to a person and the instance integrated information related to an object among the one or more instances. 6. The information processing device according to claim 5.
[0141] (Appendix 7) 7. The information processing device according to claim 1, wherein the aggregation means assigns attribute information to each of the one or more instances.
[0142] (Appendix 8) a learning unit that causes at least one of the integration means and the recognition means to learn by referring to training data that includes a plurality of pairs of a video and recognition information relating to at least one or a plurality of instances included in the video; 8. An information processing device according to any one of appendices 1 to 7, comprising:
[0143] (Appendix 9) The learning unit inputting the image included in the training data into the extraction means; At least one of parameters of an integrated model used by the integration means and parameters of a recognition model used by the recognition means is updated so that a difference between the recognition result generated by the recognition means and the recognition information included in the training data becomes small. 9. The information processing device according to claim 8.
[0144] (Appendix 10) extracting a plurality of pieces of instance information for each of one or more instances included in the input video; aggregating the plurality of instance information for each instance; generating integrated instance information by integrating the aggregated plurality of pieces of instance information for each instance; generating a recognition result for at least one of the one or more instances by referring to the generated instance integration information; Contains An information processing method comprising:
[0145] (Appendix 11) Computer, an extraction means for extracting a plurality of pieces of instance information for each of one or a plurality of instances included in the input video; aggregating means for aggregating the plurality of pieces of instance information for each instance; an integration unit that generates integrated instance information by integrating, for each instance, the plurality of pieces of instance information aggregated by the aggregation unit; and a recognition unit that generates a recognition result for at least one of the one or more instances by referring to the instance integration information generated by the integration unit. A program characterized by:
[0146] [Appendix 3] Some or all of the above-described embodiments can also be expressed as follows.
[0147] An information processing device comprising at least one processor, the processor executing an extraction process that extracts a plurality of pieces of instance information for each of one or more instances included in an input video; an aggregation process that aggregates the plurality of pieces of instance information for each instance; an integration process that generates instance integration information by integrating the aggregated plurality of pieces of instance information for each instance; and a recognition process that generates a recognition result for at least one of the one or more instances by referring to the generated instance integration information.
[0148] The information processing device may further include a memory that stores a program for causing the processor to execute the extraction process, the aggregation process, the integration process, and the recognition process. The program may be recorded on a computer-readable, non-transitory, tangible recording medium. [Explanation of symbols]
[0149] 1, 1A, 1B Information processing equipment 11 Extraction part 12 Consolidation Section 13 Integration Department 14 Recognition part 15 Output section 16 Learning Department 130 Conversion Layer 131 Integration Layer
Claims
1. For each of one or more instances included in the input video, extraction means for extracting the source information; aggregating means for aggregating the plurality of pieces of instance information for each instance; The aggregation means aggregates the plurality of instance information items, and integrates the aggregated information items for each instance. integration means for generating instance integration information; The integration unit refers to the instance integration information generated by the integration unit, and a recognition means for generating a recognition result relating to at least one of the the plurality of pieces of instance information include rectangle information that is a rectangle surrounding a person and pose information that represents a posture of the person; The aggregation means aggregating the rectangle information and the pose information based on a positional relationship between the rectangle information and the pose information; or The rectangle information and the pose information are aggregated based on the degree of overlap between the rectangle of the rectangle information and the circumscribing rectangle of the pose information. Information processing device.
2. The integration means One or more transformation layers that apply transformations to each instance, and the resulting instance and one or more integration layers for integrating stance information. The information processing device according to claim 1 .
3. The integration means assigning importance to the instance information after the conversion process by the one or more conversion layers; The information processing device according to claim 2, wherein the integration layer integrates instance information using the importance. Information processing device.
4. The integration means The one or more transformation layers serially apply transformation processes to each instance information. a plurality of conversion layers using The integration layer assigns importance to the instance information after the conversion process in the conversion layer. an integration layer that integrates instance information using the importance; The information processing device according to claim 2 , further comprising:
5. The recognition means recognizes a recognition result relating to a human behavior from among the one or more instances. The information processing apparatus according to claim 1 , wherein the information processing apparatus generates the information.
6. The recognition unit may be configured to: Different identification processes are applied to the information and the instance integration information of the object. The information processing device according to claim 5 .
7. The aggregation unit assigns attribute information to each of the one or more instances. The information processing device according to claim 1 .
8. Identification of the video and / or one or more instances contained in the video and referring to training data including a plurality of combinations of the recognition information and the integration means, A learning section that learns either The information processing device according to claim 1 , further comprising:
9. For each of one or more instances included in the input video, extracting the information; aggregating the plurality of instance information for each instance; The aggregated information of the plurality of instances is integrated for each instance. generating instance integration information; By referring to the generated instance integration information, a part of the one or more instances is integrated. generating a recognition result for at least one of the Including, the plurality of pieces of instance information include rectangle information that is a rectangle surrounding a person and pose information that represents a posture of the person; aggregating the plurality of pieces of instance information for each instance, aggregating the rectangle information and the pose information based on a positional relationship between the rectangle information and the pose information; or aggregating the rectangle information and the pose information based on the degree of overlap between the rectangle of the rectangle information and the circumscribing rectangle of the pose information. An information processing method comprising:
10. Computer, For each of one or more instances included in the input video, extraction means for extracting the source information; aggregating means for aggregating the plurality of pieces of instance information for each instance; The aggregation means aggregates the plurality of instance information items, and integrates the aggregated information items for each instance. integration means for generating instance integration information; The integration unit refers to the instance integration information generated by the integration unit, and and a recognition means for generating a recognition result relating to at least one of the It functions as the plurality of pieces of instance information include rectangle information that is a rectangle surrounding a person and pose information that represents a posture of the person; The aggregation means aggregating the rectangle information and the pose information based on a positional relationship between the rectangle information and the pose information; or The rectangle information and the pose information are aggregated based on the degree of overlap between the rectangle of the rectangle information and the circumscribing rectangle of the pose information. A program characterized by:
Citation Information
Patent Citations
Learning video selecting device, program and method for selecting, as learning video, shot video with predetermined image region masked
JP2019079357A
Program, device, and method for recognizing actions of persons using a plurality of recognition engines
JP2019144830A
Image processing device and image processing program
JP2021065617A