Information processing device and information processing method

The information processing device addresses the challenge of labeling large sequence data volumes by estimating object states and dividing sequences for pedestrian crossing intentions, enhancing labeling efficiency and accuracy.

JP7739756B2Active Publication Date: 2025-09-17AISIN CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021084941
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-19
Publication Date
2025-09-17
Estimated Expiration
2041-05-19

AI Technical Summary

Technical Problem

Conventional techniques face challenges in efficiently labeling large volumes of sequence data for machine learning, particularly in determining pedestrian crossing intentions, due to the enormous data volume and the lack of effective thinning methods beyond random thinning.

Method used

An information processing device that acquires and processes multiple frame images to estimate object states, divides sequences based on state changes, and extracts meaningful sequences for labeling, using a state estimation unit, sequence division unit, and display control unit to facilitate user input.

Benefits of technology

Enables efficient and accurate labeling of pedestrian crossing intentions by prioritizing meaningful sequences, reducing noise, and preventing biased learning data, thereby improving the accuracy of training data for machine learning applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007739756000001
    Figure 0007739756000001
  • Figure 0007739756000002
    Figure 0007739756000002
  • Figure 0007739756000003
    Figure 0007739756000003
Patent Text Reader

Abstract

To support teacher data creation for machine learning by a user using a plurality of frame images for a time series obtained by photographing a photographed area including a passage where a predetermined object can cross.SOLUTION: An information processing device acquires a plurality of frame images for a time series obtained by photographing a photographed area including a passage where a predetermined object can cross. Moreover, the information processing device estimates one of a plurality pieces of state information including a change state and a stagnation state, by using one or more last frame images together, for each frame image, about the object reflected in the frame image. Moreover, the information processing device divides, takes out and groups a first sequence from the frame image of the change state to the frame image of the stagnation state, and a second sequence from the frame image of the stagnation state to the frame image of the change state, with respect to the plurality of frame images in the time series, on the basis of the state information and a predetermined division rule, for each object reflected in the frame image.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD Embodiments of the present invention relate to an information processing device and an information processing method. [Background technology]

[0002] 2. Description of the Related Art Conventionally, there has been a technique for determining whether or not there is a pedestrian about to cross a roadway, based on an image of the area ahead of a vehicle captured by an on-board camera, for example.

[0003] Furthermore, when such a determination is made using machine learning techniques with training data, the training data must be prepared in advance. When creating the training data, for example, the user assigns a label indicating the magnitude of the intention to cross the road to each pedestrian in sequence data, which is a time series of multiple frame images. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2020-184225 [Patent Document 2] Japanese Patent Application Publication No. 2019-520655 Summary of the Invention [Problem to be solved by the invention]

[0005] However, with conventional techniques, the number of sequence data can become enormous, making it difficult to label all of the sequence data. Furthermore, even if sequence data can be thinned out, the only method available is random thinning, which leaves room for improvement.

[0006] Therefore, the problem that the present invention aims to solve is to provide an information processing device and an information processing method that can assist users in creating training data for machine learning using multiple frame images in a time series obtained by photographing a shooting area that includes a passageway that a specified object can cross. [Means for solving the problem]

[0007] An information processing device according to an embodiment includes an acquisition unit that acquires a plurality of frame images in chronological order obtained by photographing a photographing area including a passageway that a specified object can cross; a state estimation unit that, for each frame image, estimates one of a plurality of state information for the object appearing in the frame image, using one or more of the frame images immediately preceding the frame image, including a changing state in which there is a change in at least one of the movement and posture that is greater than or equal to a threshold, and a stagnant state in which there is no change in both the movement and posture that is greater than or equal to the threshold; a sequence division unit that, for each object appearing in the frame images, divides and extracts a first sequence from the frame image in the changing state to the frame image in the stagnant state and a second sequence from the frame image in the stagnant state to the frame image in the changing state based on the state information and a predetermined division rule, and groups the first sequence and the second sequence; and a display control unit that displays the grouped first sequence and the second sequence on a display unit. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram showing the overall configuration of an information processing apparatus according to this embodiment. [Figure 2] FIG. 2 is a diagram schematically showing an example of a frame image in a stagnant state in this embodiment. [Figure 3] FIG. 3 is a diagram schematically showing an example of a frame image in a changing state in this embodiment. [Figure 4] FIG. 4 is a diagram schematically illustrating an example of frame division in this embodiment. [Figure 5] FIG. 5 is a diagram schematically illustrating an example of frame extraction in this embodiment. [Figure 6] FIG. 6 is a flowchart showing the processing performed by the information processing device of this embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, an information processing apparatus and an information processing method according to embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0010] 1 is a diagram showing the overall configuration of an information processing device 1 according to this embodiment. The information processing device 1 is a computer device, and includes a storage unit 2, an input unit 3, a display unit 4, a communication unit 5, and a processing unit 6.

[0011] The storage unit 2 is a means for storing various types of information, and is realized by, for example, a RAM (Random Access Memory), a ROM (Read Only Memory), an SSD (Solid State Drive), or an HDD (Hard Disk Drive).

[0012] The input unit 3 is an information input means by the user, and is realized by, for example, a keyboard, a touch panel, a pointing device, a mouse, an input button, or the like.

[0013] The display unit 4 is a means for displaying various types of information, and is realized by, for example, an LCD (Liquid Crystal Display), an organic EL (Electro-Luminescence), or the like.

[0014] The communication unit 5 is a communication interface for communicating with an external device.

[0015] The processing unit 6 is a means for executing various types of arithmetic processing and is realized by, for example, a CPU (Central Processing Unit). The processing unit 6 includes, as functional units, an acquisition unit 61, a state estimation unit 62, a sequence division unit 63, a sequence extraction unit 64, a display control unit 65, and a control unit 66.

[0016] The acquisition unit 61 acquires various information from an external device. The acquisition unit 61 acquires, for example, a plurality of frame images in time series obtained by an imaging unit (for example, a camera mounted on a vehicle such as a passenger car) that captures an imaging area including a passageway that a predetermined object can cross. Here, the predetermined object is, for example, a person, an animal (dog, cat, bird, etc.), a car, an automobile, a bicycle, etc. In the following, an example will be described in which the predetermined object is a person, i.e., a pedestrian.

[0017] The state estimation unit 62 estimates one of the following three pieces of state information for a pedestrian captured in a frame image, for each frame image, using one or more frame images immediately preceding the frame image. (1) Change state (a state in which there is a change in at least one of the movement and posture that exceeds a threshold (a movement threshold and a posture threshold are respectively set)) (2) Stagnation (a state in which neither movement nor posture changes above the threshold) (3) Unstable state (state in which the results of image processing are unstable)

[0018] Here, Fig. 2 is a diagram schematically illustrating an example of a frame image in a stationary state in this embodiment. The frame image in Fig. 2 shows a roadway R, sidewalks W1 and W2, and a pedestrian M. Pedestrian M is walking along sidewalk W2. Therefore, there is no change in either the movement or posture of this pedestrian M that exceeds the threshold, and the pedestrian is determined to be in a "stationary state."

[0019] 3 is a diagram schematically illustrating an example of a frame image in a changing state in this embodiment. Pedestrian M has changed his / her movement (e.g., direction of travel) and posture (e.g., body orientation, facial orientation, etc.) on the sidewalk W2 toward the roadway R from the state in the immediately preceding frame image. Therefore, this pedestrian M has experienced a change in at least one of his / her movement and posture that is equal to or greater than the threshold, and is determined to be in a "changing state."

[0020] 1, the state estimation unit 62 will be described in further detail. The state estimation unit 62 includes an object detection unit 621, a posture estimation unit 622, an object recognition unit 623, and a state determination unit 624.

[0021] The object detection unit 621 sets bounding boxes for pedestrians in frame images and also associates bounding boxes between frame images. The object detection unit 621 is realized by known techniques such as "Mask R-CNN" and "Yolo (You Only Look Once)."

[0022] The pose estimation unit 622 estimates the positions of joint points of a pedestrian in a frame image. The pose estimation unit 622 is realized by a known technique such as "HRNet (High-Resolution Network) + DarkPose" or "OpenPose."

[0023] The object recognition unit 623 (semantic segmentation) performs object recognition on a pixel-by-pixel basis for the frame image. The object recognition unit 623 is realized by a known technique such as "BiSeNet (Bilateral Segmentation Network)" or "Panoptic Segmentation."

[0024] The state determination unit 624 determines that a pedestrian is in a changed state when there is a change in at least one of the movement and posture of the pedestrian between multiple frame images that is equal to or greater than a threshold value. For example, the state determination unit 624 determines that a pedestrian is in a changed state when the pedestrian changes the direction of walking on the sidewalk or when the pedestrian changes the direction of their head.

[0025] Furthermore, the state determination unit 624 determines that a pedestrian is in a stationary state when there is no change in both the movement and posture of the pedestrian between multiple frame images that is equal to or greater than a threshold value. For example, the state determination unit 624 determines that a pedestrian is in a stationary state when the pedestrian is walking straight on the sidewalk at a constant speed or when the pedestrian is standing still on the sidewalk in the same posture.

[0026] Furthermore, the state determination unit 624 determines that the state is unstable when the results of image processing for a pedestrian are unstable between multiple frame images. For example, the state determination unit 624 determines that the state is unstable when there is a significant change in SemanticSegmentation, when there is a bounding box correspondence error, or when the number of joint points calculated by the posture estimation unit 622 is equal to or less than a threshold.

[0027] For each pedestrian appearing in a frame image, the sequence division unit 63 searches for, divides, extracts, and groups the following two types of sequences for multiple frame images in chronological order based on the state information and a predetermined division rule: (11) First sequence (from a frame image in a changing state to a frame image in a stagnant state) (12) Second sequence (from a frame image in a stagnant state to a frame image in a changing state)

[0028] The sequence division unit 63 searches for, divides, and extracts (11) a first sequence and (12) a second sequence for all pedestrians captured in the frame image, and groups them.

[0029] This will be described with reference to Fig. 4. Fig. 4 is a diagram schematically showing an example of frame division in this embodiment. Fig. 4(a) shows input frames (input frame images) in chronological order with frame numbers "1" to "30".

[0030] 4(b) shows the state information determination result for pedestrian A in the input frame by the state determination unit 624. "C" represents a changing state, "S" represents a stagnant state, and "U" represents an unstable state.

[0031] 4(c) to 4(h) show sequences grouped by the sequence division unit 63. 4(c) to 4(e) show (11) the first sequence (from a frame image in a changing state to a frame image in a stagnant state). It should be noted that the sequence seq2 in FIG. 4(d) has one "S" between the first three "C"s, but it is acceptable to have one "S" between multiple "C"s, or one "C" between multiple "S"s. This is done to make the grouped sequences more meaningful, for example.

[0032] Also, (f) to (h) of FIG. 4 show (12) the second sequence (from a frame image in a stagnant state to a frame image in a changing state).

[0033] Returning to FIG. 1 , the sequence extraction unit 64 extracts sequences from the grouped one or more first sequences and one or more second sequences based on predetermined parameters. For example, the sequence extraction unit 64 extracts sequences from the grouped one or more first sequences and one or more second sequences that contain a predetermined allowable number of frame images in an unstable state. Also, for example, the sequence extraction unit 64 extracts sequences from the grouped one or more first sequences and one or more second sequences at a predetermined extraction ratio of first sequences and second sequences.

[0034] These will be described with reference to Fig. 5. Fig. 5 is a diagram schematically showing an example of frame extraction in this embodiment. Here, the sequence extraction unit 64 extracts sequences to be labeled in accordance with the following parameters: <Parameter 1> The ratio of the first sequence to the second sequence is 50% each. <Parameter 2> Unstable state tolerance "1"

[0035] FIG. 5(a) shows the sequences before extraction, including pedestrian A's sequences seq1 to seq6 and pedestrian B's sequences seq1 to seq3. FIG. 5(b) shows the sequences after extraction by the sequence extraction unit 64. (1) is pedestrian A's sequence seq1. (2) is pedestrian B's sequence seq2.

[0036] (3) is pedestrian A's sequence seq2. (4) is pedestrian A's sequence seq4. (5) is pedestrian A's sequence seq5. (6) is pedestrian B's sequence seq3.

[0037] The parameters used for sequence extraction are not limited to those described above. For example, parameters that constrain the extraction of at least one sequence for each pedestrian may be used.

[0038] 1, the display control unit 65 causes various pieces of information to be displayed on the display unit 4. The display control unit 65 causes the display unit 4 to display, for example, the first sequence and the second sequence extracted by the sequence extraction unit 64.

[0039] The control unit 66 executes processes other than those executed by the units 61 to 65. For example, the control unit 66 receives label input by the user using the input unit 3 for the first sequence and the second sequence extracted by the sequence extraction unit 64 and displayed on the display unit 4, and stores the label input in the storage unit 2.

[0040] 6 is a flowchart showing the processing by the information processing device 1 of this embodiment. First, in step S1, the acquisition unit 61 acquires a plurality of frame images (FIGS. 2 and 3) in chronological order.

[0041] Next, in step S2, the state estimation unit 62 estimates, for each frame image, whether the pedestrian shown in the frame image is in a changing state, a stagnant state, or an unstable state, using one or more frame images immediately before that frame image.

[0042] Next, in step S3, the sequence division unit 63 searches for, divides, extracts, and groups two types of sequences (first sequence and second sequence) for the multiple frame images in chronological order, based on the state information and a predetermined division rule, for each pedestrian appearing in the frame images (FIG. 4).

[0043] Next, in step S4, the sequence extraction unit 64 extracts a sequence from the grouped one or more first sequences and one or more second sequences based on predetermined parameters (FIG. 5).

[0044] Next, in step S5, the display control unit 65 causes the display unit 4 to display the first sequence and the second sequence extracted in step S4.

[0045] Next, in step S6, the control unit 66 accepts label input by the user using the input unit 3 for the first sequence and the second sequence displayed on the display unit 4 in step S5, and stores the label input in the memory unit 2.

[0046] As described above, the information processing device 1 of this embodiment performs state estimation, sequence division, and sequence extraction based on a plurality of frame images in time series, and displays the extracted sequences on the display unit 4. This allows meaningful sequences to be preferentially extracted and displayed from a huge number of sequences, making the labeling work by the user easier and more meaningful.

[0047] Furthermore, in addition to changing and stagnant states, unstable states can also be used as state information, and sequences can be extracted based on the number of frame images in the unstable state. This makes it possible, for example, to extract only sequences with few unstable frame images at the beginning of development in order to reduce noise for theory construction. Furthermore, for example, at the end of development, it is possible to extract sequences with many unstable frame images in order to intentionally increase noise to improve robustness. In other words, the labeling effort can be adjusted by adjusting the proportion of noise mixed in depending on the purpose and development phase.

[0048] Furthermore, by determining the ratio of the first sequence to the second sequence when extracting sequences, the user can extract the sequences they desire. In other words, it is possible to prevent biased learning data, such as sequences of monotonous pedestrian movements dominating the data.

[0049] Furthermore, as shown in Figure 4, by dividing the input sequence into sequences that allow overlapping frame images, it is possible to pick up all transitions in the pedestrian's crossing intention. In other words, it is possible to support labeling of a series of pedestrian movements, thereby improving the accuracy of the labels.

[0050] On the other hand, there is a conventional technology that uses machine learning or the like to estimate the start and end positions of a routine task from a video, and uses these as temporary labels, allowing the user to check the temporarily labeled sections and adjust and confirm the sections. However, because changes in crossing intentions (intentions to cross an aisle) are diverse, this conventional technology is not suitable for sequence extraction related to labeling crossing intentions as in this embodiment.

[0051] Another conventional technique reduces the burden of labeling by using constraints of physical laws, such as the requirement that an object released into the air follows a parabolic trajectory. However, because a pedestrian's crossing intention is not an event that follows physical laws, this conventional technique is not suitable for sequence extraction related to labeling crossing intention, as in this embodiment.

[0052] In the above-described embodiment, the program for executing the information processing may be provided by being recorded in an installable or executable file format on a computer-readable recording medium such as a CD-ROM, a flexible disk (FD), a CD-R, a DVD (Digital Versatile Disk), or a USB (Universal Serial Bus) memory. The program may also be provided or distributed via a network such as the Internet. The program may also be provided by being pre-installed in a ROM or the like.

[0053] The program has a modular configuration including the above-mentioned functional units, and in actual hardware, for example, a CPU (processor circuit) reads the program from a ROM or HDD and executes it, thereby loading the above-mentioned functional units into a RAM (main memory) and generating the above-mentioned functional units in the RAM (main memory). Note that some or all of the above-mentioned functional units can also be realized using dedicated hardware such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array).

[0054] Although the embodiments have been described, the above embodiments are presented as examples and are not intended to limit the scope of the invention. The novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. The above embodiments are included within the scope and spirit of the invention, and are also included in the inventions and their equivalents as defined in the claims.

[0055] For example, the application of the present invention is not limited to vehicles, but may also be to other objects such as mobile robots. [Explanation of symbols]

[0056] 1...information processing device, 2...storage unit, 3...input unit, 4...display unit, 5...communication unit, 6...processing unit, 61...acquisition unit, 62...state estimation unit, 63...sequence division unit, 64...sequence extraction unit, 65...display control unit, 66...control unit, 621...object detection unit, 622...posture estimation unit, 623...object recognition unit, 624...state determination unit, M...pedestrian, R...roadway, W1, W2...sidewalk

Claims

1. an acquisition unit that acquires a plurality of frame images in time series by capturing an image of a capturing area including a passageway through which a predetermined object can cross; a state estimation unit that estimates, for each frame image, one of a plurality of state information pieces for the object captured in the frame image, using one or more of the frame images immediately preceding the frame image, including a change state in which there is a change in at least one of movement and posture that is equal to or greater than a threshold, and a stagnant state in which there is no change in either movement or posture that is equal to or greater than the threshold; a sequence division unit that divides and extracts, for each of the objects captured in the frame images, a first sequence from the frame image in the changing state to the frame image in the stagnant state and a second sequence from the frame image in the stagnant state to the frame image in the changing state for the plurality of frame images in time series based on the state information, and groups the extracted sequences; a display control unit that displays the grouped first sequence and the grouped second sequence on a display unit.

2. the state estimation unit estimates, for each frame image, state information of the object captured in the frame image, one of the changing state, the stagnant state, and an unstable state in which a result of image processing is unstable, by using the immediately preceding one or more frame images; the information processing device further includes a sequence extraction unit that extracts a sequence including a predetermined allowable number or less of the frame images in the unstable state from among the one or more grouped first sequences and the one or more grouped second sequences; The information processing apparatus according to claim 1 , wherein the display control unit causes the display unit to display the sequence extracted by the sequence extraction unit.

3. a sequence extraction unit that extracts sequences from the grouped one or more first sequences and one or more second sequences at a predetermined extraction ratio of the first sequences and the second sequences, The information processing apparatus according to claim 1 , wherein the display control unit causes the display unit to display the sequence extracted by the sequence extraction unit.

4. An information processing method executed by an information processing device, comprising: an acquisition step of acquiring a plurality of frame images in time series obtained by photographing an imaging area including a passageway that a predetermined object can cross; a state estimation step of estimating, for each frame image, one of a plurality of state information pieces for the object shown in the frame image, using one or more of the frame images immediately preceding the frame image, including a change state in which there is a change in at least one of the movement and the posture that is equal to or greater than a threshold, and a stagnant state in which there is no change in either the movement or the posture that is equal to or greater than the threshold; a sequence division step of dividing and extracting a first sequence from the frame image in the changing state to the frame image in the stagnant state and a second sequence from the frame image in the stagnant state to the frame image in the changing state for the plurality of frame images in time series based on the state information for each of the objects shown in the frame images, and grouping the extracted sequences; a display control step of displaying the grouped first sequence and second sequence on a display unit.

Citation Information

Patent Citations

  • Object type decision device, vehicle, object type decision method, and program for object type decision

    JP2009042941A

  • Pedestrian sensor, pedestrian signal control system, pedestrian sensing method, and pedestrian sensing program

    JP2015072548A

  • Systems and methods for generating data interpretations in neural networks and related systems

    JP2019520655A

  • Information processor, and information processing method

    JP2020184225A