system

The system addresses misidentification issues by employing vanishing points and image similarity to accurately match objects across multiple camera views on a moving vehicle, enhancing object tracking and quality management.

JP2026057155APending Publication Date: 2026-04-02NIPPON SIGNAL CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing systems fail to accurately identify and manage objects captured by multiple cameras on a moving body, leading to misidentification of the same object as different objects or vice versa due to positional discrepancies.

Method used

A system that performs object identification processes using positional relationships, feature quantities, moving speed, and direction changes of a moving body, particularly utilizing vanishing points and image similarity, to accurately match objects across images captured by the same or different cameras on a moving vehicle.

Benefits of technology

Enhances object identification accuracy by correlating images from multiple cameras, ensuring consistent object tracking and quality assessment, thereby improving the management of objects captured at different times or by different cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026057155000001_ABST
    Figure 2026057155000001_ABST
Patent Text Reader

Abstract

This invention provides a means to enable identification of objects in images captured by multiple cameras mounted on a mobile vehicle. [Solution] System 1 comprises cameras 11A and 11B mounted on a railway vehicle 9 that runs on rails 8, an odometer 12 that measures the distance from the starting point of the railway vehicle 9, and a data processing device 13. The data processing device 13 identifies objects in images sequentially captured by camera 11A at predetermined time intervals while the railway vehicle 9 is moving. The data processing device 13 also identifies objects in images sequentially captured by camera 11B at predetermined time intervals while the railway vehicle 9 is moving. Furthermore, the data processing device 13 identifies objects in images captured by camera 11A and objects in images captured by camera 11B.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system for identifying an object in an image captured by a camera mounted on a moving body.

Background Art

[0002] There is a technology for extracting images of traffic facilities such as insulators, concrete poles, hangers, and overhead wires installed along a track or the like from an image captured by a camera mounted on a moving body such as a railway vehicle during the movement of the moving body.

[0003] As a patent document disclosing the above technology, for example, there is Patent Document 1.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the technology described in Patent Document 1, for each of a plurality of images continuously captured and generated by a camera mounted on a moving body during the movement of the moving body, position information indicated by latitude and longitude measured by GNSS (Global Navigation Satellite System), the moving distance from a predetermined starting point, etc. is associated and managed. Therefore, there is a problem that the same object shown in different images is managed as a different object, or different objects shown in those different images are managed as the same object.

[0006] Incidentally, a mobile vehicle may be equipped with multiple cameras with different fields of view, such as a camera that photographs the front of the vehicle and another camera that photographs the rear of the vehicle, and images of a predetermined type of object, such as traffic equipment, may be extracted from the images captured by each of these cameras. In that case, in addition to the problems mentioned above, there is also the problem that the same object appearing in different images captured by different cameras may be managed as different objects, or different objects appearing in those different images may be managed as the same object.

[0007] In view of the above circumstances, the present invention provides a means for identifying objects in images captured by multiple cameras mounted on a mobile body. [Means for solving the problem]

[0008] To solve the above-mentioned problems, the present invention provides, in a first embodiment, a system that performs a first process of identifying objects in images taken at different times by the same camera mounted on a moving body, and a second process of identifying the objects identified in the first process with respect to each of the different cameras attached to the moving body.

[0009] According to the system of the first embodiment, identification is performed between objects captured in images taken by multiple cameras mounted on a moving object.

[0010] In the system of the first embodiment described above, a configuration in which the positional relationship between a vanishing point in the image and an object in the image is used in the first processing may be adopted as a second embodiment.

[0011] According to the system of the second embodiment, the identification of objects in images taken by the same camera at different times can be achieved with higher accuracy compared to the case where vanishing points are not used.

[0012] In the system of the first embodiment described above, a third embodiment may be adopted in which feature quantities of objects in the image are used in at least one of the first and second processes.

[0013] According to the system of the third embodiment, at least one of the identification of objects in images taken at different times by the same camera mounted on a moving object, and the identification of objects in images taken by each of different cameras attached to the moving object, can be achieved with higher accuracy compared to the case where features are not used.

[0014] In the system of the first embodiment described above, a fourth embodiment may be adopted in which the moving speed of the moving body is used in at least one of the first process and the second process.

[0015] According to the system of the fourth embodiment, at least one of the identification of objects in images taken at different times by the same camera mounted on a moving object, and the identification of objects in images taken by each of different cameras attached to the moving object, can be achieved with higher accuracy compared to the case where the moving speed of the moving object is not used.

[0016] In the system of the first embodiment described above, a fifth embodiment may be adopted in which the degree of change in the direction of movement of the moving body is used in at least one of the first process and the second process.

[0017] According to the system of the fifth embodiment, at least one of the identification of objects in images taken at different times by the same camera mounted on a moving body, and the identification of objects in images taken by each of different cameras attached to the moving body, can be achieved with higher accuracy compared to the case where the degree of change in the direction of movement of the moving body is not used.

[0018] In the system of the fifth embodiment described above, a sixth embodiment may be adopted in which the moving body is a railway vehicle, and the degree of change is determined based on the shape of the rails in the image captured by a camera mounted on the moving body.

[0019] According to the system of the sixth embodiment, the degree of change in the direction of movement of the moving object is determined by the image captured by the camera.

[0020] In the system of the first aspect described above, in the second process, a configuration in which the moving speed of the position of the same object identified in the first process in a plurality of images taken by the same camera mounted on the moving body at different times is used may be adopted as the seventh aspect.

[0021] According to the system of the seventh aspect, the identification between objects in images taken by different cameras attached to the moving body is realized with high accuracy as compared with the case where the moving speed of the object in the image is not used.

Brief Description of the Drawings

[0022] [Figure 1] A diagram showing the configuration of a system according to an embodiment of the present invention. [Figure 2] A diagram showing the data configuration of a management table according to an embodiment of the present invention. [Figure 3] A diagram showing the flow of processing performed by a data processing device according to an embodiment of the present invention. [Figure 4] A diagram showing the flow of processing performed by a data processing device according to an embodiment of the present invention. [Figure 5] A diagram for explaining the processing performed by a data processing device according to an embodiment of the present invention. [Figure 6] A diagram for explaining the processing performed by a data processing device according to an embodiment of the present invention. [Figure 7] A diagram showing the flow of processing performed by a data processing device according to an embodiment of the present invention. [Figure 8] A diagram for explaining the processing performed by a data processing device according to an embodiment of the present invention. [Figure 9] A diagram showing an example of a screen displayed by a data processing device according to an embodiment of the present invention. [Figure 10] A diagram for explaining the processing performed by a data processing device according to a modification of the present invention. [Figure 11] A diagram for explaining the processing performed by a data processing device according to a modification of the present invention. [Figure 12] A diagram illustrating the processing performed by a data processing device according to one modified example of the present invention. [Figure 13] A diagram illustrating the processing performed by a data processing device according to one modified example of the present invention. [Modes for carrying out the invention]

[0023] [Embodiment] Figure 1 shows the configuration of System 1 according to one embodiment of the present invention. System 1 is a system that extracts images of specific types of objects (such as traffic facilities) from images taken around the movement path of a moving object, and determines whether the condition of the objects is good or bad based on the extracted images.

[0024] System 1 comprises cameras 11A and 11B mounted on a railway vehicle 9 (an example of a mobile vehicle) that runs on rails 8, an odometer 12 for measuring the distance traveled by the railway vehicle 9, and a data processing device 13.

[0025] Hereafter, when cameras 11A and 11B are not distinguished from each other, they will simply be referred to as camera 11. Camera 11 is a camera that continuously captures still images at a predetermined frame rate, such as 30fps (frames per second), that is, at predetermined time intervals.

[0026] Cameras 11A and 11B do not necessarily have to be the same type of camera, but for the purposes of the following explanation, we will assume that cameras 11A and 11B are the same type of camera. Cameras 11A and 11B are installed on the railway vehicle 9 so that they have different shooting directions. The horizontal field of view of camera 11A is α centered on arrow A in Figure 1, and the horizontal field of view of camera 11B is α centered on arrow B in Figure 1.

[0027] Cameras 11A and 11B are connected to the data processing unit 13 via wired or wireless communication, and sequentially transmit the captured images to the data processing unit 13.

[0028] The odometer 12 is connected to the data processing device 13 via wired or wireless communication. The odometer 12 continuously measures the distance traveled by the railway vehicle 9 from the starting point and sequentially transmits the measurement results to the data processing device 13.

[0029] The data processing device 13 is, for example, a computer and includes a memory for persistently storing data including programs, a processor for processing data according to the programs stored in the memory, a communication interface for communicating with the camera 11 and the odometer 12, a display for displaying information to the user, and an input device such as a keyboard for receiving data input from the user. These components of the data processing device 13 do not necessarily have to be housed in a single enclosure. For example, the display and input device may be connected as external devices to the computer main unit, which includes the memory, processor, and communication interface.

[0030] The data processing unit 13 continuously measures the current time, for example, using the clock of the processor.

[0031] The data processing device 13 performs the operations described below by having the processor perform data processing according to the program of this embodiment.

[0032] Figure 2 shows the data structure of the management table M stored in the memory of the data processing device 13. The data processing device 13 stores a different management table M for each camera 11 and for each section of travel of the railway vehicle 9. Figure 2 shows, as an example, the data structure of the management table M for the travel section from point X to point Y of camera 11A, but the data structure of other management tables M is similar.

[0033] Management table M is a collection of data records (hereinafter simply referred to as "records") corresponding to images captured by camera 11. Management table M has the following data fields (hereinafter simply referred to as "fields").

[0034] Time of capture: Stores data indicating the time when camera 11 captured the image (hereinafter referred to as the time of capture). Distance: This section stores data indicating the distance traveled (distance) from the starting point of the travel section to the position of the railway vehicle 9 at the time the image was taken. File name: Stores data indicating the file name of the image data representing the image captured by camera 11 (hereinafter referred to as "captured image"). Object Image: Stores data related to images of objects extracted from captured images (hereinafter referred to as "object images").

[0035] The "Object Image" field in management table M has the following subfields:

[0036] Region: Stores data indicating the region occupied by the object image in the captured image. Type: Stores data indicating the type of object represented by the object image (e.g., insulator, concrete pole, hanger, power line). Temporary ID: Stores a temporary ID, which is identification information (Identifier) ​​that identifies an object image extracted from images captured by the same camera 11. This ID: After identification has been performed between objects represented by object images extracted from images captured by the same camera 11, this ID, which is identification information used to distinguish the identified object from other objects, is stored here. Integrated ID: This ID is extracted from each captured image of different cameras 11. After identification of objects identified by this ID is performed, the integrated ID, which is identification information used to distinguish the identified objects from other objects, is stored here. Judgment Result: Stores data indicating the result of the quality judgment for the object identified by the integrated ID.

[0037] The "Object Image" field in the management table M has the above subfields, which generate as many subrecords as there are object images extracted from a single captured image, and each of these subrecords stores data corresponding to the individual object image.

[0038] When the data processing device 13 receives data indicating the mileage from the odometer 12, it temporarily stores that data in memory. In other words, the data processing device 13 avoids unnecessary consumption of memory capacity by keeping only the most recent mileage data received from the odometer 12 in memory.

[0039] Each time the data processing device 13 receives image data representing a captured image from the camera 11, it performs processing according to the flow shown in Figure 3.

[0040] The data processing device 13 first stores the received image data in memory (step S10).

[0041] Next, the data processing device 13 adds a new record to the management table M (Figure 2) relating to the section of travel on which the railway vehicle 9 is currently traveling, corresponding to the camera 11 that transmitted the image data (step S11). Hereinafter, the record added in step S11 will be referred to as the "current record".

[0042] Next, the data processing device 13 stores data indicating the current time in the "shooting time" field of the current record (step S12).

[0043] Next, the data processing device 13 stores in the "Distance" field of the current record data indicating the distance that was last received from the odometer 12 at that time (step S13).

[0044] Next, the data processing device 13 stores data indicating the file name of the image data stored in memory in step S11 in the "file name" field of the current record (step S14).

[0045] Note that the execution order of steps S12 to S14 may be changed at will.

[0046] Next, the data processing device 13 reads the image data stored in memory in step S10 from memory and recognizes a specific object (for example, traffic equipment such as insulators, concrete poles, hangers, and power lines) from the image represented by the read image data using known image recognition techniques (step S15). The image recognition technique used in step S15 may be any of the following: techniques that do not use artificial intelligence, such as feature point matching and Bag-of-Visual Words, or techniques that use neural network type (deep learning, etc.) learning models. All of these image recognition techniques use the features of objects in the image.

[0047] Next, the data processing device 13 generates a sub-record corresponding to each object recognized in step S15 within the "object image" field of the current record (step S16).

[0048] Next, for each object recognized in step S15, the data processing device 13 stores data indicating the area of ​​the captured image occupied by the image representing the recognized object (object image) in the "area" subfield of the subrecord generated in step S16 (step S17).

[0049] Next, the data processing device 13 stores data indicating the type of the recognized object in the "Type" subfield of the subrecord generated in step S16 for each object recognized in step S15 (step S18).

[0050] Next, the data processing device 13 generates a unique temporary ID to identify each object recognized in step S15, and stores the generated temporary ID in the "Temporary ID" subfield of the subrecord generated in step S16 (step S19).

[0051] Note that the execution order of steps S17 to S19 may be changed at will.

[0052] Following the process shown in Figure 3 above, the data processing device 13 uses a management table M in which data is stored in the "Shooting Time" field, "Distance" field, "File Name" field, the "Area" subfield, "Type" subfield, and "Temporary ID" subfield of the "Object Image" field, but no data is stored in the "Actual ID" subfield, "Integrated ID" subfield, and "Judgment Result" subfield of the "Object Image" field. The data processing device 13 then performs a process (hereinafter referred to as "first process") to identify objects in images taken by the same camera 11 at different times.

[0053] Figure 4 shows the flow of the first process.

[0054] The data processing device 13 first selects two records from the management table M to be processed, in order of oldest capture time (step S20). The two captured images corresponding to the records selected in step S20 are images taken by the same camera 11 (camera 11A or camera 11B) at different adjacent times.

[0055] Figure 5 schematically illustrates the captured images and object images contained in each of the two records selected by the data processing device 13 in step S20. The captured image in Figure 5(A) includes an object identified by temporary ID "T001" and an object identified by temporary ID "T002". The captured image in Figure 5(B) includes an object identified by temporary ID "T003", an object identified by temporary ID "T004", and an object identified by temporary ID "T005".

[0056] Hereinafter, an object image identified by the temporary ID "T###" (where "###" is any natural number) will be referred to as "object image T###".

[0057] The data processing device 13 selects one object image to be compared from the object images included in the captured image in Figure 5(A) (step S21). Now, let's assume that object image T001 was selected in step S21.

[0058] Next, the data processing device 13 selects one object image to be compared from the object images included in the captured image in Figure 5(B) (step S22). Now, let's assume that object image T003 was selected in step S22.

[0059] Next, the data processing device 13 generates an image (hereinafter referred to as the "enlarged / reduced image") which is obtained by scaling the object image T001 selected in step S21 with respect to the vanishing point V as the reference point, and which has the largest overlapping area with the object image T003 selected in step S22 (step S23).

[0060] The vanishing point V is a point unique to the camera 11 that captured the image, and it represents the point where the same object captured in the image taken by camera 11 converges as it moves within the camera's field of view while shrinking while the railway vehicle 9 is moving along a straight line. Furthermore, when the same object captured in the image taken by camera 11 moves within the camera's field of view while the railway vehicle 9 is moving along a straight line, the vanishing point V represents the point where the object converges as it moves in the opposite direction (shrinking direction).

[0061] Figure 6 schematically shows enlarged and reduced images of object image T001 corresponding to object image T003, which are generated in step S23.

[0062] Next, the data processing device 13 calculates the degree of overlap between the enlarged / reduced image of object image T001 generated in step S23 and object image T003 (step S24). The degree of overlap calculated in step S24 is, for example, the ratio of the area of ​​the object represented by object image T003 that overlaps with the enlarged / reduced image of object image T001 to the area of ​​the object represented by object image T003. Alternatively, in step S24, the degree of overlap may be calculated as the ratio of the area of ​​the object represented by the enlarged / reduced image of object image T001 that overlaps with object image T003 to the area of ​​the object represented by the enlarged / reduced image of object image T001.

[0063] Next, the data processing device 13 determines whether or not there are any object images included in the captured image in Figure 5(B) that were not selected in step S22 (step S25).

[0064] If there are any unselected object images in step S22 (step S25; Yes), the data processing device 13 returns to step S22, selects one unselected object image (for example, object image T004) from among the captured images in Figure 5(B), and performs the processing from step S23 onwards with respect to the selected object image.

[0065] If there are no object images that were not selected in step S22 (step S25; No), the data processing device 13 determines whether or not there are any object images among the captured images in Figure 5(A) that were not selected in step S21 (step S26).

[0066] If there are any unselected object images in step S21 (step S26; Yes), the data processing device 13 returns to step S21 and selects one unselected object image (for example, object image T002) from among the object images included in the captured image in Figure 5(A), and performs the processing from step S22 onwards with respect to the selected object image. Note that when the data processing device 13 repeats step S22 with respect to the object image newly selected in step S21, all object images included in the captured image in Figure 5(B) are treated as unselected object images.

[0067] If there are no object images that have not been selected in step S21 (step S26; No), the data processing device 13 identifies the object images based on the degree of overlap calculated in step S24 (step S27).

[0068] Here, as an example, let's assume that in step S24, the following degree of overlap (the closer to 1, the greater the degree of overlap) was calculated. Overlap degree between object image T001 and object image T003: 0.84 Overlap between object image T001 and object image T004: 0.32 Overlap between object image T001 and object image T005: 0.00 Overlap between object image T002 and object image T003: 0.00 Overlap between object image T002 and object image T004: 0.77 Overlap between object image T002 and object image T005: 0.00

[0069] The data processing device 13 makes the following determination based on the degree of overlap described above. The object image with the highest degree of overlap with object image T001 is object image T003, and since the degree of overlap is greater than or equal to a threshold (e.g., 0.6), it is determined that object image T001 and object image T003 represent the same object. The object image with the highest degree of overlap with object image T002 is object image T004, and since this overlap is above the threshold, object image T002 and object image T004 represent the same object.

[0070] Next, the data processing device 13 assigns a unique ID to the objects represented by object images T001 and T003, and to the objects represented by object images T002 and T004 (step S28). For example, the objects represented by object images T001 and T003 are assigned the unique ID "G001", and the objects represented by object images T002 and T004 are assigned the unique ID "G002".

[0071] Next, the data processing device 13 stores the original ID assigned in step S28 in the "Original ID" subrecord of the "Object Image" field of the two records selected in step S20 (step S29). Specifically, the data processing device 13 stores the original ID in the management table M as follows.

[0072] In Figure 5(A), the "G001" is stored in the "Main ID" subfield of the subrecord corresponding to the object image T001 of the record corresponding to the captured image. In Figure 5(A), the "G002" is stored in the "Main ID" subfield of the subrecord corresponding to the object image T002 of the record corresponding to the captured image. In Figure 5(B), the "G001" is stored in the "Main ID" subfield of the subrecord corresponding to the object image T003 of the record corresponding to the captured image. In Figure 5(B), the "G002" is stored in the "Main ID" subfield of the subrecord corresponding to the object image T004 of the record corresponding to the captured image.

[0073] Next, the data processing device 13 determines whether or not there are any records in the management table M to be processed that were not selected in step S20 (step S30).

[0074] If there are records in the management table M to be processed that were not selected in step S20 (step S30; Yes), the data processing device 13 returns to step S20, selects the record with the most recent shooting time from the two records selected in the last executed step S20, and the record with the next oldest shooting time, and performs the processing from step S21 onwards on those two selected records.

[0075] However, in the repeated step S28, for the subrecord in which the "Main ID" subfield contains the Main ID for the record with the older shooting time among the two records selected in step S20, the data processing device 13 does not assign a new Main ID to the object represented by the object image corresponding to that subrecord, but uses the Main ID that has already been assigned (the Main ID stored in the "Main ID" subfield).

[0076] If there are no records in the management table M to be processed that were not selected in step S20 (step S30; No), then the first processing for the management table M to be processed is considered complete.

[0077] As a result of the process following the flow shown in Figure 4 above, the "This ID" is stored in the "This ID" subfield of two management tables M, each storing data on images taken by different cameras 11, namely camera 11A and camera 11B, at the same time (hereinafter, of these management tables M, the management table M that stores data on images taken by camera 11A will be called management table M(A), and the management table M that stores data on images taken by camera 11B will be called management table M(B)).

[0078] The data processing device 13 uses the above-mentioned management table M(A) and management table M(B) to perform a process (hereinafter referred to as "second process") that identifies objects identified in the first process within images captured by different cameras 11.

[0079] Figure 7 shows the flow of the second process.

[0080] The data processing device 13 selects a predetermined number (hereinafter referred to as 4) of records from management table M(A) in order of oldest shooting time, and selects a predetermined number (hereinafter referred to as 4) of records from management table M(B) whose shooting time is around the same time as the shooting time of the records selected from management table M(A) (step S40).

[0081] Figure 8 schematically shows the captured image corresponding to the record selected in step S40, and the object image contained in that captured image. Figures 8(A1) to 8(A4) show the captured images taken by camera 11A, and Figures 8(B1) to 8(B4) show the captured images taken by camera 11B.

[0082] Hereinafter, an object identified by this ID "G###" (where "###" is any natural number) will be referred to as "Object G###".

[0083] As shown in Figure 8, the images in Figures 8(A1) to 8(A3) contain an object image representing object G101, and the images in Figures 8(A3) and 8(A4) contain an object image representing object G102. Furthermore, the images in Figures 8(B2) and 8(B3) contain an object image representing object G201, Figure 8(B3) contains an object image representing object G202, and Figures 8(B3) and 8(B4) contain an object image representing object G203.

[0084] The data processing device 13 selects one object to be compared from the objects represented by the object images included in the captured images in Figures 8(A1) to 8(A4) (step S41). Now, let's assume that object G101 was selected in step S41.

[0085] Next, the data processing device 13 selects one object to be compared from the objects represented by the object images included in the captured images in Figures 8(B1) to 8(B4) (step S42). Now, let's assume that object G201 was selected in step S42.

[0086] Next, the data processing device 13 measures the similarity between the object image with the largest area among the object images representing the object selected in step S41 (in this case, the object image representing object G101 included in Figure 8(A1)) and the object image with the largest area among the object images representing the object selected in step S42 (in this case, the object image representing object G201 included in Figure 8(B2)) using known image similarity measurement techniques (step S43). The image similarity measurement technique used in step S43 may be any of the following: techniques that do not use artificial intelligence, such as feature point matching and Bag-of-Visual Words, or techniques that use neural network type (deep learning, etc.) learning models. All of these image recognition techniques use the features of objects in the image.

[0087] Next, the data processing device 13 determines whether or not there are any objects among the objects represented by the object images included in the captured images of Figures 8(B1) to 8(B4) that were not selected in step S42 (step S44).

[0088] If there are any unselected objects in step S42 (step S44; Yes), the data processing device 13 returns to step S42, selects one unselected object (for example, object G202) from among the objects represented by the object images included in the captured images of Figures 8(B1) to 8(B4), and performs the processing from step S43 onwards with respect to the selected object.

[0089] If there are no unselected objects in step S42 (step S44; No), the data processing device 13 determines whether or not there are any objects among the objects represented by the object images included in the captured images of Figures 8(A1) to 8(A4) that were not selected in step S41 (step S45).

[0090] If there are any unselected objects in step S41 (step S45; Yes), the data processing device 13 returns to step S41 and selects one unselected object (for example, object G102) from among the objects represented by the object images included in the captured images of Figures 8(A1) to 8(A4), and then performs the processing from step S42 onward with respect to the selected object. Note that when the data processing device 13 repeats step S42 with respect to the object newly selected in step S41, it treats all objects represented by the object images included in the captured images of Figures 8(B1) to 8(B4) as unselected objects.

[0091] If there are no objects that have not been selected in step S41 (step S45; No), the data processing device 13 identifies the objects based on the similarity measured in step S43 (step S46).

[0092] Here, as an example, let's assume that in step S43, the following similarity scores (the closer to 0, the greater the degree of similarity) were measured. Similarity between object G101 and object G201: 0.28 Similarity between object G101 and object G202: 0.81 Similarity between object G101 and object G203: 0.76 Similarity between object G102 and object G201: 0.98 Similarity between object G102 and object G202: 0.87 Similarity between object G102 and object G203: 0.15

[0093] The data processing device 13 makes the following determination based on the similarity described above. The object with the smallest similarity (highest degree of similarity) to object G101 is object G201, and since its similarity is below the threshold (e.g., 0.4), it is determined that object G101 and object G201 are identical. The object with the smallest similarity (highest degree of similarity) to object G102 is object G203, and since its similarity is below the threshold, it is determined that object G102 and object G203 are identical.

[0094] Next, the data processing device 13 assigns an integrated ID to objects G101 and G201, and objects G102 and G203 to identify them (step S47). For example, assume that objects G101 and G201 are assigned the integrated ID "I001", and objects G102 and G203 are assigned the integrated ID "I002".

[0095] Next, the data processing device 13 stores the integrated ID assigned in step S47 in the "Integrated ID" subrecord of the "Object Image" field of the four records selected from management table M(A) in step S40 and the four records selected from management table M(B) in step S40 (step S48). Specifically, the data processing device 13 stores the integrated ID in management table M as follows.

[0096] In Figure 8(A1), the "I001" is stored in the "Integrated ID" subfield of the subrecord corresponding to object G101 in the record corresponding to the captured image. In Figure 8(A2), the "I001" is stored in the "Integrated ID" subfield of the subrecord corresponding to object G101 in the record corresponding to the captured image. In Figure 8(A3), the "I001" is stored in the "Integrated ID" subfield of the subrecord corresponding to object G101 in the record corresponding to the captured image. In Figure 8(A3), the "I002" is stored in the "Integrated ID" subfield of the subrecord corresponding to object G102 in the record corresponding to the captured image. In Figure 8 (A4), the "I002" is stored in the "Integrated ID" subfield of the subrecord corresponding to object G102 in the record corresponding to the captured image. In Figure 8(B2), the "I001" is stored in the "Integrated ID" subfield of the subrecord corresponding to object G201 in the record corresponding to the captured image. In Figure 8(B3), the "I001" is stored in the "Integrated ID" subfield of the subrecord corresponding to object G201 in the record corresponding to the captured image. In Figure 8(B3), the "I002" is stored in the "Integrated ID" subfield of the subrecord corresponding to object G203 in the record corresponding to the captured image. In Figure 8(B4), the "I002" is stored in the "Integrated ID" subfield of the subrecord corresponding to object G203 in the record corresponding to the captured image.

[0097] Next, the data processing device 13 determines whether or not there are any records in the management table M(A) that were not selected in step S40 (step S49).

[0098] If there are records in the management table M(A) that were not selected in step S40 (step S49; Yes), the data processing device 13 returns to step S40 and selects from the management table M(A) the three most recent records among the four records selected in the last executed step S40, and the record with the next oldest shooting time after the most recent one. From the management table M(B), it selects four records whose shooting times are around the same time as the records selected from the management table M(A), and performs the processing from step S41 onwards with respect to those selected records.

[0099] However, in the repeated step S47, for the six records with the oldest shooting time among the eight records selected in step S40 (three records from management table M(A) and three records from management table M(B)), the data processing device 13 does not assign a new integrated ID to the object corresponding to that subrecord, but uses the already assigned integrated ID (the integrated ID stored in the "integrated ID" subfield).

[0100] If there are no records in management table M(A) that were not selected in step S40 (step S49; No), then the second processing for management table M(A) and management table M(B) is considered complete.

[0101] The data processing device 13 uses known image recognition technology to determine the quality of each object identified by the integrated ID, and stores the data indicating the result of the quality determination in the "Determination Result" subfield of the management table M (Figure 2). Alternatively, instead of using image recognition technology, a user (for example, a maintenance manager of traffic facilities, etc.) may visually inspect the object image, determine the quality of the object represented by the image, and input the result into the management table M.

[0102] Figure 9 shows an example of a screen displayed by a data processing device such as the data processing device 13, using a management table M in which, as described above, the integrated ID is stored in the second process and data indicating the pass / fail judgment result is stored in the "Judgment Result" subfield. Hereinafter, the screen exemplified in Figure 9 will be referred to as the "management screen".

[0103] Area A1 of the management screen displays a selection box where the user selects a travel section. Area A2 of the management screen displays a table (hereinafter referred to as the "summary table") generated from management table M corresponding to the travel section selected by the user in the selection box in area A1. The summary table has, for example, an "integration ID" field, a "type" field, a "distance" field, and a "judgment result" field.

[0104] Area A3 of the administration screen displays multiple selection boxes and a "Filter" button for entering filtering criteria to display a summary table containing only records that meet the user's desired conditions, as shown in Area A2. By entering filtering criteria in the selection boxes in Area A3 and performing a prescribed operation such as clicking the "Filter" button, the user can display a summary report containing only records that meet the desired conditions in Area A2.

[0105] By selecting the record corresponding to the object they wish to view details for from the summary table through a predetermined operation such as clicking, the user can display a map showing the object's location in area A4 of the management screen, and an image of the object in area A5 of the management screen.

[0106] Users can find the location of an object selected from the summary table using the map displayed in area A4 of the administration screen. Additionally, users can check the status of the selected object using the image displayed in area A5 of the administration screen.

[0107] [Differentiation] The embodiments described above may be modified in various ways within the scope of the technical idea of ​​the present invention. Examples of such modifications are shown below. These modifications may be combined as appropriate.

[0108] (1) In the embodiment described above, the number of cameras 11 mounted on the railway vehicle 9 was assumed to be two, but the number of cameras 11 mounted on the railway vehicle 9 may be three or more.

[0109] (2) In the above-described embodiment, the arrangement and shooting direction of the camera 11 mounted on the railway vehicle 9 are not limited to those illustrated in Figure 1, and may be changed in various ways.

[0110] (3) In the embodiments described above, some of the processing that is performed by the data processing device 13 may be performed by a data processing device other than the data processing device 13. For example, the data processing device 13 may store image data received from the camera 11 and data indicating the distance traveled from the starting point (distance) received from the odometer 12 in memory, and while the railway vehicle 9 is stopped at a station, it may transmit this data to a server device via a communication device installed at the station, and the server device may use the data received from the data processing device 13 to perform object recognition processing according to the flow in Figure 3, a first process according to the flow in Figure 4, and a second process according to the flow in Figure 7.

[0111] (4) In the above-described embodiment, the shooting position of the captured image was determined by the distance from the starting point, but the shooting position of the captured image may be determined by means other than distance. For example, the railway vehicle 9 may be equipped with a unit that measures latitude and longitude using GNSS (Global Navigation Satellite System) (GNSS unit), and the latitude and longitude measured by the GNSS unit may be used instead of distance.

[0112] (5) In the embodiment described above, the data processing device 13, in the first process, uses the positional relationship between the vanishing point V in the captured image and the object (object image) in the captured image to identify objects in images captured by the same camera 11 at different times. Alternatively, the data processing device 13 may, in the first process, use the similarity between the images to be compared to identify objects in images captured by the same camera 11 at different times, similar to the second process.

[0113] (6) When the data processing device 13 measures the similarity between the object images to be compared in the second processing, it may also measure the similarity between images extracted from the object image (i.e., partial images of the object image) according to predetermined rules corresponding to the arrangement and shooting direction of different cameras 11.

[0114] For example, if cameras 11A and 11B are arranged as illustrated in Figure 10, the image of object O, which is located to the right of the rail 8, captured by camera 11A will be an image of object O taken from the right side as viewed from the railway vehicle 9. On the other hand, the image of object O, which is the same object O, captured by camera 11B will be an image of object O taken from the left side as viewed from the railway vehicle 9.

[0115] Figure 11 is a schematic diagram showing images of object O captured by camera 11A and camera 11B. Figure 11(A) is an image showing object O as captured by camera 11A, and Figure 11(B) is an image showing object O as captured by camera 11B. As can be seen from Figure 11, in this case, the left part of Figure 11(A) and the right part of Figure 11(B) represent the same part of object O. Therefore, instead of the data processing device 13 measuring the similarity between the entire object image of object O included in the image of Figure 11(A) and the entire object image of object O included in the image of Figure 11(B), the data processing device 13 may measure the similarity between the left part of the object image of object O included in the image of Figure 11(A) and the right part of the object image of object O included in the image of Figure 11(B).

[0116] (7) The data processing device 13 may use the type of object identified in the object recognition process according to the flow in Figure 3 in at least one of the object identification in the first process and the object identification in the second process. That is, in the first process, the data processing device 13 may calculate the degree of overlap between an enlarged or reduced image of one object image and an enlarged or reduced image of the other object image for two object images representing the same type of object. In addition, in the second process, the data processing device 13 may measure the similarity between one object image and the other object image for two object images representing the same type of object.

[0117] (8) The data processing device 13 may use the speed of movement of the railway vehicle 9 in at least one of the first and second processes. In this case, for example, the speed of movement of the railway vehicle 9 (distance traveled per unit time) measured by the odometer 12 may be used. Alternatively, the railway vehicle 9 may be equipped with a GNSS unit, and the speed of movement of the railway vehicle 9 may be calculated from the change in the position of the railway vehicle 9 over time as measured by the GNSS. Furthermore, in the second process, the data processing device 13 may use the speed of movement of the position of an object image representing the same object in multiple images taken by the same camera 11 at different times.

[0118] For example, in Figure 1, when the train car 9 moves from bottom to top, an object positioned to the right of the rail 8 is first photographed by camera 11A and then by camera 11B. The faster the train car 9 moves, the shorter the time between when the same object is captured by camera 11A and when it is captured by camera 11B.

[0119] Furthermore, the further the object moves away from the rail 8, the slower the speed at which the object moves in the image captured by camera 11A, and the slower the speed at which the object moves in the image captured by camera 11B.

[0120] Accordingly, the data processing device 13 may estimate the time between when the same object appears in the image taken by camera 11A and when it appears in the image taken by camera 11B, based on the speed of movement of the railway vehicle 9 and the speed of movement of the position of the object image of the same object in multiple images arranged in chronological order by the time of shooting by camera 11A, or the speed of movement of the position of the object image of the same object in multiple images arranged in chronological order by the time of shooting by camera 11B. The similarity between the object images appearing in the images taken by camera 11A and camera 11B is measured based on the estimated time difference, and the object represented by those object images may be identified based on the measured similarity.

[0121] (9) The data processing device 13 may use the degree of change in the direction of movement of the railway vehicle 9 in at least one of the first and second processes. The degree of change in the direction of movement refers to an index that indicates the degree to which the direction of travel changes, and is a value that can be expressed, for example, as the radius of curvature of the curve of the rail 8.

[0122] Figure 12 shows a railway vehicle 9 traveling on a curved rail 8, and Figure 13 schematically shows how the position of object Q, located to the right of the rail 8, changes in the images captured by camera 11A while the railway vehicle 9 is traveling on the rail 8 shown in Figure 12. Figure 13(A) is an image captured by camera 11A when the railway vehicle 9 is at position P1 in Figure 12, Figure 13(B) is an image captured by camera 11A when the railway vehicle 9 is at position P2 in Figure 12, and Figure 13(C) is an image captured by camera 11A when the railway vehicle 9 is at position P3 in Figure 12.

[0123] As shown in Figure 13, when the railway vehicle 9 changes direction of travel, objects in the captured images taken by the same camera 11 move in a direction corresponding to the degree of change in the direction of travel of the railway vehicle 9 within the captured image.

[0124] Therefore, for example, the data processing device 13 may estimate the movement path of the same object in the same image captured by the same camera 11 based on the degree of change in the direction of travel of the railway vehicle 9, and identify the object represented by the object image included in those captured images based on the estimated movement path.

[0125] The data processing device 13 can, but is not limited to, the following methods to determine the degree of movement of the railway vehicle 9 in the direction of travel at the time the camera 11 takes a photograph.

[0126] (a) The data processing device 13 acquires map information showing the shape of the rail 8 and the distance from the starting point to any point on the rail 8, and based on the map information, identifies the degree of change in the direction of movement of the railway vehicle 9 at a point on the rail 8 corresponding to the distance from the starting point shown by the measurement result of the odometer 12.

[0127] (b) The railway vehicle 9 is equipped with sensors such as acceleration sensors, angular velocity sensors, and geomagnetic sensors that measure information to determine the degree of change in the direction of movement of the railway vehicle 9, and the data processing device 13 determines the degree of change in the direction of movement of the railway vehicle 9 based on the information measured by these sensors.

[0128] (c) A GNSS unit is installed on the railway vehicle 9, and the data processing device 13 determines the degree of change in the direction of movement of the railway vehicle 9 based on the shape of the movement path indicated by the position of the railway vehicle 9, which is continuously measured by the GNSS unit.

[0129] (d) The railway vehicle 9 is equipped with a camera positioned so that the shooting direction is the direction of travel, and the data processing device 13 determines the degree of change in the direction of movement of the railway vehicle 9 based on the shape of the rails 8 in the image captured by the camera.

[0130] (10) The data processing device 13 may use the learning model N1 described below to identify objects in the captured images taken by the same camera 11 in the first processing.

[0131] The following describes the training data used to generate the learning model N1, the explanatory variables input to the learning model N1 during operation, and the target variable output from the learning model N1.

[0132] A series of images taken by camera 11A while the railway vehicle 9 was moving are denoted as image C1, image C2, ..., image Cn (where n is any natural number) in chronological order of capture time.

[0133] [Example 1] (Estimating the region of an object in a subsequent image from the region of an object in a preceding image.) Machine learning is performed using training data containing the following explanatory and dependent variables to generate the trained model N1. However, the following captured image Ck is one of the captured images C1 to C(n-1). (Explanatory variable) In the captured image Ck, the region occupied by the object image representing a certain object (hereinafter referred to as object R1). The speed of movement of railway vehicle 9 at the time when image Ck was taken. The degree of change in the direction of movement of the railway vehicle 9 at the time when the photographed image Ck was taken. (Dependent variable) The region occupied by the object image representing object R1 in the subsequent image C(k+1), which is the image captured after the image Ck.

[0134] During operation, the data processing device 13, in the first processing step, inputs the following information as explanatory variables to the machine learning model N1, which was trained using the aforementioned training data, in order to identify objects within the image captured by the camera 11A.

[0135] (Explanatory variable) In the captured image taken by camera 11A (hereinafter referred to as captured image Dj), the region occupied by the object image representing a certain object (hereinafter referred to as object S1) The speed of movement of railway vehicle 9 at the time when the photographed image Dj was taken. The degree of change in the direction of movement of railway vehicle 9 at the time when image Dj was taken.

[0136] The data processing device 13 inputs the above explanatory variables into the learning model N1 and obtains the following target variable output by the learning model N1 as its response.

[0137] (Dependent variable) The region estimated to be occupied by the object image representing object S1 in the subsequent image D(j+1), which is the image taken after the image Dj.

[0138] The data processing device 13 determines that if an object image exists in the region of the captured image D(j+1) estimated from the information obtained from the learning model N1, that object image represents object S1. The data processing device 13 performs a similar determination for each object image included in the image captured by the camera 11A, thereby identifying objects within the image captured by the camera 11A.

[0139] [Example 2] (Estimating the region of an object in subsequent images from the region of an object in a preceding image) Machine learning is performed using training data containing the following explanatory and dependent variables to generate the trained model N1. However, the following captured image Ck is one of the captured images C1 to C(n-2). (Explanatory variable) The region occupied by the object image representing object R1 in the captured image Ck. The speed of movement of railway vehicle 9 at the time when image Ck was taken. The degree of change in the direction of movement of the railway vehicle 9 at the time when the photographed image Ck was taken. (Dependent variable) The region occupied by the object image representing object R1 in each of the subsequent images C(k+1), C(k+2), ..., which are multiple images taken after the captured image Ck.

[0140] During operation, the data processing device 13, in the first processing step, inputs the following information as explanatory variables to the machine learning model N1, which was trained using the aforementioned training data, in order to identify objects within the image captured by the camera 11A.

[0141] (Explanatory variable) The region occupied by the object image representing object S1 in the captured image Dj taken by camera 11A. The speed of movement of railway vehicle 9 at the time when the photographed image Dj was taken. The degree of change in the direction of movement of railway vehicle 9 at the time when image Dj was taken.

[0142] The data processing device 13 inputs the above explanatory variables into the learning model N1 and obtains the following target variable output by the learning model N1 as its response.

[0143] (Dependent variable) The region estimated to be occupied by the object image representing object S1 in each of the subsequent images D(j+1), D(j+2), ..., which are multiple images taken after the captured image Dj.

[0144] The data processing device 13 determines that if there are object images in each region of the captured images D(j+1), D(j+2), ..., which are estimated based on the information obtained from the learning model N1, then those object images represent object S1. The data processing device 13 performs a similar determination for each object image included in the captured images of the camera 11A, thereby identifying objects within the captured images of the camera 11A.

[0145] [Example 3] (Estimating the region of an object in a subsequent image from the region of an object in several preceding images) Machine learning is performed using training data containing the following explanatory and dependent variables to generate the trained model N1. However, the following captured image Ck is one of the captured images C2 to C(n-1). (Explanatory variable) The region occupied by the object image representing object R1 in each of the captured image Ck and one or more captured images C(k-1), C(k-2), ... that precede the captured image Ck. The speed of movement of railway vehicle 9 at the time when image Ck was taken. The degree of change in the direction of movement of the railway vehicle 9 at the time when the photographed image Ck was taken. (Dependent variable) The region occupied by the object image representing object R1 in the subsequent image C(k+1), which is the image captured after the image Ck.

[0146] During operation, the data processing device 13, in the first processing step, inputs the following information as explanatory variables to the machine learning model N1, which was trained using the aforementioned training data, in order to identify objects within the image captured by the camera 11A.

[0147] (Explanatory variable) The region occupied by the object image representing object S1 in each of the captured image Dj taken by camera 11A and one or more captured images D(j-1), D(j-2), etc. that precede captured image Dj. The speed of movement of railway vehicle 9 at the time when the photographed image Dj was taken. The degree of change in the direction of movement of railway vehicle 9 at the time when image Dj was taken.

[0148] The data processing device 13 inputs the above explanatory variables into the learning model N1 and obtains the following target variable output by the learning model N1 as its response.

[0149] (Dependent variable) The region estimated to be occupied by the object image representing object S1 in the subsequent image D(j+1), which is the image taken after the image Dj.

[0150] The data processing device 13 determines that if an object image exists in the region of the captured image D(j+1) estimated from the information obtained from the learning model N1, that object image represents object S1. The data processing device 13 performs a similar determination for each object image included in the image captured by the camera 11A, thereby identifying objects within the image captured by the camera 11A.

[0151] [Example 4] (Estimating the region of an object in subsequent images from the region of an object in preceding images.) Machine learning is performed using training data containing the following explanatory and dependent variables to generate the trained model N1. However, the following captured image Ck is one of the captured images C2 to C(n-2). (Explanatory variable) The region occupied by the object image representing object R1 in each of the captured image Ck and one or more captured images C(k-1), C(k-2), ... that precede the captured image Ck. The speed of movement of railway vehicle 9 at the time when image Ck was taken. The degree of change in the direction of movement of the railway vehicle 9 at the time when the photographed image Ck was taken. (Dependent variable) The region occupied by the object image representing object R1 in each of the subsequent images C(k+1), c(k+2), ..., which are multiple images taken after the captured image Ck.

[0152] During operation, the data processing device 13, in the first processing step, inputs the following information as explanatory variables to the machine learning model N1, which was trained using the aforementioned training data, in order to identify objects within the image captured by the camera 11A.

[0153] (Explanatory variable) The region occupied by the object image representing object S1 in each of the captured image Dj taken by camera 11A and one or more captured images D(j-1), D(j-2), etc. that precede captured image Dj. The speed of movement of railway vehicle 9 at the time when the photographed image Dj was taken. The degree of change in the direction of movement of railway vehicle 9 at the time when image Dj was taken.

[0154] The data processing device 13 inputs the above explanatory variables into the learning model N1 and obtains the following target variable output by the learning model N1 as its response.

[0155] (Dependent variable) The region estimated to be occupied by the object image representing object S1 in each of the subsequent images D(j+1), D(j+2), ..., which are multiple images taken after the captured image Dj.

[0156] The data processing device 13 determines that if there are object images in each region of the captured images D(j+1), D(j+2), ..., which are estimated based on the information obtained from the learning model N1, then those object images represent object S1. The data processing device 13 performs a similar determination for each object image included in the captured images of the camera 11A, thereby identifying objects within the captured images of the camera 11A.

[0157] Furthermore, multiple learning models N1 from Examples 1 to 4 above may be used in combination. For example, if the object image of the object of interest is not included in any of the preceding captured images, the learning model N1 from Example 1 may be used. If the object image of the object of interest is included in one or more preceding captured images, the learning model N1 from Example 3 may be used to estimate the area occupied by the object image of that object in a subsequent captured image. Alternatively, if the object image of the object of interest is not included in any of the preceding captured images, the learning model N1 from Example 2 may be used. If the object image of the object of interest is included in one or more preceding captured images, the learning model N1 from Example 4 may be used to estimate the area occupied by the object image of that object in each of the subsequent multiple captured images. Also, the learning models N1 from Examples 1 and 3 above may be generated as a single learning model. Similarly, the learning models N1 from Examples 2 and 4 above may be generated as a single learning model.

[0158] (11) The data processing device 13 may use the learning model N2 described below to identify objects in images captured by different cameras 11 in the second processing.

[0159] The N2 learning model is a machine learning model developed using training data that includes the following information as explanatory or dependent variables.

[0160] Explanatory variables: The region occupied by the object image representing an object (hereinafter referred to as object R2) in the image captured by camera 11A in which object R2 is most prominently visible (hereinafter referred to as captured image E), the speed of movement of the railway vehicle 9 at the time captured image E was taken, and the degree of change in the direction of movement of the railway vehicle 9 at the time captured image E was taken. Target variable: The time from the time of capture of image E to the time of capture of the image from camera 11B in which object R2 is most prominently visible (hereinafter referred to as image F), and the region occupied by the object image representing object R2 in image F.

[0161] In the second processing step, the data processing device 13 inputs the following information as explanatory variables to the machine learning model N2, which was trained using the aforementioned training data, in order to identify objects in the images captured by camera 11A and camera 11B.

[0162] Explanatory variables: The region occupied by the object image (hereinafter, the object represented by this object image will be referred to as object S2) in the image captured by camera 11A (hereinafter, image D2), the movement speed of the railway vehicle 9 at the time image D2 was captured, and the degree of change in the direction of movement of the railway vehicle 9 at the time image D2 was captured.

[0163] The data processing device 13 inputs the above explanatory variables into the learning model N2 and obtains the following target variable output by the learning model N2 as its response.

[0164] Target variable: The elapsed time from the time of capture of image D2 to the time of capture of the image from camera 11B in which object S2 is presumed to be present, and the region in the image from camera 11B in which object S2 is presumed to be present that is presumed to be represented by the object image of object S2.

[0165] The data processing device 13 determines that if an object image exists in the region of the image captured by camera 11B, which is identified by the information obtained from the learning model N2, that object image represents object S2. The data processing device 13 performs a similar determination for each object image included in the image captured by camera 11A, thereby identifying objects within the images captured by camera 11A and camera 11B.

[0166] (12) In the embodiments described above, the moving body is assumed to be a railway vehicle, but the type of moving body is not limited to a railway vehicle. [Explanation of Symbols]

[0167] 1...System, 8...Rail, 9...Railway vehicle, 11A...Camera, 11B...Camera, 12...Odometer, 13...Data processing device.

Claims

1. A system that performs a first process of identifying objects in images taken at different times by the same camera mounted on a mobile body, and a second process of identifying the objects identified in the first process with respect to each of the different cameras attached to the mobile body.

2. In the first process described above, the positional relationship between the vanishing point in the image and the object in the image is used. The system according to claim 1.

3. In at least one of the first and second processes, feature quantities of objects in the image are used. The system according to claim 1.

4. In at least one of the first and second processes, the moving speed of the moving body is used. The system according to claim 1.

5. In at least one of the first and second processes, the degree of change in the direction of movement of the moving body is used. The system according to claim 1.

6. The moving object is a railway vehicle, and the degree of change is determined based on the shape of the rails in the image captured by a camera mounted on the moving object. The system according to claim 5.

7. In the second process, the movement speed of the same object identified in the first process is used in multiple images taken at different times by the same camera mounted on the moving body. The system according to claim 1.

Citation Information

Patent Citations

  • Image processing device

    JP2022033627A