Travel path estimation device, travel path estimation method, travel path estimation program, and travel path estimation system
Patent Information
- Application Number
- JP2025023738
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2026-08-27
AI Technical Summary
【0010】 本開示によれば、例えば逃走車両や迷子のような移動経路を推測したい対象が監視カメラなどに撮像されていた場合、撮像した対象の移動経路推測を、ネットワーク情報を必要とせずに可能とする。
Smart Images

Figure 2026137558000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a travel route estimation device, a travel route estimation method, a travel route estimation program, and a travel route estimation system.
Background Art
[0002] Patent Document 1 includes individual cameras installed at each of a plurality of intersections, and an investigation support device communicably connected to the individual cameras. The investigation support device includes an accumulation unit that records the imaging video of the individual cameras in association with the imaging date and time, camera information, and intersection information together with road map image information including the plurality of intersections, and information input including the date and time when an incident or the like occurred, information regarding the intersection, and feature information of a vehicle involved in the occurrence of the incident or the like. A search unit that searches for a vehicle that satisfies the feature information of the vehicle using the imaging video of the camera specified by the camera information corresponding to the intersection; a list output control unit that outputs a list of at least one candidate vehicle extracted by the search to an output unit; and a setting unit that stores, in the accumulation unit, in association with the selected candidate vehicle, the traveling direction when the selected candidate vehicle passes through the intersection in association with the selected candidate vehicle and intersection information; and a route output control unit that outputs the travel route of the selected candidate vehicle to the output unit using the traveling directions when passing through the intersections corresponding to at least two different intersection information stored in the accumulation unit in association with the same candidate vehicle. An investigation support system is described. Patent Document 1 records, in the accumulation unit, the identification information of the camera (an example of camera information) in association with the information of the imaging date and time, and also records road map image information including a plurality of intersections. And intersection camera installation data indicating the correspondence between one or more cameras installed at each intersection and the intersection is recorded. In the investigation support system of Patent Document 1, this road map image information is network information in which map data is node-linked and is data of a huge size.
Prior Art Documents
Patent Documents
[0003] [Patent Document 1] Japanese Patent Publication No. 2019-080295 [Overview of the project] [Problems that the invention aims to solve]
[0004] The purpose of this disclosure is to provide a technology that enables the estimation of the movement path of an object whose movement path we want to predict, such as a runaway vehicle or a lost child, without requiring network information, if the object has been captured on camera or similar equipment. [Means for solving the problem]
[0005] A movement path estimation device according to one aspect of the present disclosure comprises a memory and a processor, the processor working in cooperation with the memory to generate an instruction statement that includes specific information that identifies the positions of multiple cameras that have imaged an object and sequential information of the multiple cameras that have imaged the object, inputs map information and the instruction statement to a VLM (Vision-Language Model), acquires movement path information of the object output from the VLM, and displays it on a predetermined display device.
[0006] A computer-based method for estimating a movement path according to one aspect of this disclosure generates an instruction statement that includes specific information identifying the positions of multiple cameras that have captured an image of an object and sequential information of the multiple cameras that have captured the image of the object. The computer inputs the map information and the instruction statement into a Vision-Language Model (VLM), obtains the movement path information of the object output from the VLM, and displays it on a predetermined display device.
[0007] A movement path estimation program according to one aspect of this disclosure generates an instruction statement that includes specific information that identifies the positions of multiple cameras that have captured an image of an object and sequential information of the multiple cameras that have captured the image of the object, inputs the map information and the instruction statement into a VLM (Vision-Language Model), obtains the movement path information of the object output from the VLM, and instructs a computer to display it on a predetermined display device.
[0008] A movement path estimation system according to one aspect of the present disclosure comprises a plurality of surveillance cameras, a movement path estimation device, and a display device. The movement path estimation device generates an instruction statement that includes specific information identifying the positions of the plurality of surveillance cameras that have captured images of the target, and sequential information of the plurality of surveillance cameras that have captured images of the target. The movement path estimation device inputs map information and the instruction statement into a VLM (Vision-Language Model), obtains movement path information of the target output from the VLM, and displays it on the display device.
[0009] These comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or recording media, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media. [Effects of the Invention]
[0010] According to this disclosure, if, for example, a target whose movement path we want to estimate, such as a getaway vehicle or a lost person, is captured on camera, it becomes possible to estimate the movement path of the captured target without requiring network information. [Brief explanation of the drawing]
[0011] [Figure 1] Block diagram showing an overview of the travel path estimation system according to this embodiment. [Figure 2] This figure shows an example of a map image, input instruction text, and output text according to this embodiment. [Figure 3] This figure shows an example of an input instruction and output statement that uses a map image and sequence information as time information according to this embodiment. [Figure 4] This figure shows an example of input instruction and output statements that take into account the possibility that the map image and input data according to this embodiment may contain image data of multiple similar cars. [Figure 5] This figure shows an example of an input instruction statement and an output statement containing a map image and direction information according to this embodiment. [Figure 6]This figure shows an example of input instructions and output text when the map image and target are people according to this embodiment. [Figure 7] This figure shows an example of input instructions and output statements when a map image and camera number are assigned to a camera according to this embodiment. [Figure 8] This figure shows an example of a map image indicating the camera position, an input instruction, and an output statement according to this embodiment. [Figure 9] This figure shows an example of input instructions and output statements for predicting a travel path according to this embodiment. [Figure 10] Flowchart showing an example of the movement path estimation method according to this embodiment. [Modes for carrying out the invention]
[0012] (Background leading to this disclosure) According to the investigation support system of Patent Document 1, the storage unit records camera identification information (an example of camera information) in association with the date and time of image capture, and also records road map image information including multiple intersections. Furthermore, it records intersection camera installation data that shows the correspondence between one or more cameras installed at each intersection and that intersection. This road map image information is network information in which map data has been node-linked, and is a massive amount of data.
[0013] Such network information is map data broken down into nodes and links, with each node and link having various attributes. For example, it includes not only names such as road type, road name (number), and location (intersection) name, but also various information such as access restrictions by direction, road width, number of lanes, and directions for each lane. Furthermore, pedestrian map image information includes information indicating whether there is a sidewalk (or only a side road), crosswalk data, floor information such as whether it is a staircase, escalator, or elevator, and building entrance information. In other words, when using network information, the amount of data is enormous, resulting in high data costs, processing load, and processing time. Therefore, there is a need for technology that can predict a target's movement route without using network information.
[0014] Hereinafter, embodiments of the present disclosure will be described in detail with appropriate reference to the drawings. However, detailed descriptions that are more than necessary may be omitted. For example, detailed descriptions of well-known matters and redundant descriptions of substantially the same configurations may be omitted. This is to avoid making the following description unnecessarily redundant and to facilitate the understanding of those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter recited in the claims.
[0015] (This embodiment) (Regarding the movement path estimation system) First, the movement path estimation system 1 will be described. FIG. 1 is a diagram showing an example of the configuration of the movement path estimation system according to this embodiment.
[0016] The movement path estimation system 1 includes a movement path estimation device 10, a display device 14, a surveillance camera 15, a video reception and distribution device 16, a video storage device 18, and an image analysis processing device 19. The movement path estimation device 10, the display device 14, the video reception and distribution device 16, the video storage device 18, and the image analysis processing device 19 do not necessarily need to be individual devices, and all or a part thereof may be included as an integrated one device. The movement path estimation device 10, the display device 14, the surveillance camera 15, the video reception and distribution device 16, the video storage device 18, and the image analysis processing device 19 can transmit and receive information and data to and from each other via a wired cable or a communication network. Examples of the wired cable include an HDMI (registered trademark) cable, a USB cable, etc. Examples of the communication network include a wired LAN, a wireless LAN, a mobile communication network (e.g., LTE, 4G, 5G), the Internet, etc.
[0017] The movement path estimation device 10 includes a processor 11 and a memory 12. The processor 11 realizes various functions by executing a program held in the memory 12. Examples of the processor 11 include a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a controller, and an LSI (Large Scale Integration). The processor 11 may include a GPU (Graphics Processing Unit) and / or an NPU (Neural network Processing Unit). The display device 14, the monitoring camera 15, the video reception and distribution device 16, the video storage device 18, and the image analysis processing device 19 may also be provided with similar processors.
[0018] As a function, the processor 11 includes an instruction generation unit 110, an output unit 113, and a route drawing unit 112. However, when route drawing is not required, the route drawing unit 112 may not be provided.
[0019] The instruction generation unit 110 generates an instruction (prompt) to be input to the VLM 111. The VLM 111 is a Vision-Language Model, which is an AI (Artificial Intelligence) model that simultaneously processes visual information such as images and videos and language information such as text. The VLM 111 is also called a large-scale vision language model. Examples of visual information include RGB images, depth information, segmentation maps, and the like. Language information includes text data such as descriptions, captions, and questions related to images. The VLM 111 extracts the features of the input data including visual information and language information using various neural networks corresponding to each. Then, the VLM 111 performs multimodal processing on the features of the visual information and the language information. And based on the information integrated by the multimodal processing, the VLM 111 can give an appropriate answer to a question about an object in the image.
[0020] In other words, in this embodiment, map image information 121, an instruction statement including identification information of multiple surveillance cameras 15 that captured images of a target (vehicle, person, etc.) and sequence information of the multiple cameras that captured images of the target (vehicle, person, etc.), are input to the VLM 111, and the VLM 111 outputs information about the target's movement path. This makes it possible to estimate the movement path of an object whose movement path we want to estimate, such as a runaway vehicle or a lost person, if it has been captured by a surveillance camera 15 or the like, without requiring network information.
[0021] The output unit 113 acquires information such as the output shown in Figures 2 to 9, which will be described later, from the VLM 111, and instructs the route drawing unit 112 to output based on the acquired information, or causes the display device 14 to display a screen showing the acquired information.
[0022] The route drawing unit 112 is an optional function that interprets the target's estimated travel path information (text information) as shown in Figures 2 to 9 (described later), and draws the route on the display device 14 by superimposing it on the separately acquired map image information 121.
[0023] Memory 12 stores camera identification information 122 and stores computer programs and data handled by the processor 11. Memory 12 may be composed of a volatile or non-volatile storage medium and may include, for example, ROM (Read-Only Memory) and / or RAM (Random Access Memory). The display device 14, surveillance camera 15, video receiving and distribution device 16, video storage device 18, and image analysis processing device 19 may also be equipped with similar memories.
[0024] Map image information 121 is a map image of an arbitrary area. The processor 11 may obtain map image information 121 from an external server. Map image information 121 may be a map image that includes outdoor road information, or it may be indoor map image information. The map image may also include identifying information (such as numbers, IDs, or names) and location information of multiple surveillance cameras 15.
[0025] Camera identification information 122 is information that associates the location of the surveillance camera 15 with a target object on the map (for example, a store or a bridge), and is stored in memory 12 in advance. Since the target object may change (for example, the surveillance camera installed in a store may change), the processor 11 may periodically update the camera identification information 122 by referring to the map information.
[0026] The instruction text generation unit 110 generates instruction text to be input to the VLM 111. The instruction text will be described later with reference to Figures 2 to 8. The instruction text may be directly input by the user using the instruction text generation unit 110. Alternatively, the instruction text may be automatically generated by the instruction text generation unit 110 from information such as the surveillance camera 15 or pre-prepared videos and still images. In the case of automatic generation, a template may be prepared in advance, and information extracted by image processing from the surveillance camera 15 or pre-prepared videos and still images may be included in the instruction text.
[0027] The display device 14 displays the output information generated by the processor 11. The display method may include displaying on a screen or outputting via sound.
[0028] The surveillance camera 15 may be a general-purpose surveillance camera. The video receiving and distribution device 16 transmits the video from the surveillance camera 15 and the video files 17 stored in the video storage device 18 to the image analysis processing device 19. The processor 11 may receive information from multiple surveillance cameras 15 that have captured images of the target, either directly or indirectly, indicating that the target was captured (detected) and the time of capture (detection time and time period).
[0029] The image analysis processing unit 19 extracts features from images such as still images and videos and transmits the necessary information to the instruction statement generation unit 110. For example, it extracts information from the video footage of the surveillance camera 15, such as whether or not the target vehicle was captured, the time of capture, the direction the target was heading, and its orientation.
[0030] Next, we will describe the use cases with reference to Figures 2 to 8. Figure 2 shows an example of a map image, input instruction, and output statement according to this embodiment. Figure 3 shows an example of a map image, input instruction, and output statement according to this embodiment, where the sequence information is used as time information. Figure 4 shows an example of a map image, input instruction, and output statement according to this embodiment, considering the possibility that the input data may contain image data of multiple similar cars. Figure 5 shows an example of a map image, input instruction, and output statement that includes direction information according to this embodiment. Figure 6 shows an example of a map image, input instruction, and output statement when the target is a person according to this embodiment. Figure 7 shows an example of a map image, input instruction, and output statement according to this embodiment when a camera number is assigned to the camera. Figure 8 shows an example of a map image, input instruction, and output statement according to this embodiment, where the camera position is indicated.
[0031] In Figures 2 to 8, 'input' is the instruction text input to VLM111, and 'output' is the output text from VLM111. As shown in Figure 2, it is advisable to attach a map to the instruction text and indicate that "the attached image is a map." In the instruction text, the target is identified as "a car," and the targets on the map where multiple surveillance cameras 15 that captured images of the target are installed are identified as "convenience store Ikebe, supermarket Pana Ikebe, and surveillance cameras at Kamoike Ohashi." Since the locations of convenience store Ikebe, supermarket Pana Ikebe, and Kamoike Ohashi are clearly indicated in the map image, the map location information of the surveillance cameras 15 can also be identified. Although the map image itself does not show the location of the surveillance cameras 15, the camera identification information 122 stored in memory 12 associates the surveillance cameras 15 with the targets, so the instruction text generation unit 110 can refer to the camera identification information 122 and generate an instruction text that identifies the targets where the surveillance cameras 15 are installed. The camera identification information 122 may include information indicating the orientation of each surveillance camera 15. In this case, the instruction generation unit 110 may refer to the camera identification information 122 and include the orientation of the surveillance cameras 15 installed on each target object in the instruction. The sequence information of the multiple cameras that captured images of the target identifies that "the same car was captured in the order of the surveillance cameras at Convenience Store Ikebe, Supermarket Pana Ikebe, and Kamoike Ohashi Bridge." Finally, the instruction "Predict and tell me the driving path of this car" is included. As a result, the VLM 111 outputs the information shown in Figure 2. The travel path information may also be displayed and output superimposed on the map image, as shown by the arrows in Figure 2.
[0032] Next, let's explain the example in Figure 3. Note that in the map images shown in Figures 2 to 8, roads are represented in gray. In Figure 3, the instructions tell VLM111 how to read the map, stating that "roads are represented by gray straight lines and curves" and "the scale is approximately 550m, based on the length of the straight line connecting JR Kamoi Station to the Ikebe convenience store." It is helpful to identify the color and shape of objects in the map and explain what those objects represent, such as roads or stations. In addition to route estimation, you may also instruct the system to extract vehicles that are presumed to be different from the target vehicle from among the vehicles captured by multiple surveillance cameras 15.
[0033] As identifying and sequential information for the multiple cameras that captured the target, each of the multiple surveillance cameras 15 that captured the target indicates the date and time (imaging time) in which the target was captured. Specifically, it may be indicated as follows: "The same car was captured by surveillance cameras at Convenience Store Ikebe on December 1, 2024 at 15:03:00, Supermarket Pana Ikebe on December 1, 2024 at 15:04:00, Kamoike Ohashi on December 1, 2024 at 15:05:01, and Drugstore Pana Kamoi on December 1, 2024." The date and time information of the imaging should preferably be in units (hours, minutes, seconds, etc.) that indicate the order in which the multiple surveillance cameras 15 captured the image.
[0034] Furthermore, as shown in Figure 3, the instruction may include conditions for path estimation. Specifically, it may include the speed of the target vehicle, how to distinguish it from similar vehicles, and how to output if it is determined to be a different vehicle. It may also specify the format of the output text (state the conclusion first, followed by a brief explanation of the reason). As a result, VLM111 estimates and outputs a vehicle different from the target vehicle, as shown in Figure 3.
[0035] Next, let's explain the example in Figure 4. In the instruction text in Figure 4, in addition to the instruction text in Figure 3, the direction of travel of the vehicle is indicated. That is, the instruction text may include the direction of movement of the object. Here, as shown in "input" in Figure 4, the direction is indicated not by east, west, north, and south, but by the up, down, left, and right directions on the attached map using arrows. As a result, as shown in "output" in Figure 4, the output text from VLM111 also clearly indicates the direction using arrows.
[0036] As shown in Figures 3 and 4, the accuracy of the estimation is improved if vehicles other than the target vehicle are extracted first, and then the movement path estimation is performed using only the target vehicle. In other words, if VLM111 identifies information that is presumed to be erroneous among the information input to VLM111, it may generate and output movement path information without using the identified information.
[0037] Next, let's explain the example in Figure 5. In the instructions in Figure 5, the top, right, bottom, and left of the image correspond to the directions north, east, south, and west, respectively. In other words, the instructions may include interpretive information about the map image, such as directions. The conditions for estimating the travel path are also shown: "Predict the travel path of this car assuming that the car travels rationally using only roads to travel the shortest distance." As a result, the output from VLM111 also shows a travel path that meets the conditions and indicates directions using east, west, north, and south.
[0038] Next, let's explain the example in Figure 6. Note that in the map images shown in Figures 2 to 8, railway station premises are colored red. The instruction in Figure 6 instructs the system to estimate a person's movement path. Since a person's movement path may include walking within station premises as well as roads, the instruction may include interpretation information for the map image, such as "roads are represented by gray straight lines and curves, and railway station premises are colored red." As a result, the output from VLM111 will also reflect the movement path estimation, taking into account the possibility of walking within the station premises.
[0039] Next, let's explain the example in Figure 7. In the instruction text in Figure 7, each of the surveillance cameras 15 that captured the target is assigned a number as specific information. Specifically, the assignments are as follows: "The correspondence between the cameras and the location names on the map is as follows: Camera 1 - Convenience store Ikebe, Camera 2 - Supermarket Pana Ikebe, Camera 3 - Kamoike Ohashi." As a result, the output text from VLM111 is also expressed by assigning numbers to the cameras.
[0040] Next, let's explain the example in Figure 8. In the map image in Figure 8, the location of surveillance camera 15 is displayed as a circle along with the number of surveillance camera 15. As identifying information for surveillance camera 15, a number is assigned to surveillance camera 15. In this way, location information and identifying information of surveillance cameras may be added to the map image. As a result, the output statement from VLM111 will also be output based on this.
[0041] Figure 9 shows an example of input and output statements for predicting a travel path according to this embodiment. Next, the example in Figure 9 will be explained.
[0042] As shown in the input of Figure 9, the instruction text indicates how to recognize the railway and stations on the map, that surveillance cameras are installed at each station, and that the surveillance cameras record images of passengers. Furthermore, the instruction text in Figure 9 indicates that the same passenger has been recorded by surveillance cameras at Kamoi Station and Nakayama Station. The instruction text in Figure 9 then instructs the system to tell the system, in order of likelihood, which station's recording footage should be checked and when, in order to find the next recorded footage of this passenger. In other words, the instruction text generation unit 110 may generate an instruction text that predicts the movement path of an object (person or car, etc.) from the images of two or more surveillance cameras and prompts the system to answer which surveillance camera's footage should be checked. As a result, as shown in the output of Figure 9, an output text can be obtained that predicts the movement path and time of the object and indicates the surveillance camera's recording footage that should be checked.
[0043] While Figure 9 shows an example involving a railway and its passengers, the content of Figure 9 can also be applied to automobiles, pedestrians, and other similar situations.
[0044] Furthermore, in the prediction function shown in Figure 9, assumptions (conditions) appropriate to the subject being predicted may be included in the instruction statement. For example, when tracking a crime suspect, an assumption such as "they should be moving away from the scene" may be given in the instruction statement. For example, when tracking a lost child, an assumption such as "they are only moving on foot" may be given in the instruction statement. By giving such assumptions (conditions) to the instruction statement, the accuracy of the prediction can be improved or the possibilities can be narrowed down.
[0045] <Other variations> For example, if a user marks an object of interest in the surveillance camera footage and instructs the system to "predict a route" for that object, the movement path estimation device 10 may select one instruction from a pre-prepared set of instruction templates and input that instruction into the VLM 111. In addition to directly marking the footage, the user may input characteristics of an object based on eyewitness or investigation information (e.g., gender, age, height, clothing) as search conditions, and then, based on the search results, the user instructs the system to "predict a route." The movement path estimation device 10 may similarly input the selected instruction into the VLM 111. The user may select which of the multiple instruction templates to input into the VLM 111 from a selection screen, or the instruction generation unit 110 may automatically select it based on the conditions under which the object was captured, tracking conditions, etc.
[0046] Next, the movement path method and movement path program will be described. Figure 10 is a flowchart showing an example of the movement path estimation method according to this embodiment. The order of the flow may be changed unless there are constraints. The following flow may be executed by the computer running the program.
[0047] The processor 11 receives an instruction statement and map information, which includes identification information of multiple cameras that captured images of the target and sequence information of the multiple cameras that captured images of the target (ST1). This instruction statement is automatically or manually generated by the instruction statement generation unit 110 based on images and videos of the target captured by the surveillance camera 15. The processor 11 then inputs the received instruction statement and map information into the VLM 111 (ST2). The VLM 111 outputs information about the target's movement path (ST3). The processor 11 then transmits the outputted movement path information to the display device 14 (ST4). The display device 14 displays this movement path information.
[0048] (Summary of this embodiment) Based on the above description of embodiments, the following technologies are disclosed.
[0049] <Technology 1> A movement path estimation device 10 according to one embodiment comprises a memory 12 and a processor 11. The processor 11 works in cooperation with the memory 12 to generate an instruction statement that includes specific information that identifies the positions of multiple cameras (surveillance cameras 15) that have captured images of the target, and information about the order of the multiple cameras (surveillance cameras 15) that have captured images of the target. The processor 11 inputs the map information and the instruction statement to a VLM (Vision-Language Model) 111, obtains the movement path information of the target output from the VLM 111, and displays it on a predetermined display device.
[0050] This makes it possible to estimate the movement path of an object whose movement route you want to predict, such as a getaway vehicle or a lost person, if it has been captured on surveillance camera footage, without requiring network information.
[0051] <Technology 2> In the movement path estimation device 10 described in Technical 1, the processor 11 receives information on the imaging of the target and imaging time information from a plurality of cameras (surveillance cameras 15) that have imaged the target.
[0052] This makes it possible to estimate the movement path of an object whose movement route we want to predict, such as a fleeing vehicle or a lost person, if it is captured on surveillance camera footage, without requiring network information and more efficiently.
[0053] <Technology 3> In the travel path estimation device 10 described in any one of the technologies 1 to 3, the map information includes roads.
[0054] This makes it possible to estimate the movement path of an object whose movement route you want to predict, such as a runaway vehicle or a lost person, if it is captured by a surveillance camera outdoors, without requiring network information.
[0055] <Technology 4> In the travel path estimation device 10 described in any one of the technologies 1 to 3, the map information is indoor map image information.
[0056] This makes it possible to estimate the indoor movement path of an object whose movement route you want to estimate, such as a lost child, if it is captured by a surveillance camera indoors, without requiring network information.
[0057] <Technology 5> In the travel path estimation device 10 described in any one of the technologies 1 to 4, if the VLM 111 identifies information that is presumed to be erroneous among the information input to the VLM 111, it generates and outputs the travel path information without using the identified information.
[0058] This allows for more accurate prediction of the movement path of an imaged object, without being misled by misinformation, and without requiring network information.
[0059] <Technology 6> In the movement path estimation device 10 according to any one of technologies 1 to 5, the instruction statement further includes the direction of movement of the object.
[0060] This allows for more accurate estimation of the movement path of an object whose movement route we want to predict, such as a getaway vehicle or a lost person, if it is captured on surveillance camera footage, by taking direction into account without requiring network information.
[0061] <Technology 7> In the travel path estimation device 10 described in any one of the technologies 1 to 6, the instruction statement includes interpretation information of the map image.
[0062] This allows for more accurate estimation of the movement path of an object whose route you want to predict, such as a getaway vehicle or a lost person, if it is captured on surveillance camera footage, by incorporating the interpretation of a map proposed by the user, without requiring network information.
[0063] <Technology 8> In the travel path estimation device 10, according to any one of technologies 1 to 7, the travel path information is superimposed on the map image and displayed on the display device 14.
[0064] This makes it possible to predict the movement path of an object whose movement route you want to estimate, such as a getaway vehicle or a lost person, if it is captured on surveillance camera footage, without requiring network information, and in a more easily understandable output.
[0065] <Technology 9> One embodiment of the computer-based movement path estimation method involves generating an instruction statement that includes specific information identifying the positions of multiple cameras (surveillance cameras 15) that have captured images of a target, and sequential information of the multiple cameras (surveillance cameras 15) that have captured images of the target. The map information and the instruction statement are input to a VLM (Vision-Language Model) 111, the movement path information of the target output from the VLM 111 is acquired, and the information is displayed on a predetermined display device 14.
[0066] This makes it possible to estimate the movement path of an object whose movement route you want to predict, such as a getaway vehicle or a lost person, if it has been captured on surveillance camera footage, without requiring network information.
[0067] <Technology 10> A movement path estimation program according to one embodiment generates an instruction statement that includes specific information that identifies the positions of multiple cameras (surveillance cameras 15) that have captured images of a target, and sequential information of the multiple cameras (surveillance cameras 15) that have captured images of the target. The program inputs the map information and the instruction statement into a VLM (Vision-Language Model) 111, retrieves the movement path information of the target output from the VLM 111, and instructs the computer to display it on a predetermined display device 14.
[0068] This makes it possible to estimate the movement path of an object whose movement route you want to predict, such as a getaway vehicle or a lost person, if it has been captured on surveillance camera footage, without requiring network information.
[0069] <Technology 11> A movement path estimation system 1 according to one embodiment comprises a plurality of surveillance cameras 15, a movement path estimation device 10, and a display device 14. The movement path estimation device 10 generates an instruction statement that includes specific information identifying the positions of the plurality of surveillance cameras 15 that have captured images of the target, and sequential information of the plurality of surveillance cameras 15 that have captured images of the target. The device inputs the map information and the instruction statement to a VLM (Vision-Language Model) 111, obtains the movement path information of the target output from the VLM 111, and displays it on the display device 14.
[0070] This makes it possible to estimate the movement path of an object whose movement route you want to predict, such as a getaway vehicle or a lost person, if it has been captured on surveillance camera footage, without requiring network information.
[0071] While embodiments have been described above with reference to the attached drawings, this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the embodiments described above can be combined in any way without departing from the spirit of the invention. [Industrial applicability]
[0072] The technology disclosed herein can, for example, extract comparison locations that are more relevant to a target location from among multiple stores in the retail industry, and display work plans for the target location and the extracted comparison locations. [Explanation of symbols]
[0073] 1. Travel Path Estimation System 10. Movement path estimation device 11 processors 12 memory 14 Display device 15 surveillance cameras 16. Video receiving and distribution device 17 video files 18. Video storage device 19 Image Analysis Processing Device 110 Directive sentence generation section 111 VLM 112 Route drawing section 113 Output section 121 Map Image Information 122 Camera Identification Information
Claims
1. Equipped with memory and a processor, The aforementioned processor cooperates with the memory, An instruction statement is generated that includes specific information that identifies the positions of multiple cameras that captured the target, and sequential information of the multiple cameras that captured the target. The map information and the aforementioned instruction text are entered into the VLM (Vision-Language Model), The movement path information of the target output from the VLM is acquired and displayed on a predetermined display device. A device for estimating movement paths.
2. The processor receives information about the subject being captured and information about the time of capture from multiple cameras that captured the subject. A travel path estimation device according to claim 1.
3. The aforementioned map information includes roads, A travel path estimation device according to claim 1 or 2.
4. The aforementioned map information is indoor map image information. A travel path estimation device according to claim 1 or 2.
5. If the VLM identifies information that is presumed to be erroneous among the information input to the VLM, it generates and outputs the travel path information without using the identified information. A travel path estimation device according to claim 1 or 2.
6. The instruction further includes the direction of movement of the object, A travel path estimation device according to claim 1 or 2.
7. The aforementioned instruction includes information interpreting the map information, A travel path estimation device according to claim 1 or 2.
8. The aforementioned travel route information is superimposed on the map information and displayed on the display device. A travel path estimation device according to claim 1 or 2.
9. A computer-based method for estimating movement paths, An instruction statement is generated that includes specific information that identifies the positions of multiple cameras that captured the target, and sequential information of the multiple cameras that captured the target. The map information and the aforementioned instruction text are entered into the VLM (Vision-Language Model), The movement path information of the target output from the VLM is acquired and displayed on a predetermined display device. A method for estimating movement paths.
10. An instruction statement is generated that includes specific information that identifies the positions of multiple cameras that captured the target, and sequential information of the multiple cameras that captured the target. The map information and the aforementioned instruction text are entered into the VLM (Vision-Language Model), The computer is instructed to acquire the target's movement path information output from the VLM and display it on a predetermined display device. A program for estimating travel paths.
11. A movement path estimation system comprising multiple surveillance cameras, a movement path estimation device, and a display device, The aforementioned movement path estimation device is An instruction statement is generated that includes specific information that identifies the location of the multiple surveillance cameras that captured the target, and sequential information of the multiple surveillance cameras that captured the target. The map information and the aforementioned instruction text are entered into the VLM (Vision-Language Model), The movement path information of the target output from the VLM is acquired and displayed on the display device. A system for estimating travel paths.
Citation Information
Patent Citations
Investigation assist system and investigation assist method
JP2019080295A