Information processing device, information processing method, and recording medium

The information processing device accurately classifies accident videos by generating target description text and using a language model to reduce the influence of training data, improving event identification and classification accuracy.

WO2026004708A1PCT designated stage Publication Date: 2026-01-02NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/021897
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-24
Filing Date
2025-06-18
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing information processing devices fail to accurately classify videos as accident videos based on their content, particularly due to limitations in training data influencing event identification.

Method used

An information processing device and method that includes a text acquisition unit to generate target description text from a video and a text processing unit to classify the video as an accident video using a language model, reducing the influence of training data on event identification.

Benefits of technology

Accurately classifies target videos as accident videos by generating output information indicating whether the video is related to an accident, enhancing the accuracy of event identification and enabling easy recognition of accident types and scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025021897_02012026_PF_FP_ABST
    Figure JP2025021897_02012026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device is provided with a text acquisition unit and a text processing unit. The text acquisition unit acquires, from a target video, a target explanatory text that is a text describing the target video. The text processing unit generates, on the basis of the acquired target explanatory text, output information relating to the target video and including whether the target video is an accident video relating to an accident.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and recording medium

[0001] The present invention relates to an information processing device, an information processing method, and a recording medium.

[0002] For example, Patent Document 1 discloses an information processing device that allows a user to intuitively grasp the recognition results of image information. The information processing device described in Patent Document 1 includes an object information acquisition unit that acquires object information that is information related to a predetermined object included in an image, and a sentence generation unit that generates a sentence that describes the state of the object included in the image using the object information acquired by the object information acquisition unit.

[0003] Special Publication No. 2023-524608

[0004] However, the information processing device described in Patent Document 1 does not disclose a technique for classifying video to be processed (target video) based on whether it is video related to an accident (accident video).

[0005] One of the objectives of the present disclosure is to accurately classify target videos relating to various accidents as accident videos.

[0006] The information processing device of the present disclosure includes a text acquisition means for acquiring a target description text, which is text that explains the target video generated using the target video, and a text processing means for generating output information regarding the target video, including whether or not the target video is an accident video related to an accident, based on the acquired target description text.

[0007] The information processing method of the present disclosure involves one or more computers acquiring a target description text, which is text that describes a target video generated using the target video, and generating output information regarding the target video, including whether the target video is an accident video related to an accident, based on the acquired target description text.

[0008] The program in the present disclosure is a program for causing one or more computers to acquire a target description text, which is text that describes a target video generated using the target video, and generate output information regarding the target video, including whether or not the target video is an accident video related to an accident, based on the acquired target description text.

[0009] According to the present disclosure, it becomes possible to accurately classify target videos relating to various accidents as accident videos.

[0010] FIG. 1 is a block diagram showing a configuration example of a first information processing device in the present disclosure. FIG. 2 is a flowchart showing an operation example of a first information processing device in the present disclosure. FIG. 3 is a block diagram showing a physical configuration example of a first information processing device in the present disclosure. FIG. 4 is a block diagram showing a configuration example of a first information processing system including a second information processing device in the present disclosure. FIG. 5 is a flowchart showing an operation example of a second information processing device in the present disclosure. FIG. 6 is a flowchart showing an operation example of a second information processing device in the present disclosure.

[0011] Hereinafter, embodiments according to the present disclosure will be described with reference to the drawings. In all drawings, similar components are denoted by similar reference numerals, and descriptions thereof will be omitted as appropriate. In the present disclosure, the drawings relate to one or more embodiments.

[0012] First Embodiment An information processing device 100 includes a text acquisition unit 110 and a text processing unit 120, as shown in FIG.

[0013] The text acquisition unit 110 acquires a target explanation text, which is a text that explains the target video and is generated using the target video.

[0014] The text processing unit 120 generates output information about the target video, including whether or not the target video is an accident video about an accident, based on the acquired target description text.

[0015] According to this information processing device 100, it is possible to generate output information relating to a target video, including whether or not the target video is an accident video relating to an accident, from the target video via the target explanation text.

[0016] For example, if a machine learning model is used to generate output information from a target video that includes whether the target video is an accident video related to an accident, the events that can be identified as accidents may be limited depending on the training data used to train the machine learning model.

[0017] As described above, the information processing device 100 generates output information of a target video from the target video via the target explanatory text, thereby reducing the possibility that events that can be identified as accidents will be influenced by the training data. Therefore, it becomes possible to accurately classify target videos related to various accidents as accident videos.

[0018] The information processing device 100 executes information processing such as that shown in FIG.

[0019] The text acquisition unit 110 acquires a target explanation text, which is a text that explains the target video and is generated using the target video (step S110).

[0020] The text processing unit 120 generates output information about the target video, including whether or not the target video is an accident video about an accident, based on the acquired target description text (step S120).

[0021] This information processing allows for the generation of output information about the target video, including whether the target video is an accident video or not, from the target video via the target description text. As described above, this reduces the possibility that events that can be identified as accidents will be influenced by the training data. Therefore, it becomes possible to accurately classify target videos related to various accidents as accident videos.

[0022] A detailed example of this embodiment will be described below.

[0023] (Regarding target video) The target video is video that is to be processed by the information processing device 100. The target video may include, for example, at least one of video captured by a camera mounted on a vehicle and video captured of a road on which the vehicle is traveling.

[0024] The video captured by the camera mounted on the vehicle may include, for example, a video of the surroundings of the vehicle, or a video of the interior of the vehicle. The camera mounted on the vehicle may be, for example, a drive recorder, but is not limited to this.

[0025] The video of the road may include, for example, video captured by a camera installed on the road or around the road. The area around the road may include, for example, buildings around the road. Note that the camera used to capture the road is not limited to the example shown here.

[0026] (Text Acquisition Unit 110) As described above, the text acquisition unit 110 acquires the target explanation text of the target video. The target explanation text is text (sentence) that explains the target video.

[0027] For example, the text acquisition unit 110 acquires a target description text generated using the target video and a video analysis model.

[0028] The video analysis model is, for example, a machine learning model that generates explanatory text from input information including video. The input information may further include accompanying information, which is information accompanying the video. The explanatory text is text (sentence) that explains the input information. The explanatory text includes, for example, text that explains the video. A machine learning model that handles images and text in a composite manner in this way is also called a visual language model (VLM), etc.

[0029] The video analysis model is a trained machine learning model that has been trained to generate explanatory text from input information including video. The data (training data) used for this training may include, for example, training video and related correct answers. The training video may include video of the same type as the target video, i.e., at least one of video captured by a camera mounted on a vehicle and video of a road on which the vehicle travels. The training video may also include video of a different type from the target video, i.e., video captured by a camera other than the camera mounted on a vehicle, video other than video of a road on which the vehicle travels, etc.

[0030] Additionally, if the input information further includes accompanying information, the video analysis model may generate explanatory text from the video and the accompanying information. In this case, the training data may include, for example, information accompanying the training video.

[0031] For example, when target video information including a target video is input as input information, the video analysis model outputs a target description text. The target description text is text (sentence) that describes the target video information. The target description text includes, for example, text that describes the target video. The target description text may include accompanying information that accompanies the target video, as described below.

[0032] For example, if the target video is a video of an accident scene, the target description text may include a description of the accident. The description of the accident may include at least one of the details of the accident, the circumstances when the accident occurred, and the events leading up to the accident.

[0033] The text acquisition unit 110 may include a video analysis model. The text acquisition unit 110 may acquire the target explanatory text using this video analysis model. Alternatively, the text acquisition unit 110 may acquire the target explanatory text using a video analysis model provided in an external device (not shown) other than the information processing device 100. In this case, the external device may be communicably connected to the information processing device 100 via a communication network configured by wired, wireless, or a combination thereof. The text acquisition unit 110 may then transmit input information including a video to the external device and acquire the target explanatory text generated by the external device using the video analysis model.

[0034] (Regarding Accompanying Information) The target video information may further include accompanying information, which is information accompanying the target video. The video analysis model may generate explanatory text from the video and the accompanying information accompanying the video. For example, a prompt to the video analysis model may include an instruction including the accompanying information and the target video.

[0035] The accompanying information may include, for example, at least one of the following: the time the target image was captured, the vehicle's driving control information, the vehicle's location information when the target image was captured, and road information for the road on which the vehicle was traveling when the target image was captured.

[0036] The shooting time may be generated, for example, by a shooting device that generates the target video by shooting.

[0037] The driving control information may include, for example, at least one of the vehicle speed, acceleration, steering operation, braking operation, and accelerator operation. Note that the driving control information is not limited to the examples given here. For example, the driving control information may be generated by a vehicle control device (not shown) mounted on the vehicle, but the device that generates the driving control information is not limited to this.

[0038] The location information may be generated by a device (such as an imaging device or a vehicle control device) mounted on the vehicle using a GPS (Global Positioning System) function, etc. However, the device that generates the location information is not limited to this.

[0039] Road information is information relating to roads. For example, when a vehicle is traveling near an intersection, the road information may include at least one of the type of the intersection and the attributes of the road on which the vehicle is traveling. The type of intersection may include at least one of at-grade intersection, multi-level intersection, etc., but is not limited to these. The attributes of the road may include at least one of the number of lanes that make up the road and the type of the road (e.g., expressway, general road, etc.). Note that the attributes of the road are not limited to these.

[0040] The road information may be generated by a device mounted on the vehicle based on, for example, position information, map information, etc. The map information is information including road information at each position. The map information may be stored in the device mounted on the vehicle, or may be stored in an external map information providing device (not shown). If stored in the map information providing device, the device mounted on the vehicle may refer to the map information via a communication network, etc. Note that the road information is not limited to the example given here. Furthermore, the device that generates the road information is not limited to the example given here.

[0041] The text acquisition unit 110 may acquire the target description text from the target video using other general techniques, without using a machine learning model such as the video analysis model described above.

[0042] (Regarding the text processing unit 120) The text processing unit 120 generates output information based on, for example, the target explanatory text acquired by the text acquisition unit 110. The output information is information about the target video used to acquire the target explanatory text. The text processing unit 120 may output the output information to a predetermined device, such as a storage device, an external information processing device, or a display device, in a predetermined manner.

[0043] The output information includes, for example, information indicating whether the video is an accident video. An accident video is a video related to an accident. The output information regarding a target video that is an accident video may include at least one of an accident type label and a representative image.

[0044] The accident type label is, for example, information indicating a predetermined type of accident. The predetermined accident type may include, for example, at least one of a collision accident, a collision between vehicles, a collision between a vehicle and a person, a collision accident at an intersection, and an accident in the evening. Note that the types of accidents are not limited to those exemplified here.

[0045] The representative image is an image that represents the target video. For example, the representative image is an image that shows a characteristic scene in the target video. The representative image may be an image included in the target video, such as a frame image or a part of a frame image.

[0046] For example, the text processing unit 120 may use a language model to generate output information indicating whether the target video is an accident video related to an accident. The language model is a model for determining, from the explanatory text, the type of video used to generate the explanatory text.

[0047] The language model may be, for example, a machine learning model trained using training explanatory text and training including a correct answer related to the training explanatory text. The explanatory text input to the language model may be, for example, explanatory text generated by a video analysis model. For example, the language model may output output information when a target explanatory text is input. The text processing unit 120 may include, for example, a language model, and may input the target explanatory text to this language model to generate output information.

[0048] In detail, for example, the text processing unit 120 may use a language model to generate output information for a target video that is an accident video, the output information including at least one of an accident type label for identifying the type of accident and a representative image of the target video.

[0049] Note that the information included in the output information is not limited to the examples given here, and may further include, for example, target description text acquired for the target video, accompanying information for the target video, the target video, etc. The output information may include an accident type label as information indicating whether the video is an accident video. The output information may include, for example, flag information indicating whether the video is an accident video, instead of or together with the accident type label.

[0050] Furthermore, the text processing unit 120 may generate output information based on, for example, a target explanatory text and accompanying information. The language model used in this case may be a model for determining the type of video used to generate the explanatory text from the explanatory text and accompanying information. The text processing unit 120 may then input the explanatory text and accompanying information related to the target video into the language model to generate output information related to the target video.

[0051] (Example of Physical Configuration of Information Processing Device 100) The information processing device 100 physically includes a bus 1010, a processor 1020, a memory 1030, a storage device 1040, a network interface 1050, an input interface 1060, and an output interface 1070, as shown in FIG.

[0052] The bus 1010 is a data transmission path for transmitting and receiving data among the processor 1020, memory 1030, storage device 1040, network interface 1050, input interface 1060, and output interface 1070. However, the method of connecting the processor 1020 and the like to each other is not limited to bus connection.

[0053] The processor 1020 is implemented as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit).

[0054] The memory 1030 is a main storage device realized by a RAM (Random Access Memory) or the like.

[0055] The storage device 1040 is an auxiliary storage device realized by a hard disk drive (HDD), a solid state drive (SSD), a memory card, a read-only memory (ROM), or the like. The storage device 1040 stores program modules for realizing the functions of the device that includes the storage device 1040. The processor 1020 loads each of these program modules into the memory 1030 and executes them to realize the function corresponding to that program module.

[0056] The network interface 1050 is an interface for connecting a device equipped with it to a communication network.

[0057] The input interface 1060 is an interface for the user to input information, and is configured from, for example, a touch panel, a keyboard, a mouse, and the like.

[0058] The output interface 1070 is an interface for presenting information to the user, and is configured, for example, by a liquid crystal panel, an organic EL (Electro-Luminescence) panel, or the like.

[0059] In this way, the functions of the information processing device 100 can be realized by the physical components cooperating to execute a software program. Therefore, the present invention may be realized as a software program or as a storage medium on which the program is non-temporarily recorded. Note that the information processing device may be physically composed of multiple devices (e.g., computers, etc.).

[0060] (Actions and Effects) As described above, according to this embodiment, the information processing device 100 includes the text acquisition unit 110 and the text processing unit 120. The text acquisition unit 110 acquires a target description text, which is text that explains the target video and is generated using the target video. The text processing unit 120 generates output information related to the target video, including whether or not the target video is an accident video related to an accident, based on the acquired target description text.

[0061] This allows for the generation of output information about the target video, including whether the target video is an accident video related to an accident, via the target description text. As described above, this reduces the possibility that events that can be identified as accidents will be influenced by the training data. Therefore, it becomes possible to accurately classify target videos related to various accidents as accident videos.

[0062] According to this embodiment, the text acquisition unit 110 acquires a target explanatory text, which is explanatory text for target video information including a target video, using a video analysis model that generates explanatory text, which is text that explains input information including a video.

[0063] This allows for the generation of output information about the target video, including whether the target video is an accident video related to an accident, via the target description text. As described above, this reduces the possibility that events that can be identified as accidents will be influenced by the training data. Therefore, it becomes possible to accurately classify target videos related to various accidents as accident videos.

[0064] According to this embodiment, the text processing unit 120 generates output information including whether the target video is an accident video related to an accident, using a language model for determining the type of video used to generate the explanatory text from the explanatory text.

[0065] This allows for the generation of output information about the target video, including whether the target video is an accident video related to an accident, via the target description text. As described above, this reduces the possibility that events that can be identified as accidents will be influenced by the training data. Therefore, it becomes possible to accurately classify target videos related to various accidents as accident videos.

[0066] According to this embodiment, the text processing unit 120 uses a language model to generate output information for a target video that is an accident video, the output information including at least one of an accident type label for identifying the type of accident and a representative image of the target video.

[0067] This allows the user to easily recognize the outline of the accident video by using at least one of the accident type label and the representative image, making it possible to easily utilize target videos related to various accidents.

[0068] According to this embodiment, the target video information further includes accompanying information that is information accompanying the target video. The video analysis model generates explanatory text from the video and the accompanying information accompanying the video.

[0069] This allows the generation of output information regarding the target video, including whether the target video is an accident video or not, by further using the accompanying information, thereby enabling target videos regarding various accidents to be classified as accident videos with higher accuracy.

[0070] According to this embodiment, the target video includes at least one of video captured by a camera mounted on the vehicle and video captured of a road on which the vehicle is traveling. The accompanying information includes at least one of the following: a time when the target video was captured, vehicle driving control information, location information of the vehicle when the target video was captured, and road information of the road on which the vehicle was traveling when the target video was captured.

[0071] This allows for generating output information indicating whether the target video is an accident video related to an accident, via the target description text, from target video including at least one of video captured by a camera mounted on a vehicle and video captured of a road on which the vehicle travels. As described above, this reduces the possibility that events that can be identified as accidents will be influenced by the training data. Therefore, it becomes possible to accurately classify target video related to various vehicle-related accidents as accident video.

[0072] [Embodiment 2] In this embodiment, a more detailed example of the information processing device 100 described in embodiment 1 will be described. Note that in this embodiment, for simplicity of explanation, explanations that overlap with embodiment 1 will be omitted as appropriate.

[0073] 4 , the information processing system SYS includes an imaging device C mounted on a vehicle P, and an information processing device 200. The information processing device 200 includes, for example, the above-mentioned text acquisition unit 110 and text processing unit 120, an object acquisition unit 210, an output information storage unit 220, a search unit 230, and a display unit 240.

[0074] The photographing device C is an example of a device that generates a target image. The photographing device C generates a target image, for example, by photographing the surroundings of the vehicle P. The photographing device C is, for example, a camera, but is not limited to this and may be a radar device or the like.

[0075] The object acquisition unit 210 acquires an object video.

[0076] The output information storage unit 220 is a storage unit for storing output information.

[0077] The search unit 230 uses the output information to perform a search process for searching for a target video.

[0078] The display unit 240 displays the results of the search process.

[0079] The information processing device 200 executes, for example, information processing as shown in Fig. 5. This information processing may be executed repeatedly or in response to a user instruction. Note that the timing of executing this information processing is not limited to the example given here.

[0080] The target acquisition unit 210 acquires a target video (step S210).

[0081] The above-described steps S110 and S120 are executed.

[0082] The text processing unit 120 stores the output information in the output information storage unit 220 (step S220).

[0083] The information processing device 200 executes information processing such as that shown in FIG.

[0084] The search unit 230 uses the output information to perform a search process for searching for the target video (step S230).

[0085] The search unit 230 causes the display unit 240 to display the results of the search process (step S240).

[0086] (Regarding the target acquisition unit 210) The target acquisition unit 210 acquires a target video. The target acquisition unit 210 acquires, for example, target video information including the target video. In step S110, the target video or the target video information acquired by the target acquisition unit 210 may be used.

[0087] For example, the information processing device 200 and the image capture device C may be communicably connected to each other via a communication network NT configured by wired, wireless, or a combination of these. The target acquisition unit 210 may acquire target video information from the image capture device C via the communication network NT. In this case, the image capture device C may store the target video. The target acquisition unit 210 may request the image capture device C to acquire the target video information. The target acquisition unit 210 may acquire the target video information in real time.

[0088] Note that the method by which the target acquisition unit 210 acquires the target video information is not limited to the example given here. Furthermore, if the target video information includes accompanying information, the target acquisition unit 210 may acquire the accompanying information from the above-mentioned imaging device C, vehicle control device, map information providing device, etc. For example, the target acquisition unit 210 may acquire the target video information by associating the target video acquired from the imaging device C with driving control information, position information, road information, etc. that were generated at the same time as the target video (within a predetermined time period).

[0089] (Regarding the output information storage unit 220) The output information storage unit 220 is a storage unit for storing output information. For example, when the text processing unit 120 generates output information, it may output the generated output information to the output information storage unit 220. As a result, the text processing unit 120 stores the output information in the output information storage unit 220. In other words, the output information storage unit 220 is an example of an output destination to which the text processing unit 120 outputs the output information.

[0090] The output information storage unit 220 may store output information classified according to a plurality of predetermined categories. Each category may be associated with, for example, one or more accident type labels. Some or all of the plurality of categories may include some overlapping accident type labels.

[0091] (Regarding the Search Unit 230) For example, when the search unit 230 receives search conditions from a user, the search unit 230 performs a search process to search for target video using the output information. The search conditions are set using, for example, accident type labels, phrases associated with the accident type labels, text, keywords, and periods (time, period, etc.).

[0092] The phrase associated with the accident type label is a phrase indicating the meaning of the accident label, and may be, for example, the above-mentioned collision accident, collision accident between vehicles, etc. The search unit 230 may store in advance information associating the accident type label with a phrase.

[0093] The text used as a search condition may be freely input text, such as "a human-car accident that occurred during dark hours." The keywords used as a search condition may be freely input keywords, such as "collision accident" or "intersection." Examples of the time period used as a search condition include "7:00 PM to 5:00 AM" or "December 30th to January 3rd" in terms of the time of the photo or the time of the accident.

[0094] The search conditions are not limited to those exemplified here.

[0095] The search unit 230 extracts output information that matches the search conditions, for example, based on the output information and the search conditions. The search unit 230 may include, for example, a search model for extracting output information that matches the search conditions. The search model may be, for example, a machine learning model that has been trained to extract output information that matches the search conditions from information included in the output information when the search conditions are input.

[0096] As a result of the search process, the search unit 230 displays, for example, the extracted output information on the display unit 240. The display unit 240 may display, for example, a list including at least one of the accident type labels included in the extracted output information, the words associated with the accident type labels, and the representative images included in the extracted output information.

[0097] The method by which the search unit 230 outputs the results of the search process is not limited to displaying them, but may also be transmitting them to another predetermined device (not shown), for example.

[0098] (Operations and Effects) As described above, according to this embodiment, the information processing device 200 includes the search unit 230 that performs search processing to search for target video using output information.

[0099] This allows the output information including the classification results to be utilized. For example, the determination result of whether the video is an accident or not can be utilized for various purposes, such as processing insurance related to the accident, setting insurance premiums, and planning safety measures to prevent accidents.

[0100] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0101] In addition, although the flowcharts used in the above description show a sequence of steps (processes), the order of steps executed in each embodiment is not limited to the sequence shown in the flowcharts. In each embodiment, the order of steps shown in the diagrams can be changed as long as it does not cause any problems in terms of the content.

[0102] Some or all of the above-described embodiments can be described as, but are not limited to, the following supplementary notes.

[0103] 1. An information processing device comprising: a text acquisition means for acquiring a target explanatory text, which is text generated using a target video and explains the target video; and a text processing means for generating output information related to the target video, including whether or not the target video is an accident video related to an accident, based on the acquired target explanatory text. 2. The information processing device described in 1., wherein the text acquisition means acquires the target explanatory text, which is explanatory text for target video information including the target video, using a video analysis model that generates explanatory text, which is text explaining input information from input information including video. 3. The information processing device described in 1. or 2., wherein the text processing means generates the output information, including whether or not the target video is an accident video related to an accident, using a language model for determining, from explanatory text, the type of video used to generate the explanatory text. 4. The information processing device described in 3., wherein the text processing means generates, using the language model, output information for the target video, which is the accident video, including at least one of an accident type label for identifying the type of accident and a representative image of the target video. 5. The information processing device according to any one of 2. to 4., wherein the target video information further includes accompanying information that is information accompanying the target video, and the video analysis model generates the explanatory text from the video and the accompanying information accompanying the video. 6. The information processing device according to 5., wherein the target video includes at least one of video captured by a camera mounted on a vehicle and video captured of a road on which the vehicle is traveling, and the accompanying information includes at least one of a time the target video was captured, driving control information for the vehicle, position information of the vehicle when the target video was captured, and road information of a road on which the vehicle was traveling when the target video was captured. 7. The information processing device according to any one of 1. to 6., further comprising search means for performing search processing to search for the target video using the output information.8. An information processing method in which one or more computers acquire a target explanatory text, which is text that explains a target video generated using a target video, and generate output information related to the target video, including whether or not the target video is an accident video related to an accident, based on the acquired target explanatory text. 9. The information processing method described in 8., in which acquiring the target explanatory text involves acquiring the target explanatory text, which is explanatory text for target video information including the target video, using a video analysis model that generates explanatory text, which is text that explains input information from input information including video. 10. The information processing method described in 8. or 9., in which generating the output information related to the target video involves generating the output information including whether or not the target video is an accident video related to an accident, using a language model that determines, from explanatory text, the type of video used to generate the explanatory text. 11. The information processing method described in 10., in which generating the output information related to the target video involves using the language model to generate, for the target video, which is the accident video, output information including at least one of an accident type label for identifying the type of the accident and a representative image of the target video. 12. The information processing method according to any one of 9. to 11., wherein the target video information further includes accompanying information that is information accompanying the target video, and the video analysis model generates the explanatory text from the video and the accompanying information accompanying the video. 13. The information processing method according to 12., wherein the target video includes at least one of video captured by a camera mounted on a vehicle and video captured of a road on which the vehicle is traveling, and the accompanying information includes at least one of the time the target video was captured, driving control information for the vehicle, position information of the vehicle when the target video was captured, and road information of the road on which the vehicle was traveling when the target video was captured. 14. The information processing method according to any one of 8. to 13., further comprising performing a search process to search for the target video using the output information.15. A program causing one or more computers to execute the following steps: acquire a target explanatory text, which is text generated using a target video and explains the target video; and generate output information related to the target video, including whether or not the target video is an accident video related to an accident, based on the acquired target explanatory text. 16. The program described in 15., wherein acquiring the target explanatory text involves acquiring the target explanatory text, which is explanatory text for target video information including the target video, using a video analysis model that generates explanatory text, which is text explaining input information, from input information including video. 17. The program described in 15. or 16., wherein generating the output information related to the target video involves generating the output information including whether or not the target video is an accident video related to an accident, using a language model that determines, from explanatory text, the type of video used to generate the explanatory text. 18. The program described in 17., wherein generating the output information related to the target video involves using the language model to generate, for the target video, which is the accident video, output information including at least one of an accident type label for identifying the type of accident and a representative image of the target video. 19. The program according to any one of 16. to 18., wherein the target video information further includes accompanying information that is information accompanying the target video, and the video analysis model generates the explanatory text from the video and the accompanying information accompanying the video. 20. The program according to 19., wherein the target video includes at least one of video captured by a camera mounted on a vehicle and video captured of a road on which the vehicle is traveling, and the accompanying information includes at least one of the following: a time at which the target video was captured, driving control information for the vehicle, position information of the vehicle when the target video was captured, and road information for the road on which the vehicle was traveling when the target video was captured. 21. The program according to any one of 15. to 20., further comprising: using the output information to perform a search process for searching for the target video. 22. A recording medium on which the program according to any one of 15. to 21. is recorded.

[0104] This application claims priority based on Japanese Patent Application No. 2024-101082, filed on June 24, 2024, the disclosure of which is incorporated herein in its entirety by reference.

[0105] 100, 200 Information processing device 110 Text acquisition unit 120 Text processing unit 210 Object acquisition unit 220 Output information storage unit 230 Search unit 240 Display unit SYS Information processing system C Photography device P Vehicle

Claims

1. An information processing device comprising: a text acquisition means for acquiring a target explanatory text, which is text that explains a target video generated using the target video; and a text processing means for generating output information regarding the target video, including whether or not the target video is an accident video related to an accident, based on the acquired target explanatory text.

2. The information processing device according to claim 1, wherein the text acquisition means acquires the target explanatory text, which is the explanatory text of the target video information including the target video, using a video analysis model that generates explanatory text, which is text that explains input information including the video, from the input information including the video.

3. An information processing device as described in claim 1 or 2, wherein the text processing means generates the output information including whether the target video is an accident video related to an accident, using a language model for determining from the explanatory text the type of video used to generate the explanatory text.

4. The information processing device according to claim 3, wherein the text processing means uses the language model to generate output information for the target video, which is the accident video, including at least one of an accident type label for identifying the type of accident and a representative image of the target video.

5. The information processing device according to claim 2, wherein the target video information further includes accompanying information that is information accompanying the target video, and the video analysis model generates the explanatory text from the video and the accompanying information accompanying the video.

6. The information processing device according to claim 5, wherein the target image includes at least one of an image captured by a camera mounted on a vehicle and an image captured of a road on which the vehicle is traveling, and the accompanying information includes at least one of the time the target image was captured, driving control information for the vehicle, position information for the vehicle when the target image was captured, and road information for the road on which the vehicle was traveling when the target image was captured.

7. The information processing device according to claim 1 or 2, further comprising a search unit that performs a search process to search for the target video using the output information.

8. An information processing method in which one or more computers acquire a target description text, which is text that explains a target video generated using the target video, and generate output information regarding the target video, including whether or not the target video is an accident video related to an accident, based on the acquired target description text.

9. A recording medium having recorded thereon a program for causing one or more computers to acquire a target description text, which is a text that explains the target video generated using the target video, and generate output information regarding the target video, including whether or not the target video is an accident video related to an accident, based on the acquired target description text.

Citation Information

Patent Citations

  • Information display device, information display system, and control program

    JP2021067501A

  • Information processing apparatus, program, and information processing method

    JP2024012620A

  • Image recognition support apparatus and image recognition support method

    JP2024013629A

  • Signal processing device, signal processing method, program, and imaging device

    WO2020241292A1

  • Speech recognition device, acoustic model learning device, speech recognition method, and computer-readable recording medium

    WO2021166034A1