Information processing device, information processing method, and program

By automating the parameter supplementation process through an information processing device with an annotation data acquisition and video generation unit, the inefficiencies in manually updating learning models are addressed, leading to faster and more accurate vehicle situation description from camera footage.

WO2026094187A1PCT designated stage Publication Date: 2026-05-07SOFTBANK CORPORATION
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SOFTBANK CORPORATION
Filing Date
2024-10-30
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

The manual supplementation of missing parameters in video design for updating learning models is time-consuming and inefficient, especially in the context of evolving machine learning models with diverse pre-training methods and structures, necessitating the automation of parameter supplementation work in dataset creation.

Method used

An information processing device that includes an annotation data acquisition unit to identify weak words from the difference between model output and correct data, and a video generation unit to generate videos using templates with multiple categories, automatically completing parameters for video generation.

Benefits of technology

This approach reduces the burden on humans, increases processing efficiency, and accelerates the update speed of learning models, improving the accuracy of describing vehicle situations from camera footage, thereby enhancing vehicle monitoring and contributing to the automotive industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024038782_07052026_PF_FP_ABST
    Figure JP2024038782_07052026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises an annotation data acquisition unit that acquires annotation data including a text that a learning model has output for an input video and a weakness word identified based on a difference between the text and correct answer data for the input video, the learning model receiving an input of a video captured by a camera mounted on a vehicle and outputting a text representing a situation of the vehicle in the video, and a video generation unit that generates, on the basis of the annotation data, a video for causing the learning model to learn the weakness word, wherein the video generation unit generates the video using a template that is not included in the annotation data and that includes a plurality of categories used for generating the video.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] This invention relates to an information processing device, an information processing method, and a program.

[0002] Patent Document 1 describes a traffic accident situation analysis system comprising a first analysis unit and a second analysis unit, each comprising a drive recorder video acquisition unit, an accident situation classification element extraction unit, and a learning unit, wherein the classification elements of the traffic accident situation extracted from the drive recorder video include the shape of the intersection, the direction of travel of the vehicle, and the presence and color of traffic lights, the learning unit performs machine learning on the accident situation classification element extraction unit using pairs of known traffic accident situation classification elements and drive recorder videos corresponding to the known traffic accident situation classification elements as training data, and the second analysis unit comprises a case data acquisition unit that acquires case data in which the classification elements of the traffic accident situation are associated with accident situation diagrams and fault ratios, an accident situation diagram generation unit that generates accident situation diagrams corresponding to the classification elements of the traffic accident situation, a fault ratio calculation unit that calculates fault ratios corresponding to the classification elements of the traffic accident situation, and a hearing information acquisition unit. Patent Document 2 describes a system comprising: an acquisition unit that acquires from a first mobile terminal device first vehicle identification information that identifies a first vehicle in a first image captured by a first mobile terminal device and first vehicle determination information that determines the driving state of the first vehicle based on the first image; an acquisition unit that acquires from a second mobile terminal device second vehicle identification information that identifies a second vehicle in a second image captured by a second mobile terminal device and second vehicle determination information that determines the driving state of the second vehicle based on the second image; a determination unit that determines the identity of the first vehicle and the second vehicle using the first vehicle identification information and the second vehicle identification information; and a statistical processing unit that, if it is determined that the first vehicle and the second vehicle are the same, generates statistical information relating to the first vehicle using the first vehicle determination information and the second vehicle determination information. Patent Document 3 describes a situation output device characterized by comprising: a visual splendor distribution information acquisition unit that acquires visual splendor distribution information obtained by estimating the level of visual splendor within an image based on an image taken of the outside from a moving object; a visual feature extraction unit that acquires the visual trends of the moving environment of the moving object based on the visual splendor distribution information; and a situation output unit that outputs the situation of the image based on the visual trends.Patent Document 4 describes a driving support device comprising: a surrounding environment imaging means for acquiring captured images of the surrounding environment of a moving object; a risk factor determination means for recognizing surrounding conditions occurring in the surrounding environment of a moving object and determining risk factors in the surrounding environment of a moving object based on the recognized surrounding conditions by inputting the captured images into a learning model generated by machine learning using a multilayer neural network; a determination result output means for outputting the determination result of the risk factor determination means; and a determination criterion changing means for changing the determination criteria for risk factors by updating the part of the learning model that determines risk factors in the surrounding environment of a moving object based on surrounding conditions. Patent Document 5 describes an explanation text generation device comprising: an image explanation text generation unit for generating an explanation text for an image based on image features; a driving behavior explanation text generation unit for generating an explanation text for driving behavior based on image features and the explanation text generated by the image explanation text generation unit; and an explanation text generation unit for generating an image or an explanation text for driving behavior that occurs after the generation of the explanation text for the image and the explanation text for driving behavior based on the explanation text generated by the image explanation text generation unit and the explanation text generated by the driving behavior explanation text generation unit. Patent Document 6 describes an explanatory text generation device that includes: an acquisition unit that acquires material feature quantities representing each of the materials used in a process and video feature quantities extracted from each video of each process in which the process is filmed; an update unit that identifies an action on the materials included in the video of each process based on the video feature quantities of each process, and updates the material feature quantities of the identified materials according to the identified action; and a generation unit that generates text explaining the procedure of the process based on the updated material feature quantities, the identified action, and the video feature quantities. [Prior Art Documents] [Patent Documents] [Patent Document 1] Japanese Unexamined Patent Publication No. 2022-098200 [Patent Document 2] Japanese Unexamined Patent Publication No. 2021-086597 [Patent Document 3] Japanese Unexamined Patent Publication No. 2024-060029 [Patent Document 4] Japanese Unexamined Patent Publication No. 2018-173860 [Patent Document 5] Japanese Unexamined Patent Publication No. 2021-174172 [Patent Document 6] Japanese Unexamined Patent Publication No. 2023-053742.

[0003] In a conventional dataset creation process for updating a learning model that takes as input the video of a vehicle's drive recorder and outputs text indicating the vehicle's situation, procedures such as completion of annotation work, implementation of video design, manual supplementation of missing parameters, and video generation using a driving simulator were taken. In this process, especially when designing the video, it was necessary for a person to manually supplement the missing parameters, which required time and effort.

[0004] Currently, the competition in the development of machine learning models is intensifying, and various models are emerging one after another. New models generally tend to have an increasing number of parameters, making it difficult to keep up with their evolution. Furthermore, since the pre-training methods and structures of each model are different, the learning approaches also have different characteristics for each model. Regarding the dataset required for fine-tuning, it is necessary to create it efficiently according to the characteristics of the model. In particular, as an issue in video design in the dataset creation process, automation of the parameter supplementation work currently done manually is required.

[0005] According to an embodiment of the present invention, an information processing device is provided. The information processing device may include an annotation data acquisition unit that acquires annotation data including weak words identified by the difference between the text output by a learning model that takes as input the video captured by a camera mounted on a vehicle and outputs text representing the situation of the vehicle in the video, and the correct answer data of the text for the input video. The information processing device may include a video generation unit that generates a video for the learning model to learn the weak words based on the annotation data. The video generation unit may generate the video using a template including a plurality of categories used for generating the video that are not included in the annotation data.

[0006] In the information processing device, the template may include a plurality of options for each of the plurality of categories, and the video generation unit may generate the video for each variation of the plurality of options of the plurality of categories.

[0007] In the aforementioned information processing device, the template may include the number of lanes on the road on which the target vehicle travels.

[0008] In the aforementioned information processing device, the template may include the driving speed of the target vehicle.

[0009] In the aforementioned information processing device, the template may include the number of lanes on the road where other vehicles are traveling.

[0010] In the aforementioned information processing device, the template may include the driving speed of other vehicles.

[0011] In the information processing device, the template may include the direction of travel of the target vehicle and the direction of travel of other vehicles.

[0012] In the information processing device, the annotation data may include the vehicle's operation, the causative object, the object's motion, the object's absolute position, the object's relative position, the vehicle's driving intention, and the vehicle's position.

[0013] The information processing device may include a model update unit that updates the learning model using a plurality of learning data sets, including the video generated by the video generation unit based on the annotation data and the annotation data.

[0014] The information processing device may include a RAN control unit for controlling a RAN (Radio Access Network) and an AI processing unit for performing AI (Artificial Intelligence) processing, and the AI ​​processing unit may include an annotation data acquisition unit and a video generation unit. Examples of AI processing include AI processing related to RAN control (sometimes referred to as RAN control AI processing) and AI processing not related to RAN control (sometimes referred to as non-RAN control AI processing). An example of RAN control AI processing is RIC (RAN Intelligent Controller). RIC is a technology that uses AI to optimize RAN wireless resources and automate RAN operation. RICs include Non-RT RICs (Non-Real Time RICs) and Near-RT RICs (Near-Real Time RICs). Non-RT RICs are sometimes called Centralized RICs. Non-RT RICs are located within SMOs (Service Management and Orchestrations) that manage and orchestrate RANs. Non-RT RICs generate and notify policies related to RAN control and send information to Near-RT RICs. For example, a Non-RT RIC generates a learning model for RAN control by performing machine learning using data collected from the RAN and sends it to a Near-RT RIC. Near-RT RICs are sometimes called Distributed RICs. Compared to Non-RT RICs, Near-RT RICs are located closer to RAN nodes (RUs (Radio Units), DUs (Distributed Units), CUs (Central Units)) and perform tasks such as controlling RAN nodes and resources. Near-RT RICs perform processing that is more real-time than Non-RT RICs. For example, Near-RT RICs perform inference processing related to RAN control using trained models acquired from Non-RT RICs. RAN control AI processing is not limited to RICs. Non-RAN control AI processing can be so-called MEC AI applications.Non-RAN controlled AI processing includes learning and inference processing of arbitrary AIs that are not related to RAN control.

[0015] According to one embodiment of the present invention, an information processing method performed by a computer is provided. The information processing method may include an annotation data acquisition step in which a learning model, which takes a video captured by a camera mounted on a vehicle as input and outputs text representing the state of the vehicle in the video, acquires annotation data that includes weak words identified by the difference between the text output for the input video and the correct text data for the input video. The information processing method may also include a video generation step in which a video is generated to allow the learning model to learn the weak words based on the annotation data. The video generation step may generate the video using a template that includes a plurality of categories of information that are not included in the annotation data but are necessary to generate the video, which has been registered in advance.

[0016] According to one embodiment of the present invention, a program is provided for a computer to perform the following steps: an annotation data acquisition step, in which a learning model takes video footage captured by a camera mounted on a vehicle as input and outputs text representing the state of the vehicle in the video, acquires annotation data including weak words identified by the difference between the text output for the input video and the correct text data for the input video; and a video generation step, in which a video is generated to allow the learning model to learn the weak words based on the annotation data, wherein the video is generated using a template that includes a plurality of categories of information that are not included in the annotation data but are necessary to generate the video, which has been registered in advance.

[0017] It should be noted that the above summary of the invention does not enumerate all the necessary features of the present invention. Furthermore, subcombinations of these features may also constitute an invention.

[0018] A schematic diagram illustrating an example of the environment of the information processing device 100. An explanatory diagram illustrating the processing performed by the information processing device 100. An explanatory diagram illustrating the processing performed by the information processing device 100. An explanatory diagram illustrating the processing performed by the information processing device 100. An explanatory diagram illustrating the processing performed by the information processing device 100. A schematic diagram illustrating an example of the functional configuration of the information processing device 100. A schematic diagram illustrating an example of the environment of the information processing device 100. A schematic diagram illustrating an example of the hardware configuration of the computer 1200 that functions as the information processing device 100.

[0019] The present invention will be described below through embodiments, but these embodiments are not intended to limit the scope of the claims. Furthermore, not all combinations of features described in the embodiments are necessarily essential to the solution of the invention.

[0020] Figure 1 schematically shows an example of the environment of the information processing device 100. In the example shown in Figure 1, the information processing device 100 is connected to a network 50. The network 50 may include a cloud network. The network 50 may include the internet. The network 50 may include a mobile communication network. The mobile communication network may conform to any of the following mobile communication systems: LTE (Long Term Evolution) communication system, 5G (5th Generation) communication system, 3G (3rd Generation) communication system, and 6G (6th Generation) communication system or later.

[0021] For example, the information processing device 100 may be located in a cloud network. The information processing device 100 may also be located in a mobile communication network. For example, the information processing device 100 may be located in a MEC (Multi-access Edge Computing) system. The information processing device 100 may also be located in a core network.

[0022] The information processing device 100 receives video 210 captured by a camera 202 mounted on the vehicle 200 from the vehicle 200. The information processing device 100 may receive the video 210 via a mobile communication network. The video 210 may be transmitted by the camera 202 or by a communication device mounted on the vehicle 200.

[0023] Vehicle 200 may be a so-called autonomous vehicle. Camera 202 may be a drive recorder. Camera 202 may be a camera that captures images for use in controlling autonomous driving. Camera 202 may be any other type of camera. Vehicle 200 does not have to be an autonomous vehicle.

[0024] The information processing device 100 takes a video 210 as input and generates training data for updating a learning model that outputs text describing the status of the vehicle 200 in the video 210. The learning model can be anything as long as it can output text describing the status of the vehicle 200 in the video 210 in response to the input video 210. For example, the learning model may be generated by machine learning using multiple training data sets that include a video 210 in which the status of the vehicle 200 is known and text describing the status of the known vehicle. For example, the learning model may be a so-called video generation AI. If the learning model is a so-called video generation AI, the learning model may be initially tuned to output text describing the status of the vehicle 200 in response to the video 210. Usually, initial tuning alone is not sufficient, and appropriate updates to the learning model are necessary to improve accuracy.

[0025] Figure 2 is an explanatory diagram illustrating the processing performed by the information processing device 100. The information processing device 100 stores a learning model 101 that takes a video 210 as input and outputs text describing the status of the vehicle 200 in the video 210.

[0026] The information processing device 100 performs inference processing using the learning model 101 on the video 210 received from the vehicle 200. The information processing device 100 inputs the video 210 to the learning model 101 and generates text that describes the situation of the vehicle 200 in the video 210. The generated text may be referred to as the generated statement 104.

[0027] The information processing device 100 performs an evaluation / analysis 106 on the generated sentence 104. The information processing device 100 may pre-store correct data 108 for the video 210 and perform the evaluation / analysis 106 by comparing the generated sentence 104 with the correct data 108.

[0028] The correct answer data 108 is pre-registered for the video 210. The correct answer data 108 may be generated and registered by a person who views the video 210. The correct answer data 108 may also be registered by a person who views the video 210, for example, by inputting the video 210 into a learning model 101 and modifying the generated sentence 104 output from the learning model. The correct answer data 108 may be registered for the video 210 by any other method, as long as it accurately represents the state of the vehicle 200 in the video 210.

[0029] The information processing device 100 identifies weak words 110 by evaluation / analysis 106 based on the difference between the generated sentence 104 output by the learning model 101 for the video 210 and the correct answer data 108 for the video 210. For example, the information processing device 100 identifies weak words 110 as words that are included in the correct answer data 108 but not in the generated sentence 104.

[0030] The information processing device 100 performs annotation data generation 112 for the weak word 110. The information processing device 100 stores annotation categories 114 in advance and generates multiple annotation data 116 using the weak word 110 and the annotation categories 114.

[0031] Annotation category 114 includes multiple categories. For each of the multiple categories, annotation category 114 includes multiple options.

[0032] Multiple categories in annotation category 114 may include "action." "Action" represents the actions of the vehicle. Examples of "action" options include stopping, going straight (forward), reversing, turning right, turning left, accelerating, decelerating, changing lanes to the right, changing lanes to the left, etc. The "action" options may include only some of these, or may include all of them.

[0033] Multiple categories in annotation category 114 may include "causative objects." A "causative object" represents the object that caused the reason for the action. Examples of "causative object" options include pedestrians, riders (motorcycles), passenger cars, trucks, buses, trains, motorcycles, bicycles, signs, red lights, yellow lights, green lights, balls, multiple pedestrians, fire trucks, ambulances, police cars, arrow signals, obstacles, etc. The "causative object" options may include only some of these, or may include all of them.

[0034] Multiple categories in annotation category 114 may include "object motion." "Object motion" represents the action of the causative object. Examples of "object motion" options include walking, standing still, falling, running, making contact, giving a stop signal, in front of a store, running out, etc. The "object motion" options may include only some of these, or may include all of them.

[0035] Multiple categories in annotation category 114 may include "absolute object location." "Absolute object location" represents the absolute location of an object. Examples of options for "absolute object location" include roadways, shoulders (roadsides), crosswalks, sidewalks, bicycle paths, fire stations, hospitals, police stations, etc. Options for "absolute object location" may include only some of these, or may include all of them.

[0036] Multiple categories in annotation category 114 may include "Object Position (Relative)". "Object Position (Relative)" represents the relative position of an object as seen from the vehicle. Examples of "Object Position (Relative)" options include forward (in the path of travel), right front, left front, etc. The "Object Position (Relative)" options may include only some of these, or may include all of them.

[0037] Multiple categories in annotation category 114 may include "driving intention." "Driving intention" represents the planned route of the vehicle. Examples of options for "driving intention" include going straight (forward), reversing, turning right, turning left, etc. The options for "driving intention" may include only some of these, or may include all of them.

[0038] Multiple categories in annotation category 114 may include "Vehicle Location". "Vehicle Location" represents the vehicle's current location (including the vicinity of a facility). Examples of options for "Vehicle Location" include intersections, straight roads, curves, etc. Options for "Vehicle Location" may include only some of these, or all of them.

[0039] The information processing device 100 identifies a category among the multiple categories of annotation category 114 in which one of the options matches the weak word 110. By fixing the option in that category to the weak word 110 and changing the options in the other categories, it generates multiple variations of combinations of multiple options in multiple categories. For each of the multiple variations, it generates annotation data 116 using the options in the multiple categories.

[0040] As a concrete example, as illustrated in Figure 3, if the weak word 110 is a straight road, the information processing device 100 fixes the option in the "vehicle position" category to a straight road and changes the options in the categories of "action," "causative object," "object action," "object position (absolute)," "object position (relative)," and "driving intention," thereby generating multiple variations of combinations of multiple options in multiple categories. In the example shown in Figure 3, the information processing device 100 generates one variation where "action" is straight, "causative object" is an obstacle, "object action" is none, "object position (absolute)" is the roadway, "object position (relative)" is ahead, "driving intention" is straight, and "vehicle position" is a straight road. Using the multiple options in this variation, the information processing device 100 generates annotation data 116 that says, "The car is on a straight road and is trying to go straight. There is no obstacle ahead, so the car is going straight." The generation of annotation data 116 may be performed using a rule-based method with annotation categories 114. The generation of annotation data 116 may be performed using AI. For example, the generation of annotation data 116 may be performed using text generation AI. The generation of annotation data 116 may also be performed manually.

[0041] The information processing device 100 performs video design 118 for each of the multiple annotation data 116. The annotation data 116 is data that represents the status of the vehicle 200, and is not data for generating video. Therefore, the annotation data 116 alone does not have enough parameters to generate video. Thus, conventionally, people have supplemented the parameters necessary for video generation based on their experience. However, this method places a heavy burden on people and is time-consuming, which is one of the reasons why the overall processing efficiency has decreased.

[0042] In contrast, the information processing device 100 according to this embodiment has a template 120 that includes multiple categories used to generate a video, which are not included in the annotation data 116, and automatically completes the parameters necessary for video generation using the template 120.

[0043] Template 120 may include, as major categories, "Common", "Own Vehicle", and "Other Vehicle". The major categories may include others than these.

[0044] The "Common" category may include, as minor categories, "Road", "Weather", and "Time". The "Common" category may include others than these as minor categories. The minor category "Road" may include, as options, one lane, two lanes, three lanes, four lanes, and five lanes. The minor category "Road" may include others than these as options. The minor category "Weather" may include, as options, clear, cloudy, and rainy. The minor category "Weather" may include others than these as options. The minor category "Time" may include, as options, 6:00, 12:00, and 18:00. The minor category "Time" may include others than these as options.

[0045] The "Own Vehicle" category may include, as minor categories, "Speed" and "Lane". The "Own Vehicle" category may include others than these as minor categories. The minor category "Speed" may include, as options, 30 km / h, 60 km / h, and 100 km / h. The minor category "Speed" may include others than these as options. The minor category "Lane" may include, as options, right, middle, and left. The minor category "Lane" may include others than these as options.

[0046] The "Other Vehicle" category may include, as minor categories, "Speed", "Lane", and "Direction". The "Other Vehicle" category may include others than these as minor categories. The minor category "Speed" may include, as options, 0 km / h, 30 km / h, and 60 km / h. The minor category "Speed" may include others than these as options. The minor category "Lane" may include, as options, right, middle, and left. The minor category "Lane" may include others than these as options.

[0047] The information processing apparatus 100 may generate a plurality of video parameters 122 using the plurality of annotation data 116 and the template 120.

[0048] The information processing apparatus 100 may generate a plurality of video parameters 122 for each of the plurality of annotation data 116 by using the annotation data 116 and the template 120. The information processing apparatus 100 may generate variations of combinations of a plurality of options of a plurality of categories by changing the options of the plurality of categories of the template 120, and generate a plurality of video parameters 122 by combining them with the annotation data 116.

[0049] As a specific example, as illustrated in FIG. 4, for the annotation category 114 and the annotation data 116, the option of "road" in the "common" category of the template 120 is two lanes, the option of "weather" is cloudy, the option of "time" is 6:00, the option of "speed" in the "own vehicle" category is 60 km / h, the option of "lane" is left, the option of "speed" in the "other vehicle" category is 60 km / h, the option of "lane" is left, and the option of "direction" is traveling side by side to generate the video parameter 122 combined, or for the annotation category 114 and the annotation data 116, the option of "road" in the "common" category of the template 120 is two lanes, the option of "weather" is cloudy, the option of "time" is 6:00, the option of "speed" in the "own vehicle" category is 60 km / h, the option of "lane" is left, the option of "speed" in the "other vehicle" category is 60 km / h, the option of "lane" is left, and the option of "direction" is oncoming to generate the video parameter 122 combined, etc.

[0050] The information processing device 100 may generate a video 126 using the generated video parameters 122 to train the learning model 101 on weak words 110. The information processing device 100 may generate the video 126 using the video parameters 122 in a rule-based manner. For example, the information processing device 100 may generate a video 126 that reflects all elements included in the annotation category 114, annotation data 116, and template 120. The information processing device 100 may generate the video 126 from the video parameters 122 using AI. For example, the information processing device 100 generates a prompt instructing the video generation AI to generate a video 126 using all elements included in the annotation category 114, annotation data 116, and template 120, and inputs this prompt to the video generation AI to generate the video 126.

[0051] As a specific example, as illustrated in Figure 5, the information processing device 100 generates multiple videos 126 using the multiple video parameters 122 illustrated in Figure 4. For each of the multiple videos 126, the video generation 124 generates training data 128 that includes the generated video 126 and the annotation data 116 used to generate the video 126.

[0052] The information processing device 100 may perform fine-tuning 130 of the learning model 101 using the generated learning data 128. The information processing device 100 may update the learning model 101 for each of the learning data 128 so that when a video 126 is input, a generated sentence 104 based on the annotation data 116 is generated.

[0053] As described above, the annotation data 116 is data that represents the status of the vehicle 200, and not data for generating video. Therefore, the annotation data 116 alone lacks sufficient parameters for generating video. Conventionally, people have supplemented the parameters necessary for video generation based on their experience. However, this method places a heavy burden on people and is time-consuming, contributing to a decrease in overall processing efficiency. In contrast, the information processing device 100 according to this embodiment can automatically generate multiple video parameters 122 from multiple annotation data 116 using a template 120. This eliminates the need for people to supplement the parameters necessary for video generation, significantly reducing the burden on people. Furthermore, it can generate video parameters 122 at a faster rate than human supplementation, thereby increasing the update speed of the learning model 101. By increasing the update speed of the learning model 101, the accuracy of generating a generated statement 104 that accurately describes the situation of the vehicle 200 from video 210 taken by a camera 202 mounted on the vehicle 200 can be improved more quickly. This can improve the accuracy of monitoring the vehicle 200 using the generated statement 104, reduce the monitoring load on the vehicle 200, and potentially make a significant contribution to the development of the automotive industry.

[0054] In this example, the explanation uses the case where all processing is performed by the information processing device 100, but this is not the only case. The information processing device 100 may perform only video design 118 and video generation 124. The information processing device 100 may perform only video design 118, video generation 124, and fine tuning 130.

[0055] Figure 6 schematically shows an example of the functional configuration of the information processing device 100. The information processing device 100 includes a storage unit 150, a registration unit 152, a video receiving unit 154, an inference unit 156, an evaluation and analysis unit 158, an annotation data generation unit 160, an annotation data acquisition unit 162, a video generation unit 164, and a model update unit 166. It is not necessarily required that the information processing device 100 include all of these.

[0056] The registration unit 152 registers various types of data. The registration unit 152 stores the registered data in the storage unit 150. The registration unit 152 may register the learning model 101. The registration unit 152 may register the annotation category 114. The registration unit 152 may register the template 120.

[0057] The video receiving unit 154 receives the video 210. The video receiving unit 154 may receive the video 210 from the vehicle 200 via a mobile communication network. The video receiving unit 154 may also receive multiple videos 210 from a server that collects videos 210 from multiple vehicles 200. The registration unit 152 may register the correct answer data 108 for the video 210 received by the video receiving unit 154.

[0058] The inference unit 156 performs inference 102 on the video 210 received by the video receiving unit 154. The inference unit 156 generates a generated sentence 104 by inputting the video 210 into the learning model 101.

[0059] The evaluation and analysis unit 158 ​​performs an evaluation / analysis 106 on the generated sentence 104 generated by the inference unit 156. The evaluation and analysis unit 158 ​​compares the generated sentence 104 generated from the video 210 with the correct answer data 108 corresponding to the video 210. The evaluation and analysis unit 158 ​​identifies weak words 110 based on the difference between the generated sentence 104 generated from the video 210 and the correct answer data 108 corresponding to the video 210. For example, the evaluation and analysis unit 158 ​​identifies words that are included in the correct answer data 108 but not in the generated sentence 104 as weak words 110.

[0060] The annotation data generation unit 160 performs annotation data generation 112 for the weak word 110 identified by the evaluation analysis unit 158. The annotation data generation unit 160 generates multiple annotation data 116 using the weak word 110 and the annotation category 114. The annotation data generation unit 160 may identify a category among the multiple categories of the annotation category 114 in which one of the options matches the weak word 110, fix the option in that category to the weak word 110, and change the options in the other categories to generate multiple variations of combinations of multiple options in multiple categories, and generate annotation data 116 for each of the multiple variations using the options in multiple categories.

[0061] The annotation data acquisition unit 162 acquires a plurality of annotation data 116 generated by the annotation data generation unit 160.

[0062] The video generation unit 164 generates a video 126 for training the learning model 101 on weak words 110, based on a plurality of annotation data 116 acquired by the annotation data acquisition unit 162.

[0063] The video generation unit 164 performs video design 118 for each of the multiple annotation data 116. The video generation unit 164 generates multiple video parameters 122 using the multiple annotation data 116 and the template 120. The video generation unit 164 may generate multiple video parameters 122 for each of the multiple annotation data 116 using the annotation data 116 and the template 120.

[0064] As described above, the template 120 may include multiple options for each of the multiple categories. The video generation unit 164 may generate variations of combinations of multiple options for multiple categories by changing the options for multiple categories in the template 120, and generate multiple video parameters 122 by combining these with annotation data 116.

[0065] As described above, template 120 may include the number of lanes on the road on which the target vehicle 200 is traveling. Template 120 may include the speed at which the target vehicle 200 is traveling. Template 120 may include the number of lanes on the road on which other vehicles are traveling. Template 120 may include the speed at which other vehicles are traveling. Template 120 may include the direction of travel of the target vehicle 200 and the direction of travel of other vehicles.

[0066] The video generation unit 164 generates a video 126 using the multiple video parameters 122 that it has generated. The video generation unit 164 may generate the video 126 using the video parameters 122 in a rule-based manner. For example, the information processing device 100 generates a video 126 that reflects all the elements included in the annotation category 114, annotation data 116, and template 120. The video generation unit 164 may generate the video 126 from the video parameters 122 using AI. For example, the information processing device 100 generates a prompt instructing the video generation AI to generate a video 126 using all the elements included in the annotation category 114, annotation data 116, and template 120, and inputs this prompt to the video generation AI to generate the video 126. The video generation unit 164 generates training data 128 that includes the video 126 and the annotation data 116 used to generate the video 126.

[0067] The model update unit 166 updates the learning model 101 using multiple learning data sets, including the video 126 generated by the video generation unit 164 based on the annotation data 116 and the annotation data 116. The model update unit 166 may update the learning model 101 for each of the multiple learning data sets 128 such that when the video 126 is input, a generated sentence 104 based on the annotation data 116 is generated.

[0068] As described above, the information processing device 100 does not have to include all of the following: storage unit 150, registration unit 152, video receiving unit 154, inference unit 156, evaluation and analysis unit 158, annotation data generation unit 160, annotation data acquisition unit 162, video generation unit 164, and model update unit 166. For example, the information processing device 100 does not have to include the inference unit 156, evaluation and analysis unit 158, and annotation data generation unit 160. In this case, the inference unit 156, evaluation and analysis unit 158, and annotation data generation unit 160 are provided by other devices. For example, the information processing device 100 does not have to include the inference unit 156 and evaluation and analysis unit 158. In this case, the inference unit 156 and evaluation and analysis unit 158 ​​are provided by other devices. For example, the information processing device 100 does not have to include the inference unit 156. In this case, the inference unit 156 is provided by other devices. For example, the information processing device 100 does not have to include the evaluation and analysis unit 158. In this case, the evaluation and analysis unit 158 ​​is provided by another device. For example, the information processing device 100 does not need to include an annotation data generation unit 160. In this case, the annotation data generation unit 160 is provided by another device.

[0069] Figure 7 schematically shows an example of the environment of the information processing device 100. The information processing device 100 may be deployed in a mobile communication system consisting of a management infrastructure 300, a plurality of distributed infrastructures 400, and a plurality of wireless base stations 500. In this mobile communication system, the management infrastructure 300 and the plurality of distributed infrastructures 400 cooperate to control the RAN 510 and perform AI processing. Mobile communication services are provided to the vehicle 200 by the RAN 510. The RAN 510 may be a virtualized vRAN (Virtual RAN), and the mobile communication system may perform control of the vRAN. The RAN 510 may also be a physical RAN, and the mobile communication system may perform control of the physical RAN.

[0070] The distributed infrastructure 400 may be data centers located in various locations. The distributed infrastructure 400 may be composed of multiple devices. The information processing device 100 may be located on the distributed infrastructure 400. The information processing device 100 may function as a BBU (BaseBand Unit), and the wireless base station 500 may function as an RRU (Remote Radio Unit). The information processing device 100 may implement a CU. The distributed infrastructure 400 may implement a DU. The information processing device 100 may implement a UPF (User Plane Function).

[0071] The management infrastructure 300 may be a data center that manages multiple distributed infrastructures 400. The management infrastructure 300 may be composed of multiple devices. The management infrastructure 300 may be implemented on a virtualization infrastructure consisting of multiple devices. The management infrastructure 300 may be implemented by a single device. In other words, the management infrastructure 300 may be a management device.

[0072] The management infrastructure 300 may be called the Core Brain, and the distributed infrastructure 400 may be called the Regional Brain. Note that Figure 7 illustrates a case where a single-layer distributed infrastructure 400 is located below the management infrastructure 300, but this is not the only case. The distributed infrastructure 400 may have multiple layers. For example, if two layers of distributed infrastructure 400 are located below the management infrastructure 300, the management infrastructure 300 may be called the Core Brain, the lower-layer distributed infrastructure 400 may be called the Regional Brain, and the further lower-layer distributed infrastructure 400 may be called the Sub-Regional Brain.

[0073] The information processing device 100, located on the distributed infrastructure 400, may have one or more CPUs (Central Processing Units) and one or more GPUs (Graphics Processing Units). The information processing device 100 may have multiple superchips, each connected by an interconnect. This interconnect may be memory-consistent and capable of achieving high bandwidth and low latency. Thus, the information processing device 100 may have both CPU resources and GPU resources as computing resources.

[0074] The information processing device 100 may include a RAN control unit 170 for controlling the RAN 510 and an AI processing unit 180 for performing AI processing. In this case, the AI ​​processing unit 180 may include an annotation data acquisition unit 162 and a video generation unit 164. The AI ​​processing unit 180 may further include a model update unit 166. The AI ​​processing unit 180 may further include a registration unit 152, a video reception unit 154, an inference unit 156, an evaluation and analysis unit 158, and an annotation data generation unit 160.

[0075] Figure 8 schematically shows an example of the hardware configuration of a computer 1200 that functions as an information processing device 100. A program installed on the computer 1200 can cause the computer 1200 to function as one or more "parts" of the device according to this embodiment, or to cause the computer 1200 to execute operations associated with the device according to this embodiment or such one or more "parts", and / or to cause the computer 1200 to execute a process or a stage of such process according to this embodiment. Such a program may be executed by the CPU 1212 to cause the computer 1200 to execute specific operations associated with some or all of the blocks in the flowcharts and block diagrams described herein.

[0076] The computer 1200 according to this embodiment includes a CPU 1212, a GPU 1213, a RAM 1214, and a graphics controller 1216, which are interconnected by a host controller 1210. The computer 1200 also includes input / output units such as a communication interface 1222, a storage device 1224, a DVD drive 1226, and an IC card drive, which are connected to the host controller 1210 via an input / output controller 1220. The DVD drive 1226 may be a DVD-ROM drive and a DVD-RAM drive, etc. The storage device 1224 may be a hard disk drive and a solid-state drive, etc. The computer 1200 also includes legacy input / output units such as a ROM 1230 and a keyboard, which are connected to the input / output controller 1220 via an input / output chip 1240.

[0077] The CPU 1212 operates according to the programs stored in the ROM 1230 and RAM 1214, thereby controlling each unit. The graphics controller 1216 acquires the image data generated by the CPU 1212 and stores it in the frame buffer provided in the RAM 1214 or within itself, so that the image data is displayed on the display device 1218.

[0078] The communication interface 1222 communicates with other electronic devices via a network. The storage device 1224 stores programs and data used by the CPU 1212 in the computer 1200. The DVD drive 1226 reads programs or data from the DVD-ROM 1227, etc., and provides them to the storage device 1224. The IC card drive reads programs and data from the IC card and / or writes programs and data to the IC card.

[0079] The ROM 1230 stores boot programs and / or hardware-dependent programs of the computer 1200, which are executed by the computer 1200 when activated. The input / output chip 1240 may also connect various input / output units to the input / output controller 1220 via USB ports, parallel ports, serial ports, keyboard ports, mouse ports, etc.

[0080] The program is provided on a computer-readable storage medium such as a DVD-ROM 1227 or an IC card. The program is read from the computer-readable storage medium and installed on a storage device 1224, RAM 1214, or ROM 1230, which are examples of computer-readable storage media, and executed by the CPU 1212. The information processing described within these programs is read by the computer 1200, resulting in coordination between the program and the various types of hardware resources described above. The apparatus or method may be configured to realize the operation or processing of information in accordance with the use of the computer 1200.

[0081] For example, when communication is performed between a computer 1200 and an external device, the CPU 1212 may execute a communication program loaded into the RAM 1214 and, based on the processing described in the communication program, instruct the communication interface 1222 to perform communication processing. Under the control of the CPU 1212, the communication interface 1222 reads transmission data stored in a transmission buffer area provided in a recording medium such as the RAM 1214, storage device 1224, DVD-ROM 1227, or IC card, transmits the read transmission data to the network, or writes received data received from the network to a reception buffer area or the like provided on the recording medium.

[0082] Furthermore, the CPU 1212 may read all or necessary parts of a file or database stored on an external recording medium such as a storage device 1224, a DVD drive 1226 (DVD-ROM 1227), or an IC card into the RAM 1214, and perform various types of processing on the data in the RAM 1214. The CPU 1212 may then write the processed data back to the external recording medium.

[0083] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and subjected to information processing. The CPU 1212 may perform various types of processing on the data read from the RAM 1214, including various types of operations, information processing, conditional judgments, conditional branching, unconditional branching, information retrieval / replacement, etc., as described throughout this disclosure and specified by the program instruction sequence, and write the results back to the RAM 1214. The CPU 1212 may also retrieve information in files, databases, etc., within the recording medium. For example, if a plurality of entries are stored in the recording medium, each having an attribute value of a first attribute associated with an attribute value of a second attribute, the CPU 1212 may search among the plurality of entries for an entry that matches the specified condition for the attribute value of the first attribute, read the attribute value of the second attribute stored in that entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.

[0084] The program or software module described above may be stored on or near the computer 1200 in a computer-readable storage medium. Alternatively, a recording medium such as a hard disk or RAM provided within a server system connected to a dedicated communication network or the Internet can be used as a computer-readable storage medium, thereby providing the program to the computer 1200 via the network.

[0085] In this embodiment, blocks in the flowchart and block diagram may represent a stage in a process in which an operation is performed or a "part" of a device that has the role of performing an operation. A particular stage and "part" may be implemented by a dedicated circuit, a programmable circuit supplied with computer-readable instructions stored on a computer-readable storage medium, and / or a processor supplied with computer-readable instructions stored on a computer-readable storage medium. The dedicated circuit may include digital and / or analog hardware circuits, and may include integrated circuits (ICs) and / or discrete circuits. The programmable circuit may include reconfigurable hardware circuits, such as field-programmable gate arrays (FPGAs) and programmable logic arrays (PLAs), which include logical AND, logical OR, exclusive OR, negated AND, negated OR, and other logical operations, flip-flops, registers, and memory elements.

[0086] A computer-readable storage medium may include any tangible device capable of storing instructions to be executed by a suitable device, and as a result, a computer-readable storage medium having instructions stored therein will comprise a product that includes instructions that can be executed to create means for performing operations specified in a flowchart or block diagram. Examples of computer-readable storage media may include electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, etc. More specific examples of computer-readable storage media may include floppy disks (registered trademark), diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disk read-only memory (CD-ROM), digital versatile disk (DVD), Blu-ray (registered trademark) disk, memory stick, integrated circuit card, etc.

[0087] Computer-readable instructions may include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk®, Java®, C++, and conventional procedural programming languages ​​such as the C programming language or similar programming languages.

[0088] Computer-readable instructions may be provided locally or via a wide area network (WAN) such as a local area network (LAN) or the internet to a processor or programmable circuit of a general-purpose computer, a special-purpose computer, or another programmable data processing device, so that the processor or programmable circuit of the programmable data processing device, such as a computer, may execute the computer-readable instructions to generate means for performing operations specified in a flowchart or block diagram. Here, the computer may be a PC (personal computer), a tablet computer, a smartphone, a workstation, a server computer, a general-purpose computer, or a special-purpose computer, and may also be a computer system in which multiple computers are connected. Such a computer system in which multiple computers are connected is also called a distributed computing system and is a computer in a broad sense. In a distributed computing system, multiple computers execute a program collectively by each computer executing a part of the program and passing data during program execution between computers as needed.

[0089] Examples of processors include computer processors, central processing units (CPUs), processing units, microprocessors, digital signal processors, controllers, and microcontrollers. A computer may have one or more processors. In a multiprocessor system with multiple processors, each processor executes a portion of the program, and the processors collectively execute the program by passing program execution data between them as needed. For example, in the execution of multitasks, each of the multiple processors may execute a portion of each task in small chunks by switching tasks at each time slice. In this case, which part of a program each processor executes changes dynamically. Which part of a program each of the multiple processors executes may also be statically determined by multiprocessor-aware programming.

[0090] By using the invention according to this embodiment, the accuracy of the learning model 101 that can be used to monitor the vehicle 200 can be improved, and the safety level of the vehicle 200 can be improved, thereby contributing to the achievement of at least one of the Sustainable Development Goals (SDGs): Goal 7 "Affordable and Clean Energy", Goal 9 "Industry, Innovation and Infrastructure", and Goal 11 "Sustainable Cities and Communities".

[0091] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications or improvements can be made to the above embodiments. It will be clear from the claims that such modified or improved forms may also be included in the technical scope of the present invention.

[0092] It should be noted that the execution order of operations, procedures, steps, and stages in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not explicitly stated as "before" or "prior to," and that these can be performed in any order unless the output of a previous operation is used in a later operation. Even if the operation flow in the claims, specifications, and drawings is described using phrases such as "first," and "next," for convenience, this does not mean that it is mandatory to perform the operations in that order.

[0093] 50 Network, 100 Information Processing Device, 101 Learning Model, 102 Inference, 104 Generated Sentence, 106 Evaluation / Analysis, 108 Ground Response Data, 110 Weakness Words, 112 Annotation Data Generation, 114 Annotation Category, 116 Annotation Data, 118 Video Design, 120 Template, 122 Video Parameters, 124 Video Generation, 126 Video, 128 Training Data, 130 Fine Tuning, 150 Memory Unit, 152 Registration Unit, 154 Video Reception Unit, 156 Inference Unit, 158 Evaluation and Analysis Unit, 160 Annotation Data Generation Unit, 162 Annotation Data Acquisition Unit, 164 Video Generation Unit, 166 Model Update Unit, 170 RAN Control Unit, 180 AI Processing Unit, 200 Vehicle, 202 Camera, 210 Video, 300 Management Platform, 400 Distributed infrastructure, 500 wireless base stations, 510 RAN, 1200 computers, 1210 host controllers, 1212 CPUs, 1213 GPUs, 1214 RAM, 1216 graphics controllers, 1218 display devices, 1220 input / output controllers, 1222 communication interfaces, 1224 storage devices, 1226 DVD drives, 1227 DVD-ROMs, 1230 ROMs, 1240 input / output chips

Claims

1. An information processing device comprising: an annotation data acquisition unit that acquires annotation data including weak words identified by the difference between the text output for the input video and the correct text data for the input video, wherein the learning model takes video captured by a camera mounted on a vehicle as input and outputs text representing the situation of the vehicle in the video as output; and a video generation unit that generates a video for training the learning model on the weak words based on the annotation data, wherein the video generation unit generates the video using a template that includes a plurality of categories used to generate the video, which are not included in the annotation data.

2. The information processing apparatus according to claim 1, wherein the template includes a plurality of options for each of the plurality of categories, and the video generation unit generates the video for each variation of the plurality of options for the plurality of categories.

3. The information processing device according to claim 1 or 2, wherein the template includes the number of lanes on the road on which the target vehicle travels.

4. The information processing device according to any one of claims 1 to 3, wherein the template includes the driving speed of the target vehicle.

5. The information processing device according to claim 3 or 4, wherein the template includes the number of lanes on the road where other vehicles are traveling.

6. The information processing device according to claim 3 or 4, wherein the template includes the driving speed of other vehicles.

7. The information processing apparatus according to claim 5 or 6, wherein the template includes the direction of travel of the target vehicle and the direction of travel of the other vehicle.

8. The information processing apparatus according to any one of claims 1 to 7, wherein the annotation data includes the operation of the vehicle, the causative object, the motion of the object, the absolute position of the object, the relative position of the object, the vehicle's driving intention, and the position of the vehicle.

9. The information processing apparatus according to any one of claims 1 to 8, further comprising a model update unit that updates the learning model using a plurality of learning data including the video and annotation data generated by the video generation unit based on the annotation data.

10. An information processing apparatus according to any one of claims 1 to 9, comprising a RAN control unit for controlling RAN and an AI processing unit for performing AI processing, wherein the AI ​​processing unit comprises an annotation data acquisition unit and a video generation unit.

11. An information processing method performed by a computer, comprising: an annotation data acquisition step, in which a learning model that takes a video captured by a camera mounted on a vehicle as input and outputs text representing the state of the vehicle in the video acquires annotation data that includes weak words identified by the difference between the text output for the input video and the correct text data for the input video; and a video generation step, in which a video is generated to allow the learning model to learn the weak words based on the annotation data, wherein the video generation step generates the video using a template that includes a plurality of categories of information that are not included in the annotation data but are necessary to generate the video, which has been registered in advance.

12. A program for causing a computer to perform the following steps: an annotation data acquisition step, in which a learning model takes video footage captured by a camera mounted on a vehicle as input and outputs text describing the state of the vehicle in the video, acquires annotation data including weak words identified by the difference between the text output for the input video and the correct text data for the input video; and a video generation step, in which a video is generated to allow the learning model to learn the weak words based on the annotation data, wherein the video is generated using a template that has been registered in advance and includes multiple categories of information that are not included in the annotation data but are necessary to generate the video.

Citation Information

Patent Citations

  • Learning server, and assist system

    JP2019021201A

  • Image generation apparatus, image generation method, image generation program, and storage medium

    JP2023012856A

  • Method for sample analysis, electronic device, storage medium, and program product

    JP2023042582A

  • Information processing device, information processing method, and information processing program

    JP2023157645A

  • Automated annotation techniques

    US20200019799A1