Database generation system, database generation method, and program

By acquiring and integrating photographic images of the environment surrounding the moving object, detecting target objects and generating descriptive text, and establishing database associations, the latency problem in real-time judgment of moving objects is solved, and efficient information processing and autonomous decision-making are achieved.

CN121958591APending Publication Date: 2026-05-01TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2025-10-16
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies have problems with untimely judgment or response in real-time judgment of the environment around moving objects. In particular, when using pre-trained visual language models, it is difficult to quickly generate information related to objects, resulting in delays or misjudgments in the processing of moving objects.

Method used

By acquiring photographic images of the environment surrounding the moving object, the system detects the target object and generates descriptive text for the target object using a pre-trained model. It then establishes a connection between the target object and the descriptive text, generates a database, and updates the database by integrating multiple descriptive texts, thus achieving real-time information processing.

Benefits of technology

It enables mobile entities to efficiently and accurately process surrounding information in real-time environments, reduces judgment delays, improves the accuracy and real-time performance of responses, and supports autonomous decision-making by mobile entities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958591A_ABST
    Figure CN121958591A_ABST
Patent Text Reader

Abstract

The present invention provides a database generation system which enables a moving body to appropriately process on the basis of information relating to the surrounding environment. A database generation system according to the present invention is provided with: an image acquisition unit that acquires a captured image including the surroundings of a moving body; an object detection unit that detects an image of a target object included in the captured image; a persuasion acquisition unit that acquires a persuasion of the target object by applying the image of the target object to the model; and a database generation unit that associates and associates the target object with the statement and generates a database. The image acquisition unit repeatedly acquires the captured images, the utterance acquisition unit acquires utterance corresponding to each of the repeatedly acquired captured images, and the database generation unit updates the database on the basis of the acquired utterance. In the present invention, an artificial intelligence (AI) technique can be used in the processing of each function unit.
Need to check novelty before this filing date? Find Prior Art

Description

Database generation system, database generation method and program Technical Field

[0001] This invention relates to a database generation system, database generation method, and program. Background Technology

[0002] A technique is known for a mobile device acquiring information related to its surrounding environment and making various judgments based on that information. As a related technology, Patent Document 1 discloses a mobile device control apparatus that detects the parking position of a mobile device by inputting camera images of the mobile device's surroundings and user-inputted instructions into a pre-trained model containing a pre-trained Vision-Language Model (VLM), and then drives the mobile device to the parking position. This pre-trained model is pre-trained to output the parking position of the mobile device in the image corresponding to the input instructions, provided that at least an image and instructions are input.

[0003] Patent Document 1: Japanese Patent Application Publication No. 2024-31978 As with related technologies, the present invention involves a structure that inputs images and instructions into a base model (VLM) each time to obtain information about the surrounding objects (e.g., what the object is and what state it is in). Obtaining this information can be time-consuming. Therefore, especially in real-time situations where judgments are made based on images of the moving object's surroundings, there is a possibility of delayed judgment or response to the moving object.

[0004] This invention provides a database generation system, database generation method, and program that enable a mobile body to appropriately process information related to its surrounding environment.

[0005] The database generation system of the present invention comprises: an image acquisition unit that acquires photographic images of the surrounding environment including a moving object; an object detection unit that detects images of target objects contained in the photographic images; a description text acquisition unit that acquires description text of the target object by applying the image of the target object to a model; and a database generation unit that generates a database by establishing a corresponding association between the target object and the description text, wherein the image acquisition unit repeatedly acquires the photographic images, the description text acquisition unit acquires the description text corresponding to each of the repeatedly acquired photographic images, and the database generation unit updates the database based on the acquired description text.

[0006] Based on the above structure, the moving body can appropriately process information related to the surrounding environment.

[0007] When the explanatory text acquisition unit acquires multiple explanatory texts with different contents for the target object, it integrates the multiple explanatory texts, and the database generation unit uses the integrated explanatory texts to update the database.

[0008] The explanatory text acquisition unit may be configured to integrate the multiple explanatory texts based on the number of times each of the multiple explanatory texts has been acquired.

[0009] The database generation method of the present invention includes: an image acquisition step, acquiring a photographic image of the surrounding environment containing a moving object; an object detection step, detecting an image of a target object contained in the photographic image; a descriptive text acquisition step, acquiring a descriptive text of the target object by applying the image of the target object to a model; and a database generation step, establishing a corresponding association between the target object and the descriptive text to generate a database. The image acquisition step includes the step of repeatedly acquiring the photographic image, the descriptive text acquisition step includes the step of acquiring the descriptive text corresponding to each of the repeatedly acquired photographic images, and the database generation step includes the step of updating the database based on the acquired descriptive text.

[0010] The program involved in this invention enables a computer to perform the following steps: an image acquisition step, acquiring a photographic image of the surrounding environment containing a moving object; an object detection step, detecting an image of a target object contained in the photographic image; a descriptive text acquisition step, acquiring a descriptive text of the target object by applying the image of the target object to a model; and a database generation step, establishing a corresponding association between the target object and the descriptive text to generate a database. The image acquisition step includes the step of repeatedly acquiring the photographic image, the descriptive text acquisition step includes the step of acquiring the descriptive text corresponding to each of the repeatedly acquired photographic images, and the database generation step includes the step of updating the database based on the acquired descriptive text.

[0011] Effects of the Invention: The database generation system, database generation method, and program involved in this invention enable a moving body to appropriately process information related to its surrounding environment. Attached Figure Description

[0012] Figure 1 is a block diagram showing the structure of the database generation system involved in this invention.

[0013] Figure 2 is an example of a photographic image showing the detection of a target object.

[0014] Figure 3 is a diagram showing a specific example of the processing performed by the object detection unit and the explanatory text acquisition unit.

[0015] Figure 4 is a flowchart illustrating the processing flow of the database generation system. Detailed Implementation

[0016] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the drawings, the same or corresponding elements are labeled with the same symbols. For clarity, repeated descriptions are omitted as needed.

[0017] First, the problems addressed in this invention will be described in more detail. In recent years, a technique has been known to assign descriptive text to each object contained in a target image by applying it to a VLM (Visual Model). However, in situations where the state of objects surrounding a moving object (e.g., the presence or absence of objects, their position or posture) may change at any time, even if a surrounding image with accompanying descriptive text is prepared in advance, the moving object may sometimes be unable to properly refer to that image. Especially in cases where judgments are made in real time, situations may arise where the moving object's judgment or response is delayed, or where judgment or response cannot be made at all.

[0018] Furthermore, a technique is known to assign feature data (e.g., coordinate data embedded in a multidimensional space) to each object contained in a target image by applying a Virtual Model (VLM). However, in this method of using feature data, if the pre-trained VLM used to calculate the feature data is different from the actual model used, the feature data of each object contained in the target image will have different values, even when using the same target image. In this case, it is possible that even if a surrounding image with the feature data is prepared in advance, the moving object cannot utilize that surrounding image. Therefore, for example, in the case of generating map data using a VLM, once the generated map data is generated, it can only be used with the model used to generate that map data. Therefore, when using a new model, the map data needs to be generated again.

[0019] Furthermore, in existing methods for assigning descriptive text to objects, all data is processed before the text is generated when map data is generated. Therefore, the information from the descriptive text cannot be utilized in loop closures or repositioning during the map data generation process.

[0020] The purpose of this invention is to realize a database generation system, which, especially when the moving body is judged in real time, suppresses the decrease in the accuracy of the judgment or response and has high utilization.

[0021] (Database Generation System 10) Figure 1 is a block diagram showing the structure of the database generation system 10 according to the present invention. The database generation system 10 includes an image acquisition unit 1, an object detection unit 2, a descriptive text acquisition unit 3, a database generation unit 4, and a photography unit 5. Furthermore, the directional arrows shown in the figure are simply arrows indicating the direction of information (data or signals, etc.), and do not exclude the bidirectional nature of information flow.

[0022] The database generation system 10 includes a processor, a memory, and a storage device (structure not shown). The storage device stores a computer program that implements the processing according to the present invention. The processor can read the computer program from the storage device into the memory and execute the computer program. Thus, the processor performs the functions of the image acquisition unit 1, the object detection unit 2, the explanatory text acquisition unit 3, and the database generation unit 4.

[0023] Alternatively, the image acquisition unit 1, object detection unit 2, description text acquisition unit 3, and database generation unit 4 can each be implemented using dedicated hardware. Alternatively, some or all of each component can be implemented using general-purpose or special-purpose circuits, processors, or combinations thereof. These can be constructed from a single chip or from multiple chips connected via a bus. Some or all of each component can be implemented using the aforementioned circuits and programs. Furthermore, as the processor, a central processing unit (CPU), graphics processing unit (GPU), field-programmable gate array (FPGA), quantum processor (quantum computer control chip), etc., can be used.

[0024] Furthermore, when some or all of the components of the database generation system 10 are implemented by multiple information processing devices or circuits, these devices or circuits can be centrally or distributed. For example, the information processing devices or circuits can be implemented as client server systems, cloud computing systems, etc., connected via a communication network.

[0025] The database generation system 10 is, for example, a system mounted on a mobile body moving in an environment. The environment is any space where the mobile body can exist. The environment is, for example, a residence, public facility, office, shop, or hospital, but is not limited to these. In addition, the environment can be indoors or outdoors.

[0026] The mobile body has a mobility mechanism for moving in the environment. The mobility mechanism may include, for example, a drive mechanism for moving on a surface, in the air, on water, or in water. The mobile body may be, for example, a mobile robot, an autonomous vehicle, or a drone. The mobile body can be capable of autonomous movement or can be manually moved by an operator.

[0027] The following description uses an example of a database generation system 10 mounted on a mobile body, but it is not limited to this. The database generation system 10 can also be implemented by a mobile body and a database generation device. In this case, for example, the structure can be as follows: the mobile body has a camera unit 5, the database generation device has an image acquisition unit 1, an object detection unit 2, a description acquisition unit 3, and a database generation unit 4, and the mobile body sends the acquired image to the database generation device.

[0028] Image acquisition unit 1 acquires photographic images of the surrounding environment of the moving object from photography unit 5. The photographic images can be still images or moving images. Image acquisition unit 1 can repeatedly acquire photographic images. For example, image acquisition unit 1 can repeatedly acquire photographic images at predetermined time intervals. Image acquisition unit 1 can also continuously acquire photographic images. Therefore, image acquisition unit 1 can acquire photographic images in real time. Image acquisition unit 1 can repeatedly acquire photographic images based on the moving distance of the moving object, etc. For example, image acquisition unit 1 can acquire photographic images at predetermined moving distances.

[0029] The object detection unit 2 detects images of target objects contained in photographic images. The target object is an object that has been assigned a description in the database generation system 10. The target object may be, for example, an item, a person, an animal, or a building, but is not limited to these. The object detection unit 2 can use any image recognition technique to detect the image of the target object. For example, the object detection unit 2 can use artificial intelligence (AI) techniques.

[0030] For example, the object detection unit 2 detects the region of the target object using a detector that has the function of estimating the position, size, and type of the target object within the photographic image. The detector can be a detector that outputs a rectangle enclosing the object, such as in the YOLO series of partial object detection algorithms. Alternatively, a model such as Mask R-CNN or YOLOv8 can be used, which estimates mask information representing the pixels corresponding to the detected object in addition to the rectangle.

[0031] Figure 2 is an example of a photographic image P1 in which target objects are detected. In this example, the object detection unit 2 detects target objects T1 and T2 from the photographic image P1. In the figure, target objects T1 and T2 are surrounded by rectangular frames that enclose the target objects. The rectangular frames are represented by thick solid lines in the figure. Target object T1 is an "apple," and target object T2 is a "clock." In addition, both target objects T1 and T2 are placed on a table (not shown). The object detection unit 2 outputs the information of the detected target objects to the description acquisition unit 3.

[0032] Returning to Figure 1. The descriptive text acquisition unit 3 acquires the descriptive text of the target object by applying an image of the target object to the model. The descriptive text is also called a language caption or simply a caption. The descriptive text is a string describing the visible state of the observed images of various landmarks existing around the moving object.

[0033] In recent years, techniques for visual / language reasoning tasks such as image captioning using deep learning have been developed, enabling the use of pre-trained models that leverage these techniques. Image captioning is a task where a language generation model outputs descriptive text for a given image. The descriptive text acquisition unit 3 can assign meaningful information to the target object using language that is model-independent. This allows for flexible utilization of different recognition models. The meaningful information related to the target object can be used for things like planning the movement of a moving object.

[0034] The description text acquisition unit 3 can acquire description text by generating it using this pre-trained model. The model is pre-trained to output description text of the target object when an image of the target object is input. The model can be stored in the storage unit (not shown) of the database generation system 10. The description text acquisition unit 3 uses the model to acquire description text corresponding to each of the photographic images repeatedly acquired in the object detection unit 2.

[0035] As a result, multiple explanatory texts can be obtained for a single landmark. When the explanatory text acquisition unit 3 obtains multiple explanatory texts with different contents for a target object, it can integrate these multiple explanatory texts. Two methods for integrating explanatory texts are as follows:

[0036] (First Integration Method) The first integration method is to record the frequency of the generated explanatory texts and select the explanatory text with the highest frequency. The explanatory text acquisition unit 3 records the number of times each explanatory text is generated. The explanatory text acquisition unit 3 updates the explanatory text with the highest frequency each time a new explanatory text is generated, thereby successively assigning the maximum likelihood explanatory text to the landmark. In terms of implementation, for example, it is possible to consider storing the frequency in an associative array with the explanatory text as the key, and incrementing the corresponding value each time an explanatory text is generated. Moreover, in order to improve the consistency of the explanatory texts, it is also possible to consider performing syntactic analysis and extracting noun phrases.

[0037] (Second Integration Method) The second integration method is a clustering method in the embedding space. The explanatory text acquisition unit 3 uses a language encoder to embed the explanatory text generated by the model into the feature space and perform clustering. For example, it is possible to consider selecting the explanatory text corresponding to the centroid of the closest largest cluster as the representative method.

[0038] Figure 3 is a diagram illustrating a specific example of the processing performed by the object detection unit 2 and the description acquisition unit 3. As shown in the upper part of Figure 3, the object detection unit 2 acquires multiple photographic images from the image acquisition unit 1 and detects the target object T1 in each of the multiple photographic images. In Figure 3, the target objects detected by the object detection unit 2 are represented by target objects A1, A2, A3, ... in the detection order.

[0039] As shown in the middle section of Figure 3, the explanatory text acquisition unit 3 acquires explanatory texts for each of the target objects A1, A2, A3, ... The explanatory text acquisition unit 3 inputs each of the target objects A1, A2, A3, ... into a pre-trained model and acquires the explanatory texts output from the model. The explanatory texts corresponding to each of the target objects A1, A2, A3, ... are explanatory texts B1, B2, B3, ...

[0040] Explanatory texts B1 and B2 are "a red apple on a table," and explanatory text B3 is "an apple on a desk." Thus, the explanatory text acquisition unit 3 acquires multiple explanatory texts with different content. In this case, the explanatory text acquisition unit 3 integrates the multiple explanatory texts. Here, an example of the first integration method described above is used. The explanatory text acquisition unit 3 integrates the multiple explanatory texts into one based on the number of times each of the multiple explanatory texts has been acquired.

[0041] The description text acquisition unit 3 records the number of times each description text is generated. As shown in the lower part of Figure 3, the number of times "a red apple on a table" is acquired is set to the maximum. The description text acquisition unit 3 integrates multiple description texts B1, B2, B3, ... into "a red apple on a table". Therefore, the description text acquisition unit 3 assigns the description text "a red apple on a table" to the target object T1. The description text assigned to the target object can be directly used in action plans that utilize a Large Language Model (LLM).

[0042] For example, in existing monocular depth estimation techniques, explanatory text is generated and integrated after all the data used to generate map data has been processed. This technique cannot assign explanatory text sequentially, and therefore cannot be used during map data generation. In contrast, the explanatory text acquisition unit 3, by using the algorithm described above, can generate and integrate explanatory text sequentially from the acquired image.

[0043] Returning to Figure 1, the database generation unit 4 generates a database by establishing a correspondence between the target object and the explanatory text. Furthermore, the database generation unit 4 updates the database based on the acquired explanatory text. Database updates include modifications, additions, and deletions. When multiple explanatory texts in the explanatory text acquisition unit 3 are consolidated into one, the database generation unit 4 updates the database using the consolidated explanatory text.

[0044] Databases can be, for example, map data with accompanying descriptions or lists that establish corresponding associations between the coordinates of target objects and their descriptions. The database method is not limited to these. In this invention, a database is sufficient whenever at least one object contained in an image is assigned a description. Hereinafter, map data with accompanying descriptions will be used as an example of a database.

[0045] The database generation unit 4 generates and updates landmarks based on the detection results of target objects in the object detection unit 2. The database generation unit 4 uses the results obtained from the object detection unit 2 to generate landmarks in three-dimensional space. The shape representation of the landmarks is not limited to a specific shape and can utilize any method using image observation. Examples of such methods include the following. The database generation unit 4 can reconstruct landmarks using one or more of the following methods.

[0046] (1) The ellipsoid representation database generation unit 4 can use existing map generation methods that represent landmarks as three-dimensional ellipsoids. When the photography unit 5 uses an RGB camera, there are known methods to reconstruct ellipsoids from observations from multiple viewpoints as constraints. Furthermore, when using an RGB-D camera capable of depth estimation, it is also possible to directly reconstruct ellipsoids from observations from a single viewpoint.

[0047] (2) The cuboid representation database generation unit 4 can use a method to fit the three-dimensional shape of the landmark into a cuboid. Similar to ellipsoids, there are known methods for reconstruction using RGB observations from multiple viewpoints. Furthermore, the database generation unit 4 can also generate a cuboid containing a point cloud reconstructed from RGB-D data.

[0048] (3) The point cloud representation database generation unit 4 can use a method to represent an object as a three-dimensional point cloud. When using RGB-D images, the database generation unit 4 can directly reconstruct the three-dimensional point cloud by integrating the observed object and depth information. Furthermore, in recent years, monocular depth estimation techniques that estimate the depth of each pixel from RGB images have been developed, and these techniques can also be applied to this invention.

[0049] In the explanatory text acquisition unit 3, explanatory texts for landmarks can be generated sequentially during the process of generating map data with accompanying explanatory texts. Therefore, the database generation unit 4 can utilize the explanatory texts when generating map data with accompanying explanatory texts.

[0050] As one application example, there is repositioning, which involves re-identifying the self-position when the error or uncertainty of successive self-position estimations increases. Furthermore, as another application example, loop closure can be listed. Loop closure aims to ensure consistency with previous observations when re-observing previously observed objects, a task similar to repositioning. Moreover, existing methods include estimating camera position using descriptive information on a constructed map. This method can also be applied to the present invention to achieve the aforementioned tasks.

[0051] The database generation unit 4 can acquire movement information of the moving body. By using a moving device of the moving body equipped with the camera unit 5, the database generation unit 4 can acquire movement data obtained from a wheel encoder or the like as movement information. The database generation unit 4 can estimate the movement data of the camera unit 5 based on the movement information.

[0052] The camera unit 5 captures images of the environment surrounding the moving object. The camera unit 5 then outputs the captured images to the image acquisition unit 1. The camera unit 5 can capture images at predetermined time intervals or based on the movement of the moving object. For example, the camera unit 5 can capture images based on the distance the moving object has traveled. By moving alongside the moving object while capturing images, the camera unit 5 captures images of the target object from different positions. Thus, the camera unit 5 captures images of the same object from multiple viewpoints, obtaining images from each viewpoint.

[0053] The imaging unit 5 can be, for example, an RGB camera, an RGB-D sensor, a stereo camera, or a 3D LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging). It is not limited to these; any sensor capable of capturing images of the area around a moving object can be used as the imaging unit 5. A map generated in the database generation unit 4 can be used by different sensors. This is an advantage compared to sensor-dependent map representations such as image feature-based maps.

[0054] (Processing by Database Generation System 10) Referring to Figure 4, the processing performed by the database generation system 10 will be explained. Figure 4 is a flowchart showing the process performed by the database generation system 10.

[0055] First, the image acquisition unit 1 acquires a photographic image from the photography unit 5 (S1). Next, the object detection unit 2 detects an image of the target object from the acquired photographic image (S2). Next, the description text acquisition unit 3 applies the image of the target object to the model to acquire a description text of the target object (S3). The description text acquisition unit 3 stores the acquired description text in the storage unit.

[0056] Furthermore, the description text acquisition unit 3 refers to the storage unit and determines whether there are multiple description texts with different content for the target object (S4). For example, suppose the target object is "apple". The description text acquisition unit 3 confirms whether multiple description texts with different content have been assigned to "apple".

[0057] If it is determined that there are multiple explanatory texts with different content (S4 "Yes"), the explanatory text acquisition unit 3 integrates the explanatory texts (S5). For example, the explanatory text acquisition unit 3 can integrate multiple explanatory texts based on the number of times each of the multiple explanatory texts has been acquired. For example, the explanatory text acquisition unit 3 selects the explanatory text with the most acquisitions as the explanatory text to be assigned to the target object. The database generation unit 4 uses the integrated explanatory text to update the map data with accompanying explanatory texts (S6).

[0058] Furthermore, if it is determined in step S4 that there are no multiple description texts with different content (S4 "No"), the database generation unit 4 uses the description text acquired by the description text acquisition unit 3 to update the map data with accompanying description text. The database generation unit 4 establishes a corresponding association between the target object and the description text to update the map data with accompanying description text. Alternatively, if this process is performed without creating map data with accompanying description text, the database generation unit 4 establishes a corresponding association between the target object and the description text to generate map data with accompanying description text.

[0059] Next, the database generation system 10 determines whether to end the processing (S7). If it determines that the processing should end (S7 "Yes"), the processing ends; if it determines that the processing should not end (S7 "No"), it returns to step S1 and repeats the processing. The database generation system 10 may determine that the processing should end when all images of the target object detected in step S2 have been processed, and determine that the processing should not end when there are unprocessed images. In this case, the database generation system 10 may repeat the processing after step S3.

[0060] As explained above, the database generation system 10 of this invention utilizes explanatory text during the map data generation process by performing successive explanatory text generation and filtering. Therefore, the database generation system 10 can minimize the time delay from including target objects in the surrounding environment into the image to obtaining a database (map data with accompanying explanatory text) reflecting the latest situation. Thus, the database generation system 10 can identify target objects located around a moving object in real time.

[0061] Thus, especially when the mobile body makes judgments in real time, the database generation system 10 can generate a database that suppresses the decline in the accuracy of judgments or responses and has high utilization. The database generation system 10 enables the mobile body to process information appropriately based on information related to the surrounding environment.

[0062] Each functional component of the database generation system 10 described above can be implemented by hardware (e.g., hardwired electronic circuits) or by a combination of hardware and software (e.g., a combination of electronic circuits and programs that control them). For example, the present invention can also achieve arbitrary processing by having the CPU execute a computer program.

[0063] Additionally, the program contains a set of commands (or software code) used to cause the computer to perform one or more functions described in the implementation when read into a computer. The program can be stored on various types of non-transitory computer-readable media or physical storage media. By way of example, and not limitation, non-transitory computer-readable media or physical storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray disc or other optical disc storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices. Furthermore, the program can also be transmitted on various types of temporary computer-readable media or communication media. By way of example, and not limitation, temporary computer-readable media or communication media include electrical, optical, acoustic, or other forms of propagation signals.

[0064] Furthermore, the present invention is not limited to the embodiments described above, and appropriate modifications can be made without departing from the spirit of the invention. For example, in the above description, the generation of map data with accompanying explanatory text for mobile bodies such as mobile robots has been mainly described, but the scope of the present invention is not limited thereto. The present invention can also be used, for example, in map generation for estimating the position of a car. For example, by using a combination of a stereo camera more suitable for outdoor use or a 3D LiDAR and a monocular camera as the photography unit 5, landmark reconstruction and explanatory text acquisition can be achieved.

[0065] Symbol Explanation: 1-Image Acquisition Unit, 2-Object Detection Unit, 3-Descriptive Text Acquisition Unit, 4-Database Generation Unit, 5-Photography Unit, 10-Database Generation System, A1~A3, T1, T2-Target Object, B1~B3-Descriptive Text, P1-Photographed Image.

Claims

1. A database generation system, characterized in that, It includes: an image acquisition unit that acquires a photographic image of the surrounding environment containing the moving object; an object detection unit that detects an image of a target object contained in the photographic image; and a description text acquisition unit that acquires a description text of the target object by applying the image of the target object to a model. The system includes a database generation unit that establishes a corresponding association between the target object and the explanatory text to generate a database; an image acquisition unit that repeatedly acquires the photographic images; an explanatory text acquisition unit that acquires the explanatory text corresponding to each of the repeatedly acquired photographic images; and a database generation unit that updates the database based on the acquired explanatory text.

2. The database generation system according to claim 1, characterized in that, When the explanatory text acquisition unit acquires multiple explanatory texts with different contents for the target object, it integrates the multiple explanatory texts, and the database generation unit uses the integrated explanatory texts to update the database.

3. The database generation system according to claim 2, characterized in that, The explanatory text acquisition unit integrates the multiple explanatory texts based on the number of times each of the multiple explanatory texts has been acquired.

4. A database generation method, characterized in that, include: The image acquisition step involves acquiring photographic images of the surrounding environment containing the moving object; The object detection step detects images of target objects contained in the photographic image; The explanatory text acquisition step involves obtaining the explanatory text of the target object by applying an image of the target object to a model; and the database generation step involves establishing a corresponding association between the target object and the explanatory text to generate a database. The image acquisition step includes the step of repeatedly acquiring the photographic images, the explanatory text acquisition step includes the step of acquiring the explanatory text corresponding to each of the repeatedly acquired photographic images, and the database generation step includes the step of updating the database based on the acquired explanatory text.

5. A program, characterized in that, The computer performs the following steps: an image acquisition step, acquiring photographic images of the surrounding environment containing the moving object; The system includes an object detection step, which detects images of target objects contained in the photographed images; a descriptive text acquisition step, which obtains descriptive text of the target objects by applying the images of the target objects to a model; and a database generation step, which establishes a corresponding association between the target objects and the descriptive texts to generate a database. The image acquisition step includes the step of repeatedly acquiring the photographed images, the descriptive text acquisition step includes the step of acquiring the descriptive texts corresponding to each of the repeatedly acquired photographed images, and the database generation step includes the step of updating the database based on the acquired descriptive texts.

Citation Information

Patent Citations

  • Mobile object control device, mobile object control method, learning device, learning method, generation method, and program

    JP2024031978A