Control device, control method, information processing device, information processing method, and image capture control device
Converting captured images to text data for storage in a database addresses storage and privacy issues, facilitating efficient and secure analysis of imaged subjects.
Patent Information
- Application Number
- PCT/JP2025/022452
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-05
- Filing Date
- 2025-06-23
- Publication Date
- 2026-01-08
Smart Images

Figure JP2025022452_08012026_PF_FP_ABST
Abstract
Description
Control device, control method, information processing device, information processing method, and imaging control device
[0001] The present technology relates to a control device, a control method, an information processing device, an information processing method, and an imaging control device, and in particular to a technology for performing inference processing using an AI (Artificial Intelligence) model on a captured image.
[0002] For example, there is a technology that performs inference processing using an AI (Artificial Intelligence) model on sensing data such as captured image data, and there is also a technology that performs various analyses on the sensing target based on the results of such inference processing. For example, it is possible to perform inference processing such as person detection processing using an AI model on image data captured by an imaging device installed in a store, and then perform analysis such as counting the number of customers in the store based on the inference results.
[0003] As a related prior art, the following Patent Document 1 can be cited: Patent Document 1 discloses an inference device that performs inference processing on image data as target data, and that is an integrated sensor in which a signal processing unit that performs inference processing is implemented within an image sensor in which a pixel array unit is formed.
[0004] International Publication No. 2023 / 090119
[0005] Here, consider storing captured image data in a database and performing various analyses of the imaged subject based on the data stored in the database. In this case, a configuration can be adopted in which the captured image data is sent from the imaging device to the database, and an information processing device that performs the analysis retrieves the captured image data from the database and performs inference processing using an AI model and analysis processing based on the inference results.
[0006] However, storing captured images in a database as described above is undesirable because it increases the storage capacity of the database, and also because storing captured images in a database tends to increase the amount of communication data between the image capture device and the database, and between the database and the information processing device, which is also undesirable in this respect.
[0007] Furthermore, considering the risk of the captured images being leaked during communication with the database, storing the captured images in the database is undesirable from the viewpoint of privacy protection.
[0008] This technology was developed in consideration of the above circumstances, and aims to protect privacy in a system that analyzes imaged subjects using an AI model based on information related to captured images stored in a database, while reducing the storage capacity of the database and the amount of communication data between the database.
[0009] The control device according to the present technology includes an image text conversion unit that generates image text data in which the content of visual information represented by a captured image is converted into natural language by performing natural language conversion on a captured image using a natural language conversion model that converts the content of visual information represented by an image into natural language, and a storage control unit that controls storage of the image text data obtained by the image text conversion unit in a database. According to the above configuration, the stored information in the database is converted into text.
[0010] Furthermore, an information processing device according to the present technology includes a prompt receiving unit that receives a prompt for a large-scale language model that generates response information to a prompt based on accumulated information in a database in which image text data in which content of visual information from captured images is converted into natural language, and a presentation processing unit that performs processing to present to a user the response information generated by the large-scale language model in response to the prompt received by the prompt receiving unit. This makes it possible to realize a system that analyzes an imaged object based on accumulated information in which text data of captured images is accumulated, and that can present to a user appropriate analysis information in response to various input prompts in natural language.
[0011] In addition, an imaging control device according to the present technology includes an imaging control unit that inputs image text data, which is obtained by converting the content of visual information from a captured image into natural language, into a large-scale language model, causes the large-scale language model to infer imaging settings or image signal processing settings according to the content of the image text data, and controls the imaging device so that imaging operation or image signal processing based on the inferred imaging settings or image signal processing settings is performed in the imaging device.As a result, in a system that analyzes an imaging target based on accumulated information that has accumulated text data of captured images, camera operation can be automated to obtain accumulated information appropriate for analysis, such as by optimizing the angle of view of the imaging device or performing image signal processing appropriate for analysis of the target subject.
[0012] 1 is a block diagram showing an example of a schematic configuration of an information processing system as an embodiment, which is configured with a control device and an information processing device according to the present technology. FIG. 1 is a diagram illustrating an example of a plurality of imaging devices when installed for road monitoring. FIG. 2 is a diagram illustrating an example of a plurality of imaging devices when installed for home monitoring. FIG. 3 is a block diagram showing an example of a hardware configuration of a control device as an embodiment. FIG. 4 is a block diagram showing an example of a hardware configuration of an information processing device as an embodiment. FIG. 5 is a block diagram showing an example of a hardware configuration of a database and a user terminal in an embodiment. FIG. 6 is an explanatory diagram of various functions possessed by a control device as an embodiment. FIG. 7 is a diagram showing an example of a relationship between captured images and image-to-text data when used for traffic monitoring. FIG. 8 is a diagram showing an example of a relationship between captured images and image-to-text data when used for home monitoring of a specific person. FIG. 9 is an explanatory diagram of various functions possessed by an information processing device as an embodiment. FIG. 10 is a diagram showing an example of a prompt and response information when used for traffic monitoring. FIG. 11 is a diagram showing an example of a prompt and response information when used for home monitoring of a specific person. FIG. 12 is a flowchart showing an example of a processing procedure for realizing the functions of an information processing device as an embodiment. FIG. 13 is a block diagram showing an example of a schematic configuration of an imaging device having an image-to-text unit and a storage control unit. FIG. 1 is a block diagram showing an example of the schematic configuration of a control device having a function of converting sensing information other than images into natural language. FIG. 2 is an explanatory diagram of an example of converting sensing information other than images into text. FIG. 3 is a block diagram showing an example of the schematic configuration of an imaging device in the case where a plurality of types of captured images are obtained using a mixed sensor. FIG. 4 is a diagram showing an example of the configuration of a mixed sensor with in-plane mixed mounting. FIG. 5 is a diagram showing an example of the configuration of a mixed sensor with a stacked structure. FIG. 6 is a diagram showing an example of the pixel array structure of an image sensor compatible with binning. FIG. 7 is a diagram showing an example of a pixel array structure compatible with HDR imaging. FIG. 8 is an explanatory diagram of an example of the schematic configuration of a control device that controls imaging settings and image signal processing settings of an imaging device based on image text data.
[0013] Hereinafter, with reference to the accompanying drawings, embodiments of the present technology will be described in the following order: <1. Overview of the configuration of an information processing system as an embodiment> <2. Example of the hardware configuration of each device> <3. Storage control method as an embodiment> <4. Information processing method as an embodiment> <5. Modification> <6. Summary of the embodiment> <7. The present technology>
[0014] 1 is a block diagram showing a schematic configuration example of an information processing system according to an embodiment, which is configured to include a control device and an information processing device according to the present technology. As shown in the figure, the information processing system according to the embodiment includes a control device 1, an information processing device 2, an imaging device 3, a database 4, and a user terminal 5.
[0015] The control device 1, information processing device 2, database 4, and user terminal 5 are each configured as a computer device equipped with a microcomputer having a CPU (Central Processing Unit), ROM (Read Only Memory), and RAM (Random Access Memory). In this example, the control device 1 is capable of performing data communication with at least the database 4 via a network NT, which is a predetermined communication network such as the Internet. The information processing device 2 is also configured to be capable of performing data communication with at least the database 4 and the user terminal 5 via the network NT. Note that the network used by the control device 1 for communication with the database 4 and the network used by the information processing device 2 for communication with the database 4 and the user terminal 5 do not need to be the same.
[0016] The control device 1 is also capable of performing data communication with the imaging device 3 .
[0017] The imaging device 3 obtains a captured image of a subject. In this specification, "capturing" broadly refers to obtaining image data capturing the subject. The term "image data" here refers collectively to data consisting of multiple pixel data. The concept of pixel data broadly encompasses not only data indicating the intensity of light received from the subject, but also data indicating, for example, the distance to the subject, polarization information about the subject, and temperature information. In other words, the "image data" (captured image data) obtained by "capturing" includes data such as a grayscale image indicating the intensity of light received for each pixel, a distance image indicating the distance to the subject for each pixel, a polarization image indicating the polarization information of incident light for each pixel, and a thermal image indicating temperature information for each pixel. Furthermore, "image data" also includes data such as an event image obtained by an event-based vision sensor (EVS) having an event sensor with a two-dimensional array of event detection pixels that detect changes in the amount of received light as events. This event image data can be described as image data indicating whether an event has occurred for each pixel, and can be expressed as image data capturing the movement of the subject.
[0018] Here, it is assumed that the imaging device 3 is configured to obtain a captured image using the above-described gradation image. Specifically, it is assumed that the imaging device 3 is configured to obtain an RGB image as the captured image, that is, to obtain an image showing the luminance values of R (red), G (green), and B (blue) for each pixel.
[0019] In this example, each imaging device 3 outputs a captured image as a moving image to the control device 1 .
[0020] The information processing system in this example includes a plurality of imaging devices 3. In this example, surveillance is assumed as the use of the imaging devices 3. For example, Fig. 2 illustrates an example of a plurality of imaging devices 3 installed for road surveillance, and Fig. 3 illustrates an example of a plurality of imaging devices 3 installed for in-home surveillance.
[0021] 2, it is conceivable to install a plurality of sets of imaging devices 3 with different imaging directions so that the monitoring target can be captured from different directions, for example, from the front to the rear. Also, in this case, it is conceivable to arrange the imaging devices 3 so that the imaging fields of view of adjacent imaging devices 3 overlap, so that the monitoring target can be easily tracked.
[0022] Similarly, in the home monitoring application shown in FIG. 3 , it is possible to install a set of imaging devices 3 with different imaging directions so that the monitoring target can be captured from different directions. Note that the installation manner of the imaging devices 3, including the number of imaging devices 3 to be installed, may be determined appropriately depending on the application and is not limited to a specific installation manner. For example, for a simple traffic volume analysis application, it is possible to install multiple imaging devices 3 with the same imaging direction, such as installing multiple imaging devices 3 with imaging directions perpendicular to the direction in which the road extends. Furthermore, for a home monitoring application, for example, simply determining whether the user is at home or not, it is possible to install a single imaging device 3 only in a specified room such as the living room, or in each room.
[0023] In FIG. 1 , the control device 1 includes an image-to-text unit 14. The image-to-text unit 14 converts the content of visual information from captured images obtained by the imaging device 3 into natural language. The technology for converting the content of visual information from images into natural language is known as image captioning technology, and can be performed using an AI (Artificial Intelligence) model. Hereinafter, text data obtained by such natural language conversion processing as image captioning will be referred to as "image-to-text data." In this example, the image-to-text unit 14 performs natural language conversion processing using images captured by multiple imaging devices 3 as input. Specifically, it uses an AI model (a natural language conversion model MT, described below) that has been trained to convert the aspects of the objects captured in these multiple captured images into natural language. Details of the natural language conversion performed by the image-to-text unit 14 will be described later.
[0024] The control device 1 also performs a process of storing the image-to-text data obtained by the image-to-text unit 14 in the database 4. This eliminates the need to store the captured images themselves in the database 4, making it possible to reduce the storage capacity of the database 4 and the amount of communication data between the database 4. It also reduces the risk of the captured images being leaked to the outside, thereby ensuring privacy protection. Hereinafter, the image-to-text data stored in the database 4 will be referred to as "stored information 4a" as shown in the figure.
[0025] The information processing device 2 is a device for presenting to a user analytical information about an imaged subject captured by the imaging device 3. The user here refers to a person who receives an analytical service for an imaged subject using the information processing system. The information processing device 2 is a computer device intended to be used by a provider of the analytical service, and the user terminal 5 is a computer device intended to be used by a user.
[0026] The information processing device 2 has a large language model (LLM) ML. As is well known, the large language model ML, such as GPT (Generative Pre-trained Transformers) or BERT (Bidirectional Encoder Representations from Transformers), is an AI model capable of executing various natural language processing tasks, which generates response information according to the content of an input prompt. For example, when a question sentence is input as a prompt, the LLM generates answer information to the question by performing a data search process according to the question content and a sentence generation process according to the search results. In addition to the task of generating answer information to such a question, the LLM can also execute various other tasks for generating response information according to the content of the input prompt, such as a task of creating a computer program that realizes processing specified by the prompt and a task of generating an image or music that satisfies conditions specified by the prompt.
[0027] The information processing device 2 according to the embodiment accepts a prompt input from the user terminal 5. In the information processing device 2, the large-scale language model ML generates response information to the prompt based on the stored information 4a in the database 4. For example, if the imaging device 3 is installed for road monitoring purposes, in response to an input of a prompt such as "I would like to know the traffic volume at X intersection yesterday," the large-scale language model ML may search a group of image-to-text data stored as the stored information 4a and generate response information such as "The traffic volume at X intersection yesterday was X vehicles." Alternatively, in response to an input of a prompt such as "What dangerous incidents occurred in the past three days?", the large-scale language model ML may generate response information such as "In the past three days, X vehicles crossed the crosswalk at speeds above the speed limit." Furthermore, if the imaging device 3 is used for monitoring a home, in response to a prompt input such as "What dangerous incidents have occurred in the past month?", the imaging device 3 may generate response information such as "In the past month, there have been xx incidents of xx stumbling when standing up." Alternatively, if the imaging device 3 is used for monitoring a warehouse, in response to a prompt input such as "What are the incidents of intrusions between 10 PM and 6 AM in the past week?" the imaging device 3 may generate response information such as "In the past week, there have been x incidents of humans or animals intrusions into this warehouse between 10 PM and 6 AM." The information processing device 2 performs a process of presenting the response information generated by the large-scale language model ML as described above to the user who input the prompt. It is also possible to play the corresponding video together with the presentation of the response information.
[0028] In this way, in the information processing system of the embodiment, analysis of the imaged subject by the imaging device 3 is performed in response to input prompts from the user, and the analysis results are presented to the user. At this time, image-to-text data of past events can be stored in the database 4, and therefore, according to the information processing system of the embodiment, it is possible to perform post-event analysis from various perspectives specified in the prompts based on the stored information of image-to-text data of such past events, thereby realizing a highly functional analysis system.
[0029] 1, the number of user terminals 5 is one, but there may be a plurality of user terminals 5. In other words, it is also possible that there are a plurality of users who will receive services provided by the information processing system.
[0030] 2. Example of Hardware Configuration of Each Device> With reference to FIGS. 4 to 6, an example of the hardware configuration of the control device 1, the information processing device 2, the database 4, and the user terminal 5 shown in FIG. 1 will be described.
[0031] 4 is a block diagram showing an example of the hardware configuration of the control device 1. As shown in the figure, the control device 1 includes a CPU 11. The CPU 11 executes various processes in accordance with programs stored in a ROM 12 or programs loaded from a storage unit 19 into a RAM 13. The RAM 13 also stores data necessary for the CPU 11 to execute various processes as appropriate.
[0032] The CPU 11, ROM 12, and RAM 13 are connected to one another via a bus BS. The bus BS is also connected to the image-to-text unit 14. The image-to-text unit 14 performs natural language translation of the captured image using an AI model as a natural language translation model MT that translates the content of visual information from an image into natural language, thereby generating image-to-text data in which the content of the visual information from the captured image is translated into natural language.
[0033] The following models can be used as natural language processing models (MT): NIC (Neural Image Captioning): A model announced by Google in 2014 that combines a CNN (Convolutional Neural Network) and LSTM (Long Short Term Memory). It extracts feature vectors from images and passes them to LSTM to generate text. Show and Tell: An improved version of NIC, announced by Google in 2015. While the structures of CNN and LSTM are the same, improvements to the training and evaluation methods have enhanced image captioning performance. Show, Attend, and Tell: A model announced by the University of Montreal in Canada in 2015 that incorporates a mechanism called self-attention in addition to CNN and LSTM. Self-attention aims to generate more accurate captions by weighting the parts of an image's feature vector that are relevant to text generation.・Transformer-based models: A model with multiple layers of self-attention, announced by Google in 2017. Originally developed in the field of natural language processing, it has also been applied to image captioning. There are also models that generate text from images using only the Transformer without using a CNN.
[0034] AI models that perform image captioning (image captioning models) are generally trained using supervised learning techniques. Specifically, machine learning is performed using pairs of image and text captions as training data. Image captioning models are broadly divided into two parts: a part that extracts image features and a part that generates text. The part that extracts image features uses a model such as a CNN or Transformer to convert the image into a numerical vector. This vector represents information such as the content and shape of the image. The part that generates text uses a model such as an RNN or Transformer to sequentially generate text based on the image feature vector. At this time, the next word is predicted using information such as the generated words and context. When training an image captioning model, an index called a loss function is used to measure the difference between the model output and the correct text, and the model parameters are updated to minimize this difference. Common loss functions include cross-entropy loss and BLUE score.
[0035] The type of natural language model MT is not particularly limited, and it is possible to adopt not only existing models but also models to be developed in the future.
[0036] As described above, the image-to-text unit 14 in this example performs natural language processing using images captured by multiple imaging devices 3 as input, and the natural language model MT uses an AI model trained to natural languageize the aspects of the objects captured in the multiple images. Because the image-to-text unit 14 is configured to generate image-to-text data using images captured by multiple imaging devices 3 as input, it becomes possible to recognize events that would not be recognized if only images captured by a single imaging device 3 were input, such as recognizing that a subject captured by one imaging device 3 is approaching another subject captured by another imaging device 3. This allows the amount of information in the image-to-text data stored in the database 4 to be increased, thereby improving the usefulness of the stored information 4a.
[0037] Furthermore, in this embodiment, the natural language model MT of the image-to-text unit 14 is a model trained to naturalize the state of the imaged object. For the image-to-text unit 14, naturalizing changes in the state of the imaged object from images is highly difficult and requires resources and processing time. On the other hand, the large-scale language model ML, which analyzes the imaged object by referring to the accumulated information 4a, can relatively easily estimate changes in the state of the imaged object from the time-series information of text data indicating the state of the imaged object. Therefore, by performing the process of naturalizing the state of the imaged object as the image-to-text process (image captioning process) as described above, the processing difficulty of the image captioning process is reduced, and the time and resources required for image captioning can be reduced. In this case, using natural language is advantageous because it enables high-speed processing using the LMM in cases such as "detecting changes," "summarizing state changes," and "answering the final (latest) state," and also reduces the amount of input data.
[0038] Here, in the image captioning process by the image text conversion unit 14, the language of the image text data to be generated can be selected as appropriate depending on the destination of the service, etc., but from the standpoint of efficiency, it is conceivable to perform image captioning in English, i.e., to have the image text conversion unit 14 perform natural language conversion in English. By using natural language conversion in English, tokens can be saved compared to, for example, natural language conversion in Japanese, and since there is less variation in expression than in Japanese, searchability using the large-scale language model ML can be improved, thereby improving the accuracy of response information presented to the user.
[0039] In the control device 1, an input / output interface (I / F) 15 is also connected to the bus BS. An input unit 16 consisting of operators and operation devices is connected to the input / output interface 15. For example, the input unit 16 may be various operators and operation devices such as a keyboard, a mouse, keys, a dial, a touch panel, a touch pad, a remote controller, etc. An operation by a user is detected by the input unit 16, and a signal corresponding to the input operation is interpreted by the CPU 11.
[0040] A display unit 17, such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) panel, and an audio output unit 18, such as a speaker, are connected integrally or separately to the input / output interface 15. The display unit 17 is used to display various types of information, and may be, for example, a display device provided in the housing of the computer device, or a separate display device connected to the computer device.
[0041] The display unit 17 displays images for various image processing, moving images to be processed, etc. on the display screen based on instructions from the CPU 11. The display unit 17 also displays various operation menus, icons, messages, etc., i.e., a GUI (Graphical User Interface), based on instructions from the CPU 11.
[0042] The input / output interface 15 may be connected to a storage unit 19 configured with a hard disk drive (HDD) or solid-state memory, or a communication unit 20 configured with a modem or the like.
[0043] The communication unit 20 performs communication processing via a transmission path such as the Internet, and communication with various devices via wired / wireless communication, bus communication, and the like.
[0044] A drive DR is also connected to the input / output interface 15 as required, and a removable recording medium RM such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory is appropriately loaded therein.
[0045] The drive DR can read data files, such as programs used for various processes, from the removable recording medium RM. The read data files are stored in the storage unit 19, and images and sounds contained in the data files are output on the display unit 17 and the audio output unit 18. Furthermore, the computer programs, etc. read from the removable recording medium RM are installed in the storage unit 19 as needed.
[0046] In a computer device having the above-described hardware configuration, for example, software for the processing of this embodiment can be installed via network communication by the communication unit 20 or via a removable recording medium RM. Alternatively, the software may be stored in advance in the ROM 12, the storage unit 19, etc. The CPU 11 performs processing operations based on various programs, thereby executing the information processing and communication processing required by the control device 1.
[0047] The control device 1 is not limited to being configured as a single computer device as shown in Fig. 4, but may be configured as a system of multiple computer devices. The multiple computer devices may be systemized using a LAN (Local Area Network) or the like, or may be located in a remote location using a VPN (Virtual Private Network) using the Internet or the like. The multiple computer devices may include computer devices as a server group (cloud) available through a cloud computing service.
[0048] 5 is a block diagram showing an example of the hardware configuration of the information processing device 2. In the following description, parts that are the same as parts that have already been described will be given the same reference numerals and description thereof will be omitted.
[0049] The difference from the control device 1 shown in Fig. 4 is that an AI processing unit 24 is provided instead of the image-to-text unit 14. The AI processing unit 24 generates response information to input prompts using the large-scale language model ML described above.
[0050] Hereinafter, the CPU included in the information processing device 2 will be referred to as a CPU 21 in order to distinguish it from the CPU 11 included in the control device 1 .
[0051] The information processing device 2 is not limited to being configured as a single computer device as shown in FIG. 5, but may be configured as a system of multiple computers.
[0052] 6 is a block diagram showing an example of the hardware configuration of the database 4 and the user terminal 5. The difference from the control device 1 shown in FIG. 4 is that the image-to-text unit 14 is omitted. Here, to distinguish them from the CPU 11 provided in the control device 1, the CPU provided in the database 4 and the CPU provided in the user terminal 5 will be referred to as CPU 41 and CPU 51, respectively.
[0053] Of the database 4 and the user terminal 5, the database 4 in particular is not limited to being configured by a single computer device as shown in FIG. 6, but may also be configured by a system of multiple computer devices.
[0054] 7 is an explanatory diagram of various functions of the control device 1 according to the embodiment. As shown in the figure, the control device 1 includes the image-to-text unit 14 described above, and the CPU 11 has functions as a storage control unit F11, an input control unit F12, an image analysis unit F13, and an imaging control unit F14.
[0055] The storage control unit F11 controls the storage of the image-to-text data obtained by the image-to-text unit 14 in the database 4. Specifically, the storage control unit F11 causes the communication unit 20 shown in FIG.
[0056] The input control unit F12 controls so that only some frame images of the captured moving image are input to the image-to-text unit 14. By inputting only some frame images to the image-to-text unit 14, the number of images to be processed by the image-to-text unit 14 is reduced, which reduces the processing load on the image-to-text unit 14 and reduces the amount of image-to-text data stored in the database 4.
[0057] Specifically, the input control unit F12 may perform control so that frame images are input to the image-to-text unit at predetermined time intervals, such as every second or every few minutes.
[0058] Alternatively, some of the frame images input to the image-to-text unit 14 may be frame images in which there has been a change in the imaged subject. In this case, the input control unit F12 performs image analysis processing on the captured images to determine whether there has been a change in the imaged subject, and controls so that the frame images in which there has been a change are input to the image-to-text unit 14. In this case, the image analysis processing for determining whether there has been a change in the imaged subject may involve detection of a difference between frames. Alternatively, when determining whether there has been a specific change in the imaged subject, such as a person turning around or raising their hand, the image analysis processing may involve image recognition processing using an AI model.
[0059] By inputting only frame images in which a change has occurred in the imaged subject as described above, it becomes possible to subject only frame images in which a specific change has occurred in the imaged subject to image text conversion processing, such as frame images in which a specific subject such as a person or vehicle has framed in, or frame images in which that person, vehicle, etc. has begun to move. This makes it possible to prevent frame images that are not useful for analysis, such as frame images in which the subject desired to be analyzed does not exist, or frame images in which the subject desired to be analyzed exists but is not moving, from being subjected to image text conversion processing indiscriminately, thereby reducing the processing load on the image text conversion unit 14 and reducing the amount of image text data stored in the database 4, while ensuring the usefulness of the stored information 4a.
[0060] The image analysis unit F13 and the imaging control unit F14 have functions for controlling the imaging device 3 so that imaging settings or image signal processing settings suitable for a specific subject in a captured image are performed. Specifically, the image analysis unit F13 performs processing to detect a specific subject in the captured image as image analysis processing for the captured image input from the imaging device 3. Furthermore, the imaging control unit F14 controls the imaging settings or image signal processing settings in the imaging device 3 so that they are set according to the specific subject detected by the image analysis unit F13.
[0061] The image analysis unit F13 may detect, for example, a moving object as a specific subject. In this case, the image analysis unit F13 performs processing to detect inter-frame differences in captured images as image analysis processing for detecting moving objects. Alternatively, the specific subject may be a person, a vehicle, an animal, or a subject identified based on shape characteristics such as a specific body part of a person or animal, such as a face or hand, or a vehicle license plate. In this case, the image analysis unit F13 may perform processing to detect the specific subject in the captured image using an image recognition AI model trained to enable class identification of the specific subject.
[0062] Furthermore, in the imaging control unit F14, imaging settings according to a specific subject include, for example, a setting to zoom in on the specific subject, a setting to adjust the exposure to the specific subject, etc. Furthermore, image signal processing settings according to a specific subject include, for example, a setting to emphasize the edges of the specific subject, a setting to enhance the noise reduction effect, etc. When a specific subject is detected in a captured image by the image analysis unit F13, the imaging control unit F14 controls the imaging device 3 so that imaging settings and image signal processing settings according to such a specific subject are performed.
[0063] The functions of the image analysis unit F13 and the imaging control unit F14 described above make it possible to control imaging settings and image signal processing settings so as to increase the amount of information about the object to be converted into text, thereby improving the usefulness of the stored information 4a in the database 4.
[0064] Here, the image analysis unit F13 may perform scene recognition processing on the captured image and determine an object corresponding to the recognized scene as a specific subject. In this case, for example, an AI model that performs scene recognition processing is used as the image analysis unit F13. Examples of scene recognition include scenes related to the behavior of a target subject, such as a person or a vehicle. For example, in the case of a person, scenes such as sitting, standing up, raising a hand, and eating can be considered, and in the case of a vehicle, scenes such as a vehicle stopping or departing can be considered. In the process of determining an object corresponding to the recognized scene as a specific subject, for example, in a meal scene, a person's mouth or hands can be considered to be determined as a specific subject, and in a scene where a vehicle is stopped, the license plate portion can be considered to be determined as a specific subject (e.g., for managing vehicles parked in a parking lot). Here, the type of object to be determined as a specific subject depending on the scene may be determined appropriately depending on the purpose of the service, etc.
[0065] In addition, there is a possibility that important (meaningful) objects, important movements of objects, or important actions are captured near a person's hands. For example, if a user later asks, "I seem to have left my smartphone somewhere, please tell me where it is," the system can detect the smartphone from the image near the hand, extract the last scene, and use that as a candidate location to answer the question.
[0066] Previously, humans had to specify rules to determine where important information was in an image, such as "near the hands," "near the mouth," or "car license plate," but in the age of LLM, it is possible to input an image and have the parts containing important information automatically selected (by learning from many examples). This means there is no longer any need to specify specific subjects in advance.
[0067] As described above, by configuring the image analysis unit F13 to determine a specific subject according to the scene recognized by the scene recognition process, it becomes possible to detect notable objects in the scene as specific subjects, and images with increased information content regarding notable subjects according to the scene can be input to the image text conversion unit 14, thereby improving the usefulness of the stored information 4a in the database 4.
[0068] An example of the relationship between the captured image and the converted image data input to the image-to-text unit 14 will now be described with reference to Figures 8 and 9. Figure 8 shows an example of the relationship between the captured image and the converted image data in the case of traffic monitoring, and Figure 9 shows an example of the relationship between the captured image and the converted image data in the case of monitoring the inside of a specific person such as an elderly person.
[0069] 8 and 9, time information is associated with the image-to-text data, and the image-to-text data associated with this time information is stored in the database 4. The time information here may be time information indicating the timing at which the target event for image captioning occurred, and for example, information indicating the capture time of the captured image input as the target image for image captioning may be associated. Alternatively, information indicating the time at which the image captioning process was performed may be associated.
[0070] 4. Information Processing Method as an Embodiment Next, a description will be given of the analysis-side processing using the stored information 4a in the database 4. Fig. 10 is an explanatory diagram of various functions of an information processing device 2 as an embodiment. The information processing device 2 includes an AI processing unit 24 having the large-scale language model ML described above, and the CPU 21 functions as a prompt receiving unit F21 and a presentation processing unit F22.
[0071] The prompt receiving unit F21 receives a prompt for the large-scale language model ML, and the presentation processing unit F22 performs processing to present to the user response information generated by the large-scale language model ML in response to the prompt received by the prompt receiving unit F21. The prompt receiving unit F21 receives a prompt input from the user terminal 5, and the presentation processing unit F22 performs processing to display the response information generated by the large-scale language model ML in response to the prompt based on the stored information 4a on the display unit 17 (see FIG. 6) of the user terminal 5. Note that the response information may be presented, for example, on a web page.
[0072] Examples of prompts and response information will be described with reference to Figures 11 and 12. Figure 11 shows examples of prompts (Figure 11A) and response information (Figure 11B) in the case of traffic monitoring applications, and Figure 12 shows examples of prompts (Figure 12A) and response information (Figure 12B) in the case of in-home monitoring of a specific person such as an elderly person.
[0073] 11 shows an example of response information generated by the large-scale language model ML in response to a prompt input requesting a search for a specific event at a specific location on a specific date and time. In this case, the large-scale language model ML may generate not only text-based response information, but also image-based response information such as the map image shown in the figure. It may also generate sound-based response information and present it to the user (by having the audio output unit 18 present information audibly).
[0074] FIG. 12 shows an example of response information generated by the large-scale language model ML in response to a prompt input to confirm whether the person being monitored has performed a specific action at a specific time.
[0075] Here, because it is a large-scale language model ML, it is possible to respond to prompt inputs that ask for more detailed information. For example, in the example of Fig. 12, it is possible to generate response information to a prompt input such as "Did you drink yesterday?" after the response information is presented.
[0076] In addition to the task of searching for a specific event as described above, various other tasks can be considered as tasks for the large-scale language model ML, such as creating a summary of the "status" over a specific period of time, such as one month.
[0077] The information processing device 2 having the above-described prompt receiving unit F21 and presentation processing unit F22 can realize a system that analyzes an image capture target based on accumulated information 4a that stores text data of captured images, and can present appropriate analysis information to a user in response to various input prompts in natural language. In particular, a highly functional analysis system can be realized that can retrospectively analyze past events from various perspectives specified in the prompts, based on accumulated information 4a about past events.
[0078] 13 is a flowchart showing an example of a processing procedure for realizing the functions of the information processing device 2 according to the embodiment described above. In this example, the processing shown in FIG. 13 is executed by the CPU 21 in the information processing device 2 based on a program stored in, for example, the ROM 12 or the storage unit 19.
[0079] 13, the CPU 21 performs a prompt receiving process in step S101 to receive a prompt input from the user terminal 5.
[0080] In the next step S102, the CPU 21 issues a prompt to the large-scale language model ML to generate response information. In the next step S103, the CPU 21 displays the response information generated by the large-scale language model ML in the process of step S102 on the display unit 17 of the user terminal 5 as a presentation process of the response information.
[0081] After executing the process of step S103, the CPU 21 ends the series of processes shown in FIG.
[0082] In the above example, the AI processing unit 24 having the large-scale language model ML is provided in a device that accepts prompts and processes the presentation of response information. However, it is also possible that the AI processing unit 24 is provided in a device separate from the device that accepts prompts and processes the presentation of response information.
[0083] 5. Modifications Note that the embodiment is not limited to the specific examples described above, and various modified configurations are possible. For example, although the above example illustrates a configuration in which the image-to-text unit 14 and the storage control unit F11 are provided in devices external to the imaging device 3, it is also possible to provide the image-to-text unit 14 and the storage control unit F11 in the imaging device 3.
[0084] 14 is a block diagram showing an example of the schematic configuration of an imaging device 3A having an image-to-text unit 14 and a storage control unit F11. As shown in the figure, the imaging device 3A includes an imaging unit 31, a control unit 32, and a communication unit 33, as well as the image-to-text unit 14.
[0085] The imaging unit 31 is configured as an image sensor such as a CCD (Charge Coupled Device) image sensor or a CMOS (Complementary Metal Oxide Semiconductor) image sensor, and obtains a captured image as a grayscale image (an RGB image in this example).
[0086] The image captured by the imaging unit 31 is input to the image-to-text unit 14 .
[0087] The control unit 32 is configured with, for example, a microcomputer equipped with a CPU, ROM, and RAM, and performs overall control of the imaging device 3A by performing processing in accordance with programs stored in the ROM and programs loaded into the RAM.
[0088] The communication unit 33 performs wired or wireless data communication with an external device. In particular, the communication unit 33 in this example has a network communication function and is capable of performing data communication with the database 4 via the network NT.
[0089] In the imaging device 3A, the function of the storage control unit F11 is realized by software processing of the control unit 32. This makes it possible to realize an imaging device that handles everything from capturing images to controlling the storage of image-text data, and makes it possible to eliminate the need to provide a control device 1 separate from the imaging device.
[0090] Although not shown, it is also possible to consider a configuration in which the input control unit F12, the image analysis unit F13, and the imaging control unit F14 are provided in the imaging device 3A.
[0091] Although not specifically mentioned above, the natural language model MT of the image text conversion unit 14 can also be a model that performs natural language conversion on sensing information other than images, in addition to natural language conversion on images.
[0092] Fig. 15 is a block diagram showing an example of the schematic configuration of a control device 1A having a function of naturalizing information sensed by a microphone 40. Note that Fig. 15 does not show the input control unit F12, image analysis unit F13, and imaging control unit F14 described above. As shown, the control device 1A includes an image-to-text unit 14A instead of the image-to-text unit 14. The image-to-text unit 14A has a natural-language model MT that naturalizes not only images but also information sensed by the microphone 40 (sound information).
[0093] 16 shows an example of image-to-text data generated by the image-to-text unit 14 A. In this way, the image-to-text unit 14 A can obtain image-to-text data that expresses not only visual phenomena of the imaged subject but also auditory phenomena (e.g., the notation of braking sounds in the figure) in natural language.
[0094] In addition to sound sensing information, sensing information other than images can be thought of in a variety of ways, such as sensing information about temperature, humidity, wind speed, air pressure, and acceleration (vibrations during an accident, etc.). By including sensing information other than images as a target for natural language conversion, it becomes possible to accumulate text data of events identified from sensing information other than images in the database 4. This makes it possible to increase the amount of information that can be used for analysis, making it possible to handle more multifaceted analysis.
[0095] Furthermore, in the above description, an RGB image is assumed as the captured image to be input to the image-to-text unit 14, but a configuration in which a captured image other than an RGB image is input to the image-to-text unit 14 is also possible. For example, as a gradation image other than an RGB image, a monochrome image, an infrared image (IR, SWIR, NIR, etc.), a multispectral image (a wavelength analysis image with a bandwidth narrower than at least the wavelength bandwidths of R, G, and B), etc. may be input to the image-to-text unit 14. Furthermore, an example of a gradation image is an image captured by photon counting imaging (PCI) using a single photon avalanche diode (SPAD) element. Alternatively, captured images other than gradation images may include distance images obtained by various distance measurement methods such as a stereo method, a dToF (direct Time of Flight) method, an iToF (indirect ToF) method, an STL (Structured Light) method, a thermal image obtained by a thermal camera, a polarized image generated using a polarization sensor, an event image obtained by an EVS, etc. The captured images to be converted into text may be selected from an appropriate type of image depending on the monitoring application, etc.
[0096] It is also conceivable to adopt a configuration in which a plurality of types of captured images are input to the image-to-text conversion unit 14. In other words, the image-to-text conversion unit 14 is configured to generate image-to-text data using a plurality of types of captured images as input. For example, it is conceivable to adopt a configuration in which, in addition to captured images as gradation images such as RGB images, captured images as distance images, captured images as polarized images, and captured images as thermal images are input to the image-to-text conversion unit 14.
[0097] In this way, by configuring the image-to-text unit 14 to generate image-to-text data using multiple types of captured images as input, it becomes possible for the image-to-text unit 14 to recognize events that would not be recognized if only one type of captured image were input. Therefore, it is possible to increase the amount of information in the image-to-text data, and to improve the usefulness of the stored information 4a.
[0098] Here, it is conceivable that the multiple types of captured images are each obtained by a different imaging device. For example, if an RGB image and a distance image are input to the image text conversion unit 14, the RGB image may be obtained by an imaging device equipped with an RGB image sensor, and the distance image may be obtained by an imaging device equipped with a distance measuring sensor.
[0099] Alternatively, multiple types of captured images may be obtained by a single imaging device equipped with a so-called hybrid sensor. Fig. 17 is a block diagram showing an example of the schematic configuration of an imaging device 3B in which multiple types of captured images are obtained by a hybrid sensor. The imaging unit 31B is configured as a hybrid sensor in which multiple types of pixels are formed to obtain different types of captured images.
[0100] 18 and 19 show configuration examples of a pixel array unit included in the imaging unit 31B as a mixed sensor. Fig. 18 shows a configuration example of an in-plane mixed sensor. Fig. 18A illustrates a pixel array structure in which, in addition to R, G, and B pixels, white pixels (W) and IR pixels are two-dimensionally arranged according to a predetermined rule. Fig. 18B illustrates a pixel array structure in which, in addition to R, G, and B pixels, pixels for ranging (D) are two-dimensionally arranged according to a predetermined rule. Fig. 18C illustrates a pixel array structure in which, in addition to R, G, and B pixels, pixels for event detection (E) are two-dimensionally arranged according to a predetermined rule.
[0101] 19A shows an example of a stacked structure, in which FIG. 19A is an example in which an IR pixel is provided below a pixel array section for obtaining an RGB image, and FIGS. 19B and 19C are examples in which a pixel array section for event detection and a pixel array section for ranging are provided below a pixel array section for obtaining an RGB image, respectively.
[0102] 17, when the imaging unit 31B as a combined sensor includes pixels for distance measurement using the dToF method, iToF method, or STL method, or IR pixels, it is assumed that the subject will be irradiated with light in a wavelength band other than the visible light band. For this reason, in this case, the imaging device 3B is provided with a light emitting unit 35 that emits light to irradiate the subject.
[0103] In the imaging device 3B, multiple types of captured images obtained by the imaging unit 31B, which serves as a mixed sensor, are input to the image-to-text unit 14B. The image-to-text unit 14B has a natural language model MT that has been trained to perform image captioning in response to input of multiple different types of captured images. The image-to-text data obtained by the image-to-text unit 14B is input to the control unit 32 and stored in the database 4 via the communication unit 33 by the function of the storage control unit F11.
[0104] In the above example, an imaging device equipped with an imaging unit 31B as a mixed sensor has an image-to-text unit 14B and an accumulation control unit F11, but a system configuration in which a device separate from the imaging device equipped with the imaging unit 31B has the image-to-text unit 14B and the accumulation control unit F11 can also be adopted.
[0105] To increase the amount of information, for example, for distance images or infrared images, it is possible to adopt a configuration in which images captured when the light emitting unit 35 is turned on and images captured when it is turned off are input to the image-to-text unit 14B. Alternatively, it is possible to adopt a configuration in which a plurality of types of captured images obtained by emitting light at different wavelengths are input to the image-to-text unit 14B.
[0106] Here, the multiple types of captured images may be captured images with different brightness levels, such as short-exposure images and long-exposure images obtained by HDR (High Dynamic Range) imaging. Alternatively, the multiple types of captured images may be captured images with different resolutions, obtained by using an image sensor that supports binning.
[0107] Fig. 20 shows an example of a pixel array structure of an image sensor that supports binning, and Fig. 21 shows an example of a pixel array structure for obtaining a plurality of types of captured images with different exposure times within the same frame period for HDR imaging. In Fig. 20, Fig. 20A shows an example of a 2 × 2 pixel-sharing type pixel array structure, and Fig. 20B shows an example of a 3 × 3 pixel-sharing type pixel array structure.
[0108] Also, in Figure 21, Figure 21A is an example in which pixel units of "RGGB" in a Bayer array are 2 x 2 size pixel units each consisting of a single R pixel, G pixel, G pixel, and B pixel, and pixel units consisting of only short exposure pixels (S) and pixel units consisting of only long exposure pixels (L) are arranged alternately in both the horizontal and vertical directions. Also, Figures 21B and 21C show examples in which short exposure pixels and long exposure pixels are arranged in R pixel groups, G pixel groups, G pixel groups, and B pixel groups within a pixel unit in which the pixel units of "RGGB" in a Bayer array are pixel units of a size of 4x4 or more, each consisting of a plurality of R pixels, G pixels, G pixels, and B pixels. Figure 21B shows an example in which the size of each of the R, G, and B pixel groups within the pixel unit is 2x2 (pixel unit size is 4x4), and Figure 21C shows an example in which the size of each of the R, G, and B pixel groups within the pixel unit is 3x3 (pixel unit size is 9x9).
[0109] In the above example of HDR imaging, pixels with different exposure times are arranged, but it is also possible to obtain multiple types of images with different brightness by arranging pixels with different sizes. Furthermore, while the above example shows two types of brightness, it is also possible to use three or more types of brightness. In this case, three or more types of pixels with different exposure times or sizes are arranged in the pixel array unit.
[0110] Furthermore, as explained above with reference to the image analysis unit F13 and the imaging control unit F14, it has been mentioned that the imaging settings and image signal processing settings of the imaging device 3 are controlled based on the image analysis results of the captured image, but it is also possible to perform control based on the image text data obtained by the image text conversion unit 14 rather than control based on the image analysis results of the captured image.
[0111] 22 is an explanatory diagram of a schematic configuration example of a control device 1B that controls the imaging settings and image signal processing settings of the imaging device 3 based on the image-to-text data. As shown in the figure, the control device 1B includes an image-to-text unit 14 and a storage control unit F11, as well as an imaging control unit F15 and a large-scale language model for control tasks 50.
[0112] The control task large-scale language model 50 is a large-scale language model trained to be able to use image-to-text data as input and infer imaging settings or image signal processing settings for the imaging device 3 according to the content of the image-to-text data. Specifically, the control task large-scale language model 50 estimates a subject (interesting subject) that is useful for analysis from the content of the image-to-text data, and performs processing to estimate imaging settings or image signal processing settings that will enable more detailed information about the useful subject to be obtained. For example, if the content of the image-to-text data indicates the presence of a vehicle that is likely to cause an accident, the control task large-scale language model 50 estimates that the license plate of the vehicle is the interesting subject, and estimates, as suitable imaging settings, imaging settings that zoom in on the license plate and image signal processing settings that perform edge enhancement processing on the license plate area.
[0113] The imaging control unit F15 controls the imaging device 3 so that imaging operations or image signal processing according to imaging settings or image signal processing settings estimated (inferred) by the large-scale language model for control tasks 50 are executed in the imaging device 3. For example, in the above example, if the large-scale language model for control tasks 50 outputs, as inference results, imaging settings for zooming in on the useful subject and image signal processing settings for performing edge enhancement processing on the useful subject portion, the imaging control unit F15 controls the imaging device 3 so that the zoom-in operation and edge enhancement processing are performed.
[0114] The function of the imaging control unit F15 may be realized by software processing of the CPU 11 included in the control device 1B.
[0115] By configuring the system to have the imaging control unit F15 as described above, in a system that analyzes an imaged subject based on accumulated information 4a of image text data, it is possible to automate camera operations to obtain accumulated information 4a appropriate for analysis, for example by optimizing the angle of view of the imaging device 3 or by performing image signal processing suitable for analyzing the target subject.
[0116] While the above example illustrates a configuration in which a device having an imaging control unit F15 includes the large-scale language model for control tasks 50, it is also possible for the large-scale language model for control tasks 50 to be included in a device separate from the device having the imaging control unit F15. Furthermore, while the above example illustrates an example in which the imaging control unit F15 is provided in a device having the image-to-text unit 14 and the storage control unit F11, it is also possible for the imaging control unit F15 to be provided in a cloud server such as the information processing device 2. Furthermore, in the case of an imaging device having an image-to-text unit 14, such as the imaging device 3A illustrated in FIG. 14, it is also possible for the imaging control unit F15 to be provided in the imaging device. In this case, the large-scale language model for control tasks 50 can be provided in the imaging device, or it can also be provided in a device external to the imaging device.
[0117] Although the image-to-text data generated by the image-to-text conversion unit 14 is stored in the database 4, it is also possible to perform processing to suppress information duplication in the stored information 4a. For example, when multiple imaging devices 3 are used as illustrated in FIG. 7 , it is possible to adopt a configuration in which the image-to-text conversion unit 14 performs image captioning for each imaging device 3. In such a case, if the fields of view of the imaging devices 3 overlap, information duplication occurs in the image-to-text data. For example, it is possible to perform processing to eliminate such information duplication. Such information duplication suppression processing may involve, for example, determining whether or not there is information duplication in the image-to-text data for each imaging device 3 as a preprocessing step before storage in the database 4. If there is information duplication, only one image-to-text data is stored as is, and text corresponding to the duplicated information is deleted from the other image-to-text data before storing it in the database 4. Alternatively, it is also possible to perform such information duplication suppression processing after storage in the database 4. The former suppression process before storage may be performed in a device equipped with the image-to-text conversion unit 14, while the latter suppression process after storage may be performed in a computer device serving as the database 4.
[0118] In the above, we have mentioned that the “state” of the captured image is converted into text by image caption processing and the large-scale language model ML estimates (analyzes) the “changes” in the state. However, in this case, the information processing device 2 may obtain frame-by-frame caption information from the database 4, convert it into natural language expressing the “changes” in the state of the captured image, and store the converted information in the database 4. In this case, two processing timings are possible. The first is to analyze the “changes” in the scenes and store them in the database 4 when a user later inputs a request from the user terminal 5, such as “Please summarize the scenes from the past month.” The second is to analyze the “changes” in the scenes, for example, at regular intervals, without relying on a prompt input from the user, and store the results in the database 4. With the latter second method, it is also possible to perform the analysis in response to instructions other than a human prompt input, or to perform the analysis triggered by the results of image captioning. The analysis of the “changes” and their storage in the database 4 may also be performed based on a user program. Specifically, in this case, the user's program will be a program that realizes the process of combining the "accumulated data" and the "prompt" and sending them to the large-scale language model ML (information processing device 2), and storing the results in the database 4.
[0119] Furthermore, when the method of storing scene "changes" in advance as described above is employed, since the "changes" are verbalized for each imaging device 3, it is possible to quickly respond to a user's inquiry by picking up information from any of the imaging devices 3. For example, in an indoor monitoring system using multiple imaging devices 3, if a user makes a request such as "Please summarize and explain in 100 characters or less any unusual events that have occurred in this room in the past month," or "Please observe your mother who lives at your parents' house, and if there have been any changes in her behavior in the past month, please summarize them in about one minute and report them by voice," it is possible to quickly obtain information on "changes" related to images from the imaging devices 3 installed in "this room" or in "mother's" room, thereby enabling a prompt response to the request.
[0120] 6. Summary of the Embodiments As described above, the control device (1, 1A, 1B, imaging device 3A, 3B) according to the embodiment includes an image text conversion unit (14, 14A, 14B) that generates image text data that converts the content of visual information from the captured image into natural language by naturalizing the content of the visual information from the captured image using a natural language conversion model that converts the content of visual information from an image into natural language, and a storage control unit (F11) that controls the storage of the image text data obtained by the image text conversion unit in a database. According to the above configuration, the information stored in the database is converted into text. Therefore, privacy protection can be achieved while reducing the storage capacity of the database and the amount of communication data between the database.
[0121] Furthermore, the control device according to the embodiment includes an input control unit (F12) that controls so that only some frame images of the captured moving images are input to the image-to-text unit, thereby reducing the number of images to be processed by the image-to-text unit, thereby reducing the processing load on the image-to-text unit and reducing the amount of image-to-text data stored in the database.
[0122] Furthermore, in the control device according to the embodiment, the input control unit controls the input of frame images to the image-to-text unit at regular time intervals, thereby reducing the processing load on the image-to-text unit and the amount of data stored in the database, while allowing image-to-text data to be stored in the database in a time-lapse manner at predetermined time intervals, such as every few seconds or minutes.
[0123] Furthermore, in the control device according to the embodiment, the input control unit performs image analysis processing on the captured image to determine whether or not there has been a change in the captured subject, and controls frame images in which a change has occurred to be input to the image-to-text conversion unit. This makes it possible to select only frame images in which a change in the captured subject has occurred, such as frame images in which a specific subject, such as a person or vehicle, has entered the frame or frame images in which that person or vehicle has begun to move, as the subject of the image-to-text conversion process. This prevents frame images that are not useful for analysis, such as frame images in which the subject desired for analysis does not exist or in which the subject desired for analysis exists but is stationary, from being subjected to image-to-text conversion processing indiscriminately. This reduces the processing load on the image-to-text conversion unit and the amount of image-to-text data stored in the database, while ensuring the usefulness of the information stored in the database.
[0124] Furthermore, in the control device according to the embodiment, the image-to-text conversion unit converts the state of the imaging target into natural language. For the image-to-text conversion unit, converting changes in the state of the imaging target from images into natural language is highly difficult and requires resources and processing time. On the other hand, a large-scale language model makes it relatively easy to estimate changes in the state of the imaging target from time-series information in text data indicating the state of the imaging target. For this reason, as described above, the image-to-text conversion process converts the state of the imaging target into natural language. This reduces the processing difficulty of the image-to-text conversion process and reduces the time and resources required for image-to-text conversion.
[0125] Furthermore, in the control device according to the embodiment, the image text conversion unit converts the image into natural language in English, which reduces the number of tokens required compared to converting the image into natural language in Japanese, and also reduces the variation in expression compared to Japanese, which improves the ease of searching using a large-scale language model and improves the accuracy of the response information presented to the user.
[0126] Furthermore, the control device according to the embodiment includes an image analysis unit (F13) that performs image analysis processing on the captured image to detect a specific subject in the captured image, and an imaging control unit (F14) that controls the imaging settings or image signal processing settings of the imaging device that obtains the captured image so that they correspond to the specific subject detected by the image analysis unit. Examples of imaging settings that correspond to the specific subject include zooming in on the specific subject and adjusting the exposure to the specific subject. Examples of image signal processing settings that correspond to the specific subject include emphasizing the edges of the specific subject and enhancing the noise reduction effect. This configuration makes it possible to control the imaging settings and image signal processing settings so as to increase the amount of information about the object to be converted into text, thereby improving the usefulness of the information stored in the database.
[0127] In addition, in the control device according to the embodiment, the image analysis unit detects a moving object as a specific subject based on the inter-frame difference of the captured image. This allows images of the moving object with increased information content, such as zoomed-in images or edge-enhanced images of the moving object, to be input to the image-to-text unit in response to analysis of the moving object. This improves the usefulness of the information stored in the database.
[0128] Furthermore, in the control device according to the embodiment, the image analysis unit performs scene recognition processing on the captured image and determines an object corresponding to the recognized scene as the specific subject. This makes it possible, for example, to determine the object in the person's hand or the object being held in their mouth as the specific subject in a scene where a person is eating something, thereby detecting a noteworthy object in the scene as the specific subject. Therefore, it is possible to input an image with an increased amount of information about a noteworthy subject corresponding to the scene to the image-to-text unit, thereby improving the usefulness of the information stored in the database.
[0129] Furthermore, in the control device according to the embodiment, the image-to-text unit generates image-to-text data using images captured by multiple imaging devices as input. This allows the image-to-text unit to recognize events that would not be recognized if only images captured by a single imaging device were input, such as recognizing that a subject captured by one imaging device is approaching another subject captured by another imaging device. This increases the amount of information in the image-to-text data, thereby improving the usefulness of the information stored in the database.
[0130] In addition, in the control device (image capture device 3B) of the embodiment, the image text conversion unit (image capture device 14B) generates image text data using multiple types of captured images as input. For example, different types of captured images, such as grayscale images such as RGB images, distance images, polarized images, and short-exposure and long-exposure images obtained by HDR imaging, are input to generate image text data. This enables the image text conversion unit to recognize events that would not be recognized if only one type of captured image were input. This increases the amount of information in the image text data, improving the usefulness of the information stored in the database.
[0131] Furthermore, in the control device (1A) according to the embodiment, the image-to-text unit (14A) performs natural language processing on the image as well as on sensing information other than the image. This allows for textual data of events identified from sensing information other than the image, such as sound, temperature, and acceleration, to be stored in a database. This increases the amount of information available for analysis, enabling more multifaceted analysis.
[0132] In a control method according to an embodiment, a computer device uses a natural language model that converts the content of visual information from images into natural language to generate image-text data that converts the content of visual information from the captured images into natural language, and controls the storage of the image-text data in a database. This control method can also achieve the same effects and advantages as the control device according to the embodiment described above.
[0133] Furthermore, an information processing device (same as the second embodiment) includes a prompt receiving unit (same as the first embodiment) that receives prompts from a large-scale language model that generates response information to prompts based on information stored in a database that stores image text data in which the content of visual information from captured images is converted into natural language, and a presentation processing unit (same as the first embodiment) that performs processing to present to a user the response information generated by the large-scale language model in response to the prompt received by the prompt receiving unit. This makes it possible to realize a system that analyzes an image subject based on stored information in which text data from captured images is stored, and that can present to a user appropriate analysis information in response to various input prompts in natural language.
[0134] An information processing method as an embodiment is an information processing method in which a computer device receives a prompt from a large-scale language model that generates response information to the prompt based on information stored in a database that stores image-text data in which the content of visual information from captured images is converted into natural language, and presents the response information generated by the large-scale language model in response to the received prompt to a user. This information processing method can also achieve the same functions and effects as the information processing device as the above-mentioned embodiment.
[0135] An imaging control device (control device 1B) according to an embodiment includes an imaging control unit (F15) that inputs image text data, which is a natural language version of the content of visual information from a captured image, into a large-scale language model, causes the large-scale language model to infer imaging settings or image signal processing settings according to the content of the image text data, and controls the imaging device so that imaging operations or image signal processing based on the inferred imaging settings or image signal processing settings are performed in the imaging device. This makes it possible to automate camera operations in a system that analyzes an imaging target based on accumulated information that has accumulated text data of captured images, by, for example, optimizing the angle of view of the imaging device or performing image signal processing suitable for analyzing the target subject, thereby obtaining accumulated information suitable for analysis.
[0136] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0137] <7. The Present Technology> The present technology may also have the following configurations. (1) A control device including: an image text generation unit that generates image text data that naturalizes the content of visual information from a captured image by performing natural language generation on a captured image using a natural language generation model that naturalizes the content of visual information from the image; and an accumulation control unit that controls the accumulation of the image text data obtained by the image text generation unit in a database. (2) The control device according to (1), including an input control unit that controls so that only some frame images of a captured moving image are input to the image text generation unit. (3) The control device according to (2), in which the input control unit controls so that frame images are input at regular intervals to the image text generation unit. (4) The control device according to (2), in which the input control unit performs image analysis processing on the captured image to determine whether or not there is a change in the captured object, and controls so that frame images that have undergone the change are input to the image text generation unit. (5) The control device according to any of (1) to (4), in which the image text generation unit naturalizes the state of the captured object. (6) The control device according to any of (1) to (5), wherein the image-to-text unit performs natural language conversion in English. (7) The control device according to any of (1) to (6), comprising: an image analysis unit that performs image analysis processing on the captured image to detect a specific subject in the captured image; and an imaging control unit that controls imaging settings or image signal processing settings of an imaging device that obtains the captured image to be settings corresponding to the specific subject detected by the image analysis unit. (8) The control device according to (7), wherein the image analysis unit performs processing to detect a moving object as the specific subject based on an inter-frame difference of the captured image. (9) The control device according to (7), wherein the image analysis unit performs scene recognition processing on the captured image and determines an object corresponding to the recognized scene as the specific subject. (10) The control device according to any of (1) to (9), wherein the image-to-text unit generates the image-to-text data using images captured by a plurality of imaging devices as input.(11) The control device according to any one of (1) to (9), wherein the image-to-text unit generates the image-to-text data using a plurality of types of captured images as input. (12) The control device according to any one of (1) to (11), wherein the image-to-text unit performs natural language translation on images as well as natural language translation on sensing information other than images. (13) A control method in which a computer device performs the natural language translation on captured images using a natural language translation model that naturalizes content of visual information from images, thereby generating image-to-text data that naturalizes content of visual information from the captured images, and controls the accumulation of the image-to-text data in a database. (14) An information processing device comprising: a prompt receiving unit that receives a prompt from a large-scale language model that generates response information to a prompt based on information stored in a database in which image-to-text data that naturalizes content of visual information from captured images is stored; and a presentation processing unit that performs processing to present to a user the response information generated by the large-scale language model in response to the prompt received by the prompt receiving unit. (15) An information processing method in which a computer device receives a prompt from a large-scale language model that generates response information to the prompt based on information stored in a database in which image text data in which content of visual information from a captured image has been converted into natural language, and presents the response information generated by the large-scale language model in response to the received prompt to a user. (16) An imaging control device comprising: an imaging control unit that inputs image text data in which content of visual information from a captured image has been converted into natural language into a large-scale language model, causes the large-scale language model to infer imaging settings or image signal processing settings according to the content of the image text data, and controls the imaging device so that an imaging operation or image signal processing according to the inferred imaging settings or image signal processing settings is performed in the imaging device.
[0138] REFERENCE SIGNS LIST 1, 1A, 1B Control device 2 Information processing device 3, 3A, 3B Imaging device 4 Database 4a Stored information 5 User terminal NT Network ML Large-scale language model 11, 21, 41, 51 CPU 14, 14A, 14B Image-to-text unit MT Natural language model 17 Display unit BS Bus DR Drive RM Removable recording medium 24 AI processing unit F11 Storage control unit F12 Input control unit F13 Image analysis unit F14 Imaging control unit F21 Prompt reception unit F22 Presentation processing unit 31, 31B Imaging unit 32 Control unit 33 Communication unit 35 Light-emitting unit 50 Large-scale language model for control task F15 Imaging control unit
Claims
1. A control device comprising: an image text conversion unit that generates image text data that converts the content of visual information from a captured image into natural language by using a natural language conversion model that converts the content of visual information from an image into natural language, and a storage control unit that controls the storage of the image text data obtained by the image text conversion unit in a database.
2. The control device according to claim 1, further comprising an input control section that controls so that only some frame images of a moving image captured are input to the image text conversion section.
3. The control device according to claim 2, wherein the input control unit controls so that frame images are input to the image-to-text conversion unit at regular intervals.
4. The control device according to claim 2, wherein the input control unit performs image analysis processing on the captured image to determine whether or not there has been a change in the subject being imaged, and controls the frame image in which the change has occurred to be input to the image text conversion unit.
5. The control device according to claim 1, wherein the image text conversion unit converts the state of the imaged object into natural language.
6. The control device according to claim 1, wherein the image-to-text conversion unit converts the image into natural language in English.
7. The control device according to claim 1, comprising: an image analysis unit that performs image analysis processing on the captured image to detect a specific subject within the captured image; and an imaging control unit that controls the imaging settings or image signal processing settings of the imaging device that obtains the captured image to be settings that correspond to the specific subject detected by the image analysis unit.
8. The control device according to claim 7, wherein the image analysis unit performs processing to detect a moving object as the specific subject based on the inter-frame difference of the captured image.
9. The control device according to claim 7, wherein the image analysis unit performs scene recognition processing on the captured image and determines an object according to the recognized scene as the specific subject.
10. The control device according to claim 1, wherein the image-to-text conversion unit generates the image-to-text data using images captured by a plurality of imaging devices as input.
11. The control device according to claim 1, wherein the image-to-text unit receives a plurality of types of captured images as input and generates the image-to-text data.
12. The control device according to claim 1, wherein the image text conversion unit converts images into natural language as well as sensing information other than images into natural language.
13. A control method in which a computer device uses a natural language model that converts the content of visual information from images into natural language, performs the natural language conversion on a captured image, generates image text data in which the content of the visual information from the captured image is converted into natural language, and controls the storage of the image text data in a database.
14. An information processing device comprising: a prompt receiving unit that receives prompts from a large-scale language model that generates response information to prompts based on information stored in a database that stores image text data in which the content of visual information from captured images has been converted into natural language; and a presentation processing unit that performs processing to present to a user the response information generated by the large-scale language model in response to the prompt received by the prompt receiving unit.
15. An information processing method in which a computer device receives a prompt from a large-scale language model that generates response information to the prompt based on information stored in a database that stores image text data in which the content of visual information from captured images has been converted into natural language, and then presents the response information generated by the large-scale language model in response to the received prompt to a user.
16. An imaging control device comprising an imaging control unit that inputs image text data, which is the content of visual information from a captured image converted into natural language, into a large-scale language model, causes the large-scale language model to infer imaging settings or image signal processing settings according to the content of the image text data, and controls the imaging device so that the imaging operation or image signal processing according to the inferred imaging settings or image signal processing settings is performed in the imaging device.
Citation Information
Patent Citations
Story video production method and story video production system
JP2020512759A
Information generation method, device, computer device, storage medium, and computer program
JP2023545543A
Low Latency Captioning System
JP2024521232A