Information processing device and information processing system

The information processing system addresses the challenge of integrating data from different sensor modalities by using an arithmetic unit to optimize converters for aligning intermediate representations, allowing a single model to be applied across sensors without extensive retraining, thereby enhancing processing efficiency and robustness.

WO2025127023A1PCT designated stage expired Publication Date: 2025-06-19SONY SEMICON SOLUTIONS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/043569
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-12-10
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing information processing systems face challenges in seamlessly integrating and processing data from different sensor modalities without the need for extensive manual engineering and retraining of neural network models.

Method used

The system employs an arithmetic unit that converts data from different sensors into intermediate representations, allowing a model learned from one sensor to be applied to another, without requiring significant retraining, by optimizing converters to align intermediate representations.

Benefits of technology

This approach enables robust processing across different sensor modalities by allowing the same model to be used for different sensors, reducing the need for extensive retraining and manual engineering, and improving the efficiency of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024043569_19062025_PF_FP_ABST
    Figure JP2024043569_19062025_PF_FP_ABST
Patent Text Reader

Abstract

[Problem] To perform training to facilitate switching of a sensor modality, or, to execute information processing using training results. [Solution] This information processing device comprises a calculation unit. The calculation unit: acquires first data output by a first sensor and second data output by a second sensor that acquires information of a type different from that of the first sensor; inputs the first data to a first converter to acquire a first intermediate expression; inputs the second data to a second converter to acquire a second intermediate expression; and trains the second converter on the basis of the first intermediate expression and the second intermediate expression.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device and information processing system

[0001] The present disclosure relates to an information processing device and an information processing system.

[0002] Neural network models trained by various machine learning techniques are used in various fields. The trained models may be used to process and analyze sensor data acquired from various sources, e.g., various sensors. This analysis typically involves a specific processing algorithm tailored to each sensor modality, and the trained models are trained to infer the results of this specific processing algorithm.

[0003] Techniques for integrating sensor data from different modalities require the development of both custom algorithms to align and combine the data, as well as preprocessing techniques. This development often requires domain expertise and extensive manual engineering. Furthermore, because the intermediate data representations differ, it is virtually impossible to change the sensor and use the same pre-trained model on the same processor. Therefore, incorporating a new sensor or changing the sensor modality likely requires extensive retraining of the neural network model using a combined (labeled) dataset to adapt to the new input data, making robust processing difficult.

[0004] Japanese Patent Publication No. 2023-062217

[0005] Therefore, one non-limiting problem that the embodiments of the present disclosure aim to solve is information processing that performs learning to facilitate switching between sensor modalities or utilizes the learning results. The problem that the embodiments of the present disclosure aim to solve can also be, as some further non-limiting examples, a problem corresponding to the effects described in the embodiments. In other words, a problem that corresponds to at least one of the effects described in the description of the embodiments of the present disclosure can be the problem that the present disclosure aims to solve.

[0006] According to one embodiment, an information processing device includes a calculation unit that acquires first data output by a first sensor and second data output by a second sensor that acquires a different type of information from the first sensor, inputs the first data to a first converter to acquire a first intermediate representation, inputs the second data to a second converter to acquire a second intermediate representation, and performs training of the second converter based on the first intermediate representation and the second intermediate representation. This embodiment enables the information processing device to use a model trained using a data set of the first sensor for data acquired from the second sensor, thereby enabling processing even when only the sensor is replaced without changing the configuration of the information processing device.

[0007] The first intermediate representation may be an intermediate representation that can execute a first process when input to a first model. The intermediate representation can also be referred to as a latent representation, for example. The above embodiment can be implemented by training a converter that results in the same intermediate representation for outputs from a first sensor and a second sensor having the same characteristics.

[0008] The first model may be a model trained to perform the first processing upon input of the first intermediate representation, which may be, for example, image processing or signal processing such as object detection, classification, or segmentation.

[0009] The first intermediate representation may be an intermediate representation that can execute a second process when input to a second model. The first intermediate representation can also be used as an input to a second model that executes a second process different from the first process. According to the above embodiment, the second intermediate representation can be input to the second model to obtain a result of executing the second process. In other words, it is possible to replace the sensor or the model that performs the process.

[0010] The second model may be a model trained to execute the second process when the first intermediate representation is input. Like the first model, the second model may also be a model trained using a dataset of outputs from a first sensor.

[0011] The calculation unit may acquire information of the same object in the same environment using the first sensor and the second sensor, respectively, to acquire the first data and the second data, input the first data to the first converter to acquire the first intermediate representation, input the second data to the second converter to acquire the second intermediate representation, calculate a loss between the first intermediate representation and the second intermediate representation, update parameters of the second converter, and optimize the second converter. By acquiring information of the same object in the same environment using sensors that acquire different types of information, multiple types of sensory information that may have the same characteristics can be acquired. By optimizing the converter so that the differences between the intermediate representations of the acquired information are reduced, it is possible to refine the converter so that sensor outputs with the same characteristics are converted into the same intermediate representation.

[0012] The calculation unit may input the first data to a third converter to obtain a third intermediate representation, may input the second data to a fourth converter to obtain a fourth intermediate representation, and may perform learning of the fourth converter based on the third intermediate representation and the fourth intermediate representation. In this way, the converter is not limited to being of one type.

[0013] The third intermediate representation may be an intermediate representation that can be input to a third model to perform a third process. For example, a converter different from the converter that obtains data to be input to the first model can be used for a third model that performs a third process different from the first process.

[0014] The third model may be a model trained to execute the third process when the third intermediate representation is input. Like the first model, the third model may also be a model trained using a dataset of outputs from a first sensor.

[0015] The calculation unit may acquire information of the same object in the same environment using the first sensor and the second sensor, respectively, to acquire the first data and the second data, may input the first data to the third converter to acquire the third intermediate representation, may input the second data to the fourth converter to acquire the fourth intermediate representation, may calculate a loss between the third intermediate representation and the fourth intermediate representation, and may update parameters of the fourth converter to optimize the fourth converter. Multiple types of sensory information that are likely to have the same characteristics can also be acquired for the third model. By optimizing the converter so that the differences between the intermediate representations of the acquired information are reduced, it is possible to refine converters that convert sensor outputs with the same characteristics into the same intermediate representation.

[0016] The first converter and the second converter may be domain adaptation models, and the domain adaptation models can be used to align data in different domains, such as sensory information, into an equivalent intermediate representation.

[0017] The first converter and the second converter may be the same model. Converters that convert data acquired from different types of sensors into intermediate representations may use the same model. In this case, the model associated with the converter may be optimized so that only the second intermediate representation changes, without changing the first intermediate representation. Furthermore, if the first intermediate representation changes, subsequent re-learning of the first model and / or the second model may be performed.

[0018] According to one embodiment, an information processing device includes a calculation unit that acquires second data output by a second sensor that acquires a type of information different from that of a first sensor, converts the second data to acquire a second intermediate representation, inputs the second intermediate representation to a first model that executes a first process when it receives the first intermediate representation obtained by converting the first data output by the first sensor, and acquires a result of executing the first process on the second data. This information processing device can perform processing (including inference) using a converter acquired by any of the methods described above.

[0019] The calculation unit may input the second intermediate representation to a second model that receives the first intermediate representation and performs a second process, and obtain a result of performing the second process on the second data. As described above, a converter can convert the second intermediate representation into a first intermediate representation having the same processing characteristics.

[0020] The calculation unit may convert the second data to obtain a fourth intermediate representation, input the fourth intermediate representation to a third model that performs a third process upon receiving a third intermediate representation obtained by converting the first data, and obtain a result of performing the third process on the second data. As above, different converters may be used for models that perform different processes.

[0021] According to one embodiment, an information processing system includes a first sensor, a second sensor that acquires a different type of information from the first sensor, and an information processing device that optimizes any of the converters described above. The information processing device optimizes a converter that converts second data acquired from the second sensor into a first intermediate representation based on a first intermediate representation obtained by converting first data acquired from the first sensor and a second intermediate representation obtained by converting second data acquired from the second sensor. The information processing system includes multiple sensors that acquire different types of information and an information processing device that optimizes the converters, and the converters can be trained by the information processing device.

[0022] The transformation to obtain the first intermediate representation and the transformation to obtain the second intermediate representation may be models related to domain application.

[0023] The transformation to obtain the first intermediate representation and the transformation to obtain the second intermediate representation may be performed using the same model.

[0024] According to one embodiment, an information processing device includes a first sensor, a second sensor that acquires a different type of information from the first sensor, and any of the above-described information processing devices capable of using a model for the different sensors. The information processing device inputs a second intermediate representation obtained by converting second data acquired from the second sensor into a first model that executes a first process when a first intermediate representation obtained by converting first data acquired from the first sensor is input, and acquires a result of executing the first process on the second data. An information processing system can include multiple sensors that acquire different types of information and an information processing device having a converter trained by the above-described information processing system or information processing device. This information processing system can realize processing (including inference) using information acquired from the different sensors.

[0025] 1 is a schematic block diagram showing an example of the overall configuration of an information processing system according to an embodiment. FIG. 2 is a block diagram showing an example of the configuration of each system in an information processing system according to an embodiment, which registers or downloads an artificial intelligence (AI) model or an AI application via the operation of an information processing device on the cloud side. FIG. 3 is a flowchart showing an example of at least a portion of the processing of an information processing system according to an embodiment. FIG. 4 is a block diagram showing an example of the internal configuration of a camera as an imaging device according to an embodiment. FIG. 5 is a schematic diagram showing an example of the configuration of an image sensor as an imaging device according to an embodiment. FIG. 6 is a diagram showing an outline of an information processing system according to an embodiment. FIG. 7 is a flowchart showing the processing of an information processing device according to an embodiment. FIG. 8 is a diagram showing an outline of an information processing system according to an embodiment. FIG. 9 is a diagram showing an outline of an information processing system according to an embodiment.

[0026] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. The drawings are used for explanation purposes, and the shape, size, and size ratio of each component in an actual device do not necessarily have to be the same as those shown in the drawings. Furthermore, since the drawings are simplified, components necessary for implementation other than those shown in the drawings are also assumed to be appropriately provided.

[0027] (First embodiment)

[0028] [Configuration of information processing system]

[0029] (1) Overall configuration of the information processing system

[0030] 1 shows an example of a schematic configuration of an information processing system 100 that constitutes an imaging device system according to the first embodiment. That is, the information processing system 100 can be an example of a system to which the present invention is applied.

[0031] 1, an information processing system 100 according to the first embodiment includes at least a cloud server 1, a user terminal 2, cameras 3 as a plurality of imaging devices, a fog server 4, and a management server 5. Here, at least the cloud server 1, the user terminal 2, the fog server 4, and the management server 5 are configured to be able to communicate with each other via a network 6, such as the Internet.

[0032] The cloud server 1, the user terminal 2, the fog server 4, and the management server 5 are all configured as information processing devices equipped with a microcomputer having a CPU, a ROM (Read Only Memory), and a RAM (Random Access Memory).

[0033] The camera 3 as an imaging device includes an image sensor, such as a CCD (Charge Coupled Device) image sensor or a CMOS (Complementary Metal Oxide Semiconductor) image sensor. These image sensors constitute an imaging unit (see reference numeral 41 in FIG. 4). The camera 3 captures an image of a subject and obtains image information (captured image information) as digital data. The camera 3 also has a function of performing AI-based processing on the captured image. Examples of this processing include image recognition processing and image detection processing.

[0034] In the following description, various types of processing on images, such as image recognition processing and image detection processing, will be referred to simply as "image processing." For example, various types of processing on images using AI or an AI model will be referred to as "AI image processing."

[0035] The multiple cameras 3 are configured to be able to communicate data with the fog server 4. For example, various data such as processing result information showing the results of image processing using AI is transmitted from the cameras 3 to the fog server 4. In addition, the cameras 3 receive various data from the fog server 4.

[0036] Here, the information processing system 100 is assumed to be used in the following manner, for example. First, the fog server 4 or the cloud server 1 generates analytical information about the subject based on the processing result information obtained by image processing of the multiple cameras 3. This generated analytical information can be viewed by the user via the user terminal 2.

[0037] In this case, the multiple cameras 3 are used as surveillance cameras. For example, they can be used as surveillance cameras for monitoring the interior of stores, offices, homes, etc., or as surveillance cameras for monitoring the exterior of parking lots, city streets, etc. Examples of surveillance cameras for monitoring the exterior include traffic surveillance cameras for monitoring traffic conditions. They can also be used as surveillance cameras for monitoring production lines for FA (Factory Automation), IA (Industrial Automation), etc. Furthermore, they can also be used as surveillance cameras for monitoring the interior or exterior of automobiles, trains, etc.

[0038] Furthermore, when used as a surveillance camera in a store, multiple cameras 3 can be placed at predetermined locations within the store. Using multiple cameras 3 allows the user to check the demographics of customers (gender, age, etc.) and their behavior (traffic flow) within the store. In this case, information on the demographics of customers, information on their traffic flow within the store, and information on the congestion status at the cash registers (for example, waiting times at the cash registers) can be generated as analysis information.

[0039] In addition, for traffic monitoring camera applications, multiple cameras 3 can be placed at various locations near the road. Using multiple cameras 3 allows the user to recognize information such as the license plate number (vehicle number), color, and model of the vehicle of passing vehicles.

[0040] In this case, information such as the license plate number, vehicle color, and model can be generated.

[0041] Furthermore, if the camera is intended for use as a surveillance camera in a parking lot, the camera 3 can be placed in a position where it can monitor parked vehicles. For example, the camera 3 can be used to monitor whether there is a suspicious person behaving suspiciously around the vehicle. Furthermore, if a suspicious person is found, a notification device can be provided that notifies the presence of the suspicious person and their attributes (gender or age group), etc.

[0042] Furthermore, if the camera is used as a surveillance camera to monitor available spaces in towns or parking lots, it can notify users of the locations of spaces where they can park their cars.

[0043] For example, in the above-mentioned store monitoring application, the fog server 4 is arranged in the store to be monitored together with the multiple cameras 3. In other words, the fog server 4 is arranged for each monitored store.

[0044] In this way, if a fog server 4 is provided for each monitored object such as a store, the cloud server 1 does not need to directly receive data transmitted from multiple cameras 3 at the monitored object, thereby reducing the processing load on the cloud server 1.

[0045] In addition, when there are multiple stores to be monitored and all of the stores belong to the same chain, it is preferable to place a fog server 4 for each of the stores, rather than for each individual store. In other words, it is not limited to placing one fog server 4 for each monitored object, but it is possible to place one fog server 4 for each of multiple monitored objects.

[0046] Furthermore, if the cloud server 1 or the multiple cameras 3 have processing capabilities, the cloud server 1 or the multiple cameras 3 can have the functions of the fog server 4. As a result, in the information processing system 100, the fog server 4 can be omitted, and the multiple cameras 3 can be directly connected to the network 6, allowing the cloud server 1 to directly receive data transmitted from the multiple cameras 3.

[0047] The above various devices are roughly classified into cloud-side information processing devices and edge-side information processing devices.

[0048] The cloud-side information processing devices include the cloud server 1 and the management server 5. The cloud-side information processing devices are a group of devices that provide services that are expected to be used by multiple users.

[0049] The edge-side information processing devices include the camera 3 and the fog server 4. The edge-side information processing devices are a group of devices prepared by users who use cloud services and placed in the environment.

[0050] However, both the cloud-side information processing device and the edge-side information processing device may be placed in an environment prepared by the same user.

[0051] The fog server 4 may be an on-premise server.

[0052] (2) Registration of AI models and AI applications

[0053] As described above, in the information processing system 100, AI image processing is performed in the camera 3, which is an edge-side information processing device. Then, in the cloud server 1, which is a cloud-side information processing device, advanced application functions are realized using the result information of the AI ​​image processing on the edge side. The result information of the AI ​​image processing is, for example, result information of image recognition processing using AI.

[0054] Here, various methods for registering application functions in the cloud server 1 which is an information processing device on the cloud side or the cloud server 1 including the fog server 4 are as follows.

[0055] FIG. 2 shows an example of the configuration of each device in the information processing system 100 that registers or downloads an AI model or AI application via a marketplace function provided in an information processing device on the cloud side.

[0056] 2, the fog server 4 is not shown, but the fog server 4 may be provided. In this case, the fog server 4 may take on part of the functions of the edge side.

[0057] The above-mentioned cloud server 1 and management server 5 are information processing devices that constitute the cloud environment.

[0058] Furthermore, camera 3 is an information processing device that constitutes the edge environment.

[0059] The camera 3 can be constructed as a device equipped with a control unit that performs overall control of the camera 3. The camera 3 can also be constructed as a device equipped with an image sensor IS that has an arithmetic processing unit that performs various processes including AI image processing on captured images. In other words, the camera 3, which is an edge-side information processing device, may be equipped with an image sensor IS that is another edge-side information processing device inside.

[0060] In addition, the user terminals 2 used by users who use various services provided by the cloud-side information processing device include an application developer terminal 2A, an application user terminal 2B, an AI model developer terminal 2C, etc.

[0061] The application developer terminal 2A is used by users who develop applications used in AI image processing. The application user terminal 2B is used by users who use applications. The AI ​​model developer terminal 2C is used by users who develop AI models used in AI image processing.

[0062] In addition, the application developer terminal 2A may be used by a user who develops applications that do not use AI image processing.

[0063] The cloud-side information processing device has prepared training data sets for AI learning. A user developing an AI model communicates with the cloud-side information processing device using an AI model developer terminal 2C and downloads these training data sets.

[0064] In this case, the training dataset may be provided for a fee. For example, an AI model developer may purchase the training dataset by registering personal information in a marketplace (electronic market) provided as a function on the cloud, thereby enabling the purchase of various functions and materials registered in the marketplace.

[0065] After developing an AI model using a training dataset, the AI ​​model developer uses the AI ​​model developer terminal 2C to register the developed AI model in the marketplace. As a result, an incentive may be paid to the AI ​​model developer when the AI ​​model is downloaded.

[0066] Furthermore, a user who develops an application downloads an AI model from the marketplace using the application developer terminal 2A and develops an application using this AI model (hereinafter simply referred to as an "AI application"). At this time, as described above, an incentive may be paid to the AI ​​model developer.

[0067] A user who develops an application registers the developed AI application in the marketplace using the application developer terminal 2A. In this way, an incentive may be paid to the user who developed the AI ​​application when the AI ​​application is downloaded.

[0068] A user who uses an AI application uses an application user terminal 2B to perform operations to deploy the AI ​​application and AI model from the marketplace to a camera 3 as an edge-side information processing device that the user manages. At this time, an incentive may be paid to the AI ​​model developer.

[0069] This allows the camera 3 to perform AI image processing using AI applications and AI models. Specifically, it is possible to not only capture images but also detect customers and vehicles through AI image processing.

[0070] Here, "deploying" an AI application or an AI model refers to installing the AI ​​application or the AI ​​model in a target (device) as an execution subject so that the target can use the AI ​​application or the AI ​​model. Furthermore, "deploying" also includes installing the AI ​​application or the AI ​​model in a target as an execution subject so that at least a part of the program as the AI ​​application can be executed.

[0071] In addition, camera 3 may be configured to be able to extract attribute information of customers from images captured by camera 3 using AI image processing.

[0072] This attribute information is transmitted from the camera 3 via the network 6 to an information processing device on the cloud side.

[0073] Cloud applications are deployed on the cloud-side information processing device. Each user can use the cloud applications via network 6. The cloud applications include an application that analyzes the movement of customers using their attribute information and captured images. Such cloud applications are uploaded by application developers and users.

[0074] A user of the application uses the cloud application for flow analysis using the application user terminal 2B. This allows the user to analyze the flow of customers visiting their own store and view the analysis results. Viewing the analysis results refers to viewing the flow of customers displayed graphically on a map of the store, for example.

[0075] The results of the flow line analysis may also be displayed in the form of a heat map, and the density of customers visiting the store may be presented, allowing the analysis results to be viewed.

[0076] In addition, the information may be displayed in a manner that is categorized according to the attribute information of the customers.

[0077] In the cloud-side marketplace, AI models optimized for each user may be registered. For example, images captured by a camera 3 installed in a store managed by a certain user are uploaded and stored in an information processing device on the cloud side as appropriate.

[0078] In the cloud-side information processing device, each time a certain number of uploaded captured images are accumulated, a re-learning process for the AI ​​model is performed, and the AI ​​model is updated and re-registered in the marketplace.

[0079] The AI ​​model re-learning process may be made available as an option to users on a marketplace, for example.

[0080] For example, an AI model retrained using dark images from a camera 3 placed inside a store is deployed to that camera 3. This makes it possible to improve the recognition rate, etc. of image processing for images captured in dark places. Also, an AI model retrained using bright images from a camera 3 placed outside the store is deployed to that camera 3. This makes it possible to improve the recognition rate, etc. of image processing for images captured in bright places.

[0081] In other words, a user of the application can always obtain optimized processing result information by deploying the updated AI model again to the camera 3.

[0082] The re-learning process for the AI ​​model will be explained later.

[0083] Furthermore, if personal information is included in information (such as captured images) uploaded from the camera 3 to the cloud-side information processing device, the data may be uploaded with the privacy information deleted from the viewpoint of privacy protection. The data with the privacy information deleted may be made available to users who develop AI models and users who develop applications.

[0084] The AI ​​model developer terminal 2C is an information processing device used by the AI ​​model developer.

[0085] In addition, the software developer terminal 7 is an information processing device used by the developer of the AI ​​application.

[0086] FIG. 3 is a flowchart showing an example of the flow of the above-mentioned process.

[0087] The information processing device on the cloud side corresponds to the cloud server 1, management server 5, etc. shown in FIG.

[0088] The AI ​​model developer browses a list of datasets registered in the marketplace using an AI model developer terminal 2C having a display unit that may include an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) panel, etc. When the AI ​​model developer selects a desired dataset, in response to this selection, the AI ​​model developer terminal 2C transmits a download request for the selected dataset to an information processing device on the cloud side (step S21).

[0089] The cloud-side information processing device receives the request (step S1), and then performs processing to transmit the requested data set to the AI ​​model developer terminal 2C (step S2).

[0090] The AI ​​model developer terminal 2C performs a process to receive the data set (step S22), which enables the AI ​​model developer to develop an AI model using the data set.

[0091] After the AI ​​model developer finishes developing an AI model, the AI ​​model developer performs an operation to register the developed AI model in the marketplace. This operation involves, for example, specifying the name of the AI ​​model and the address where the AI ​​model is located. As a result, the AI ​​model developer terminal 2C transmits a request to register the AI ​​model in the marketplace to the information processing device on the cloud side (step S23).

[0092] The cloud-side information processing device receives the registration request (step S3). The cloud-side information processing device performs registration processing for the AI ​​model (step S4). The cloud-side information processing device can, for example, display the AI ​​model on a marketplace. This allows users other than the AI ​​model developer to download the AI ​​model from the marketplace.

[0093] For example, an application developer who wishes to develop an AI application uses the application developer terminal 2A to browse a list of AI models registered in the marketplace. In response to an operation by the application developer, the application developer terminal 2A transmits a download request for the selected AI model to the cloud-side information processing device (step S31). The operation here is, for example, an operation to select one of the AI ​​models on the marketplace.

[0094] The cloud-side information processing device accepts the request (step S5) and transmits the AI ​​model to the application developer terminal 2A (step S6).

[0095] The application developer terminal 2A receives the AI ​​model (step S32), which enables the application developer to develop an AI application that uses an AI model developed by another person.

[0096] After completing development of an AI application, the application developer performs an operation to register the AI ​​application in the marketplace. This operation involves, for example, specifying the name of the AI ​​application and the address where the AI ​​model is located. As a result, the application developer terminal 2A sends a registration request for the AI ​​application to the cloud-side information processing device (step S33).

[0097] The cloud-side information processing device accepts the registration request (step S7). The cloud-side information processing device registers the AI ​​application (step S8). The cloud-side information processing device can, for example, display the AI ​​application on a marketplace. This allows users other than the application developer to select and download the AI ​​application on the marketplace.

[0098] (3) System Functionality Overview

[0099] In the first embodiment, a service using the information processing system 100 is assumed in which a user as a customer can select a function type for AI image processing of multiple cameras 3. For example, an image recognition function, an image detection function, or the like may be selected as the function type, or a more specific type may be selected so as to perform an image recognition function, an image detection function, or the like for a specific subject.

[0100] For example, as a business model, a service provider sells cameras 3 and fog servers 4 with AI image recognition functions to users, and has them install the cameras 3 and fog servers 4 in locations to be monitored. The service provider then develops a service that provides users with the above-mentioned analytical information.

[0101] In this case, each customer has different requirements for the system, such as store surveillance, traffic surveillance, etc. Therefore, the AI ​​image processing function of the camera 3 can be selectively set to obtain analytical information that corresponds to the customer's desired use.

[0102] In the first embodiment, the management server 5 has a function for selectively setting the AI ​​image processing function of such a camera 3.

[0103] In addition, the functions of the management server 5 may be provided by the cloud server 1 or the fog server 4.

[0104] (4) Configuration of the imaging device

[0105] FIG. 4 shows an example of the internal configuration of the camera 3.

[0106] 4, the camera 3 includes an imaging optical system 31, an optical system driving unit 32, an image sensor IS, a control unit 33, a memory unit 34, and a communication unit 35. The image sensor IS, the control unit 33, the memory unit 34, and the communication unit 35 are connected to each other via a bus 36, which enables data communication between them.

[0107] The imaging optical system 31 includes lenses such as a cover lens, a zoom lens, and a focus lens, as well as an iris mechanism. Light (incident light) from the subject is guided by the imaging optical system 31, and the light is collected on the light receiving surface of the image sensor IS.

[0108] The optical system driving unit 32 collectively refers to the driving units for the zoom lens, focus lens, and diaphragm mechanism of the imaging optical system 31. Specifically, the optical system driving unit 32 has actuators and actuator driving circuits for driving the zoom lens, focus lens, and diaphragm mechanism, respectively.

[0109] The control unit 33 is configured with, for example, a microcomputer having a CPU, ROM, and RAM, and performs overall control of the camera 3 by the CPU executing various processes in accordance with programs stored in the ROM or programs loaded into the RAM.

[0110] Furthermore, the control unit 33 issues drive instructions to the optical system drive unit 32 to drive the zoom lens, focus lens, diaphragm mechanism, etc. In response to these drive instructions, the optical system drive unit 32 moves the focus lens and zoom lens, opens and closes the diaphragm blades of the diaphragm mechanism, etc.

[0111] The control unit 33 also controls the writing and reading of various data to and from the memory unit 34 .

[0112] The memory unit 34 includes a non-volatile storage device such as a hard disk drive (HDD), a solid state drive (SSD), a flash memory device, etc. The memory unit 34 is used as a storage destination (recording destination) for the image data output from the image sensor IS.

[0113] Furthermore, the control unit 33 performs various data communications with external devices via the communication unit 35. The communication unit 35 in the first embodiment is capable of data communications with at least the fog server 4 (or the cloud server 1) shown in FIG.

[0114] The image sensor IS is configured as, for example, a CCD type, a CMOS type, or the like image sensor.

[0115] It should be noted that in the present disclosure, the image sensor IS is not limited to the above-mentioned devices such as CCD and CMOS, but may be, for example, a sensor including ToF (Time of Flight) pixels, a sensor for acquiring other distance images, a sensor for acquiring X-ray images, a sensor for acquiring temperature images, a sensor for acquiring ultrasonic images, an infrared sensor, an ultraviolet sensor, or other sensors for acquiring various types of information. Furthermore, the image sensor IS is not limited to one unit, and one or more types of sensors may be combined to form a plurality of image sensors IS.

[0116] The image sensor IS includes an imaging unit 41, an image signal processing unit 42, an internal sensor control unit 43, an AI image processing unit 44, a memory unit 45, and a communication interface (hereinafter referred to as a communication I / F 46). These are connected via a bus 47 and are capable of mutual data communication.

[0117] The imaging unit 41 includes a pixel array unit in which a plurality of pixels are arranged two-dimensionally, and a readout circuit. The pixels include photoelectric conversion elements such as photodiodes. The readout circuit reads out electrical signals obtained by photoelectric conversion from each pixel in the pixel array unit. The imaging unit 41 outputs the obtained electrical signals as captured image signals.

[0118] The readout circuit performs, for example, CDS (Correlated Double Sampling) processing, AGC (Automatic Gain Control) processing, etc. on the electrical signal obtained by photoelectric conversion, and further performs A / D (Analog / Digital) conversion processing.

[0119] The image signal processing unit 42 performs pre-processing, synchronization processing, YC generation processing, resolution conversion processing, codec processing, etc. on the captured image signal as digital data after A / D conversion processing.

[0120] In the pre-processing, the captured image signal is subjected to a clamping process for clamping the R, G, and B black levels to a predetermined level, a correction process between the R, G, and B color channels, and the like.

[0121] In the image synchronization process, color separation is performed so that the image data for each pixel contains all the color components R, G, and B. For example, in the case of an image sensor that uses a Bayer color filter, demosaic processing is performed as the color separation process.

[0122] In the YC generation process, a luminance (Y) signal and a color (C) signal are generated (separated) from the R, G, and B image data. In the resolution conversion process, resolution conversion is performed on image data that has undergone various signal processes.

[0123] In codec processing, the image data that has undergone the various processes described above is encoded for recording or communication, and a file is generated. In codec processing, moving image file formats such as MPEG-2 (Moving Picture Experts Group) and H.264 can be generated. Still image files can be generated in compressed formats such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), and GIF (Graphics Interchange Format), or in uncompressed formats such as TIFF (Tagged Image File Format) and raw data.

[0124] The sensor control unit 43 issues instructions to the imaging unit 41 and controls the execution of imaging operations. Similarly, the sensor control unit 43 also controls the image signal processing unit 42 to execute processing.

[0125] The AI ​​image processing unit 44 performs image recognition processing as AI image processing on the captured image. The image recognition function using AI can be realized using a programmable arithmetic processing device such as a CPU, FPGA (Field-Programmable Gate Array), ASIC (Application Specific Integrated Circuit), or DSP (Digital Signal Processor).

[0126] The image recognition functions that can be realized in the AI ​​image processing unit 44 can be switched by changing the algorithm of the AI ​​image processing. In other words, the type of AI image processing function can be switched by switching the AI ​​model used for the AI ​​image processing. The types of AI image processing functions are, for example, as follows: Class identification Semantic segmentation Person detection Vehicle detection Target tracking Optical character recognition (OCR)

[0127] Of the above function types, class identification is a function that identifies the class of a target. This "class" is information that represents the category of an object. For example, classes distinguish between "person," "car," "airplane," "ship," "truck," "bird," "cat," "dog," "deer," "frog," "horse," etc.

[0128] Target tracking is a function for tracking a subject that has been targeted. In other words, target tracking is a function for obtaining historical information about the position of the subject.

[0129] The memory unit 45 is used as a storage destination for various data such as captured image data obtained by the image signal processing unit 42. In the first embodiment, the memory unit 45 is also used for temporary storage of data used by the AI ​​image processing unit 44 in the AI ​​image processing process.

[0130] In addition, the memory unit 45 stores information on AI applications and AI models used in the AI ​​image processing unit 44.

[0131] Information about the AI ​​application and the AI ​​model may be deployed in the memory unit 45 as a container or the like using the container technology described below. Information about the AI ​​application and the AI ​​model may also be deployed using microservice technology. By deploying the AI ​​model used for AI image processing in the memory unit 45, it is possible to change the function type of the AI ​​image processing and change to an AI model whose performance has been improved by re-learning.

[0132] Although the above-described first embodiment has been described based on examples of AI models and AI applications used for image recognition, the present technology is not limited thereto and may also be applied to programs executed using AI technology.

[0133] Furthermore, if the capacity of the memory unit 45 is small, information on the AI ​​application and AI model may be expanded into memory outside the image sensor IS as a container using container technology, such as the memory unit 34, and then only the AI ​​model may be stored in the memory unit 45 within the image sensor IS via the communication I / F 46 described below.

[0134] The communication I / F 46 is an interface for communicating with the control unit 33, memory unit 34, etc., which are external to the image sensor IS. The communication I / F 46 communicates to acquire from the outside the program executed by the image signal processing unit 42, the AI ​​application used by the AI ​​image processing unit 44, the AI ​​model, etc. This information is stored in the memory unit 45 provided in the image sensor IS. As a result, the AI ​​model, etc., is stored in part of the memory unit 45 provided in the image sensor IS and can be used by the AI ​​image processing unit 44.

[0135] The AI ​​image processing unit 44 performs predetermined image recognition processing using the AI ​​application or AI model obtained in this manner, thereby recognizing the subject according to the purpose.

[0136] The recognition result information of the AI ​​image processing is output to the outside of the image sensor IS via the communication I / F 46.

[0137] That is, the communication I / F 46 of the image sensor IS outputs not only the image data output from the image signal processing unit 42 but also the recognition result information of the AI ​​image processing.

[0138] The communication I / F 46 of the image sensor IS can output either the image data or the recognition result information.

[0139] For example, when using the re-learning function of the AI ​​model described above, the captured image data used for the re-learning function is uploaded from the image sensor IS to an information processing device on the cloud side via the communication I / F 46 and the communication unit 35.

[0140] In addition, when inference is performed using an AI model, the recognition result information of the AI ​​image processing is output from the image sensor IS to another information processing device outside the camera 3 via the communication I / F 46 and the communication unit 35.

[0141] The image sensor IS may have various structures. In the first embodiment, the structure of the image sensor IS having a two-layer stacked structure will be described.

[0142] FIG. 5 shows an example of the configuration of an image sensor IS as an imaging device.

[0143] As shown in FIG. 5, the image sensor IS is formed by a semiconductor device in which two semiconductor chips, dies D1 and D2, are stacked together to form a single semiconductor chip.

[0144] The die D1 has the function of the imaging unit 41 shown in Fig. 4. The die D2 has the functions of an image signal processing unit 42, an internal sensor control unit 43, an AI image processing unit 44, a memory unit 45, and a communication I / F 46.

[0145] The die D1 and the die D2 each have terminals on their opposing surfaces. The terminals are formed, for example, using copper (Cu) as a wiring material. That is, the die D1 and the die D2 are electrically connected by Cu-Cu bonding that joins the terminals.

[0146] The technology in the present disclosure is not limited to the above-described embodiment, and various modifications are possible without departing from the spirit of the present disclosure.

[0147] For example, the solid-state imaging devices according to two or more of the embodiments may be combined.

[0148] (Second embodiment)

[0149] A more specific description will be given of AI processing including AI image processing in the first embodiment described above. In this embodiment, a detailed description will be given of processing when sensors of various modalities are used as the image sensor IS.

[0150] A configuration according to this embodiment is shown, for example, in Fig. 2. In Fig. 2, the AI ​​model developer terminal 2C or the software developer terminal 7 may be an information processing device that trains the model and converter according to this embodiment, and the application developer terminal 2A and the application user terminal 2B may be information processing devices that execute processes such as estimation and analysis using the trained model according to this embodiment.

[0151] Furthermore, the multiple cameras 3 are equipped with sensors that acquire various different types of information as described above. As another example, the cameras 3 are not limited to optical sensing. The sensors included in the cameras 3 may also be sensors that acquire information other than light.

[0152] Examples of sensors included in the multiple cameras 3 include, but are not limited to, sensors that acquire RGB information, sensors that acquire grayscale or brightness, sensors that acquire raw data, sensors that acquire depth information, sensors that acquire event information, sensors that acquire polarization information, multispectral sensors, etc.

[0153] Each of the application developer terminal 2A, the application user terminal 2B, the AI ​​model developer terminal 2C, and / or the software developer terminal 7 may include, for example, a processor such as a CPU or a GPU (Graphics Processing Unit) and a processing circuit as a computing unit. These terminals may also include, for example, temporary and / or non-temporary storage circuits and storage devices as a storage unit. At least a portion of the storage unit may be provided in the cloud server 1 or the management server 5, and each information processing device may acquire data by accessing these servers.

[0154] 2, the camera 3 is connected to each terminal via the cloud, but this is not limiting. The camera 3 may be connected to the AI ​​model developer terminal 2C or the software developer terminal 7 during learning in a manner that allows direct data transmission and reception, and may be connected to the application developer terminal 2A or the application user terminal 2B during inference in a manner that allows direct data transmission and reception.

[0155] Examples of processes that a trained model may perform include, but are not limited to, person detection, detection of the orientation of a person (partially or entirely), detection of two-dimensional barcodes, image classification, counting of objects (e.g., people or animals), object detection, semantic segmentation, and feature point detection.

[0156] 6 is a diagram showing an outline of processing according to one embodiment. An information processing device, for example, an AI model developer terminal 2C or a software developer terminal 7, optimizes processing using a first model that executes a first processing step, based on information output from a first sensor 3A.

[0157] The first model 22A is, for example, a model that outputs a result of executing a first process on first data output from the first sensor 3A. The first model 22A may be a model of any format. The first model 22A may be, for example, a model that has been trained by machine learning.

[0158] More specifically, the first model 22A is a model trained to output the result of executing the first processing when the first intermediate representation obtained by converting the first data by the first converter 21A is input. In other words, the first intermediate representation is an intermediate representation (latent representation) that, when input to the first model 22A, can output the result of executing the first processing on the first data.

[0159] An information processing device (e.g., an AI model developer terminal 2C or a software developer terminal 7) that performs learning of the converter converts the second data output from the second sensor 3B into a second intermediate representation using the second converter 21B, and performs learning (optimization) of the second converter 21B based on this second intermediate representation.

[0160] Here, the first converter 21A and the second converter 21B may be models formed based on neural network models such as CNN and transformers, or may be models related to domain application.

[0161] By training the second converter 21B in this manner, when both the first intermediate representation, which is an intermediate representation of the first data acquired from the first sensor 3A, and the second intermediate representation, which is an intermediate representation of the second data acquired from the second sensor 3B, are input into the first model 22A, the result of executing the first processing can be obtained.

[0162] As described above, the first model 22A is a model that realizes the first processing on the output of the first sensor 3A, but according to this embodiment, it is also possible to use the same first model 22A to perform the first processing on the second data, which is the output of the second sensor 3B that outputs a different type of information from the first sensor 3A.

[0163] Although two sensors are shown, similar processing can be performed with sensors that acquire three or more different types of information. The same applies to the embodiments described below.

[0164] 7 is a flowchart showing information processing according to this embodiment. The first model 22A (and the first converter 21A) have been trained in advance using a data set of the first data. This training may be performed based on any method. This training may be supervised, semi-supervised, or unsupervised.

[0165] The calculation unit of the information processing device acquires first data output by the first sensor 3A and second data output by the second sensor 3B (S100). For example, the first sensor 3A and the second sensor 3B acquire information of the same object in the same environment and output the first data and the second data, respectively. The calculation unit acquires information of the same object in the same environment, for example, captured image data, as the first data and the second data.

[0166] As described above, the first data and the second data are different types of information, such as an RGB image and a depth image. It is difficult to obtain the same processing results by simply inputting such different types of information into the same model. By optimizing the converter, it is possible to perform the same processing on the different types of information using the same model.

[0167] Next, the calculation unit obtains a first intermediate representation and a second intermediate representation (S102). For example, the calculation unit inputs the first data to the first converter 21A to obtain the first intermediate representation, and inputs the second data to the second converter 21B to obtain the second intermediate representation.

[0168] Next, the calculation unit obtains a loss between the first intermediate representation and the second intermediate representation (S104). This loss may be calculated using any method for calculating an error. For example, any norm or KL divergence may be used as the loss.

[0169] Next, the calculation unit optimizes the converter by updating the parameters of the second converter 21B based on the loss calculated in S104 (S106). The processes from S102 to S106 may be repeated as needed, and the processes from S100 to S106 may be repeated as needed. The optimization can be completed based on a general termination condition, such as when a predetermined number of epochs have been processed or when the evaluation value has fallen below a predetermined threshold.

[0170] The calculation unit transmits the parameters of the second converter 21B optimized in S106 to a server such as the cloud server 1 or the management server 5, for output (S108), and completes the process. As another example, the calculation unit may store the parameters in a memory unit within an information processing device such as the AI ​​model developer terminal 2C or the software developer terminal 7, instead of outputting them.

[0171] The second converter 21B optimized in this way makes it possible to convert the second data into a second intermediate representation having a distribution similar to that of the intermediate representation of the first data obtained by acquiring information of the same subject in the same environment. Therefore, by inputting the second intermediate representation into the first model 22A, which accepts as input the first intermediate representation, which is an intermediate representation of the first data, it becomes possible to acquire the results of appropriately executing the first process using the second data.

[0172] That is, by using the converter, the application developer terminal 2A, the application user terminal 2B, or the software developer terminal 7 can realize the first process by having the calculation unit acquire the second data and convert the second data into a second intermediate representation, and inputting this second intermediate representation into the first model 22A. The first model 22A is a model trained to execute the first process when it receives the first intermediate representation obtained by converting the first data output from the first sensor 3A. In this way, it becomes possible to apply a model optimized for the first data to the second data via the converter.

[0173] For example, the application developer terminal 2A, the application user terminal 2B, or the software developer terminal 7 can achieve appropriate processing by changing the converter for obtaining the intermediate representation calculated as the front stage of the model for input from different sensors, without changing the configuration of the model executed in the calculation unit.

[0174] (Third embodiment)

[0175] 8 is a diagram showing an outline of an information processing system according to an embodiment. The information processing system can use the same converter 21 instead of providing a converter for each sensor. In this case, the first data is converted to a first intermediate representation by the converter 21, and the second data is converted to a second intermediate representation by the same converter 21. As above, the converter 21 can be a model that realizes domain application.

[0176] The processing flow is generally the same as in Fig. 7. The calculation unit can optimize the converter 21 by updating the parameters related to the second intermediate representation so that the first intermediate representation remains unchanged.

[0177] The converter 21 may be configured to obtain an intermediate representation based on data automatically input in the input layer without considering the type of sensor, for example.

[0178] The converter 21 may be configured, for example, to execute input from different neurons for each sensor in the input layer to obtain an intermediate representation.

[0179] The converter 21 may have a configuration in which, for example, a neuron that inputs the type of sensor exists, and the type of connected sensor is explicitly given to the input layer. As an application of this, the converter 21 may also have a configuration in which, in addition to inputting image data, etc., a one-hot vector indicating the sensor can be input.

[0180] In this way, it is not necessary to prepare a converter 21 for each individual sensor, but it may be implemented as a single converter. In this case, it is possible to use a model that performs appropriate processing on input from the sensor by simply switching the sensor connection without any processing on the calculation unit side.

[0181] In the embodiment described below, the processing by this converter 21 will be explained, but of course it can also be applied to the case where there are multiple converters as shown in FIG. 6 within the scope of no contradiction.

[0182] (Fourth embodiment)

[0183] Fig. 9 is a diagram showing an outline of an information processing system according to one embodiment. The information processing system includes a converter 21 for executing the first model 22A of Fig. 8, and can also perform different processing based on the intermediate representation converted by this converter 21.

[0184] The calculation unit can input the first intermediate representation and / or the second intermediate representation output from the converter 21 into the second model 22B and obtain the results of executing a second process that is different from the first process.

[0185] The second model 22B is a model trained to perform a second process on the first data. The first data can be converted into a first intermediate representation by the converter 21, thereby generating input data for the second model 22B.

[0186] The converter 21 is the same in Figures 8 and 9. That is, for both the data acquired from the first sensor 3A and the second sensor 3B, the intermediate representation acquired through the converter 21 can be used to acquire the result of a first process using a first model, or the result of a second process using a second model 22B. In other words, it is possible to replace only the model following the converter 21 with a model that realizes another process.

[0187] More specifically, the calculation unit inputs the second intermediate representation into a second model 22B that has been trained to execute the second processing when the first intermediate representation is input, thereby enabling the second processing to be executed on the output of the second sensor 3B.

[0188] As described above, according to this embodiment, it is possible not only to replace the sensor but also to change the model that executes the processing. This configuration makes it possible to realize more robust analysis using various types of sensors by using a model trained using output data from a specific sensor.

[0189] However, the implementation is not limited to this.

[0190] FIG. 10 is a diagram showing another aspect of this embodiment. As shown in this figure, the information processing system can use a third converter 21C to obtain a third intermediate representation of the first data, and a third model 22C that has been trained to be able to obtain the results of a third process by inputting this third intermediate representation. That is, the intermediate representation for the first model 22A and the intermediate representation for the third model 22C that performs a different process may be different even for the same sensor. The third intermediate representation is an input representation of the first data for the third model 22C.

[0191] In this case, the calculation unit can be configured to convert the first data acquired from the first sensor 3A into a third intermediate representation using the third converter 21C, convert the second data acquired from the second sensor 3B into a fourth intermediate representation using the fourth converter 21D, and perform learning of the fourth converter 21D based on this third intermediate representation and the fourth intermediate representation to perform optimization.

[0192] In this way, when the model for executing the process is changed, it is also possible to change the converter. Of course, the converter may be one that is shared by both the first data and the second data, as shown in Fig. 9 etc.

[0193] The calculation unit of the AI ​​model developer terminal 2C can, for example, acquire first data and second data obtained using a first sensor 3A and a second sensor 3B, respectively, which are information on the same object in the same environment, input the first data into a third converter 21C to obtain a third intermediate representation, input the second data into a fourth converter 21D to obtain a fourth intermediate representation, and calculate the loss of these intermediate representations, thereby updating and optimizing the parameters of the fourth converter 21D.

[0194] The calculation unit of the application developer terminal 2A, the application user terminal 2B or the software developer terminal 7 can, for example, convert the data output from the second sensor 3B using the fourth converter 21D to obtain an intermediate representation, and input this intermediate representation into the third model 22C, thereby realizing the third processing using the third model 22C that has been trained to execute the third processing using the output of the first sensor 3A on the output from the second sensor 3B.

[0195] Although the above-described embodiments have been described mainly as information processing devices, they can also operate as information processing systems including sensors.

[0196] An information processing system as a converter learning system includes sensors that acquire different types of information, and an information processing device acquires intermediate representations for each of these sensors, and the converter can be optimized from these intermediate representations. The optimized converter makes it possible to perform converter learning to realize processing such as analysis using the same model for outputs from different sensors acquired inside or outside the information processing system.

[0197] As in the above, this converter may correspond to a different converter for each sensor, which is, for example, a model related to domain application, or the same converter may correspond to each sensor.

[0198] According to all embodiments from the second embodiment onward, it is possible to implement appropriate processing for outputs from sensors that produce different types of output without retraining the model that implements the processing, i.e., without incurring the cost of retraining. For example, it is possible to shorten the training time. Furthermore, when supervised learning is realized for a certain sensor (labels are prepared), it is possible to apply unsupervised data (data without labels) from other sensors to the same trained model.

[0199] This modular system of interchangeable sensors and domain adaptation allows for seamless integration of diverse sensor data within the same trained neural network processing pipeline. The modular system and ability to utilize common input feature generation reduces the cost of hardware development and production.

[0200] It also provides an efficient and flexible approach to processing sensor data, allowing users to adapt devices to different sensors without significant changes to the implementation or time-consuming retraining. This also saves time and resources. For example, it allows users to effectively omit pre-processing before inputting each sensor into the model.

[0201] By using this information processing device or system, it is easy to make changes to the modular design, such as incorporating additional components such as different mechanisms for sensor compatibility, different architectures, different algorithms, etc. To achieve a similar effect, a trained neural network model may be prepared for each sensor, and in this case, the information processing device can be adapted by changing the model depending on the sensor, but on the other hand, as described above, retraining for each sensor requires time and resources.

[0202] This system allows for greater flexibility in sensor replacement, allowing users to seamlessly change sensors without changing the system contents, which is advantageous in scenarios where different sensor modalities are required for different applications or environments.

[0203] Furthermore, processing suited to output data from a certain sensor can be learned using a data set from that sensor. Furthermore, since it is possible to align intermediate representations of data output from other sensors, it becomes possible to achieve highly accurate processing using the same trained model for those other sensors. From another perspective, even for a sensor for which it is difficult to acquire a large amount of data for learning, if there is an abundance of data from other sensors for learning the model, it becomes possible to achieve that processing for that sensor by learning a model that processes using a sensor with an abundant amount of data available.

[0204] Furthermore, when developing a new sensor or applying a new sensor, it is possible to execute the processing using the same model by learning a model that aligns the intermediate representation, without having to learn a new model to execute the processing.

[0205] The above-described embodiment may be modified as follows.

[0206] (1) An information processing device comprising: a calculation unit that acquires first data output by a first sensor and second data output by a second sensor that acquires a different type of information from the first sensor; inputs the first data into a first converter to acquire a first intermediate representation; inputs the second data into a second converter to acquire a second intermediate representation; and executes learning of the second converter based on the first intermediate representation and the second intermediate representation.

[0207] (2) The information processing device according to (1), wherein the first intermediate representation is an intermediate representation that can execute a first process when input to a first model.

[0208] (3) The information processing device according to (2), wherein the first model is a model trained to execute the first processing when the first intermediate representation is input.

[0209] (4) The information processing device according to (2) or (3), wherein the first intermediate representation is an intermediate representation that can execute a second process when input to a second model.

[0210] (5) The information processing device according to (4), wherein the second model is a model trained to execute the second processing when the first intermediate representation is input.

[0211] (6) The information processing device according to any one of (1) to (5), wherein the calculation unit acquires information of the same object in the same environment using the first sensor and the second sensor, respectively, to acquire the first data and the second data, inputs the first data to the first converter to acquire the first intermediate representation, inputs the second data to the second converter to acquire the second intermediate representation, calculates a loss between the first intermediate representation and the second intermediate representation, updates parameters of the second converter, and optimizes the second converter.

[0212] (7) The information processing device according to any one of (1) to (6), wherein the calculation unit inputs the first data to a third converter to obtain a third intermediate representation, inputs the second data to a fourth converter to obtain a fourth intermediate representation, and performs learning of the fourth converter based on the third intermediate representation and the fourth intermediate representation.

[0213] (8) The information processing device according to (7), wherein the third intermediate representation is an intermediate representation that can execute a third process when input to a third model.

[0214] (9) The information processing device according to (8), wherein the third model is a model trained to execute the third processing when the third intermediate representation is input.

[0215] (10) The information processing device according to any one of (7) to (9), wherein the calculation unit acquires information of the same object in the same environment using the first sensor and the second sensor, respectively, to acquire the first data and the second data, inputs the first data to the third converter to acquire the third intermediate representation, inputs the second data to the fourth converter to acquire the fourth intermediate representation, calculates a loss between the third intermediate representation and the fourth intermediate representation, updates parameters of the fourth converter, and optimizes the fourth converter.

[0216] (11) The information processing device according to any one of (1) to (10), wherein the first converter and the second converter are models related to domain application.

[0217] (12) The information processing device according to (11), wherein the first converter and the second converter are of the same model.

[0218] (13) An information processing device including a calculation unit that acquires second data output by a second sensor that acquires a type of information different from that of a first sensor, converts the second data to acquire a second intermediate representation, and inputs the second intermediate representation to a first model that executes a first process when a first intermediate representation obtained by converting the first data output by the first sensor is input, and acquires a result of executing the first process on the second data. The converter that acquires the intermediate representation may be a model optimized by any of the information processing devices of (1) to (12).

[0219] (14) The information processing device described in (13), wherein the calculation unit inputs the second intermediate representation to a second model that executes a second process when the first intermediate representation is input, and obtains a result of executing the second process on the second data.

[0220] (15) The information processing device according to (13) or (14), wherein the calculation unit converts the second data to obtain a fourth intermediate representation, inputs the fourth intermediate representation to a third model that executes a third process when a third intermediate representation obtained by converting the first data is input, and obtains a result of executing the third process on the second data.

[0221] (16) An information processing system comprising: a first sensor; a second sensor that acquires a type of information different from that of the first sensor; and an information processing device according to any one of (1) to (11), wherein the information processing device optimizes a converter that converts second data acquired from the second sensor into a first intermediate representation based on a first intermediate representation obtained by converting first data acquired from the first sensor and a second intermediate representation obtained by converting second data acquired from the second sensor.

[0222] (17) The information processing system according to (16), wherein the transformation for obtaining the first intermediate representation and the transformation for obtaining the second intermediate representation are models related to domain application.

[0223] (18) The information processing system according to (17), wherein the transformation to obtain the first intermediate representation and the transformation to obtain the second intermediate representation are performed using the same model.

[0224] (19) An information processing system comprising: a first sensor; a second sensor that acquires a type of information different from that of the first sensor; and an information processing device according to any one of (12) to (14), wherein the information processing device inputs a second intermediate representation obtained by converting second data acquired from the second sensor into a first model that executes a first process when a first intermediate representation obtained by converting first data acquired from the first sensor is input, and acquires a result of executing the first process on the second data.

[0225] The aspects of the present disclosure are not limited to the above-described embodiments and include various conceivable modifications, and the effects of the present disclosure are not limited to the above-described contents. The components in each embodiment may be appropriately combined and applied. In other words, various additions, modifications, and partial deletions are possible within the scope of the conceptual idea and intent of the present disclosure, which is derived from the content defined in the claims and their equivalents.

[0226] 100: Information processing system, 1: Cloud server, 2: User terminal, 2A: Application developer terminal, 2B: Application user terminal, 2C: AI model developer terminal, 21: Converter, 21A: First converter, 21B: Second converter, 21C: Third converter, 21D: Fourth converter, 22A: First model, 22B: Second model, 22C: Third model, 3: Camera, 31: Imaging optical system, 32: Optical system driving unit, 33: Control unit, 34: Memory unit, 35: Communication unit, 36: Bus, IS: Image sensor, 41: Imaging unit, 42: Image signal processing unit, 43: Sensor internal control unit, 44: AI image processing unit, 45: Memory unit, 46: Communication I / F, 47: Bus, D1, D2: Die, 3A: First sensor, 3B: Second sensor, 4: Fog server, 5: Management server, 6: Network, 7: Software developer terminal

Claims

1. An information processing device comprising: a calculation unit that acquires first data output by a first sensor and second data output by a second sensor that acquires a different type of information from the first sensor; inputs the first data to a first converter to acquire a first intermediate representation; inputs the second data to a second converter to acquire a second intermediate representation; and performs learning of the second converter based on the first intermediate representation and the second intermediate representation.

2. The information processing device according to claim 1, wherein the first intermediate representation is an intermediate representation capable of executing a first process when inputted into a first model.

3. The information processing device according to claim 2, wherein the first model is a model trained to execute the first processing when the first intermediate representation is input.

4. The information processing device according to claim 3, wherein the first intermediate representation is an intermediate representation capable of executing a second process when inputted into a second model.

5. The information processing device according to claim 4, wherein the second model is a model trained to execute the second processing when the first intermediate representation is input.

6. The information processing device according to claim 1, wherein the calculation unit: acquires information of the same object in the same environment using the first sensor and the second sensor, respectively, to acquire the first data and the second data; inputs the first data to the first converter to acquire the first intermediate representation; inputs the second data to the second converter to acquire the second intermediate representation; calculates a loss between the first intermediate representation and the second intermediate representation, updates parameters of the second converter, and optimizes the second converter.

7. The information processing device according to claim 1, wherein the calculation unit inputs the first data to a third converter to obtain a third intermediate representation, inputs the second data to a fourth converter to obtain a fourth intermediate representation, and performs learning of the fourth converter based on the third intermediate representation and the fourth intermediate representation.

8. The information processing device according to claim 7, wherein the third intermediate representation is an intermediate representation capable of executing a third process when inputted into a third model.

9. The information processing device according to claim 8, wherein the third model is a model trained to execute the third processing when the third intermediate representation is input.

10. The information processing device according to claim 7, wherein the calculation unit: acquires information of the same object in the same environment using the first sensor and the second sensor, respectively, to acquire the first data and the second data; inputs the first data to the third converter to acquire the third intermediate representation; inputs the second data to the fourth converter to acquire the fourth intermediate representation; calculates a loss between the third intermediate representation and the fourth intermediate representation, updates parameters of the fourth converter, and optimizes the fourth converter.

11. The information processing device according to claim 1, wherein the first converter and the second converter are models related to domain application.

12. The information processing device according to claim 11, wherein the first converter and the second converter are of the same model.

13. An information processing device comprising a calculation unit, which acquires second data output by a second sensor that acquires a different type of information from a first sensor, converts the second data to acquire a second intermediate representation, inputs the second intermediate representation to a first model that executes a first processing upon receiving a first intermediate representation obtained by converting the first data output by the first sensor, and acquires a result of executing the first processing on the second data.

14. The information processing device according to claim 13, wherein the calculation unit inputs the second intermediate representation to a second model that executes a second process when the first intermediate representation is input, and obtains a result of executing the second process on the second data.

15. The information processing device according to claim 13, wherein the calculation unit converts the second data to obtain a fourth intermediate representation, inputs the fourth intermediate representation to a third model that executes a third process when a third intermediate representation obtained by converting the first data is input, and obtains a result of executing the third process on the second data.

16. An information processing system comprising: a first sensor; a second sensor that acquires a type of information different from that of the first sensor; and an information processing device according to claim 1, wherein the information processing device optimizes a converter that converts second data acquired from the second sensor into a first intermediate representation based on a first intermediate representation acquired by converting first data acquired from the first sensor and a second intermediate representation acquired by converting second data acquired from the second sensor.

17. The information processing system of claim 16, wherein the transformation to obtain the first intermediate representation and the transformation to obtain the second intermediate representation are models related to domain application.

18. The information processing system of claim 17, wherein the transformation to obtain the first intermediate representation and the transformation to obtain the second intermediate representation are performed using a same model.

19. An information processing system comprising: a first sensor; a second sensor that acquires a different type of information from the first sensor; and the information processing device according to claim 12, wherein the information processing device inputs a second intermediate representation obtained by converting second data acquired from the second sensor into a first model that executes a first processing upon inputting a first intermediate representation obtained by converting first data acquired from the first sensor, and acquires a result of executing the first processing on the second data.

Citation Information

Patent Citations

  • Inference device, inference method, and inference program

    JP2021022079A

  • Method and apparatus for providing artificial intelligence-based energy saving model for IoT device

    KR1020250009846A