Signal processing device, signal processing method, and program

The signal processing device addresses the challenge of real-time processing by selectively outputting AI model inference results based on designation information, enhancing processing efficiency and reducing data transfer burdens.

WO2025134692A1PCT designated stage expired Publication Date: 2025-06-26SONY SEMICON SOLUTIONS CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/041565
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-11-25
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing systems face challenges in performing real-time processing of recognition results from AI models deployed on edge devices, due to the need to analyze and select necessary information from the received results.

Method used

A signal processing device with an inference processing unit that performs inference processing using an AI model, receives designation information to select specific inference results, and outputs these results based on an output instruction, allowing for partial output of inference results rather than all of them.

Benefits of technology

This solution enables easy real-time processing by allowing only necessary inference results to be output, reducing data transfer load and facilitating processing even with low data transfer rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024041565_26062025_PF_FP_ABST
    Figure JP2024041565_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A signal processing device according to the present technology comprises: an inference processing unit that performs inference processing using an AI model; a control unit that receives designation information designating at least part of an inference result obtained as a result of the inference processing and gives an output instruction regarding the inference result on the basis of the designation information; and an output unit that outputs the inference result on the basis of the output instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Signal processing device, signal processing method, and program

[0001] The present technology relates to the technical fields of a signal processing device, a signal processing method, and a program that perform inference using an AI model.

[0002] A technology for obtaining a recognition result by inputting predetermined input data into an AI model and performing predetermined recognition processing has become widespread. It is possible to deploy an AI model on an edge device by making it smaller and lighter (see, for example, Patent Document 1 below). By deploying the AI ​​model on the edge device, the recognition processing can be performed on the edge device, and only the recognition result can be output to a downstream device such as a server device, thereby reducing the amount of data involved in data communication.

[0003] International Publication No. 2018 / 051809

[0004] In edge devices where AI models are deployed, such as image sensors, interfaces are constructed so that predetermined recognition results and image data are output for each captured frame from the start of the image capture operation. Therefore, downstream devices that receive the recognition results from the AI ​​model analyze the received recognition results and select the necessary information before using them for various processes, making real-time processing difficult.

[0005] The present technology has been developed in view of such problems, and aims to facilitate real-time processing.

[0006] The signal processing device according to the present technology includes an inference processing unit that performs inference processing using an AI model, a control unit that receives designation information that designates at least a part of the inference results obtained as a result of the inference processing and issues an output instruction for the inference results based on the designation information, and an output unit that outputs the inference results based on the output instruction. This makes it possible to output a part of the inference results obtained by the inference processing, rather than necessarily outputting all of the inference results to a subsequent stage.

[0007] 1 is a block diagram showing a schematic configuration of a signal processing system in the present embodiment. FIG. 1 is an explanatory diagram showing first input image data and a first inference result. FIG. 2 is a table showing an example of a first inference result. FIG. 2 is an explanatory diagram showing second input image data and an inference result. FIG. 3 is a table showing an example of a second inference result. FIG. 3 is a diagram showing an example of the structure of a MIPI-compliant packet. FIG. 4 is a diagram showing an example of an inference result output when a camera device is installed. FIG. 5 is a diagram showing an example of an inference result output when an AI model is deployed. FIG. 6 is a diagram showing an example of an inference result output when the system is in operation. FIG. 7 is a diagram showing an example in which the size of data stored in the data section differs for each packet. FIG. 8 is a diagram showing an example in which the size of data stored in the data section is unified. FIG. 9 is a flowchart showing an example of processing executed by an image sensor as a pre-stage device. FIG. 10 is a flowchart showing an example of processing executed by a signal processing device. FIG. 11 is a block diagram showing an example of a signal processing system according to a modified example.

[0008] Hereinafter, with reference to the accompanying drawings, an embodiment of an information processing device according to the present technology will be described in the following order: <1. Configuration of signal processing system> <2. Data transfer> <3. Processing flow> <4. Modification example> <5. Application example> <6. Summary> <7. Present technology>

[0009] 1 is a block diagram showing a schematic configuration example of a signal processing system 1 according to an embodiment of the present technology. The signal processing system 1 is, for example, a camera device 1A configured to have each unit included in a single camera housing.

[0010] The camera device 1A is configured with multiple AI processing devices that perform inference processing using an AI (Artificial Intelligence) model. Note that inference processing using an AI model will be simply referred to as "AI processing." In other words, an AI processing device refers to a device that performs AI processing. Furthermore, AI processing on images will be referred to as "AI image processing."

[0011] The multiple AI processing devices provided in the camera device 1A are, for example, a signal processing device 2 and a pre-stage device 3 that is a device in the pre-stage of the signal processing device 2.

[0012] The pre-stage device 3 is, for example, an image sensor 3A that receives light from a subject and performs photoelectric conversion to obtain a predetermined image.

[0013] The signal processing device 2 is configured as a single board on which an ISP (Image Signal Processor) capable of executing AI image processing is mounted, for example.

[0014] The camera device 1A includes a subsequent-stage device 4 that is a device subsequent to the signal processing device 2. The subsequent-stage device 4 is a device that identifies inference results required outside the camera device 1A from among various inference results obtained by AI image processing in the previous-stage device 3 or the signal processing device 2 located in the previous stage, and transmits the inference results to the previous-stage device 3 or the signal processing device 2.

[0015] The subsequent device 4 is, for example, a microcomputer such as an AP (Application Processor). The subsequent device 4 may be mounted on a different board from the signal processing device 2, or may be mounted on the same board as the signal processing device 2.

[0016] The image sensor 3A as the pre-stage device 3 includes, for example, an image sensor unit 31, a development processing unit 32, a control unit 33, a first AI processing unit 34, and a data transfer unit 35.

[0017] The image sensor unit 31 includes a pixel array unit 31a as a light receiving unit and a readout circuit (not shown). The pixel array unit 31a is configured by a two-dimensional array of pixels that perform photoelectric conversion to output a signal according to the amount of received light.

[0018] The readout circuit performs, for example, CDS (Correlated Double Sampling) processing, AGC (Automatic Gain Control) processing, etc. on the electrical signal obtained by photoelectric conversion, and further performs A / D (Analog / Digital) conversion processing.

[0019] Image data output from the image sensor 3A and based on the digital signal read out by the readout circuit is referred to as RAW image data.

[0020] The development processing unit 32 performs preprocessing, synchronization processing, YC generation processing, resolution conversion processing, codec processing, and other processes on the RAW image data after A / D conversion processing. Preprocessing includes clamping the R, G, and B black levels of the captured image signal to a predetermined level, and correction processing between the R, G, and B color channels. Synchronization processing involves color separation processing to ensure that the image data for each pixel contains all R, G, and B color components. For example, in the case of an image sensor using a Bayer array color filter, demosaic processing is performed as the color separation processing. YC generation processing generates (separates) a luminance (Y) signal and a color (C) signal from the R, G, and B image data. Resolution conversion processing involves executing resolution conversion processing on image data that has undergone various signal processes.

[0021] In codec processing, the image data that has undergone the various processes described above is subjected to encoding processing and file generation for recording or communication, for example. In codec processing, it is possible to generate moving image file formats such as MPEG-2 (Moving Picture Experts Group) and H.264. It is also possible to generate still image file formats such as JPEG (Joint Photographic Experts Group), TIFF (Tagged Image File Format), and GIF (Graphics Interchange Format).

[0022] The image data generated by the development processing unit 32 is used as input image data to be input to the AI ​​model.

[0023] Such a development processing section 32 is provided in the image sensor 3A as, for example, an ISP (Image Signal Processor).

[0024] The control unit 33 causes the first AI processing unit 34 to execute AI image processing using the first AI model M1. Here, the AI ​​image processing by the first AI processing unit 34 is referred to as "first AI image processing."

[0025] The control unit 33 also receives, as instruction information, information for selecting at least a portion of the first inference results obtained by the first AI image processing from a downstream device, specifically, the downstream device 4. Based on the received instruction information, the control unit 33 causes the first AI processing unit 34 to supply a predetermined first inference result to the data transfer unit 35. Note that the control unit 33 may cause all of the first inference results obtained by the first AI processing unit 34 to be supplied to the data transfer unit 35, and may instruct the data transfer unit 35 to transfer a portion of the first inference results based on the instruction information to the downstream device.

[0026] The first AI processing unit 34 inputs the first input image data G1 to the first AI model M1 based on instructions from the control unit 33, and obtains a first inference result. The first input image data G1 is, for example, image data generated by the development processing unit 32.

[0027] The first input image data G1 input to the first AI model M1 may be image data that has undergone various development processes by the development processing unit 32, may be RAW image data, or may be cropped image data in which a specified area has been further cut out by the development processing unit 32.

[0028] The first AI model M1, for example, executes a process of estimating an area in which a person appears in the first input image data G1 (see FIG. 2). Specifically, the first AI model M1 performs object recognition on the input image data and outputs classification information, likelihood information, and coordinate information of the detected object.

[0029] 2 shows an image in which a person recognized from the first input image data G1 is surrounded by a bounding box drawn with a dashed line, and FIG. 3 shows an overview of the first inference result obtained by the first AI image processing using the first AI model M1.

[0030] Based on an instruction from the control unit 33, the first AI processing unit 34 supplies, for example, only the coordinate information out of the classification information, likelihood information, and coordinate information to the data transfer unit 35.

[0031] The first AI processing unit 34 is provided in the image sensor 3A as, for example, a DSP (Digital Signal Processor).

[0032] The data transfer unit 35 transfers the first inference result selected based on an instruction from the control unit 33, for example, coordinate information, to the downstream signal processing device 2.

[0033] Data transfer between the image sensor 3A and the signal processing device 2 is performed in accordance with, for example, MIPI (Mobile Industry Processor Interface).

[0034] Transfer data conforming to MIPI includes a packet header, a data body, and a packet footer. The packet header includes a data type field that specifies the data type of each packet.

[0035] MIPI transfer data includes packets that store RAW image data by specifying "RAW10" or the like in the data type field, and packets that store the first inference result by specifying "embedded data" in the data type field.

[0036] The signal processing device 2 includes a receiving unit 21 , a development processing unit 22 , a control unit 23 , a second AI processing unit 24 , and a data transfer unit 25 .

[0037] The receiving unit 21 receives the first inference result and image data conforming to MIPI from the upstream image sensor 3 A. The image data received by the receiving unit 21 from the image sensor 3 A may be, for example, RAW image data or the first input image data G1 input to the first AI model M1.

[0038] The development processing unit 22 is provided as an ISP or the like, and performs development processing on RAW image data received from the receiving unit 21. Note that execution of development processing is not essential when image data to which development processing has been applied is received from the receiving unit 21. In this case, the development processing unit 22 may function as a processing unit that performs image processing on image data other than RAW image data.

[0039] The image data generated by the development processing unit 22 is used for AI image processing by the second AI processing unit 24.

[0040] The control unit 23 changes the parameters used in the development processing performed by the development processing unit 22, for example, based on the first inference result received from the receiving unit 21. For example, if the inference accuracy of the first AI image processing performed in the upstream image sensor 3A is low, the control unit 23 changes the parameters for the development processing in order to improve the inference accuracy of the second AI image processing performed in the second AI processing unit 24 of the signal processing device 2.

[0041] Specifically, when a RAW image is captured in backlight and AI image processing is performed to detect a subject of a specific color, parameters for color correction are changed to restore the color changed by backlight to its original color. Also, when it is estimated that the inference accuracy of the first AI image processing has decreased because the brightness of the entire image is too high, parameters for development processing that affect the brightness of the image are changed.

[0042] The control unit 23 causes the second AI processing unit 24 to perform AI image processing using the second AI model M2. Here, the AI ​​image processing by the second AI processing unit 24 is referred to as "second AI image processing."

[0043] The control unit 23 also receives, as instruction information, information for selecting at least a portion of the second inference results obtained by the second AI image processing from a downstream device, specifically, the downstream device 4. Based on the received instruction information, the control unit 23 causes the second AI processing unit 24 to supply a predetermined second inference result to the data transfer unit 25.

[0044] The second AI processing unit 24 is provided as a DSP or the like, and inputs the second input image data G2 into the second AI model M2 based on instructions from the control unit 23, thereby obtaining a second inference result.

[0045] The second input image data G2 input to the second AI model M2 is image data that has been subjected to various development processes by the development processing unit 22, but may also be cropped image data in which a predetermined area has been cut out based on the coordinate information obtained as the first inference result. The second input image data G2 may also be RAW image data.

[0046] The second AI model M2 identifies a face area and performs face recognition processing based on the second input image data G2, which is a cropped image in which a person is captured (see FIG. 4). Specifically, the second AI model M2 outputs ID (identification) information, likelihood information, and coordinate information for the face area of ​​the person identified as a face recognition result for the input image data.

[0047] 4 shows the second input image data G2 input to the second AI model M2, which is a cropped image cut out from the first input image data G1, with the face region used for face recognition surrounded by a bounding box indicated by a dashed line. Also, FIG. 5 shows an overview of the second inference results obtained when the second AI image processing is performed on each person detected in the first AI image processing.

[0048] The second input image data G2 does not necessarily have to be a cropped image. For example, the first input image data G1 and the second input image data G2 may be the same image.

[0049] Based on an instruction from the control unit 23, the second AI processing unit 24 supplies, for example, only the ID information out of the ID information, likelihood information, and coordinate information to the data transfer unit 25.

[0050] The data transfer unit 25 transfers the second inference result selected based on an instruction from the control unit 23, for example, ID information about the detected person, to the subsequent device 4 at the subsequent stage.

[0051] Data transfer between the signal processing device 2 and the subsequent device 4 is performed in accordance with, for example, MIPI.

[0052] The ID information transferred in the data transfer from the signal processing device 2 to the subsequent device 4 is stored in a packet in which "embedded data" is specified in the data type field and transferred.

[0053] The subsequent device 4 includes at least a control unit 41. The control unit 41 identifies an inference result to be output from the camera device 1A and notifies the image sensor 3A and the signal processing device 2. As a result, a specific first inference result or second inference result is output from the camera device 1A via the image sensor 3A or the signal processing device 2.

[0054] These inference results may be output from the subsequent device 4. In this case, the subsequent device 4 may include a receiving unit 42 and a transmitting unit 43, and the receiving unit 42 may receive the inference result selected from the image sensor 3A or the signal processing device 2, and the transmitting unit 43 may output the inference result to the outside of the camera device 1A.

[0055] When outputting image data from the camera device 1A, the image data is received from the image sensor 3A or the signal processing device 2 by the receiving unit 42, and the image data is output to the outside of the camera device 1A by the transmitting unit 43. Note that the image data received from the image sensor 3A or the signal processing device 2 may be subjected to a predetermined development process, in which case a development processing unit 44 may be provided in the subsequent device 4.

[0056] The receiving unit 42 receives MIPI-compliant packets from the signal processing device 2 .

[0057] When image data is included in the packet and development processing is required, the development processing unit 44 receives the RAW image data and the like from the receiving unit 42 and performs the necessary development processing. Note that the development processing in the development processing unit 44 may use parameters adjusted by the control unit 41 in order to suitably perform the processing content in the subsequent device 4.

[0058] The control unit 41 may perform processing using the inference results received from the image sensor 3A and the signal processing device 2.

[0059] For example, consider a case where the camera device 1A is used as a device that authenticates a photographed person and determines whether to unlock a door based on the authentication result. The control unit 41 of the subsequent device 4 performs a process of matching the detected person with a person permitted to enter the room by referring to a database or the like based on the second inference result obtained from the signal processing device 2, and determines whether to output a command to unlock the door based on the matching result.

[0060] This eliminates the need to output not only image data but also inference results from the camera device 1A to the outside, thereby ensuring privacy protection.

[0061] Furthermore, in the subsequent device 4, AI image processing may be performed using image data obtained by development by the development processing unit 44. For example, an AI processing unit may be provided in the subsequent device 4, and the AI ​​image processing or the like may be performed in the AI ​​processing unit.

[0062] <2. Data Transfer> As described above, data transfer between the image sensor 3A, the signal processing device 2, and the subsequent device 4 is performed using, for example, MIPI-compliant packets.

[0063] A MIPI-compliant packet comprises a packet header PH, a data section DS, and a packet footer PF (see FIG. 6).

[0064] The packet header PH further includes a data ID, a word count (WC), and an ECC (Error Correction Code). The data ID further includes a 2-bit virtual channel VC and a 6-bit data type. In Figure 6, only the data type of the packet header PH is shown. Note that "EBD" in the figure indicates embedded data as a data type.

[0065] 6 can be output from the camera device 1A to an external device. Here, the image sensor 3A transfers only information instructed by the control unit 41 of the subsequent device 4 to the subsequent signal processing device 2. Similarly, the signal processing device 2 outputs only information instructed by the control unit 41 from the subsequent device 4 or the camera device 1A to the external device.

[0066] The image sensor 3A and the signal processing device 2 can change the data to be transferred to the subsequent stage depending on the situation.

[0067] For example, when the camera device 1A is installed, as shown in Fig. 7, the control unit 41 instructs the image sensor 3A and the signal processing device 2 to transfer only information related to the image quality of the captured image. In the example shown in Fig. 7, only the information on the image capture settings and the result information of the analysis of the image after development processing by the control unit 41 are output from the camera device 1A.

[0068] In a device disposed outside the camera device 1A, it becomes easy to use this information to perform optimal imaging settings according to the installation location of the camera device 1A.

[0069] Next, when deploying the AI ​​models after installing the camera device 1A, the control unit 41 instructs the image sensor 3A and the signal processing device 2 to transfer only information about the first AI model M1 and information about the second AI model M2, as shown in Fig. 8. In the example shown in Fig. 8, the control unit 41 outputs only the first input image data G1, the first inference result, the second input image data G2, and the second inference result from the camera device 1A.

[0070] A device located outside the camera device 1A can use this information to generate and adjust an AI model to be deployed to the camera device 1A.

[0071] When the camera device 1A is used, the control unit 41 instructs the signal processing device 2 to transfer only the second inference result, as shown in FIG.

[0072] During operation, a device located outside the camera device 1A can use the second inference result to determine whether to unlock the door, etc. Furthermore, since the image data is not output to a device outside the camera device 1A, privacy can be appropriately protected.

[0073] As described above, the information stored in a MIPI-compliant packet differs depending on the information specified by the control unit 41. Therefore, the size of the data stored in the data section DS of the packet also differs.

[0074] In transferring packets, differences in the data size of the data section DS may be allowed, or the data size may be unified.

[0075] For example, as shown in FIG. 10, even if the packet size differs due to the data size of the data section DS being different, the packet may be transferred to the subsequent stage as is.

[0076] Alternatively, as shown in Fig. 11, when the data size of the data section DS is different, padding data PD may be used to fill the data section DS to a predetermined size. In this case, the size of the transferred packets is unified.

[0077] 12 shows an example of processing executed by the image sensor 3A of the camera device 1A. In step S101, the image sensor unit 31 of the image sensor 3A performs exposure control and readout processing to perform image capture processing. As a result, RAW image data is output as digital data from the image sensor unit 31.

[0078] The development processing unit 32 performs the development processing described above in step S102, thereby obtaining, for example, RGB image data.

[0079] In step S103, the first AI processing unit 34 performs first AI image processing on the developed image data as first input image data G1.

[0080] In step S104, the first AI processing unit 34 determines whether an inference result has been obtained as a result of the first AI image processing, for example, whether the target subject has been detected.

[0081] If the first AI processing unit 34 determines that the target subject has been detected (step S104: Yes), the process proceeds to step S105, where it acquires a first inference result.

[0082] On the other hand, if the first AI processing unit 34 determines that the target subject has not been detected (step S104: No), it avoids the processing of step S105 and proceeds to step S106.

[0083] In step S106, the control unit 33 determines whether information to be transferred from the image sensor 3 A is designated. If it is determined that transfer information is designated (step S106: Yes), the control unit 33 instructs the data transfer unit 35 in step S107 to select only the designated information and store it in a MIPI-compliant packet.

[0084] On the other hand, if it is determined that transfer information has not been specified (step S106: No), the control unit 33 instructs the data transfer unit 35 to select all information that can be output and store it in a MIPI-compliant packet in step S108. Note that if it is determined in step S104 that the first inference result cannot be acquired, information other than the first inference result, such as imaging setting information, is stored in the packet.

[0085] After executing either step S107 or step S108, the data transfer unit 35 performs processing to transfer the generated packet to the signal processing device 2 at the subsequent stage in step S109.

[0086] After completing the process of step S109, the image sensor 3A executes the process of step S101 again. That is, the image sensor 3A executes the series of processes shown in Fig. 13 for each captured frame.

[0087] 13 shows an example of processing executed by the signal processing device 2 of the camera device 1A. In step S201, the receiving unit 21 of the signal processing device 2 acquires a first inference result from the upstream image sensor 3A. Note that in step S201, the receiving unit 21 may acquire information other than the first inference result, such as imaging setting information, and the received information may not include the first inference result.

[0088] In step S202, the control unit 23 performs processing to change the parameters used in the development processing. Note that this processing is not essential and may be performed, for example, when the second input image data G2 for performing highly accurate inference cannot be appropriately generated.

[0089] The development processing unit 22 performs development processing in step S203. Note that, instead of performing development processing on the RAW image data in step S203, the development processing unit 22 may perform image processing on image data other than RAW image data. Furthermore, when RAW image data is input to the second AI model M2 as the second input image data G2, the processing of step S203 is not essential.

[0090] In step S204, the second AI processing unit 24 performs second AI image processing on the image data after the image processing as second input image data G2.

[0091] In step S205, the second AI processing unit 24 determines whether an inference result has been obtained as a result of the second AI image processing, for example, whether the target subject has been detected.

[0092] If the second AI processing unit 24 determines that the target subject has been detected (step S205: Yes), it proceeds to step S206 and acquires a second inference result.

[0093] On the other hand, if the second AI processing unit 24 determines that the target subject could not be detected (step S205: No), it avoids the processing of step S206 and proceeds to step S207.

[0094] In step S207, the control unit 23 determines whether or not information to be transferred has been designated from the signal processing device 2. If it is determined that transfer information has been designated (step S207: Yes), the control unit 23 instructs the data transfer unit 25 in step S208 to select only the designated information and store it in a MIPI-compliant packet.

[0095] On the other hand, if it is determined that no transfer information has been specified (step S207: No), the control unit 23 instructs the data transfer unit 25 in step S209 to select all information that can be output and store it in a MIPI-compliant packet.

[0096] After executing either step S208 or step S209, the data transfer unit 25 performs processing in step S210 to transfer (transmit) the generated packet to the subsequent stage device 4 or a device external to the camera device 1A.

[0097] After completing the process of step S210, the signal processing device 2 executes the process of step S201 again.

[0098] 4. Modified Example In the above description, a camera device 1A equipped with one image sensor 3A has been given as an example of the signal processing system 1. In this modified example, the signal processing system 1 is a camera device 1B equipped with multiple image sensors 3B, 3C, and 3D, as shown in FIG.

[0099] The image sensors 3B, 3C, and 3D have the same configuration as the image sensor 3A. In Fig. 14, the configuration of the image sensors 3C and 3D is omitted.

[0100] The control unit 33 included in the image sensors 3B, 3C, and 3D is supplied with instruction information regarding the transfer data from the control unit 41 of the subsequent device 4.

[0101] The data transfer units 35 of the image sensors 3B, 3C, and 3D store transfer data corresponding to the instruction information in packets conforming to the MIPI standard and transfer the packets to the signal processing device 2.

[0102] The signal processing device 2 receives the transferred data from each of the image sensors 3B, 3C, and 3D at the receiving unit 21, and performs second inference processing at the second AI processing unit 24.

[0103] The signal processing device 2 also selects transfer data from various information generated by the image sensors 3B, 3C, 3D and the signal processing device 2 based on instructions from the control unit 23, and the selected data is transmitted from the data transfer unit 25 to a subsequent subsequent device 4 or a device external to the camera device 1B.

[0104] Note that image sensors 3B, 3C, and 3D may be different types of image sensors. For example, image sensor 3B may be an RGB sensor that generates a color image, image sensor 3C may be a ToF (Time of Flight) sensor that generates a distance image, and image sensor 3D may be a thermosensor that generates a temperature image.

[0105] The information generated by these image sensors 3B, 3C, and 3D, such as various image data and inference results, can be appropriately selected by the signal processing device 2, allowing it to be processed appropriately in subsequent processing.

[0106] In the above example, the type of data to be transferred is selected for each phase of the camera device 1A, such as during installation or operation, but different data may be selected and transferred for each frame.

[0107] For example, various possible uses are possible, such as outputting the inference results of image sensor 3D as a thermal sensor mainly from camera device 1B until a person is detected, and then outputting the inference results of image sensor 3B as a color image sensor from camera device 1B when a subject whose surface temperature is above a predetermined temperature is detected.

[0108] The above example shows the transfer of packet data conforming to MIPI as the transfer data, but it is also possible to transfer data conforming to SPI (Serial Peripheral Interface) or data conforming to a parallel interface.

[0109] 5. Application Examples In the above-described examples, face authentication is performed on a person captured by the camera device 1A (1B), and whether or not to unlock a door is determined based on the result of the face authentication. However, the application of the present technology is not limited to this.

[0110] For example, even in the same face recognition system, a person's posture is detected by AI image processing using the image sensor 3A (3B), and only information about a person who is approaching the camera device 1A for face recognition is transferred to the downstream signal processing device 2. The signal processing device 2 performs face detection processing as AI image processing, and transfers the face detection result information to the downstream subsequent device 4. At this time, the signal processing device 2 may identify how light hits the subject based on the recognition result of the image sensor 3A and optimize parameters for the development processing. The downstream device 4 or an external device to the camera device 1A may further perform AI image processing and face recognition processing to identify the subject. At this time, the downstream device 4 may optimize parameters for the development processing in the same way as the signal processing device 2, and the optimization processing may be performed by receiving parameters from the signal processing device 2.

[0111] In addition to facial recognition systems, a vehicle license plate recognition system may also be considered. Specifically, a vehicle is detected by AI image processing using the image sensor 3A, the signal processing device 2 detects the license plate of the detected vehicle, and the subsequent device 4 or an external device to the camera device 1A performs a process of identifying the vehicle and its owner by verifying the vehicle's license plate information. The process of verifying the vehicle's license plate information may further determine whether the verification result matches a specified vehicle, and if so, output information indicating that the vehicle to be detected has been captured by the camera device 1A as a surveillance camera. As with the facial recognition system described above, parameter optimization processing may be performed for the development process in each device. Furthermore, for example, the parameters of the development process may be changed depending on the body color of the detected vehicle.

[0112] Alternatively, the signal processing system 1 may be a system that reads two-dimensional codes, etc. For example, an image sensor 3A performs processing to detect an object. A signal processing device 2 at a subsequent stage detects a two-dimensional code attached or printed on the detected object. A further subsequent stage 4 performs processing to read the detected two-dimensional code and verifies whether an appropriate two-dimensional code has been assigned to the object. The development processing in each device may be performed by optimizing parameters as appropriate, similar to the face authentication system and license plate recognition system described above.

[0113] Alternatively, a dog may be detected using the image sensor 3A, and if a dog is detected, a person close to the dog may be detected using the downstream signal processing device 2, and if these are detected, the downstream device 4 may execute a process to make an audio announcement, taking into account that the person in question is visually impaired.

[0114] The example given here is merely an example, and various other configurations are possible.

[0115] 6. Summary As described in the various examples above, the signal processing device 2 in the signal processing system 1 includes an inference processing unit (second AI processing unit 24) that performs inference processing (e.g., second AI image processing) using an AI model (second AI model M2), a control unit 23 that receives designation information specifying at least a portion of the inference results (second inference results) obtained as a result of the inference processing and issues an output instruction for the inference results based on the designation information, and an output unit (data transfer unit 25) that outputs the inference results based on the output instruction. This allows for the inference results obtained by the inference processing to be output to a subsequent stage, rather than necessarily being output in their entirety. Therefore, only the inference results designated according to the processing content of the subsequent stage can be output, eliminating the need to analyze the recognition results in the subsequent stage (e.g., subsequent-stage device 4), and facilitating real-time processing. Furthermore, outputting only the necessary inference results reduces the data transfer load and enables support for slow data transfer rates.

[0116] As described with reference to FIG. 1 and other figures, the signal processing device 2 includes a receiving unit 21 that receives the result of a first inference process (e.g., first AI image processing) using a first AI model M1 from a pre-stage device 3 (e.g., image sensor 3A) as a first inference result. The control unit 23 may issue an output instruction based on designation information that specifies predetermined information from the first inference result and a second inference result obtained as an inference result of a second inference process (e.g., second AI image processing), which is an inference process performed by an inference processing unit (second AI processing unit 24). This makes it possible to select and output only the inference results necessary for processing in a subsequent stage (e.g., subsequent stage device 4) from multiple inference results obtained by multiple inference processes. Therefore, even if the processing in the subsequent stage is changed, it is possible to output an appropriate inference result that meets the changed requirements, thereby reducing the processing burden on the subsequent stage. Furthermore, by distributing the inference process across multiple devices, such as performing person detection in the pre-stage device 3 and face detection in the subsequent stage signal processing device 2, it is possible to perform inference processing with high inference accuracy even in low-performance edge-side devices. Furthermore, using multiple devices makes it possible to perform appropriate inference processing that meets the requirements of various edge devices. It should be noted that the "pre-stage device 3" referred to here includes not only an independent unit device but also a microcomputer that performs pre-stage processing, a board on which a predetermined IC (Integrated Circuit) is mounted, etc. In other words, the pre-stage device 3 and the signal processing device 2 may be included in a single housing. In this case, the pre-stage device 3 may be an image sensor 3A, and the signal processing device 2 may be a microcomputer. In other words, the pre-stage device 3 and the signal processing device 2 may be provided within a single device.

[0117] As described with reference to FIGS. 1 to 3 , the receiving unit 21 of the signal processing device 2 receives image data (e.g., RGB image data or RAW image data) along with the first inference result from the preceding device 3 (e.g., the image sensor 3A), and the inference processing unit (second AI processing unit 24) may perform a second inference process (e.g., second AI image processing) using the second AI model M2 based on the image data. The image data received by the signal processing device 2 may be, for example, image data input to the first AI model M1 or RAW image data captured by the preceding device 3. Receiving image data from the preceding device 3 enables the signal processing device 2 to perform inference processing on the image. For example, person detection may be performed in the preceding first AI model M1 and face detection may be performed in the following second AI model M2. Therefore, even an edge-side device with limited processing power can perform advanced inference processing.

[0118] As described with reference to FIG. 1 etc., the signal processing device 2 includes an image processing unit (e.g., development processing unit 22) that performs image processing on image data to generate second input image data G2 to be input to the second AI model M2, and the inference processing unit (second AI processing unit 24) may perform second inference processing by inputting the second input image data G2 to the second AI model M2. This allows the signal processing device 2 to generate an input image to be input to the second AI model M2 anew. Therefore, it is possible to generate second input image data G2 that is appropriate as image data to be input to the second AI model M2, and it is possible to improve the accuracy of the inference processing (e.g., second AI image processing) in the second AI model M2.

[0119] 1 etc., the receiving unit 21 in the signal processing device 2 may receive RAW image data as image data, and the image processing unit (e.g., the development processing unit 22) may perform development processing as image processing to generate the second input image data G2. Receiving the RAW image data by the signal processing device 2 makes it easy for the signal processing device 2 to generate appropriate second input image data G2 to be input to the second AI model M2.

[0120] As described with reference to FIG. 1 and other figures, the control unit 23 of the signal processing device 2 may supply parameters based on the first inference result to the image processing unit (e.g., the development processing unit 22), and the image processing unit may perform image processing based on the parameters. This makes it possible to change the image processing as the development processing in the downstream signal processing device 2 depending on the first inference result of the inference processing (e.g., the first AI image processing) in the upstream device 3 (e.g., the image sensor 3A). For example, the upstream device 3 may receive likelihood information of the first inference result, and if the likelihood information is lower than a predetermined value, the parameters used in the development processing may be changed to suitably perform the second inference processing (e.g., the second AI image processing) by the second AI model M2. Furthermore, if the inference processing in the upstream device 3 fails to detect the subject to be detected, the parameters used in the development processing to generate the second input image data G2 to be input to the downstream second AI model M2 may be changed to obtain an appropriate second inference result in the second AI model M2. Note that an output instruction may be issued to the upstream device 3 so that information related to parameter adjustment is transferred from the upstream device 3 to the signal processing device 2. As a result, for example, information on the brightness of the image and information on how the light hits the subject's face are appropriately transferred from the pre-stage device 3 to the signal processing device 2. This makes it possible to optimize the AI ​​image processing in the signal processing device 2.

[0121] As described with reference to FIG. 14 and other figures, the signal processing device 2 may include a plurality of pre-stage devices 3 (e.g., image sensors 3B, 3C, and 3D), the receiving unit 21 receiving the first inference results for each of the plurality of pre-stage devices 3, and the control unit 23 issuing an output instruction based on the second inference result and designation information indicating predetermined information from the plurality of first inference results. For example, the plurality of pre-stage devices 3 may perform an inference process for detecting a person (first AI image processing) and an inference process for inferring a posture (first AI image processing). Then, the downstream signal processing device 2 may perform facial recognition based on these detection results, for example, by estimating the area where a person is detected and the area where a face is captured based on the posture inference result, and inputting a cropped image of that area into the second AI model M2. In this way, using various inference results can reduce the amount of computation required for the inference process (e.g., second AI image processing) in the downstream signal processing device 2, thereby lowering the specifications required for the signal processing device 2. Furthermore, by configuring the subsequent stage device 4, which is located further downstream of the signal processing device 2, to be able to transmit not only the second inference result but also the first inference result, it becomes possible to select and output a wide range of inference results required for the inference processing of the subsequent stage device 4. Therefore, it becomes possible to appropriately perform the inference processing in the subsequent stage device 4.

[0122] As described with reference to FIG. 1 and other figures, the pre-stage device 3 for the signal processing device 2 may be an image sensor 3A. This allows a first inference process (e.g., first AI image processing) and a second inference process (e.g., second AI image processing) to be performed as inference processes using an image captured by the image sensor. This makes it possible to perform various inference processes without transmitting image data to a stage downstream of the signal processing device 2, which is preferable from the perspective of privacy. In particular, if the signal processing device 2 is a microcomputer provided within the housing of the camera device 1A (1B), it is possible to output only the inference result (the first inference result or the second inference result) without outputting the image outside the camera, thereby making it possible to preferably protect privacy.

[0123] As described with reference to FIG. 1 and other figures, the output unit (data transfer unit 25) of the signal processing device 2 may output the inference results (the first inference result and the second inference result) in a format conforming to MIPI. Utilizing MIPI, which is designed on the assumption that image data will be stored, is suitable for transmitting image data to a downstream device. At that time, by transmitting only the specified inference result, the burden of data transfer can be reduced.

[0124] 10 and other figures, the packet size of the MIPI-compliant format may be variable in data transfer in the signal processing device 2. This eliminates the need to worry about the data volume of the inference results when transmitting the inference results (first inference result and second inference result) as embedded data.

[0125] 11 and the like, the packet size of the MIPI-compliant format may be fixed in data transfer in the signal processing device 2. This allows the receiving side to receive packet data of a fixed size when transmitting the inference result (the first inference result or the second inference result) as EBD, thereby making it possible to improve the efficiency of the receiving process.

[0126] As described with reference to FIG. 1 and other figures, the control unit 23 of the signal processing device 2 may receive designated information from the subsequent device 4, which receives output data from the output unit (data transfer unit 25). It is possible to output only necessary information from the inference results (first inference result or second inference result) in response to instructions from the subsequent device 4. This makes it possible to achieve data transfer that is suitable for the subsequent device 4. Note that the term "subsequent device 4" here includes not only an independent, stand-alone device but also a microcomputer or the like that performs subsequent processing. That is, the subsequent device 4 and the signal processing device 2 may be included in a single housing. For example, the preceding device 3 may be an image sensor 3A, and the signal processing device 2 and subsequent device 4 may each be separate microcomputers. That is, the preceding device 3, signal processing device 2, and subsequent device 4 may be provided within a single camera device 1A (1B).

[0127] The signal processing method of the present technology is a method executed by a computer device, and includes an inference process (e.g., a second AI image process) using an AI model (a second AI model M2), a process of receiving designation information that designates at least a portion of the inference result (a second inference result) obtained as a result of the inference process, a process of issuing an output instruction for the inference result based on the designation information, and a process of outputting the inference result based on the output instruction.

[0128] The program of the present technology is a program executed by a processing device, and includes a function of performing inference processing (e.g., second AI image processing) using an AI model (second AI model M2), a function of receiving specification information that specifies at least a portion of the inference result (second inference result) obtained as a result of the inference processing, a function of issuing an output instruction for the inference result based on the specification information, and a function of outputting the inference result based on the output instruction.

[0129] Such a signal processing method and program can also provide the various functions and effects described above.

[0130] Such programs can be pre-recorded on a hard disk drive (HDD) or a read-only memory (ROM) in a microcomputer having a central processing unit (CPU) as a built-in recording medium in a computer or other device. Alternatively, the programs can be temporarily or permanently stored (recorded) on removable recording media such as a flexible disk, a compact disk read-only memory (CD-ROM), a magneto optical (MO) disk, a digital versatile disc (DVD), a Blu-ray Disc (registered trademark), a magnetic disk, a semiconductor memory, or a memory card. Such removable recording media can be provided as so-called packaged software. In addition to being installed on a personal computer or the like from a removable recording medium, such programs can also be downloaded from a download site via a network such as a local area network (LAN) or the Internet.

[0131] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0132] Furthermore, the above-described examples may be combined in any manner, and even when various combinations are used, the various effects described above can be obtained.

[0133] <7. The Present Technology> The present technology may also have the following configuration. (1) A signal processing device comprising: an inference processing unit that performs inference processing using an AI model; a control unit that receives designation information that designates at least a portion of the inference result obtained as a result of the inference processing and issues an output instruction for the inference result based on the designation information; and an output unit that outputs the inference result based on the output instruction. (2) The signal processing device according to (1) above, comprising: a receiving unit that receives a result of a first inference processing using a first AI model from a preceding device as a first inference result, wherein the control unit issues the output instruction based on the designation information that designates a second inference result as the inference result of a second inference processing that is the inference processing by the inference processing unit and predetermined information from the first inference result. (3) The signal processing device according to (2) above, wherein the receiving unit receives image data together with the first inference result from the preceding device, and the inference processing unit performs the second inference processing using a second AI model based on the image data. (4) The signal processing device according to (3) above, further comprising an image processing unit that generates second input image data to be input to the second AI model by performing image processing on the image data, wherein the inference processing unit performs the second inference processing by inputting the second input image data to the second AI model. (5) The signal processing device according to (4) above, wherein the receiving unit receives RAW image data as the image data, and the image processing unit generates the second input image data by performing development processing as the image processing. (6) The signal processing device according to any of (4) to (5) above, wherein the control unit supplies parameters based on the first inference result to the image processing unit, and the image processing unit performs the image processing based on the parameters. (7) The signal processing device according to any of (2) to (6) above, wherein a plurality of the pre-stage devices are provided, wherein the receiving unit receives the first inference result for each of the plurality of pre-stage devices, and the control unit issues the output instruction based on the second inference result and the designation information that designates predetermined information from the plurality of first inference results. (8) The signal processing device according to any one of (2) to (7), wherein the preceding device is an image sensor.(9) The signal processing device according to any one of (1) to (8), wherein the output unit outputs the inference result in a format conforming to MIPI. (10) The signal processing device according to (9), wherein a packet size of the MIPI-compliant format is variable. (11) The signal processing device according to (9), wherein the packet size of the MIPI-compliant format is fixed. (12) The signal processing device according to any one of (1) to (11), wherein the control unit receives the designation information from a subsequent device that receives output data from the output unit. (13) A signal processing method, wherein a computer device executes the following processes: inference processing using an AI model; receiving designation information that designates at least a part of the inference result obtained as a result of the inference processing; issuing an output instruction for the inference result based on the designation information; and outputting the inference result based on the output instruction. (14) A program that causes a processing device to execute the following functions: a function of performing inference processing using an AI model; a function of receiving designation information that designates at least a portion of the inference result obtained as a result of the inference processing; a function of issuing an output instruction for the inference result based on the designation information; and a function of outputting the inference result based on the output instruction.

[0134] 2 Signal processing device 21 Receiving unit 22 Development processing unit (image processing unit) 23 Control unit 24 Second AI processing unit (inference processing unit) 25 Data transfer unit (output unit) 3 Preceding device 3A Image sensor (preceding device) 3B Image sensor (preceding device) 3C Image sensor (preceding device) 3D Image sensor (preceding device) 4 Subceding device G2 Second input image data M1 First AI model M2 Second AI model (AI model)

Claims

1. A signal processing device comprising: an inference processing unit that performs inference processing using an AI model; a control unit that receives designation information that designates at least a portion of the inference result obtained as a result of the inference processing, and issues an output instruction for the inference result based on the designation information; and an output unit that outputs the inference result based on the output instruction.

2. A signal processing device as described in claim 1, further comprising a receiving unit that receives the result of a first inference process using a first AI model from a preceding device as a first inference result, and the control unit issues the output instruction based on a second inference result as the inference result of a second inference process, which is an inference process by the inference processing unit, and on the designation information that designates predetermined information from among the first inference result.

3. The signal processing device according to claim 2, wherein the receiving unit receives image data together with the first inference result from the pre-stage device, and the inference processing unit performs the second inference processing using a second AI model based on the image data.

4. A signal processing device as described in claim 3, further comprising an image processing unit that generates second input image data to be input to the second AI model by performing image processing on the image data, and the inference processing unit performs the second inference processing by inputting the second input image data to the second AI model.

5. The signal processing device according to claim 4, wherein the receiving unit receives RAW image data as the image data, and the image processing unit generates the second input image data by performing a development process as the image processing.

6. The signal processing device according to claim 4, wherein the control unit supplies parameters based on the first inference result to the image processing unit, and the image processing unit performs the image processing based on the parameters.

7. A signal processing device as described in claim 2, wherein a plurality of the front-stage devices are provided, the receiving unit receives the first inference result from each of the plurality of front-stage devices, and the control unit issues the output instruction based on the second inference result and the designation information indicating predetermined information from among the plurality of first inference results.

8. The signal processing device according to claim 2, wherein the preceding device is an image sensor.

9. The signal processing device according to claim 1, wherein the output unit outputs the inference result in a format conforming to MIPI.

10. The signal processing device according to claim 9, wherein the packet size of the MIPI-compliant format is variable.

11. The signal processing device according to claim 9, wherein the packet size of the MIPI-compliant format is fixed.

12. The signal processing device according to claim 1, wherein the control section receives the designation information from a downstream device that receives output data from the output section.

13. A signal processing method in which a computer device executes the following processes: an inference process using an AI model; a process of receiving designation information that designates at least a portion of the inference result obtained as a result of the inference process; a process of issuing an output instruction for the inference result based on the designation information; and a process of outputting the inference result based on the output instruction.

14. A program that causes a processor to execute the following functions: performing inference processing using an AI model; receiving designation information that designates at least a portion of the inference result obtained as a result of the inference processing; issuing an output instruction for the inference result based on the designation information; and outputting the inference result based on the output instruction.

Citation Information

Patent Citations

  • Inference processing apparatus, imaging apparatus, inference processing method, and program

    JP2022172661A

  • Information processor, method for processing information, information processing program, and information processing system

    JP2023134172A

  • Information processing device, control method, and storage medium

    WO2021240732A1

  • Image sensor, information processing method, and program

    WO2023218936A1