Signal processing device and signal processing method
The signal processing device employs a stacked autoencoder and supervised learning to generate approximate feature equations, reducing input data and computational costs for inference, thereby enhancing processing efficiency.
Patent Information
- Application Number
- JP2022099697
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-08-14
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Increasing the amount of information in input data for inference increases computational costs, which hampers the efficiency of machine-learned learning devices.
A signal processing device utilizing a stacked autoencoder for pre-training and control line association learning to generate an approximate feature equation, followed by supervised learning to reduce the amount of input data for inference processing.
Reduces computational costs by dimensionally compressing data using approximate feature expressions, enabling efficient inference processing.
Smart Images

Figure 2025119072000001_ABST
Abstract
Description
[Technical Field]
[0001] The present technology relates to a signal processing device and a method thereof, and in particular to a technology for making a predetermined inference from input data using a machine-learned learner. [Background technology]
[0002] Various techniques have been proposed for making predetermined inferences from input data using machine-learned learners. For example, in the field of image recognition, a technique has been proposed for recognizing target subjects such as people, animals, vehicles, etc. from input image data using a machine learning device with a neural network such as a convolutional neural network (CNN). In addition to image data, there are also techniques for inferring the melody of a song or the attributes of a speaker (e.g., gender, age, etc.) using audio data as input data.
[0003] As a related prior art, the following Patent Document 1 can be cited. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-66995 Summary of the Invention [Problem to be solved by the invention]
[0005] Increasing the amount of information in the input data for inference makes it easier for the learning device (artificial intelligence model) to capture various features of the input data, thereby improving the accuracy of inference. However, increasing the amount of information in the input data increases the computational costs involved in inference.
[0006] This technology was developed in consideration of these circumstances, and aims to reduce the computational costs for inference by reducing the amount of input data to an inference processing unit that has a machine-learned learning device and performs predetermined inference based on input data. [Means for solving the problem]
[0007] The signal processing device according to the present technology includes an approximate feature equation generation unit that has a first learning device including a stacked autoencoder, and that performs control line association learning on the first learning device after pre-training of the stacked autoencoder using a predetermined type of data as learning input data, thereby obtaining an approximate feature equation in an intermediate layer of the stacked autoencoder that can be regarded as an equation that indicates the features of the predetermined type of data, and an inference processing unit that has a second learning device that has undergone machine learning as supervised learning using the approximate feature equation as learning input data, and that performs predetermined inference using the feature approximation equation as input data. According to the above configuration, inference processing is performed on approximate feature expressions obtained by dimensionally compressing data of a predetermined type using a layered autoencoder.
[0008] Moreover, a signal processing method according to the present technology is a signal processing method including: an approximate feature equation generation step of providing the predetermined type of data as input data to the approximate feature equation generation unit having a first learning device including a stacked autoencoder, in which control line association learning has been performed for the first learning device after pre-training of the stacked autoencoder using data of a predetermined type as learning input data, and obtaining an approximate feature equation that can be regarded as an equation that shows features of the predetermined type of data in an intermediate layer of the stacked autoencoder; and an inference processing step of performing a predetermined inference using the feature approximation formula as input data by an inference processing unit having a second learning device that has undergone machine learning as supervised learning using the approximate feature formula as learning input data. Such a signal processing method also provides the same effect as the signal processing device according to the present technology described above. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a block diagram illustrating an example of the configuration of a signal processing device according to an embodiment of the present technology. [Figure 2] FIG. 1 is an explanatory diagram of an SAE (Stacked Auto-Encoder). [Figure 3] FIG. 10 is an explanatory diagram of an approximate feature formula. [Figure 4] FIG. 10 is an explanatory diagram of a modification of an approximate feature formula. [Figure 5] FIG. 10 is a diagram showing an image of an approximate feature formula as a transformation formula. [Figure 6] This is a diagram to explain a specific learning method for the second AI learning device. [Figure 7] FIG. 1 is an explanatory diagram of description rules for input to SAE and approximate feature expressions output from SAE. [Figure 8] FIG. 10 is a block diagram showing an example of the internal configuration of an image reference processing unit and an image processing unit as a comparative example. [Figure 9] FIG. 10 is a diagram showing an image captured by a tilted camera. [Figure 10] FIG. 1 is a diagram showing the relationship between IMU quaternions and image inputs. [Figure 11] FIG. 10 is a diagram illustrating an example of a lattice point mesh. [Figure 12] FIG. 10 is an explanatory diagram of coordinate transformation of a lattice point mesh. [Figure 13] FIG. 10 is a diagram for explaining the relationship between a segment matrix and a lattice point mesh. [Figure 14] FIG. 10 is an explanatory diagram of a segment search in the embodiment. [Figure 15] FIG. 10 is an explanatory diagram of triangular interpolation for determining reference coordinates for each segment position. [Figure 16] FIG. 10 is a diagram illustrating an example of triangular interpolation. [Figure 17] FIG. 10 is an explanatory diagram of remesh data. [Figure 18] FIG. 10 is an image diagram of determining reference coordinates in pixel position units from remesh data. [Figure 19] 10A and 10B are explanatory diagrams of interpolation processing for obtaining a stabilized output image. [Figure 20] FIG. 10 is a diagram for explaining an example of the internal configuration of a lattice point mesh generating and shaping unit in a comparative example. [Figure 21] 10A and 10B are diagrams for explaining machine learning for realizing geometric modulation processing on approximate surface data. [Figure 22] FIG. 10 is an explanatory diagram of quadrilateral elements of a lattice point mesh. [Figure 23] FIG. 10 is a diagram illustrating an approximate curve of a lattice point mesh when the lattice point mesh is assumed to be one-dimensional. [Figure 24] FIG. 10 is a diagram illustrating an approximate curve of a segment matrix when the segment matrix is assumed to be one-dimensional. [Figure 25] FIG. 10 is a diagram illustrating an example of the configuration of an image reference processing unit as an embodiment using each learning device that has been trained. [Figure 26] FIG. 2 is a block diagram illustrating an example of the internal configuration of an image processing unit according to an embodiment. [Figure 27] 10A to 10C are explanatory diagrams of thinning-out stabilization processing in the embodiment. [Figure 28] FIG. 10 is an explanatory diagram of machine learning of a high-frequency estimation learner. [Figure 29] FIG. 1 is an explanatory diagram of a 1 / 8-dimensional compression encoder. [Figure 30] FIG. 10 is an explanatory diagram of machine learning of a high frequency correction learner. [Figure 31] FIG. 10 is an explanatory diagram of machine learning of a high frequency correction learner when inter-frame rotation information is used. [Figure 32] FIG. 10 is an explanatory diagram illustrating an example of multi-layered high-frequency estimation. [Figure 33] FIG. 10 is an explanatory diagram of a machine learning method for realizing low-frequency correction processing. [Figure 34] FIG. 10 is a diagram illustrating a configuration at the time of inference related to low-frequency correction processing. [Figure 35]10A and 10B are explanatory diagrams of a modified example relating to generation of packed data in which a composite image and high-frequency compressed data are packed; [Figure 36] FIG. 2 is a block diagram showing an example of the internal configuration of an image analysis processing unit. [Figure 37] FIG. 10 is a diagram showing an example of segmentation analysis. [Figure 38] FIG. 10 is an explanatory diagram of distance estimation from segmentation analysis information. [Figure 39] FIG. 1 is an explanatory diagram of a learning method related to segmentation analysis. [Figure 40] FIG. 1 is an explanatory diagram of a learning method related to optical flow analysis. [Figure 41] FIG. 10 is an explanatory diagram of a learning method for estimating an extended depth image. [Figure 42] FIG. 10 is an explanatory diagram of the redefinition of the extended depth analysis. [Figure 43] FIG. 1 is an explanatory diagram of an extended depth analysis learning machine with a universal interface specification. [Figure 44] FIG. 10 is an explanatory diagram of an example of input corresponding to stereo depth analysis. [Figure 45] FIG. 1 is an explanatory diagram of an example of input corresponding to iToF depth analysis. [Figure 46] FIG. 10 is an explanatory diagram of an example of input corresponding to dToF depth analysis. [Figure 47] FIG. 10 is an explanatory diagram of an example of input corresponding to monocular depth analysis. [Figure 48] 10A and 10B are diagrams for explaining SLAM performance when using an extended depth image obtained by an extended depth analysis learning device with a universal interface specification. [Figure 49] FIG. 2 is a block diagram showing an example of the internal configuration of an image recognition processing unit. [Figure 50] FIG. 1 is an explanatory diagram of machine learning of a Delta Pose estimator. [Figure 51] FIG. 10 is an explanatory diagram of a modified Delta Pose estimator. [Figure 52]FIG. 10 is an explanatory diagram of machine learning of a modified Delta Pose estimator. [Figure 53] 10A and 10B are explanatory diagrams of experimental results relating to the estimation accuracy of an image-derived quaternion and an image-derived translational velocity by a Delta Pose estimator as a modified example. [Figure 54] 10A and 10B are diagrams illustrating experimental results regarding tracking performance using a modified Delta Pose estimator. [Figure 55] FIG. 10 is an explanatory diagram of an evaluation result by RMSE for the translation amount estimated using a modified Delta Pose estimator. [Figure 56] FIG. 2 is a block diagram showing an example of the internal configuration of an image synchronization processing unit. [Figure 57] FIG. 10 is an explanatory diagram of a learning method for a delay analysis layer. [Figure 58] FIG. 10 is an explanatory diagram of a learning method for the score analysis layer. [Figure 59] FIG. 10 is a diagram showing experimental results (in the case of +1 frame delay) relating to the accuracy of delay / advance determination by a Delta Phase estimator. [Figure 60] FIG. 10 is a diagram showing experimental results (in the case of −1 frame delay) relating to the accuracy of delay / advance determination by the Delta Phase estimator. [Figure 61] FIG. 2 is a block diagram showing an example of the internal configuration of a posture control processing unit. [Figure 62] FIG. 10 is an explanatory diagram of a learning method of a centrifugal force removal correction learning device. [Figure 63] FIG. 10 is an explanatory diagram of a learning method for a state machine correction learning device. [Figure 64] FIG. 10 is an explanatory diagram of a learning method for an effect correction learning device. [Figure 65] FIG. 10 is an explanatory diagram of a learning method of a stabilizer braking correction learning device. [Figure 66] FIG. 10 is an explanatory diagram of a learning method for a rotation prediction learning device. [Figure 67] 10A and 10B are diagrams showing experimental results relating to performance evaluation of a rotation prediction learning device. [Figure 68] FIG. 1 is an explanatory diagram illustrating an example of a configuration for realizing fusion SLAM to which a Delta Pose estimator and a Delta Phase estimator are applied. [Figure 69] FIG. 1 is an explanatory diagram of an example of a signal processing device that uses an integrated sensor specialized for AR. [Figure 70] FIG. 10 is a diagram for explaining a configuration example of a stacked logic unit and an AP layer to be provided in the subsequent stage of a small AR sensor. [Figure 71] FIG. 1 is an explanatory diagram of memory control of a rotary memory to which elements of garbage collection are applied. [Figure 72] FIG. 10 is a diagram for explaining an example of a machine learning technique when the generation process of a memory lookup table and a memory release table is realized as AI processing. [Figure 73] FIG. 10 is a diagram showing an example of a configuration for realizing memory control using a trained memory lookup table generation learning device and a memory release table generation learning device. [Figure 74] FIG. 10 is an explanatory diagram of an analysis method for an approximate feature expression. [Figure 75] FIG. 10 is a diagram illustrating an analysis method for an approximate feature expression. [Figure 76] FIG. 10 is a diagram for explaining combination elements including a blend ratio in a DNN estimator. [Figure 77] FIG. 10 is a diagram for explaining a method for searching for an optimal combinatorial solution by game mining. [Figure 78] FIG. 10 is an explanatory diagram of inference using a combination as an optimal solution. [Figure 79] FIG. 10 is an explanatory diagram of the effect of game mining (the effect in the first scene). [Figure 80] FIG. 10 is an explanatory diagram of the effect of game mining (effect in the second scene). DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present technology will be described in the following order with reference to the accompanying drawings. <1. Overall configuration of signal processing device> <2. Approximate feature formula> <3. Image reference processing section and image processing section (stabilization processing and image correction processing)> [3-1. Configuration of Image Reference Processing Unit and Image Processing Unit as Comparative Example] [3-2. Stabilization processing methods] [3-3. Image reference processing unit as an embodiment] [3-4. Image Processing Unit as an Embodiment] [3-5. Modifications Related to Image Processing Unit] <4. Image analysis processing unit> [4-1. Configuration and processing of image analysis processing unit] [4-2. Modifications related to extended depth analysis] <5. Image Recognition Processing Section (Delta Pose Estimation)> [5-1. Configuration and processing of image recognition processing unit] [5-2. Modified Example of Image Recognition Processing Unit] <6. Image synchronization processing section (Delta Phase estimation)> <7. About the attitude control processing unit> <8. Variations> [8-1. Fusion SLAM using Delta Pose and Delta Phase estimators] [8-2. Memory control variations in stabilization processing] [8-3. Approximate feature equation analysis method] [8-4. Game Mining] [8-5.Other] <9. Summary of embodiments> <10. This Technology>
[0011] <1. Overall configuration of signal processing device> 1 is a block diagram showing an example configuration of a signal processing device 1 according to an embodiment of the present technology. Here, the signal processing device 1 is illustrated as being applied to a portable computer device such as a smartphone, a tablet terminal, or a notebook PC (personal computer).
[0012] As shown in the figure, the signal processing device 1 includes an IMU (Inertial Measurement Unit) 2, an image sensor 3, a ToF (Time of Flight) sensor 4, an image synchronization processing unit 5, an attitude control processing unit 6, an image reference processing unit 7, an image processing unit 8, an image analysis processing unit 9, an image recognition processing unit 10, and an AR (Augmented Reality) processing unit 11.
[0013] In the signal processing device 1 of the embodiment, the AR processing unit 11 performs self-position estimation using SLAM (Simultaneous Localization and Mapping) and display processing of virtual images to realize AR content, based on IMU information (acceleration and angular velocity information) obtained based on the detection signal of the IMU 2, the captured image obtained by the image sensor 3, and the depth image obtained by the ToF sensor 4.
[0014] Furthermore, the signal processing device 1 performs signal processing for electronic image stabilizer (EIS) on the image captured by the image sensor 3. In this specification, the processing for electronic image stabilizer is referred to as "stabilization processing."
[0015] The IMU 2 has motion sensors that detect the motion of the signal processing device 1, specifically, an acceleration sensor 2a and an angular velocity sensor 2b in this example. In this example, the acceleration sensor 2a is configured as a triaxial acceleration sensor that can detect accelerations acting on the X-axis, Y-axis, and Z-axis that define a three-dimensional space. The angular velocity sensor 2b is configured as a triaxial angular velocity sensor that can detect angular velocities based on the X-axis, Y-axis, and Z-axis. The information detected by the IMU 2, that is, the information detected by the acceleration sensor 2a and the angular velocity sensor 2b in this example, is supplied to an image synchronization processing unit 5 and an image analysis processing unit 9.
[0016] The image sensor 3 is, for example, a CMOS (Complementary Metal Oxide Semiconductor) or CCD (Charge Coupled Device) image sensor, and includes a pixel array in which a plurality of pixels each having a light receiving element are arranged two-dimensionally, and obtains a captured image by photoelectrically converting light received by the light receiving element in each pixel. In this example, the image sensor 3 is configured to capture a color image, and obtains an RGB (red, green, blue) image as the captured image. The captured image (RGB image) obtained by the image sensor 3 is supplied to an image processing unit 8.
[0017] The ToF sensor 4 has a pixel array section in which multiple pixels each having a light receiving element are arranged two-dimensionally, and by performing distance measurement using the ToF method for each pixel, it obtains a depth image (distance image) which is an image that shows the distance to the subject for each pixel.
[0018] Here, in this specification, "image" broadly means information in which values are indicated for multiple pixels each having a fixed position on a two-dimensional plane, and is a concept that broadly includes, for example, gradation images in which the brightness value of each pixel is indicated, such as the RGB image described above, as well as the depth image described above, polarization images in which a value representing the polarization state of incident light is indicated for each pixel, and thermal images in which a value representing temperature is indicated for each pixel.
[0019] The image reference processing unit 7 and the image processing unit 8 perform processing related to the stabilization processing. Specifically, the image reference processing unit 7 performs processing to calculate reference coordinates CR used for rendering a stabilized image (i.e., a shake-corrected image) based on the attitude-controlled IMU information obtained by the attitude control processing unit 6, that is, the reference coordinates CR indicating the value at which coordinate position in the input coordinate system a pixel position in the output coordinate system should refer to.
[0020] The image processing unit 8 performs stabilization processing to obtain a stabilized image based on the input image (an RGB image in this example) from the image sensor 3 and the reference coordinates CR calculated by the image reference processing unit 7. The image processing unit 8 also performs predetermined correction processing, such as noise reduction, on the stabilized image. The stabilized image obtained by the image processing unit 8 after the correction processing is supplied to the AR processing unit 11.
[0021] In this embodiment, the calculation process of the reference coordinates CR by the image reference processing unit 7 and the correction process of the stabilized image by the image processing unit 8 are performed using AI (Artificial Intelligence), the details of which will be explained later.
[0022] The image analysis processing unit 9 performs processing to generate an extended depth image based on the depth image obtained by the ToF sensor 4, the thinned image obtained by the thinning stabilizer processing unit 47 described later in the image processing unit 8, and the detection information of the IMU 2. The term "extended depth image" as used herein refers to an image in which the dynamic range of distance has been extended from the depth image obtained by the ToF sensor 4. The dynamic range of distance corresponds to the range of distance measurable by a distance measuring sensor. For example, the term "extended" here refers to changing a distance image of a distance measuring sensor whose measurable distance range is 50 cm to 5 m into a distance image of a distance measuring sensor whose measurable distance range is 50 cm to 10 m. In this case, the dynamic range of distance can be expressed as being extended from "a range of 50 cm to 5 m" to "a range of 50 cm to 10 m." The ToF method receives reflected light projected onto a subject and calculates distance based on the time difference between emitting and receiving the projected light, so the range of distance measurement is limited to the range that the projected light reaches, and the range of distance measurement is relatively short due to restrictions on increasing the amount of projected light for safety reasons, etc. For this reason, an extended depth image is generated from the depth image captured by the ToF sensor 4, enabling distance information to be obtained for subjects that are further away.
[0023] In the image analysis processing unit 9, the generation of the extended depth image is performed using various types of AI, the details of which will be explained later. The extended depth image obtained by the image analysis processing unit 9 is supplied to the AR processing unit 11.
[0024] The image recognition processing unit 10 performs a Delta Pose estimation process based on a segment approximation feature equation obtained by a segmentation analysis learning device 61a (described later) in the image analysis processing unit 9 and detection information by the IMU 2, to generate an image-derived quaternion. The image-derived quaternion is a quaternion that is estimated from the movement of a subject in an image and represents the posture of the signal processing device 1. The Delta Pose estimation process by the image recognition processing unit 10 is also performed using AI, but details will be explained later.
[0025] The image synchronization processing unit 5 performs a Delta Phase estimation process based on the image-derived quaternion obtained by the image recognition processing unit 10 and the detection information by the IMU 2, and synchronizes the IMU information with the image to obtain image-synchronized IMU information. Normally, it is desirable for the IMU to be synchronized with the image, but there are few hardware configurations that guarantee synchronization between the IMU and the image in the architecture of modern smartphones, etc. For this reason, in this embodiment, processing is performed to synchronize the IMU information with the image. The delta phase estimation process by the image synchronization processing unit 5 is also performed using AI, but details will be explained later.
[0026] The attitude control processor 6 performs various processes for attitude control on the image-synchronized IMU information obtained by the image synchronization processor 5. For example, the image-synchronized IMU information is processed to remove centrifugal force components and to correct effects. In this embodiment, the processing of the attitude control processing unit 6 is also performed using AI, but details will be explained again later.
[0027] The attitude-controlled IMU information obtained by the attitude control processing unit 6 is supplied to the AR processing unit 11. The attitude-controlled IMU information obtained by the attitude control processing unit 6 is also supplied to the image reference processing unit 7, and is used in the calculation process of the reference coordinate CR described above.
[0028] The AR processing unit 11 performs self-position estimation of the signal processing device 1 using SLAM and display processing of virtual images to realize AR content, based on the attitude-controlled IMU information supplied from the attitude control processing unit 6, the stabilized RGB image supplied from the image processing unit 8, and the extended depth image supplied from the image analysis processing unit 9.
[0029] <2. Approximate feature formula> In this embodiment, the AI machine learning machine uses a machine learning machine having a network structure of a DNN (Deep Neural Network), such as an SAE (Stacked Auto-Encoder). In the following explanation, various blocks will appear regarding DNN, but detailed descriptions of specific embodiments such as detailed input / output configurations and firing functions will be omitted because they are not the essence of this technology. The specific number of layers and number of inputs of the DNN exemplified below are merely examples for the purpose of explanation and are not limited to these configurations.
[0030] FIG. 2 is an explanatory diagram of SAE. As shown in FIG. 2, the SAE has a part that performs convolution processing and pooling, and a fully connected layer at the subsequent stage. SAE undergoes a pre-training process. This pre-training process is a type of unsupervised learning (also called semi-supervised learning) that trains the model so that the output matches the input. Then, supervised learning (called fine tuning) is performed in the subsequent fully connected layer, enabling the generation of a recognition algorithm.
[0031] Here, pre-training is generally intended to reduce the dimensions, but by performing machine learning to match input and output, it has the function of self-teaching and learning the target feature representation. In this specification, the information indicating the features of the target, which is obtained in the intermediate layer of the SAE that has undergone such pre-training and has learned the feature representation of the target, is referred to as an "approximate feature formula."
[0032] FIG. 3 is an explanatory diagram of the approximate feature formula. Figure 3A shows an example of encoding and decoding processes for an approximate feature equation in a pre-trained SAE. Here, 16 pieces of data are used as an example of input data as discretized digital data. Specifically, the data indicates the y value for each of 16 positions in the x coordinate. This discretized data is input to the 16-tap input in SAE, and as pre-training, it is appropriately compressed in the intermediate layer, and machine learning is performed to restore the 16-tap input in the output layer. At this time, the features of the input data are in a state where they have been dimensionally compressed in the intermediate layer, and it can be interpreted that an approximate equation has been generated in the intermediate layer. However, it is difficult to imagine that an approximate equation is generated by SAE from this explanation alone.
[0033] Therefore, a modified form of control line association is used as shown in FIG. 3B. After the pre-training in Figure 3A, a control line is set as shown in Figure 3B, and supervised learning (Fine Tuning) is performed to output one of 16 input Taps as output O according to the address number of the x coordinate given as the control line value C. In this supervised learning, an integer value is used to select one of 16 values for the control line value C.
[0034] By performing this type of control line association learning, it becomes possible to output, as output O, any data input to any input 16Tap that corresponds to the address specified by the control line value C. Furthermore, in the previous Figure 3A, discretized digital data is output on the output side as well as the input data, but by performing control line association learning as in Figure 3B, even if any decimal point type within the range of 1 to 16 is taken for the control line value C, intermediate data interpolated from the input I data can be obtained as output O, and smooth output data can be obtained.
[0035] In this way, after pre-training, in the SAE that has undergone control line association learning, it can be interpreted that a mathematical formula (approximation formula) is formed based on the input I to determine the value of the output O corresponding to the control line value C. The key points here are that SAE converts input data into an approximate feature equation through the interaction of some of the intermediate and output layers, and that control line association learning outputs a smooth output O for any value C within the range of the value C used in learning, and also provides a kind of interpolation function.
[0036] Strictly speaking, the intermediate layer alone does not form a complete mathematical formula, but rather a mathematical formula is formed that includes part of the synaptic data on the output layer side. However, since the autoencoder stacked at the subsequent stage plays the same role in the processing on the output layer side, we will not go into detail about this, and in this case, for the sake of convenience, we will treat the data encoded and transformed in the intermediate layer of SAE as an approximate feature formula.
[0037] In this embodiment, the approximate feature equation obtained as described above is used as input data, and the AI at the subsequent stage performs inference processing. As described above, the approximate feature formula is obtained by compressing the dimension of the input data in the intermediate layer of SAE, and by configuring an AI to perform inference using the approximate feature formula as input data, it is possible to reduce the amount of data input to the AI. Furthermore, by reducing the amount of input data, it is possible to reduce the computational cost of inference processing by the AI.
[0038] In this embodiment, a DNN learning device such as an autoencoder is also used for the subsequent AI. Specifically, a learning device that performs supervised learning and simultaneously transforms and decodes mathematical formulas by connecting autoencoders with fully connected layers is used. Furthermore, in this embodiment, particularly in the posture control processing unit 6 and the image reference processing unit 7, the approximate feature equation obtained by the SAE of the learning device as the AI of the subsequent stage is used as input data of the AI of the subsequent stage (see Figures 25, 61, etc.). For the sake of explanation, the "later AI" will be referred to as the "second AI" and the "further later AI" will be referred to as the "third AI."
[0039] At this time, the second AI undergoes supervised learning (Fine Tuning) so that it can perform predetermined inferences, and therefore the approximate feature equation obtained by the SAE of the second AI is transformed from the input approximate feature equation.
[0040] The modification of the approximate feature formula will be described with reference to FIG. In Figure 4, the SAE of the learning device that provides the approximate feature equation to the second AI as input data in the stage before the second AI is denoted as the first SAE, and the SAE that the second AI has is denoted as the second SAE. As an illustrative example, the inference of the second AI is assumed to be the inference resulting from rolling the input data by a predetermined angle. In this case, the machine learning involves first pre-training the second SAE, and then performing supervised learning as fine tuning using the results of the roll rotation by the predetermined angle as training data. The training data in this case can be obtained by rolling the discretized input data using rule-based processing.
[0041] By carrying out the above-described machine learning, an algorithm can be obtained in the learning device as the second AI that rotates the approximate feature formula input from the first SAE and decodes and outputs it as discretized data. This can be interpreted as the approximate feature formula being obtained as a transformed formula, so to speak, by transforming (rotating in this case) the approximate feature formula input from the first SAE in the intermediate layer of the second SAE. FIG. 5 shows an image of the approximate characteristic equation as a transformation equation obtained in the intermediate layer of the second SAE.
[0042] Referring to Figure 6, a specific learning method for the second AI learning device will be explained. For the second AI, machine learning involves control line association learning. Specifically, assuming that the input data is data indicating the y value for each of 16 x coordinate addresses, the x coordinate address is specified by the control line value, and supervised learning is performed using correct answer data corresponding to the control line value as training data. Specifically, the training data used in this case is correct answer data obtained when the y value at the x address specified by the control line is rotated by a predetermined angle. During inference, as shown in the lower part of the figure, by specifying the x-coordinate address on the control line, the y value at the specified x-coordinate address can be decoded and output as the output. In this case, inference is also based on an approximate feature equation, so even if the input to the control line is a decimal input, an appropriate output can be obtained.
[0043] When a third AI or later AIs are connected in tandem to the second AI or later, the learning units of those AIs also undergo SAE pre-training and control line association learning so that they are capable of predetermined inference, just like the learning unit of the second AI. As a result, the intermediate layer of the SAE of each AI from the third AI onwards obtains approximate feature equations that are further modified from the approximate feature equations that are input as modified equations.
[0044] In the following description, the description rules for the SAE are defined as shown in Figure 7. That is, input data to the SAE is represented by an arrow pointing to the left side of the learning device block containing the SAE. Also, the output from the hidden layer of the SAE is represented by a dotted line pointing to the bottom side of the learning device block containing the SAE.
[0045] This technology is effective in applications where relatively low resolution is not required and input data can be expressed by an approximate formula, and in the signal processing device 1 of the embodiment, the AI configuration that performs inference using the approximate feature formula as input as described above employs the same technique of transforming the feature formula in each block of the image synchronization processing unit 5, attitude control processing unit 6, image reference processing unit 7, image processing unit 8, image analysis processing unit 9, and image recognition processing unit 10, thereby achieving a significant reduction in the computational cost of inference processing using AI. Here, only the image processing unit 8 has a large number of high-frequency components that are difficult to express with an approximate formula, but this problem is solved by using it within a limited range of the number of pixels.
[0046] <3. Image reference processing section and image processing section (stabilization processing and image correction processing)> First, the image reference processing unit 7 and the image processing unit 8 will be described. The image reference processing unit 7 and the image processing unit 8 mainly perform processing to realize a stabilization function for the image captured by the image sensor 3. The image processing unit 8 also has a function to perform predetermined correction processing, specifically, in this example, correction processing such as 3DNR (three-dimensional noise reduction) processing and flicker correction processing, on the stabilized image, which is a captured image that has undergone stabilization processing.
[0047] [3-1. Configuration of Image Reference Processing Unit and Image Processing Unit as Comparative Example] Here, the image reference processing unit 7 performs AI processing using approximate feature equations for various processes to obtain the reference coordinates CR, and in this example, the results of rule-based processing of the corresponding processes are used as training data in machine learning to realize various AI processes using these approximate feature equations. Therefore, before describing the image reference processing unit 7, we will describe an image reference processing unit 7' as a comparative example, which performs rule-based processing to realize each function of the image reference processing unit 7.
[0048] Furthermore, the image processing unit 8 employs a configuration for reducing the computational costs for the process for obtaining a stabilized image based on the reference coordinates CR, and also employs a configuration for similarly reducing the computational costs for the correction process for the stabilized image. However, to facilitate understanding of the effects of these computational cost reductions, the following description will be preceded by a description of the image processing unit 8, and will then be given of an image processing unit 8' as a comparative example.
[0049] FIG. 8 is a block diagram showing an example of the internal configuration of the image reference processing unit 7' and the image processing unit 8'. In the following description, the coordinate system of the input image to be subjected to stabilization processing, i.e., the coordinate system of the captured image obtained by the image sensor 3 in this example, will be referred to as the "input coordinate system." Additionally, the coordinate system of the output image of the stabilization processing, i.e., the stabilized output image, will be referred to as the "output coordinate system."
[0050] In the stabilization process, electronic image stabilization involves cutting out a portion of the input image to obtain a stabilized output image, so it is assumed that the number of pixels in the input image is greater than the number of pixels in the output image. Specifically, in this example, the input image is a 4k image (number of horizontal pixels = approximately 4000, number of vertical pixels = approximately 2000), and the output image is a 2k image (number of horizontal pixels = approximately 2000, number of vertical pixels = approximately 1000).
[0051] 8, the image reference processing unit 7′ includes a lattice point mesh generating / shaping unit 21′, a segment matrix generating unit 22, a segment searching unit 23, a remesh data generating unit 24, and a pixel coordinate interpolating unit 25. By including these blocks, the image reference processing unit 7′ generates reference coordinates CR based on IMU information based on the detection signal of the IMU 2, specifically, the quaternions of acceleration and angular velocity in this example. The reference coordinates CR are information indicating which position value in the input coordinate system should be used as the value of each pixel position in the output coordinate system when cutting out an output image from an input image. In other words, they are information indicating which position value in the input coordinate system should be referenced for each pixel position in the output coordinate system.
[0052] The specific processing of each of the above blocks in the image reference processing section 7' will be explained later.
[0053] The image processing unit 8' includes a memory control unit 41', a buffer memory 42, a cache memory 43, a stabilizer interpolation processing unit 44, and a correction processing unit 45'. The buffer memory 42 is a memory that sequentially buffers one frame of input image, and the memory control unit 41 ′ controls writing and reading of image data to and from the buffer memory 42 . The cache memory 43 is a memory used to cut out an output image from an input image, and the memory control unit 41′ controls writing of image data to the cache memory 43. The memory control unit 41′ obtains image data corresponding to the cut-out range from the image data buffered in the buffer memory 42, and writes the image data to the cache memory 43.
[0054] Furthermore, based on the reference coordinates CR input from the image reference processing unit 7', the memory control unit 41' performs processing to release (make overwritable) the memory area for data that has been determined to be unnecessary for generating a stabilized output image, from the image data of the current frame stored in the buffer memory 42. Then, the memory control unit 41' starts writing the data of the input image of the next frame into the released memory area before the start of the stabilization processing of the input image of the next frame. This makes it possible to ensure that the image data of the next frame is already buffered in the buffer memory 42 before the start of the stabilization processing of the next frame.
[0055] The stabilizer interpolation processing unit 44 reads out image data for multiple pixels (for example, image data for 4 x 4 = 16 pixels in the case of Lanczos2 interpolation) including the pixel in the input coordinate system indicated by the reference coordinate CR and its surrounding pixels for each pixel position in the output coordinate system from the image data (image data of the input image) cached in the cache memory 43, and performs interpolation processing for each pixel position in the output coordinate system using a method described below based on the image data for multiple pixels read out, to determine the value of each pixel position in the output coordinate system. This results in a stabilized output image.
[0056] The correction processing unit 45' performs predetermined correction processing, specifically, in this example, correction processing such as 3DNR processing and flicker correction processing, on the stabilized output image obtained by the stabilizer interpolation processing unit 44.
[0057] [3-2. Stabilization processing methods] The stabilization processing method employed in this embodiment will be described with reference to FIGS. In the stabilization process, the influence of the tilt and movement of the camera (device equipped with the image sensor 3) is removed from the captured image.
[0058] Figure 9 shows the image captured by the tilted camera. The tilted state here means that the camera is tilted in the roll direction and the horizontal and vertical directions are not maintained. In this case, the image data obtained by capturing the image shows the subject tilted, as shown in Figure 9B. By applying stabilization processing to such image data and rotating the image in the same direction as the tilt of the camera, the image data shown in Fig. 9C can be obtained. This image data in Fig. 9C is the same as an image captured with a camera in a straight position (with no tilt in the roll direction), as shown in Fig. 9D. Rotation is performed in this way to compensate for the tilt of the orientation, and this is achieved by rotating the pixel range to be cut out based on the orientation information when cutting out an output image smaller than the input image size from the input image.
[0059] The relationship between the IMU quaternion and image input is shown in Figure 10. When capturing images while moving the camera, the IMU quaternion will change even during one frame. If IMU data is acquired every several lines, for example, IMU quaternions (represented by r0, r1, r2, and r3 in the figure) are also acquired every few lines, as shown. Here, four IMU quaternions are acquired during one frame period indicated by the vertical synchronization signal Vsync, but this is merely an example for explanatory purposes. In this case, IMU quaternion r0 corresponds to the image in the upper 1 / 4 of the frame, IMU quaternion r1 to the next 1 / 4 of the image, IMU quaternion r2 to the next 1 / 4 of the image, and IMU quaternion r3 to the final 1 / 4 of the image. Here, the "virtual line L1" in the figure indicates a virtual line corresponding to an IMU quaternion of the same value.
[0060] Conventionally, on the premise that IMU data is acquired multiple times within one frame period as described above, multiple virtual lines L1 are assumed, each corresponding to the same IMU quaternion value, and reference coordinates CR are fitted to each pixel position in the output coordinate system according to these virtual lines L1, and the input image is cut out based on the fitted reference coordinates CR to obtain a stabilized output image.
[0061] However, it has been found that sufficient stabilization performance cannot be obtained with such stabilization processing using the virtual line L1.
[0062] Therefore, in this embodiment, a method using a lattice point mesh as shown in FIG. 11 is adopted. The grid point mesh has a plurality of grid points (represented by triangle marks in the drawing) arranged in both the horizontal and vertical directions. In a lattice point mesh, a plurality of lattice point rows, each consisting of a plurality of lattice points arranged horizontally, are arranged vertically. Alternatively, this can be said as a plurality of lattice point columns, each consisting of a plurality of lattice points arranged vertically, are arranged horizontally. In the grid point mesh, each grid point row corresponds to the virtual line L1 shown in Fig. 10, and each grid point row is associated with an IMU quaternion based on IMU data acquired at a timing corresponding to the row position. In other words, the IMU quaternion value associated with each grid point is the same for each grid point row.
[0063] Note that the figure shows an example in which the number of lattice points in each lattice point row of the lattice point mesh is 6, i.e., the number of divisions in the horizontal direction is 5, and the number of lattice points in each lattice point column is 5, i.e., the number of divisions in the vertical direction is 4, but the number of divisions in the horizontal and vertical directions of the lattice point mesh are not limited to these values.
[0064] The position of each grid point in the grid point mesh is managed as a position in the input coordinate system to correspond to the timing of IMU data acquisition. The image reference processing unit 7' converts the positions of the grid points in the input coordinate system into positions in the output coordinate system.
[0065] FIG. 12 is an explanatory diagram of coordinate transformation of a lattice point mesh. To convert the positions of grid points to positions in the output coordinate system, we simply apply the same changes to the grid point mesh as to the changes that the input image undergoes. Specifically, as shown in Figure 12, first, the grid point mesh undergoes lens distortion removal processing in response to the lens distortion removal processing that is applied to the input image, and then it is rotated to the same orientation as the camera. This is the result of the conversion to the output coordinate system.
[0066] The stabilization process of this example uses a lattice point mesh converted into the output coordinate system as described above and a segment matrix as shown in FIG. 13A. The segment matrix represents the position of each segment (indicated by a ● mark in the figure) when the output image (the image frame of the output image obtained by stabilization processing) is divided into a predetermined number of segments. In this example, the size of one segment is assumed to be, for example, 64 pixels x 64 pixels.
[0067] FIG. 13B shows the lattice point mesh and the segment matrix that have been coordinate-transformed into the output coordinate system, superimposed on each other in the output coordinate system. The size of the lattice point mesh is larger than the size of the segment matrix because, as described above, the size of the input image is larger than the size of the output image. By converting the lattice point mesh into the output coordinate system, it becomes possible to identify the positional relationship between the position of each segment (● mark) in the segment matrix and each lattice point in the lattice point mesh, as shown in the figure.
[0068] The image reference processing unit 7' determines the reference coordinates CR for each segment based on the positional relationship between each segment and the grid point in the output coordinate system. For this purpose, first, the image reference processing unit 7' performs a segment search as shown in FIG. The segment search is a process for determining in which square of the lattice point mesh the segment position indicated by the ● mark is located for each segment that constitutes the segment matrix. Specifically, the image reference processing unit 7' identifies the segment position contained within each square in the lattice point mesh by inside / outside determination. This inside / outside determination identifies which square in the lattice point mesh each segment position is located within. The reference coordinate CR at each segment position can be calculated based on the IMU quaternions at each of the four grid points that contain that segment position. In the following explanation, it is assumed that each grid point in the grid point mesh is associated with reference coordinate CR information calculated from the corresponding IMU quaternion. Hereinafter, the reference coordinate CR associated with each grid point in this way will be referred to as the "reference coordinate of each grid point."
[0069] The image reference processing unit 7' determines in which square of the lattice point mesh each segment position is located by inside / outside determination (segment search), and then calculates the reference coordinate CR for each segment position by triangular interpolation as shown in Figure 15. Specifically, this triangular interpolation uses the coordinates of the segment position, the coordinates of three of the four grid points of the square that contains the segment position in the grid point mesh, and information on the grid point reference coordinates associated with those grid points. This triangular interpolation may be performed, for example, in a manner as shown in FIG.
[0070] By calculating the reference coordinate CR of each segment position using triangular interpolation, remesh data such as that shown in Figure 17 can be obtained. This remesh data is data that indicates the reference coordinate CR of each position at the segment granularity in the output coordinate system. In the figure, the reference coordinate CR of each position at the segment granularity, i.e., the reference coordinate CR calculated for each segment position, is represented by a ◆ mark.
[0071] The image reference processing unit 7' calculates the reference coordinates CR for each pixel position in the output image based on the remesh data described above.
[0072] FIG. 18 is an image diagram of determining the reference coordinates CR in pixel position units from the remesh data, and in the figure, the reference coordinates CR in pixel position units are represented by ■ marks. In this example, the reference coordinates CR are calculated by linear interpolation (bilinear interpolation) using remesh data (reference coordinates CR at the segment granularity). Specifically, the reference coordinates CR are calculated by bilinear interpolation using the reference coordinates CR of each of the four corner points of the segment that contains the target pixel position. The reason why triangular interpolation is not used in this case is that bilinear interpolation is lighter than triangular interpolation, and once the data is converted into remesh data, sufficient accuracy can be obtained with bilinear interpolation. However, if triangular interpolation is implemented as a hardware circuit within an LSI (Large Scale Integrated circuit), it is more advantageous from the perspective of circuit size to use this block to triangularly interpolate all pixels than to provide a separate bilinear interpolation circuit.
[0073] By determining the reference coordinate CR for each pixel position of the output image, it is possible to identify which position in the input coordinate system should be referenced for each pixel position. However, since the reference coordinate CR is calculated by interpolation processing based on the remesh data as described above, it may be a value that includes decimals rather than an integer unit (i.e., a pixel unit in the input image). For this reason, when rendering the output image based on the reference coordinate CR, an interpolation processing is performed using image data for multiple pixels, including the pixel in the input coordinate system indicated by the reference coordinate CR and its surrounding pixels, as described above as the stabilizer interpolation processing unit 44.
[0074] FIG. 19 is an explanatory diagram of the interpolation process for obtaining a stabilized output image. 8, pixel values for multiple pixels required for rendering each output pixel are sequentially read out from the input image (pixel values) buffered in the buffer memory 42 and stored in the cache memory 43 under the control of the memory control unit 41'. Specifically, the pixel values for multiple pixels required for rendering each output pixel are data for an area made up of multiple pixels including a pixel that includes the position in the input coordinate system indicated by the reference coordinate CR for that output pixel and the pixels surrounding that pixel (see the area Ar surrounded by a thick frame in the figure). For the sake of explanation, the pixel that includes the position in the input coordinate system indicated by the reference coordinate CR will be referred to as the "reference pixel Pr." Furthermore, the pixel area required for rendering, including this reference pixel Pr and its surrounding pixels, will be referred to as the "reference area Ar." The reference area Ar is an area of m pixels x m pixels (m is a natural number greater than or equal to 3) centered on the reference pixel Pr. Note that in the drawings, the reference area Ar is an area of 3 pixels x 3 pixels = 9 pixels centered on the reference pixel Pr, but this is an example for explanation purposes and does not limit the size of the reference area Ar.
[0075] The stabilizer interpolation processor 44 calculates the value of the position indicated by the reference coordinate CR for the output pixel to be processed by interpolation using the values of each pixel in the reference area Ar. This interpolation process uses, for example, a Lanczos filter. Specifically, a hybrid filter that blends a Lanczos filter with a Gaussian filter to prevent aliasing can be used. This hybrid filter is effective for interpolation in RAW format, where the image format is RGGB, and is used to prevent aliasing, particularly in the high-frequency band. The stabilizer interpolation processor 44 performs such an interpolation process sequentially for each output pixel, thereby obtaining a stabilized output image.
[0076] FIG. 20 is a diagram for explaining an example of the internal configuration of the lattice point mesh generating and shaping unit 21′ shown in FIG. FIG. 20 shows an example of the internal configuration of the lattice point mesh generating / shaping unit 21', as well as an image diagram that schematically shows the process of shaping the lattice point mesh.
[0077] The lattice point mesh generating / shaping unit 21' performs processing for generating a lattice point mesh, rotating the lattice point mesh for conversion to the output coordinate system described above (see Figure 12), and shaping the lattice point mesh, and as shown in the figure, is equipped with a lattice point mesh generator 31, a lens distortion corrector 32, a projector 33, a rotator 34, a free curvature perspective projector 35, a scan controller 36, a mesh masker 37, and a lattice point reference coordinate calculator 38.
[0078] The lattice point mesh generator 31 generates a lattice point mesh (see FIG. 11).
[0079] The lens distortion corrector 32 performs lens distortion correction processing on the lattice point mesh based on the lens parameters.
[0080] The projector 33 projects the lattice point mesh after lens distortion correction processing by the lens distortion corrector 32 onto the virtual celestial sphere. As a projection method, for example, central projection or equidistant projection can be adopted (the image in the figure shows an example of central projection). In this example, the projector 33 performs projection processing using either central projection or equidistant projection based on the projection parameters.
[0081] The rotator 34 rotates the lattice point mesh projected onto the virtual celestial sphere by the projector 33 based on the IMU quaternion. This rotation has the effect of rotating the mesh in the same direction as the camera, as described above. For the rotation, information indicating the amount of rotation in the IMU quaternion (rotation amount parameter) is used.
[0082] The free curvature perspective projector 35 projects (reprojects) the lattice point mesh rotated by the rotator 34 onto a plane by free curvature perspective projection based on the projection parameters. By employing free curvature perspective projection, it is possible to impart a desired lens effect to the reprojected lattice point mesh, thereby enabling the creation of an output image. The projection parameters are parameters for specifying the mode of such a lens effect. The scanning controller 36 performs affine transformation processing for appropriate scaling and offset changes on the grid point mesh projected onto the plane. The scanning controller 36 performs these scaling and offset changes based on predetermined parameters, such as predetermined scale / offset parameters.
[0083] The mesh masker 37 inputs the lattice point mesh processed by the scan controller 36 and performs a process of masking the outer edge portion of the lattice point mesh. In the lattice point mesh, the data of the masked outer edge portion is treated as invalid data.
[0084] Information on the position coordinates of each lattice point in the area that has been unmasked by the mesh masker 37 (information on the position coordinates in the input coordinate system) is output as lattice point coordinate information to the segment search unit 23 shown in FIG.
[0085] The grid point reference coordinate calculator 38 calculates the reference coordinates of each grid point in the grid point mesh (each grid point reference coordinate described above) based on the IMU quaternion.
[0086] Returning to the explanation of Figure 8. The segment search unit 23 performs the above-mentioned segment search (inside / outside determination: see FIGS. 13 and 14) based on the segment matrix generated by the segment matrix generation unit 22 and the lattice point coordinate information supplied from the lattice point mesh generation / shaping unit 21. As a result, for each segment position in the segment matrix, four lattice points that include that segment position are identified.
[0087] The remesh data generation unit 24 generates remesh data (see FIG. 17) by performing the above-mentioned triangular interpolation (see FIGS. 15 and 16) for each segment position based on the information on each lattice point reference coordinate supplied from the lattice point mesh generation / shaping unit 21' and the information on the segment search result by the segment search unit 23. As mentioned above, the remesh data can be said to be reference coordinates CR at the segment granularity. The remesh data generating unit 24 outputs the generated remesh data to each pixel coordinate interpolating unit 25 .
[0088] Each pixel coordinate interpolation unit 25 determines a reference coordinate CR for each pixel position of the output image based on the remesh data. As described above, the reference coordinate CR for each pixel position is determined by performing bilinear interpolation based on the remesh data.
[0089] The reference coordinates CR obtained by each pixel coordinate interpolation unit 25 are supplied to a memory control unit 41' and a stabilizer interpolation processing unit 44 in the image processing unit 8'.
[0090] The memory control unit 41' performs a process of writing data of the reference area Ar (see FIG. 19) corresponding to each pixel position of the output image from the buffer memory 42 to the cache memory 43 based on the reference coordinates CR. As described above, the memory control unit 41' also performs the process of releasing the buffer memory 42 based on the reference coordinates CR.
[0091] The stabilizer interpolation processing unit 44 obtains the pixel value at the position indicated by the reference coordinate CR by performing interpolation processing using the data of the reference area Ar stored in the cache memory 43 as described above. The stabilizer interpolation processor 44 performs such interpolation processing sequentially for each pixel position of the output image, thereby obtaining a stabilized output image.
[0092] Here, in the stabilization processing method described above, when determining the reference coordinate CR for each pixel position in the output image, instead of using only one-dimensional information such as the virtual line L1 as in the conventional method to achieve alignment with the output coordinate system, two-dimensional information such as a lattice point mesh is used to achieve alignment with the output coordinate system. This makes it possible to increase the accuracy of the reference coordinates CR, and improve the performance of the stabilization process.
[0093] [3-3. Image reference processing unit as an embodiment] As can be seen by referring to Figure 20, the stabilization processing method assumed in the embodiment involves performing geometric modulation processing on the lattice point mesh, such as lens distortion correction processing by the lens distortion corrector 32 (i.e., deformation of the lattice point mesh), rotation by the rotator 34, and warp processing by the scan controller 36 (i.e., translation and scaling of the lattice point mesh).
[0094] Geometric modulation processing for an object consisting of multiple elements is generally performed as a coordinate transformation process for each element, but performing coordinate transformation for each element individually increases the computational cost required for the geometric modulation processing, which is undesirable.
[0095] Therefore, in this embodiment, AI processing using the aforementioned approximate feature formula as input is applied to various geometric modulation processes for the lattice point mesh. Specifically, a method is adopted in which the lattice point mesh is converted into an approximate feature formula, and various geometric modulation processes are performed in the state of the approximate feature formula.
[0096] FIG. 21 is a diagram for explaining machine learning for realizing geometric modulation processing on the lattice point mesh that has been approximated as described above. First, in the learning environment, a lattice point mesh generation and shaping unit 21' is used to generate training data. The lattice point mesh generation and shaping unit 21' includes a lens distortion corrector 32, a projector 33, a rotator 34, a free curvature perspective projector 35, a scanning controller 36, a mesh masker 37, and a lattice point reference coordinate calculator 38.
[0097] In addition, the learning environment uses a lattice point mesh approximation curved surface generation unit 39, a lens distortion correction learner 32b, a projection learner 33b, a rotation learner 34b, a free curvature perspective projection learner 35b, a scanning control learner 36b, and a mesh mask learner 37b. A DNN-based machine learning machine is used for the lattice point mesh approximation curved surface generation unit 39, lens distortion correction learner 32b, projection learner 33b, rotation learner 34b, free curvature perspective projection learner 35b, scan control learner 36b, and mesh mask learner 37b. Specifically, in this example, a CNN-based machine learning machine is used.
[0098] First, for the learning device serving as the lattice point mesh approximation unit 39, pre-training for SAE is performed using the lattice point mesh (coordinate data for each lattice point) generated by the lattice point mesh generator 31 as input data. After this pre-training, control line association learning is performed. As can be seen from FIG. 11 above, the lattice point mesh is two-dimensional data, so in this case, control line association learning is performed by sequentially assigning x and y address values indicating the position of each lattice point (position in the input coordinate system) to the control line.
[0099] In Figure 21, for the lens distortion correction learner 32b, projection learner 33b, rotation learner 34b, free curvature perspective projection learner 35b, scan control learner 36b, and mesh mask learner 37b, machine learning (including the pre-training of FineTuning:SAE mentioned above) is performed using the outputs (processing results) of the lens distortion corrector 32, projector 33, rotator 34, free curvature perspective projector 35, scan controller 36, and mesh masker 37, respectively, as training data.
[0100] Specifically, for the lens distortion correction learning device 32b, machine learning is performed using the approximate feature equation of the lattice point mesh obtained in the intermediate layer of SAE in the lattice point mesh approximation curved surface generation unit 39 as learning input data and the output of the lens distortion corrector 32 as training data. By performing such machine learning, an algorithm is generated in the learning device serving as the lens distortion correction learning device 32b to apply a correction (deformation) similar to the correction process by the lens distortion corrector 32 to the approximate feature equation of the input lattice point mesh. Therefore, when realizing geometric modulation processing as lens distortion correction processing for a lattice point mesh, there is no need to perform coordinate transformation processing on the coordinate data of each lattice point, and the computational costs for geometric modulation processing can be reduced.
[0101] For the projection learner 33b, machine learning is performed using the approximate feature equation (transformation equation) of the lattice point mesh obtained in the intermediate layer of SAE in the lens distortion correction learner 32b, which is the learner placed immediately before, as learning input data, and the output of the projector 33 as training data. By performing such machine learning, an algorithm is generated in the learning device as the projection learning device 33b to apply a geometric modulation (deformation) to the input approximate feature equation in the same way as when the projection process is performed by the projector 33. Therefore, when implementing geometric modulation processing as a projection process onto a lattice point mesh, it is not necessary to perform coordinate transformation processing on the coordinate data of each lattice point, and it is possible to reduce the computational cost for the geometric modulation processing. Furthermore, the input lattice point mesh data is dimensionally compressed as an approximate feature formula, which reduces the amount of input data. This also reduces the computational cost.
[0102] For the rotation learner 34b, the free curvature perspective projection learner 35b, the scan control learner 36b, and the mesh mask learner 37b, machine learning is performed using the output of the corresponding one of the rotator 34, the free curvature perspective projector 35, the scan controller 36, and the mesh masker 37 as training data, and the approximate feature equation obtained in the intermediate layer of the SAE in the learner placed immediately before as learning input data. As a result, in the learning devices such as the rotation learning device 34b, the free curvature perspective projection learning device 35b, the scan control learning device 36b, and the mesh mask learning device 37b, algorithms are generated to apply geometric modulation to the input approximate feature equation in the same manner as when rotation processing by the rotator 34, reprojection processing by the free curvature perspective projector 35, translation / enlargement / reduction transformation processing by the scan controller 36, and mask processing by the mesh mask device 37 are performed. Therefore, it is not necessary to perform coordinate transformation processing for each grid point in the rotation processing, reprojection processing, and translation / enlargement / reduction transformation processing, which reduces the calculation cost.Furthermore, the calculation cost for these rotation processing, reprojection processing, and translation / enlargement / reduction transformation processing is also reduced by using the approximate feature equation as input data.
[0103] Here, in this example, each of the learners, i.e., the lens distortion correction learner 32b, the projection learner 33b, the rotation learner 34b, the free curvature perspective projection learner 35b, the scan control learner 36b, and the mesh mask learner 37b, is capable of learning an algorithm for geometric modulation of the input approximate feature equation for each different parameter setting for the corresponding geometric modulation process. Specifically, in this case, lens distortion correction learning device 32b learns an algorithm for geometric modulation of the input approximated surface data for each different lens parameter setting. For example, if two types of lens parameter settings, "A" and "B," are possible, learning for lens parameter A is performed using the output of lens distortion corrector 32 in a state where lens parameter A is set as training data, so that an algorithm for lens parameter A is generated. Learning for lens parameter B is performed using the output of lens distortion corrector 32 in a state where lens parameter B is set as training data, so that an algorithm for lens parameter B is generated. At this time, lens distortion correction learning device 32b stores the algorithm generated for each parameter setting in a manner that allows identification of which parameter setting the algorithm corresponds to.
[0104] The projection learner 33b, rotation learner 34b, free curvature perspective projection learner 35b, scan control learner 36b, and mesh mask learner 37b also learn for each parameter setting in a similar manner, and the algorithms generated for each parameter setting by this learning are stored in a manner that allows identification of which parameter setting the algorithm corresponds to.
[0105] In this example, the above-mentioned segment matrix and remesh data are also handled by the approximate feature equation, thereby further reducing the calculation cost. For this purpose, in the learning environment of this example, as shown in Figure 21, a segment matrix generation unit 22, a segment search unit 23, a remesh data generation unit 24, and a pixel coordinate interpolation unit 25 are provided, as well as a lattice point mesh segment matrix conversion learner 26b, a remesh learner 27b, and a remesh extension decoding learner 28b. In this example, a fully connected layer is used as each of the lattice point mesh segment matrix conversion learner 26b, the remesh learner 27b, and the remesh extension decoding learner 28b.
[0106] The lattice point mesh segment matrix conversion learner 26b performs machine learning using, as learning input data, the approximate feature equation (approximate feature equation of the lattice point mesh) obtained in the intermediate layer of SAE in the mesh mask learner 37b, and as training data, the segment search processing results by the segment search unit 23. The data resulting from the segment search processing here is assumed to be data indicating, for each segment position (segment number) in the segment matrix, the number of the quadrilateral element in the lattice point mesh that contains that segment position.
[0107] FIG. 22 is an explanatory diagram of quadrilateral elements of a lattice point mesh. As shown in the figure, a quadrilateral element of a lattice point mesh refers to a grid portion surrounded by four adjacent lattice points (four lattice points that are adjacent in each of the row, column, and diagonal directions), and the number of the quadrilateral element refers to the number assigned to each grid.
[0108] In the lattice point mesh segment matrix conversion learner 26b, an algorithm for realizing domain conversion from the approximate feature formula of the lattice point mesh to the approximate feature formula of the segment matrix is generated by the above machine learning. The principle of such domain transformation will be explained using a one-dimensional example with reference to Figures 23 and 24. Figure 23 shows an example of an approximate curve of a lattice point mesh when the lattice point mesh is assumed to be one-dimensional. The horizontal axis represents the element number, and the vertical axis represents the coordinate. As can be seen from this one-dimensional example, the approximated surface data of the lattice point mesh can be expressed as equivalent to the relationship between the numbers and coordinates of the quadrilateral elements of the lattice point mesh. 24 shows an example of an approximation curve of a segment matrix when the segment matrix is assumed to be one-dimensional. The horizontal axis represents the segment number, and the vertical axis represents the number of a quadrilateral element in the lattice point mesh. The approximation curve of the segment matrix can be expressed as equivalent to the relationship between the number of a quadrilateral element in the lattice point mesh and the segment number. Comparing FIG. 23 with FIG. 24, it can be seen that the approximate feature formula (approximate surface) of the segment matrix has a deformable correlation with the approximate feature formula (approximate surface) of the lattice point mesh.
[0109] In the lattice point mesh segment matrix conversion learner 26b, machine learning is performed using the approximate feature formula from the mesh mask learner 37b described above as learning input data and the segment search processing result by the segment search unit 23 as training data, thereby generating an algorithm that converts the approximate feature formula of the lattice point mesh input from the mesh mask learner 37b into an approximate feature formula of the segment matrix (a formula relating the numbers of the quadrangular elements of the lattice point mesh to the segment numbers).
[0110] The remesh learner 27b performs machine learning using the approximate feature equation of the segment matrix obtained in the intermediate layer of SAE in the lattice point mesh segment matrix conversion learner 26b, the approximate feature equation of the lattice point mesh obtained in the intermediate layer of SAE in the mesh mask learner 37b, and each lattice point reference coordinate output from each lattice point reference coordinate calculator 38 as learning input data, and the output of the remesh data generation unit 24 (reference coordinates CR at segment granularity) as training data. By this machine learning, the remesh learner 27b generates an algorithm for obtaining an approximate feature equation corresponding to the relational expression between each segment position (number) in the segment matrix and the reference coordinate CR corresponding to that segment position.
[0111] The remesh extension decoding learner 28b performs machine learning as control line association learning using the approximate feature equation obtained in the intermediate layer of the remesh learner 27b as learning input data, coordinate data (shown as x, y in the figure because the output coordinate system is two-dimensional) for indicating each pixel position of the output image as control line input, and the output of each pixel coordinate interpolation unit 25 (i.e., the reference coordinate CR for each pixel position of the output image) as teacher data.
[0112] By performing such learning, the remesh extension decoding learner 28b generates an algorithm for decoding the reference coordinate CR of the pixel position specified by the x, y coordinate data from the approximate feature equation input from the remesh learner 27b.
[0113] FIG. 25 is a diagram showing an example of the configuration of the image reference processing unit 7 according to the embodiment, which is configured to include each learning device that has been trained by the above-described method. As shown in the figure, the image reference processing unit 7 of this embodiment includes a lattice point mesh generating and shaping unit 21, which includes a lattice point mesh generator 31 and a lattice point reference coordinate calculator 38, as well as a lattice point mesh approximation curved surface unit 39, a lens distortion correction learner 32a, a projection learner 33a, a rotation learner 34a, a free curvature perspective projection learner 35a, a scan control learner 36a, and a mesh mask learner 37a. Here, lens distortion correction learner 32a, projection learner 33a, rotation learner 34a, free curvature perspective projection learner 35a, scanning control learner 36a, and mesh mask learner 37a represent lens distortion correction learner 32b, projection learner 33b, rotation learner 34b, free curvature perspective projection learner 35b, scanning control learner 36b, and mesh mask learner 37b, respectively, which have been considered to have completed training.
[0114] The image reference processing unit 7 is also provided with a lattice point mesh segment matrix conversion learner 26a, a remesh learner 27a, and a remesh extension decoding learner 28a. These lattice point mesh segment matrix conversion learner 26a, remesh learner 27a, and remesh extension decoding learner 28a represent the trained lattice point mesh segment matrix conversion learner 26b, remesh learner 27b, and remesh extension decoding learner 28b, respectively.
[0115] In the lattice point mesh generating / shaping unit 21 , the lattice point mesh approximating curved surface unit 39 receives as input coordinate data for each lattice point of the lattice point mesh generated by the lattice point mesh generator 31 . Here, the lattice point mesh data may be common to each frame, and therefore the approximation surface data generation process by the lattice point mesh approximation unit 39 may be performed at least once (the generated approximation surface data may be stored in memory and read out sequentially for each frame).
[0116] As shown in the figure, lens distortion correction learning device 32a receives as input data an approximate feature equation obtained in the intermediate layer of lattice point mesh approximation curved surface generation unit 39, and projection learning device 33a receives as input data an approximate feature equation that has been subjected to geometric modulation processing as a lens distortion correction process in lens distortion correction learning device 32a. Rotation learning device 34a receives as input data an approximate feature equation that has been subjected to geometric modulation processing as a projection process in projection learning device 33a, and free curvature perspective projection learning device 35a receives as input data an approximate feature equation that has been subjected to geometric modulation processing as a rotation process in free curvature perspective projection learning device 35a. Scanning control learning device 36a receives as input data an approximate feature equation that has been subjected to geometric modulation processing as a reprojection process in free curvature perspective projection learning device 35a, and mesh mask learning device 37a receives as input data an approximate feature equation that has been subjected to geometric modulation processing as an affine transformation process in scanning control learning device 36a.
[0117] As can be understood from the above description, in this example, lens distortion correction learner 32a, projection learner 33a, rotation learner 34a, free curvature perspective projection learner 35a, scan control learner 36a, and mesh mask learner 37a can learn algorithms for geometric modulation processing on approximate feature equations for different parameter settings. In this case, lens distortion correction learner 32a, projection learner 33a, rotation learner 34a, free curvature perspective projection learner 35a, scan control learner 36a, and mesh mask learner 37a have a function of switching algorithms so as to use an algorithm corresponding to the parameter setting from among the multiple learned algorithms.
[0118] The lattice point mesh segment matrix conversion learning device 26a is provided with the approximate feature equation of the lattice point mesh that has been subjected to mask processing by the mesh mask learning device 37a as input data.
[0119] The remesh learner 27a is given, as input data, an approximate feature equation of the segment matrix obtained by the conversion process in the lattice point mesh segment matrix conversion learner 26a, each lattice point reference coordinate from the lattice point reference coordinate calculator 38, and the approximate feature equation from the mesh mask learner 37a.
[0120] The remesh extension decode learner 28a is provided with, as input data, an approximate feature equation (an approximate feature equation corresponding to the relational equation between the segment position and the reference coordinate CR) for the remesh data obtained by the remesh learner 27a and x, y coordinate data for specifying the pixel position of the output image. As a result, the remesh extension decode learner 28a decodes the reference coordinate CR of the pixel position specified by the x, y coordinate data from the approximate feature equation input from the remesh learner 27a and outputs it.
[0121] As can be understood from the above explanation, the image reference processing unit 7 according to the embodiment applies AI processing technology that uses an approximate feature formula as input to processes such as shaping and rotation of lattice point meshes, lattice point mesh segment matrix conversion for determining reference coordinates CR for each pixel position, and remesh data generation processing, thereby reducing the calculation costs of various processes for realizing stabilization processing.
[0122] Regarding grid points, approximately 65 x 65 grid points are required to achieve sufficient accuracy to prevent faults in the final output image. However, performing various processes for geometric modulation using rule-based processing for such a large number of grid points requires a significant amount of computation. In contrast, an approach such as the embodiment in which various processes for geometric modulation of the grid point mesh are realized using AI processing with approximate feature equations as input can significantly reduce computational costs compared to rule-based processing. For example, while conventional rule-based projection calculations require 13 DSP (Digital Signal Processor) instruction cycles per grid point, the AI method used in the embodiment achieves an average of 0.1 cycles per grid point, achieving a 130-fold increase in computational speed.
[0123] Furthermore, adopting an AI approach like the embodiments contributes to reducing the burden on designers compared to rule-based processing. For example, in a processing system involving distortion correction or complex rotation control, it is extremely difficult for a designer to mathematically determine the transformation formula. However, by adopting an AI approach like the embodiments, once the input and output data are prepared, it is possible to easily obtain a mathematical transformation algorithm by expanding each warp process using mathematical processing. Furthermore, the mathematical formula can be calculated much faster and with lower computational costs than a rule-based processing approach that searches quantized digital data one by one. Thus, the AI approach of the embodiments is an effective approach suitable for deriving transformation formulas that are difficult to design even when a designer can observe geometric correlations.
[0124] Here, in the image reference processing unit 7 as an embodiment, a mesh mask learning device 37a is provided to mask the outer edge portion of the lattice point mesh, thereby making it possible to suppress deterioration in image quality of the stabilized output image. When an approximate feature equation is used in shaping or rotating a grid point mesh, the approximate feature equation for the grid point mesh is, in principle, equivalent to finding an equation for a curved surface using the so-called least squares method, and therefore the accuracy of estimating the grid points tends to decrease at the outer edges of the grid point mesh. Therefore, if the grid point mesh is used as is, the accuracy of the grid point mesh segment matrix conversion process in the subsequent stage will decrease, resulting in deterioration of image quality at the outer edges of the stabilized output image. Therefore, a mesh masker 37 is provided to mask the outer edges of the grid point mesh, thereby preventing such image quality deterioration.
[0125] [3-4. Image Processing Unit as an Embodiment] Next, the image processing unit 8 included in the signal processing device 1 according to the embodiment will be described. FIG. 26 is a block diagram showing an example of the internal configuration of the image processing unit 8. As shown in FIG. Compared to the image processing unit 8' shown in Figure 8 as a comparative example, the image processing unit 8 differs in the part that obtains a stabilized output image from an input image and the part that performs correction processing on the stabilized output image.
[0126] In the following description, parts that are similar to parts that have already been described will be given the same reference numerals and description thereof will be omitted.
[0127] In Figure 26, the image processing unit 8 includes a memory control unit 41, a buffer memory 42, and a cache memory 43, as well as a four-corner extraction unit 46, a thinning stabilizer processing unit 47, a high-frequency estimation learning unit 48a, a synthesis unit 49, 8x8 buffers 50 and 51, a 1 / 8-dimensional compression encoder 52a, a delay unit 53, and a high-frequency correction learning unit 45a. In comparison with the image processing unit 8' shown in Figure 8 as a comparative example, as surrounded by dashed lines in the figure, the thinning stabilizer processing unit 47, high-frequency estimation learning unit 48a, and synthesis unit 49 correspond to the stabilizer interpolation processing unit 44, and the 8x8 buffers 50, 51, 1 / 8-dimensional compression encoder 52a, delay unit 53, and high-frequency correction learning unit 45a correspond to the correction processing unit 45'.
[0128] Here, in the image processing unit 8' as a comparative example, Lanczos interpolation is performed to interpolate one pixel from a plurality of pixels, such as 3×3 pixels, as stabilized interpolation processing to obtain a stabilized output image (see FIG. 19). This method loads an image made up of a plurality of pixels, such as 3×3 pixels, to output one pixel, and therefore the calculation cost is high when implemented in an LSI. Therefore, in this embodiment, thinning-out stabilization processing is adopted, which refers to processing in which, instead of referencing the corresponding pixel value in the input coordinate system for each pixel position in the output coordinate system, the corresponding pixel value in the input coordinate system is referenced in units of pixel groups consisting of a predetermined number of pixels in the output coordinate system.
[0129] The thinning-out stabilization process in this embodiment will be described with reference to FIG. Here, an example is shown in which the original resolution of the output image (FHD resolution in this example) is thinned out to 1 / 16. In this case, 4x4 pixels in the output image are treated as one pixel group, and for each pixel group, 16 pixels are thinned out to 1 pixel. Specifically, one pixel value is calculated from the input image based on the reference coordinates CR of the four corner pixels in the pixel group of the output image. For example, pixel values are obtained for four pixel positions in the input coordinate system specified by the integer values of the reference coordinates CR of the four corners, and the average of these four pixel values is used as the pixel value of the 4x4 pixel group. The thinning stabilization process is not based on a single reference coordinate CR, such as the upper left corner of a 4x4 pixel group, but rather on the reference coordinate CR of all four corners, as described above. This is because it takes into account that the amount of warping from a pixel position in the output coordinate system to the corresponding position in the input coordinate system can differ even within a 4x4 pixel group.
[0130] By performing the thinning-out stabilization process described above, it is possible to significantly reduce the amount of calculation required for obtaining a stabilized output image based on the reference coordinates CR and the input image. However, in this case, the stabilized output image is an image that has been thinned out (thinned out to 1 / 16 in this example) relative to the resolution of the original output image, and therefore has a reduced resolution.
[0131] Therefore, in this embodiment, a reduction in resolution of the stabilized output image is suppressed by synthesizing high-frequency components with the thinned image obtained by the thinning stabilization process. Specifically, based on the 6x6 pixel input image shown in Figure 27, the high frequency components of the pixel group (4x4 pixels) in the output image are estimated, and the estimated high frequency components are synthesized with the thinned image, thereby restoring the original resolution.
[0132] Based on the above explanation, the configuration shown in FIG. 26 for implementing the thinning stabilization process and the process related to the synthesis of high frequency components will be described. First, the corner extraction unit 46 inputs the reference coordinates CR (reference coordinates CR for each pixel of the output image) output from the image reference processing unit 7 (remesh extension decode learning device 28a) shown in Figure 25, and extracts and outputs the reference coordinates CR of the four corner pixels for each pixel group of 4x4 pixels in the output image. In this example, pixel values are buffered for each 8x8 pixel in the 8x8 buffers 50, 51 described below and output in sequence, so the four-corner extraction unit 46 sequentially outputs the reference coordinates CR of the four corners of the four pixel groups included in each pixel unit consisting of 8x8 pixels in the output image. In the figure, the reference coordinates CR of the four corners are expressed as "CR×4".
[0133] The memory control unit 41 receives the reference coordinates CR of the four corners of each pixel group in the output image from the four corner extraction unit 46 in order. The memory control unit 41 is similar to the memory control unit 41' in that it controls buffering of the input image in the buffer memory 42 and release of the memory based on the input reference coordinates CR. Based on the reference coordinates CR of the four corners input from the four corner extraction unit 46, the memory control unit 41 reads out the pixel values of a group of pixels (6 x 6 pixels in this example) in the input image corresponding to the reference coordinates CR of the four corners from the buffer memory 42 and writes them to the cache memory 43. Here, the group of pixels in the input image corresponding to the reference coordinates CR of the four corners means the group of pixels on the input image that contains each position in the input coordinate system indicated by the reference coordinates CR of the four corners, and in this example, as illustrated in Figure 27 above, it corresponds to a group of pixels consisting of 6 x 6 pixels that contains each of the four positions in the input coordinate system indicated by the reference coordinates CR of the four corners. Hereinafter, the pixel group in the input image corresponding to the reference coordinates CR of the four corners will be referred to as the "input image target pixel group."
[0134] Note that the image size of the input image target pixel group is not limited to the 6x6 pixels shown as an example, and other image sizes such as 8x8 pixels can be set depending on the scale ratio of the output image to the input image.
[0135] The pixel values of the input image target pixel group that are sequentially output from the memory control unit 41 to the cache memory 43 are also input to the high frequency estimation learning unit 48a.
[0136] The thinning stabilization processing unit 47 performs the above-mentioned thinning stabilization processing for each pixel group based on the reference coordinates CR of the four corners of each pixel group input in order from the four-corner extraction unit 46 and the pixel values of the input image target pixel group written to the cache memory 43. Specifically, in this example, the pixel values of four pixel positions in the input coordinate system specified by the integer values of the reference coordinates CR of the four corners are obtained, and the average of these four pixel values is set as the pixel value of the 4×4 pixel group. In addition, the thinning stabilization process can also be considered to be a process in which the value of a pixel position in the input image identified from the integer value of one of the four corner reference coordinates CR, such as the upper left corner, is calculated as the pixel value of a 4x4 pixel group.
[0137] The thinned image (stabilized) obtained by the thinning and stabilization processing unit 47 is input to the synthesis unit 49. In this embodiment, the thinned image obtained by the thinning stabilizer processing unit 47 is supplied to the image analysis processing unit 9, which will be described later.
[0138] Here, the image reference processing unit 7 shown in FIG. 25 was described as calculating the reference coordinates CR for each pixel of the output image. However, when performing the thinning stabilization process described above, only the reference coordinates CR for the four corners of each pixel group are required as reference coordinates CR. Therefore, in the image reference processing unit 7 described above, the remesh extension decoding learner 28a can be configured to calculate only the reference coordinates CR for the four corners of each 4×4 pixel group. In this case, the amount of calculation required to calculate the reference coordinates CR can be reduced to one-quarter. Furthermore, in this case, there is no need to provide the four-corner extraction unit 46 in the image processing unit 8, thereby reducing the amount of calculation accordingly.
[0139] In Figure 26, the high-frequency estimation learning device 48a is configured with a machine learning device that supports deep learning, such as an autoencoder, and estimates (infers) the high-frequency components of the pixel group (4 x 4 pixels) of the output image using the reference coordinates CR of the four corners output by the four-corner extraction unit 46 and the pixel values of the input image target pixel group (6 x 6 pixels) output by the memory control unit 41 as input data.
[0140] The synthesis unit 49 receives the thinned image output by the thinning stabilizer processing unit 47 and the high frequency components estimated by the high frequency estimation learning unit 48a, and synthesizes them.
[0141] FIG. 28 is a diagram illustrating machine learning by the high-frequency estimation learning device 48a. Machine learning for high-frequency estimation is performed as supervised learning, using the difference value between the correct pixel value of a pixel group (4x4 pixels) and the pixel value of one thinned pixel as the teacher. This difference value is calculated by the high frequency correct value calculation unit 55 in the figure. Specifically, the high frequency correct value calculation unit 55 inputs the correct pixel value of the pixel group (4 × 4 pixels) and the pixel value of one pixel after thinning out of the pixel group obtained by the thinning-out stabilizer processing unit 47, and obtains the difference value therebetween (the difference value of each pixel of the 4 × 4) as correct data (teaching data) of the high frequency component.
[0142] In the figure, high-frequency estimation learning device 48b represents high-frequency estimation learning device 48a before learning. As shown in the figure, for high-frequency estimation learning device 48b, machine learning is performed using reference coordinates CR of the four corners of the pixel group and pixel values of 6 × 6 pixels of the input image target pixel group as input data, and difference values of 4 × 4 pixels obtained by high-frequency correct value calculation unit 55 as training data. As a result, high-frequency estimation learning device 48b generates an algorithm for estimating high-frequency components of a pixel group using the reference coordinates CR of the four corners of the pixel group and the pixel values of the 6×6 pixels of the input image target pixel group as input data.
[0143] Next, a correction processing system for the stabilized output image in the image processing unit 8 will be described. Specifically, the components are 8×8 buffers 50 and 51, a 1 / 8 DNN encoder 52a, a delay unit 53, and a high frequency correction learning unit 45a shown in FIG.
[0144] Here, we will use a method for correcting stabilized output images that is suitable for performing correction processing primarily targeting high frequency components. We will also assume that the correction processing used here is image correction processing that utilizes correlation between frames, such as 3DNR or flicker correction. Basically, when the current frame is the T frame, the synthesized image of the T frame obtained by the synthesis unit 49 and the high-frequency components of the T-1 frame obtained by the high-frequency estimation learning unit 48a are input to the high-frequency correction learning unit 45a, and this high-frequency correction learning unit 45a infers a processed result image for high-frequency correction processing that utilizes inter-frame correlation such as 3DNR.
[0145] In this example, an approximate feature equation is used for the high frequency components input to the high frequency correction learning device 45a to reduce the calculation cost. The approximate feature equation for the high frequency components is generated by the 1 / 8-dimensional compression encoder 52a. In addition, in order to provide this 1 / 8-dimensional compression encoder 52a, an 8×8 buffer 50 that buffers 8×8 pixels of high-frequency components is provided in the image processing unit 8. Furthermore, in order to correspond to the provision of the 8×8 buffer 51 in the system of high-frequency components, an 8×8 buffer 50 is provided in the system of the composite image obtained by the composition unit 49. By providing these 8×8 buffers 50 and 51, the high frequency correction learning device 45a in this case is provided with an 8×8 pixel T frame composite image and an 8×8 pixel T−1 frame high frequency component as input data. That is, the high frequency correction learning device 45a in this case performs high frequency correction for each pixel unit consisting of 8×8 pixels.
[0146] First, the 1 / 8-dimensional compression encoder 52a will be described with reference to FIG. The 1 / 8-dimensional compression encoder 52b in the figure represents the 1 / 8-dimensional compression encoder 52a before learning. The 1 / 8-dimensional compression encoder 52b is configured as a machine learning machine that has at least SAE. Here, the compression ratio is determined by the number of inputs in SAE and the number of neurons in the hidden layer. In this example, for the purpose of 1 / 8 compression, the number of neurons in the input layer is set to 64 and the number of neurons in the hidden layer is set to 8.
[0147] Pre-training is performed on such a 1 / 8-dimensional compression encoder 52b, with input data being high-frequency components of 8×8 pixels (high-frequency components of four pixel groups) obtained in the 8×8 buffer 51. After this pre-training, control line association learning is performed, in which the x and y coordinate values for specifying pixel positions are set as the control line values, although this is not shown in the figure. As a result, an approximate feature equation indicating the features of the high frequency components of the input 8×8 pixels can be obtained in the intermediate layer of the SAE in the trained 1 / 8-dimensional compression encoder 52a.
[0148] Because the high-frequency components for each pixel group obtained by the high-frequency estimation learning unit 48a are mostly zero, the 1 / 8-dimensional compression encoder 52a described above can compress the data relatively efficiently. However, some image degradation occurs, making this a lossy compression method. While this compressed data tends to be less efficient than when variable-length coding such as entropy coding is used, it allows access to local coordinates without reverse coding, making it a compact compression technique suitable for reducing the load on image processing and bus transfers. Since the compressed data itself is a DNN feature, it can be treated as data that has undergone dimensionality compression processing for subsequent machine learning, which also contributes to reducing the amount of calculations.
[0149] The machine learning of the high frequency correction learning device 45a will be described with reference to FIG. The high frequency correction learning device 45b in the figure represents the high frequency correction learning device 45a before machine learning. As shown in the figure, the high-frequency correction learning device 45b is provided with, as learning input data, a composite image (T frame composite image) of 8×8 pixels obtained by an 8×8 buffer 50 and an approximate feature equation of high-frequency components in a T−1 frame output from a 1 / 8-dimensional compression encoder 52a and delayed by one frame by a delay device 53. Then, the correct values of the 8×8 pixels, that is, the correct values of the target correction processing such as 3DNR or flicker correction processing, are provided as training data, and machine learning is performed as supervised learning. As a result, the trained high-frequency correction learning device 45a receives as input the composite image of the T frame and the high-frequency components (approximate feature equation) of the T-1 frame, and generates an algorithm that infers the processing result image of the desired correction processing such as 3DNR or flicker correction processing (specifically, image processing that is expected to improve performance by improving frame correlation through stabilization).
[0150] 26, the high frequency correction learning device 45a outputs a processed image of the high frequency correction process for each pixel unit (8 × 8 pixels in this example). The processed image output in this manner is supplied to the AR processing unit 11 shown in FIG. 1 as a stabilized RGB image and is used for self-position estimation by SLAM, etc.
[0151] In the image processing unit 8 of this embodiment, only the high-frequency components on the T-1 frame side are compressed to 1 / 8 dimensions and buffered to reduce the frame buffer capacity, and the high-frequency correction learning unit 45a directly inputs and processes the compressed fixed-length data without decompressing it. Conventionally, fixed-length compressed coded data of about 1 / 8 size is difficult to adopt in practice due to severe image degradation, but in this case, only the high-frequency components are compressed and the T frame is baseband, so there is not much image degradation and it contributes to improving machine learning accuracy.
[0152] In the image processing unit 8, the thinning stabilization process is classified as a compression technique that reduces the amount of calculation compared to the method of pixel interpolation on a pixel-by-pixel basis as in the comparative example, and the resulting resolution is lossy compression that involves degradation inherent to compression. However, this is a suitable approach for processing that does not require very high resolution data, such as AR processing as in this embodiment, and that requires reduced power consumption.
[0153] For confirmation, in the image processing unit 8, the high-frequency correction learning device 45a corresponds to an example of the "inference processing unit" recited in the claims, and the 1 / 8-dimensional compression encoder 52a corresponds to an example of the "approximate feature equation generation unit."
[0154] [3-5. Modifications Related to Image Processing Unit] Modifications of the image processing unit 8 will be described with reference to FIGS. In the above, it is assumed that image stabilization processing will be performed, so in the high frequency correction processing in the high frequency correction learning device 45a, the image of the T frame and the image of the T-1 frame are used as is without rotation. However, in cases where image rotation may occur between frames, such as when stabilization processing is not performed, processing should be performed in the high frequency correction processing to cancel the rotation between frames in order to increase the correlation between frames.
[0155] In this case, for learning related to high frequency correction processing, as shown in Figure 31, a high frequency correction learning device 45b' is prepared that can input inter-frame rotation information representing the amount of rotation between the T frame and the T-1 frame, and for machine learning of this high frequency correction learning device 45b', supervised learning is performed using the composite image of the T frame, the approximate feature equation of the high frequency component in the T-1 frame, and the inter-frame rotation information as learning input data, and the processing result image of the desired high frequency correction processing such as 3DNR or flicker correction processing as training data. This makes it possible to realize high frequency correction processing that increases the correlation between frames in response to cases where image rotation may occur between frames, thereby improving the accuracy of high frequency correction processing.
[0156] In addition, in the above, an example was given of the high-frequency estimation learning device 48a estimating the high-frequency components of the difference with the thinned image, but for high-frequency estimation, it is also possible to set up frequency layers and gradually increase the resolution to create multiple layers, if necessary.
[0157] FIG. 32 shows an example of multi-layered high frequency estimation. In the example shown in this figure, in addition to the previously described high frequency estimation learning device 48a, i.e., a learning device that estimates high frequency components to increase the resolution of a 1 / 16 thinned image to the original resolution (1x), there is also provided a high frequency estimation learning device 48a' that estimates high frequency components to increase the resolution to 1 / 4x, and a high frequency estimation learning device 48a' that estimates high frequency components to obtain super-resolution images such as 4x super-resolution. In this case, the resolution of the training data is adjusted so that a high-frequency image based on the target frequency can be estimated in the learning device of each layer.
[0158] In this case as well, the image of the estimated high frequency components can be compressed at a predetermined compression rate and treated as fixed-length compressed code data, as exemplified above in the 1 / 8-dimensional compression encoder 52a.
[0159] Furthermore, although high-frequency correction processing has been mentioned above as image processing for stabilized images, low-frequency correction processing can also be performed. Low-frequency correction processing here refers to processing suitable for image correction that primarily targets low-frequency components, such as shading correction (correction of color unevenness in an imager) and white balance adjustment (adjusting RGB gain according to the light source color).
[0160] FIG. 33 is an explanatory diagram of a machine learning method for realizing low-frequency correction processing. In the figure, high frequency estimation learning device 48Ab has the same configuration as high frequency estimation learning device 48b described above, but is given a different reference numeral because the teaching data is different from that of high frequency estimation learning device 48b. As in the case of high-frequency estimation learning device 48b, high-frequency estimation learning device 48Ab also uses the reference coordinates CR of the four corners and the pixel values of the input image target pixel group (6×6 pixels) as learning input data. In this case, the high-frequency corrected value calculation unit 55 receives as input the thinned image obtained by the thinning stabilizer processing unit 47, as well as a low-frequency corrected image (image at the original resolution) obtained by performing a predetermined low-frequency correction process, such as shading correction, on the correct values of a pixel group (4 × 4 pixels) in the output image using the low-frequency image processing unit 56. The high-frequency corrected value calculation unit 55 calculates a difference image between the thinned image and the low-frequency corrected image, and uses the difference image as training data for the high-frequency estimation learning unit 48Ab. Machine learning is performed on the high-frequency estimation learner 48Ab using the learning input data and teacher data described above. As a result, the trained high-frequency estimation learner 48Ab (hereinafter referred to as "48Aa") receives the reference coordinates CR of the four corners and the pixel values of the input image target pixel group (6 x 6 pixels) as inputs, and generates an algorithm for estimating high-frequency components as the difference between the thinned image and a low-frequency corrected image obtained by performing low-frequency correction processing on the original resolution image.
[0161] FIG. 34 illustrates an example of the configuration at the time of inference. As shown in the figure, a low-frequency image processing unit 57 is provided that performs low-frequency correction processing on the thinned image obtained by the thinning stabilizer processing unit 47, and the thinned image that has been subjected to low-frequency correction processing by this low-frequency image processing unit 57 and the high-frequency components (high-frequency image) estimated by the high-frequency estimation learning unit 48Aa are input to the synthesis unit 49, which then synthesizes them.
[0162] According to the above configuration, the high-frequency estimation learning device 48Aa estimates the difference between the original resolution image that has been subjected to low-frequency correction processing and the thinned image as a high-frequency component, so that the synthesized image obtained by combining this high-frequency component with the thinned image can be an image that is approximately equivalent to an image obtained by applying low-frequency correction processing to the original resolution image. Furthermore, with the above configuration, the low-frequency correction process only needs to be performed on the thinned image, thereby reducing the calculation cost.
[0163] The high-frequency correction process described above has a relatively high computational cost, and in some cases, it may be undesirable to process it on the LSI side, especially when real-time performance is not required. For example, in a smartphone environment, it is more advantageous to process it on a processor in the AP (application) layer (e.g., a GPU (Graphics Processing Unit)) at a later stage, rather than incorporating high-frequency correction functionality into the LSI. One possible implementation for addressing such use cases is to provide a function that packs high-frequency compressed data of T-1 frames for T frame images and transmits it as metadata to a later stage, without providing logic for high-frequency correction processing in the LSI.
[0164] A specific configuration example will be described with reference to FIG. 35A, in this case as well, the system includes a configuration for obtaining an image of T frame and high-frequency compressed data of T-1 frame, specifically, a thinning stabilizer processing unit 47, a high-frequency estimation learning unit 48a, a synthesis unit 49, 8×8 buffers 50 and 51, a 1 / 8-dimensional compression encoder 52a, a delay unit 53, and a high-frequency correction learning unit 45a. Furthermore, the system includes a packing processing unit 58, which packs the synthesized image of T frame (RGB image in this example) output from the 8×8 buffer 50 and the high-frequency compressed data of T-1 frame (1 / 8 fixed-length compressed code data in this example) output from the delay unit 53, as shown in FIG. 35B, and outputs the packed data to a subsequent stage. Based on such packed data, a processor such as a GPU at a later stage can perform high frequency correction processing similar to that of the high frequency correction learning device 45a.
[0165] As explained above with reference to FIG. 31, when inter-frame rotation information is used in the high frequency correction process, a configuration can be adopted in which the inter-frame rotation information is output to a subsequent stage together with the above-mentioned packing data.
[0166] <4. Image analysis processing unit> [4-1. Configuration and processing of image analysis processing unit] FIG. 36 is a block diagram showing an example of the internal configuration of the image analysis processing unit 9. As described above, the image analysis processing unit 9 generates an extended depth image, which is an image in which the dynamic range of distance is extended from the depth image obtained by the ToF sensor 4.
[0167] The image analysis processing unit 9 performs optical flow analysis, which is information about the movement of a subject in an image, segmentation analysis, which is analysis of image segmentation (semantic segmentation), and depth analysis, which analyzes the depth of an image.
[0168] For confirmation, an example of segmentation analysis is shown in Figure 37. Figure 37A shows an example in which segments are analyzed as the ground and segments as the sky, and Figure 37B shows an example in which segments are analyzed as buildings existing on the ground, along with segments of the ground and sky.
[0169] The optical flow can be understood as a change in contrast between frames.
[0170] In this example, segmentation analysis and optical flow analysis are performed using a thinned image (1 / 16 thinned image) obtained by the image processing unit 8 described above (see FIG. 36). Also in this example, a thinned depth image obtained by thinning the depth image obtained by the ToF sensor 4 (with a resolution equivalent to that of a 1 / 16 thinned image, for example) is used for depth analysis. By using thinned images, we aim to reduce the computational costs associated with segmentation analysis, optical flow analysis, and depth analysis.
[0171] In segmentation analysis, optical flow analysis, and depth analysis, the analysis information is encoded into approximate feature equations, and by using these approximate feature equations and IMU information, it becomes possible to estimate long distances that are difficult to measure with a ToF sensor in extended depth analysis. This improves the performance of applications that use depth information in the subsequent stage, specifically the AR processing unit 11 in this example.
[0172] In FIG. 36, the image analysis processing unit 9 includes a segmentation analysis learning device 61a, an optical flow analysis learning device 62a, a depth image mathematical encoder 63, a delay device 64, a thinning processing unit 65, and an extended depth analysis learning device 66a.
[0173] The segmentation analysis learning device 61a is configured with a fully connected layer that also functions as an autoencoder, and by performing machine learning described later, it has the function of inferring the segmentation analysis results for the thinned image input from the image processing unit 8, and also obtains an approximate feature equation (hereinafter referred to as the "segment approximate feature equation") that indicates the characteristics of the segmentation analysis results in the intermediate layer of the SAE (fully connected layer autoencoder 61-2 described later) in the fully connected layer.
[0174] The optical flow analysis learning device 62a is configured with a fully connected layer that also functions as an autoencoder, and receives as input a thinned image (T frame) from the image processing unit 8 and a thinned image of T-1 frame obtained by delaying the thinned image by one frame using a delay device 64. By performing machine learning described later, the optical flow analysis learning device 62a has the function of inferring an optical flow analysis result using the thinned image of T frame and the thinned image of T-1 frame as input, and also obtains an approximate feature equation (hereinafter referred to as an "optical flow approximate feature equation") that indicates the features of the optical flow analysis result in the intermediate layer of the SAE (fully connected layer autoencoder 62-2 described later) in the fully connected layer.
[0175] The depth image mathematical equation encoder 63 is configured with a fully connected layer that also functions as an autoencoder, and receives as input a thinned depth image obtained by thinning the depth image from the ToF sensor 4 using a thinning processing unit 65. The depth image mathematical equation encoder 63 performs at least pre-training using the thinned depth image as input, thereby obtaining an approximate feature equation (hereinafter referred to as a "thinned depth image approximate feature equation") that indicates the features of the thinned depth image in the intermediate layer of the SAE.
[0176] The extended depth analysis learner 66a is configured with a fully connected layer that also functions as an autoencoder, and performs machine learning, as described below, to infer an extended depth image using as input data a segment approximation feature equation, an optical flow approximation feature equation, a thinned depth image approximation feature equation, and IMU information (in this example, triaxial acceleration and triaxial angular velocity information) detected by the IMU 2. The extended depth analysis learner 66a also has a function to decode the extended depth image, and outputs the distance value of the corresponding pixel according to the value of the pixel address (pixel address in the extended depth image) specified by the control line.
[0177] Here, the inference of the extended depth image in the extended depth analysis learning device 66a can be considered as follows. For example, even if the only input information is IMU information, if there is a gravity reference, the general position of the ground can be predicted, and in an environment where the lens parameters are known, the depth can be estimated up to the Earth's horizon (because the direction perpendicular to the horizon coincides with the direction of depth change). In addition, if an image captured by the image sensor 3 is available and segmentation analysis information that can understand the image composition is available, objects such as buildings on the ground can be recognized, and depth can be estimated to a certain extent from image recognition (see Figure 38). Furthermore, more specific depth can be estimated from short-distance depth information from the ToF sensor 4 and the amount of translation of the image (optical flow analysis information), which results in a general depth estimation outside the range of distance measurement possible by the ToF sensor 4. In this example, the depth estimation using the IMU, segmentation analysis information of the captured image, and optical flow analysis information can be understood as estimating depth based on multiple types of input information, just as humans empirically perceive depth using not only visual information but also other information such as the sense of balance in the ear, for example, being able to correctly perceive depth using only an image from one eye.
[0178] Note that when estimating an augmented depth image, variations in the amount of translation from the optical flow may occur, and it is also possible that the amount of translation from the optical flow may be misdetected due to insufficient contrast in the sky or buildings. For this reason, it is not always possible to estimate an ideal augmented depth image. The design concept incorporates the idea that any minor misdetections will be removed by RANSAC (Random Sampling Consensus) in the subsequent SLAM processing.
[0179] Here, in the image analysis processing unit 9, the segment approximation feature equation obtained by the segmentation analysis learning device 61a is used for Delta Pose estimation in the image recognition processing unit 10 shown in FIG. 1, but this point will be explained again later.
[0180] The learning method of each learning device in the image analysis processing unit 9 will be described with reference to FIGS. FIG. 39 is an explanatory diagram of a learning method related to segmentation analysis, FIG. 40 is an explanatory diagram of a learning method related to optical flow analysis, and FIG. 41 is an explanatory diagram of a learning method related to estimation of an extended depth image. In each of these figures, the learning devices before learning are represented by the letter "b" at the end of the code.
[0181] 39, the segmentation recognizer 61r is an AI recognizer using CNN or the like that performs image recognition processing as semantic segmentation. Note that for semantic segmentation, various methods have been proposed in the computer vision field, and a training set can be prepared using a general tool, so details will not be mentioned here. In addition, in FIG. 39, the segmentation analysis learning device 61b is configured as a learning device having an SAE 61-1 and a fully connected layer autoencoder 61-1 at the subsequent stage.
[0182] For the segmentation analysis learning device 61b, the thinned image obtained by thinning the original resolution image by the thinning processing unit 67 is used as learning input data. In addition, the training data is the result of thinning the segmentation recognition result of the original resolution image obtained by the segmentation recognizer 61r using the original resolution image as input by the thinning processing unit 67.
[0183] As a specific learning procedure, first, pre-training is performed on the SAE 61-1 using a thinned image as input data. Through this pre-training, an approximate feature equation showing the characteristics of the thinned image is obtained in the intermediate layer of the SAE 61-1 in response to the thinned image being input to the SAE 61-1, and the approximate feature equation obtained in the intermediate layer is input to the fully connected layer autoencoder 61-2. After pre-training, machine learning is performed using the thinned image as input data and the segmentation recognition result by the segmentation recognizer 61r thinned by the thinning processing unit 67 as training data. As a result, in the segmentation analysis learning device 61a that has been trained, an algorithm is generated that uses the thinned image as input data to infer the segmentation analysis results for the thinned image, and an approximate feature equation (segment approximate feature equation) that indicates the characteristics of the segmentation analysis results is obtained in the intermediate layer of the fully connected layer autoencoder 61-2.
[0184] Next, in FIG. 40, the optical flow analysis learning device 62b uses a learning device having two SAEs 62-1, one for obtaining an approximate feature equation for the thinned image of T frame and the other for obtaining an approximate feature equation for the thinned image of T-1 frame, as shown in the figure, and a fully connected layer autoencoder 62-2 to which the approximate feature equations output from these two SAEs 62-1 are input. In this case, as the learning input data, a thinned image of T frames obtained by thinning the original resolution image by the thinning processing unit 67 is used as the input to one SAE 62-1, and a thinned image of T-1 frames obtained by thinning the original resolution image by the thinning processing unit 67 and then delaying it by the delay unit 64 is used as the input to the other SAE 62-1. The training data uses the optical flow information detected by the optical flow detection unit 62r, which is thinned out by a thinning processing unit 67. As shown in the figure, the optical flow detection unit 62r receives the original resolution image and the original resolution image delayed by a delay unit 64, and detects the optical flow.
[0185] In this case as well, the learning procedure first involves pre-training the SAE 62-1. As a result, when thinned images of T and T-1 frames are input to each SAE 62-1, an approximate feature equation indicating the characteristics of the thinned image of T frame is obtained in the intermediate layer of one SAE 62-1, and an approximate feature equation indicating the characteristics of the thinned image of T-1 frame is obtained in the intermediate layer of the other SAE 62-1. These approximate feature equations are then input to the fully connected layer autoencoder 62-2. After the pre-training, machine learning is performed using the optical flow detection results obtained by the optical flow detection unit 62r, which have been thinned out by the thinning processing unit 67, as training data. As a result, in the trained optical flow analysis learning device 62a, an algorithm is generated that uses the thinned images of T frame and T-1 frame as input data to infer the optical flow analysis results for the thinned images, and an approximate feature equation (optical flow approximate feature equation) that indicates the features of the optical flow analysis results is obtained in the intermediate layer of the fully connected autoencoder 62-2.
[0186] In Figure 41, machine learning for estimating an extended depth image uses a stereo camera 68, a stereo depth analysis unit 69, a ToF pseudo-degradation processing unit 70, a depth image mathematical encoder 63, and a trained segmentation analysis learner 61a and optical flow analysis learner 62a, as shown. Since distance measurement using the stereo method using the stereo camera 68 has a wider range of distances that can be measured compared to distance measurement using the ToF method, in the machine learning here, the distance measurement results are used as training data for the extended depth analysis learner 66b.
[0187] The left and right images (RGB images) captured by the stereo camera 68 are input to a stereo depth analysis unit 69, which obtains distance measurement results (depth images) using the stereo method. These distance measurement results are thinned out by a thinning processing unit 65 and used as training data for an extended depth analysis learning device 66b.
[0188] The distance measurement results obtained by the stereo depth analysis unit 69 are input to a ToF pseudo-degradation processing unit 70, which gives the degradation that would occur if distance measurement were performed using the ToF method. Specifically, the dynamic range of distance is reduced to the dynamic range of the ToF method (i.e., distance values outside the range of distance measurable by the ToF method are invalidated). The distance measurement results subjected to such ToF pseudo-degradation processing are then thinned out by the thinning processing unit 65 and provided to the input of the depth image mathematical formula encoder 63. This makes it possible to obtain an approximate feature formula equivalent to the approximate feature formula for the depth image obtained by the ToF sensor 4 as the depth image approximate feature formula input from the depth image mathematical formula encoder 63 to the extended depth analysis learning device 66b.
[0189] Further, a thinned image obtained by thinning out either the left or right image captured by the stereo camera 68 by the thinning processing unit 67 is input to the segmentation analysis learning device 61a. Furthermore, the thinned image itself is input to the optical flow analysis learning device 62a as a thinned image of T frame, and a thinned image obtained by delaying the thinned image by a delay device 64 is input as a thinned image of T-1 frame.
[0190] The extended depth analysis learner 66b performs machine learning using, as learning input data, the approximate feature equation for the ToF depth image obtained by the depth image mathematical equation encoder 63, the segment approximate feature equation obtained by the segmentation analysis learner 61a, the optical flow approximate feature equation obtained by the optical flow analysis learner 62a, and IMU information from the IMU 2, and also uses, as training data, the ranging results by the stereo method thinned out by the thinning processor 65 as described above. In this case, machine learning involves control line association learning by sequentially assigning x and y pixel coordinate values to the control lines in order to realize the decoding function described above. By performing this type of machine learning, the trained extended depth analysis learner 66a generates an algorithm that infers an extended depth image based on the various inputs described above and decodes and outputs the distance value at the pixel specified by the control line.
[0191] In the above, an example of inferring an extended depth image, which is an extended image of a depth image, was given as an example of inferring an extended image. However, it is also possible to infer extended images other than extended depth images, such as inferring an extended image with an extended brightness dynamic range compared to the original image for a gradation image such as an RGB image.
[0192] [4-2. Modifications related to extended depth analysis] Here, the extended depth analysis can be redefined as shown in FIG. First, the lowest layer is estimation based on IMU gravity, followed by depth analysis from segmentation analysis information, followed by depth analysis using translational surveying based on optical flow analysis for the long-distance side and depth analysis using iToF (Indirect ToF) for the short-distance side, and then, at a higher layer, depth analysis using a fusion of these. Because optical flow analysis cannot be performed under conditions where translational motion does not occur, depth analysis based on stereo images, depth analysis using dToF (direct ToF), and depth analysis based on changes in focus lens wobbling are located at higher layers to ensure consistently stable depth analysis. However, each method has its own drawbacks: stereo methods can easily produce false positives due to false peaks, dToF has low resolution, and wobbling can only estimate depth from high-frequency components. Therefore, the top layer is depth analysis based on 3D MAP data (3D point cloud data) to enable highly accurate depth analysis. Here, we will mainly assume a use case of SLAM using iToF sensors, and below we will explain various examples targeting the combined use of RGB / iToF / dToF sensors (sensor fusion).
[0193] A wide variety of variations can be realized as a learning machine for inferring augmented depth images. For example, it is possible to select whether or not to include depth images using the ToF method for input, or to support stereo image input. The following describes an extended depth analysis learning device 66aA of a multipurpose interface (hereinafter referred to as a "universal interface") that can infer extended depth images in response to various inputs.
[0194] 43, the extended depth analysis learning device 66aA is configured to be able to handle up to eight input systems, including two control lines. Specifically, the device is configured to be able to handle a total of eight input systems, including the mode control line shown in the figure, two RGB images RGB_0 and RGB_1, ToF, IMU information, a segment approximation feature formula, a depth image approximation feature formula, and a control line for specifying pixel addresses in decoding.
[0195] Examples of input variations are shown in Figures 44 to 47. 44 shows an example of input corresponding to stereo depth analysis, in which left and right RGB images obtained by stereo imaging are input to RGB_0 and RGB_1 as shown, and IMU information, a segment approximation feature equation, and a depth image approximation feature equation are also input. In this case, the depth image is expanded by the stereo method, and input of a depth image by the ToF method is not necessary, so the ToF input system is not used.
[0196] 45 shows an example of input corresponding to iToF depth analysis, and FIG. 46 shows an example of input corresponding to dToF depth analysis. In these cases, the ToF input system receives depth images obtained by the iToF method and depth images obtained by the dToF method, respectively. In these cases, RGB images of the T frame and the T-1 frame are input to RGB_0 and RGB_1, respectively, to enable depth estimation based on optical flow analysis information.
[0197] Figure 47 shows an example of input corresponding to mono depth analysis, i.e., obtaining an extended depth image from a depth image estimated from a monocular RGB image. In this case, to enable depth estimation based on optical flow analysis information, RGB images of the T frame and T-1 frame are input to RGB_0 and RGB_1, respectively, and the ToF input system is not used.
[0198] In order to enable inference of an extended depth image in response to the various input variations described above, the extended depth analysis learner 66aA performs control line association learning using mode control lines. Specifically, to enable response to the input variations illustrated in, for example, FIGS. 44 to 47, first to fourth machine learning processes are performed as described below. That is, as the first machine learning process, supervised learning is performed in which a first value is input to the mode control line and the learning input data is the input data of FIG. 44. As the second machine learning process, supervised learning is performed in which a second value is input to the mode control line and the learning input data is the input data of FIG. 45. Furthermore, as the third machine learning process, supervised learning is performed in which a third value is input to the mode control line and the learning input data is the input data of FIG. 46. As the fourth machine learning process, supervised learning is performed in which a fourth value is input to the mode control line and the learning input data is the input data of FIG. 47. By performing this type of control line association learning, it is possible to realize an inference device with universal I / F specifications that, for the trained extended depth analysis learning device 66aA, can infer an extended depth image corresponding to the input shown in Figure 44 by inputting a first value into the mode control line, can infer an extended depth image corresponding to the input shown in Figure 45 by inputting a second value into the mode control line, can infer an extended depth image corresponding to the input shown in Figure 46 by inputting a third value into the mode control line, and can infer an extended depth image corresponding to the input shown in Figure 47 by inputting a fourth value into the mode control line.
[0199] FIG. 48 is a diagram for explaining SLAM performance when using an extended depth image obtained by the extended depth analysis learning device 66aA with the universal interface specifications as described above. As experimental results related to SLAM performance, Figure 48A shows the results for SLAM using stereo depth images as ground truth data, Figure 48B shows the results for SLAM using iToF depth images, Figure 48C shows the results for SLAM using augmented depth images obtained by the aforementioned Mono (single-eye) depth analysis (see Figure 47), and Figure 48D shows the results for SLAM using augmented depth images obtained by the aforementioned iToF depth analysis (see Figure 45). The experimental results shown here include translation, rotation (quaternion), number of tracks, and the number of aggregated 3D points (number of three-dimensional point clouds) calculated by SLAM.
[0200] In Fig. 48, it is shown that posture estimation cannot be performed normally with the iToF method in Fig. 48B due to a problem with the measurable distance range, while the Mono (single-eye) depth analysis in Fig. 48C and the iToF depth analysis in Fig. 48D both produce results close to Ground Truth. Furthermore, a comparison between these monocular depth analyses and iToF depth analyses shows that iToF depth analyses have slightly better performance.
[0201] <5. Image Recognition Processing Section (Delta Pose Estimation)> [5-1. Configuration and processing of image recognition processing unit] FIG. 49 is a block diagram showing an example of the internal configuration of the image recognition processing unit 10. As shown in the figure, the image recognition processing unit 10 includes a Delta Pose estimator 71a. The Delta Pose estimator 71a is configured with a DNN learner such as a CNN learner, and uses the segment approximation feature equation output from the segmentation analysis learner 61a of the image analysis processing unit 9 as input data to infer an image-derived quaternion.
[0202] Here, the reason why the image-derived quaternion is obtained in the image recognition processing unit 10 is to enable bias removal and phase adjustment of the IMU in the subsequent image synchronization processing unit 5. As will be described later, the image-derived quaternion is used in the image synchronization processing unit 5 to obtain a phase shift and a gyro bias by comparing it with the IMU quaternion.
[0203] FIG. 50 is a diagram illustrating machine learning of the Delta Pose estimator 71a. In the figure, Delta Pose estimator 71b represents Delta Pose estimator 71a before learning. As shown in the figure, Delta Pose estimator 71b has a CNN 71-1 and a fully connected layer 71-2 at the subsequent stage. The purpose is to apply the CNN 71-1 using the segment approximation feature equation as learning input data, and to estimate image-derived quaternions in the fully connected layer 71-2. The quaternions that serve as the teacher in this case are acquired by the posture estimation processing unit 72 in the drawing through SLAM analysis of the left and right images captured by the stereo camera 68. For confirmation, since the segment approximation feature equation is generated based on the thinned image as described above, the resolution corresponds to the thinned resolution (1 / 16 in this example). The CNN 71-1 is provided with a segment approximation feature equation as learning input data, and the fully connected layer 71-1 is provided with quaternions as training data obtained by the pose estimation processing unit 72, and machine learning is performed as supervised learning for the Delta Pose estimator 71b. As a result, an algorithm is generated in the trained Delta Pose estimator 71a that uses the segment approximation feature equation as input data to estimate (infer) image-derived quaternions.
[0204] [5-2. Modified Example of Image Recognition Processing Unit] FIG. 51 is an explanatory diagram of a modified Delta Pose estimator 71Aa. The Delta Pose estimator 71Aa differs from the Delta Pose estimator 71a in that it infers not only an image-derived quaternion but also an image-derived translational velocity. As the image-derived translational velocity, velocities for the three axes X, Y, and Z are inferred. In this case, the input data is the segment approximation feature equation and a quaternion calculated from the detection information of the IMU 2 (hereinafter referred to as "IMU quaternion"). The reason for using the IMU quaternion as input is to improve the accuracy of translational surveying.
[0205] Here, SLAM also performs pose estimation, but the Delta Pose estimator 71Aa in this example is not a PTAM (Parallel Tracking and Mapping) type self-position estimator that analyzes movement based on so-called feature points, but rather estimates image-derived quaternions and translational velocity from the movement (contrast change) of image segments, with the aim of providing a fail-safe when tracking is lost in the PTAM type SLAM processing performed later. Pose estimation by the Delta Pose estimator 71Aa eliminates the concept of tracking stoppage due to feature point loss. Instead, estimation is always performed based on changes in image contrast, contributing to IMU image synchronization and stable SLAM operation. If high-precision SLAM is used to attempt IMU image synchronization, if tracking stops, IMU image synchronization will be lost until it is restored, and SLAM recovery processing cannot be guaranteed. Furthermore, in smartphone environments, IMU phase shifts can easily occur due to temperature increases, CPU load, and CPU interrupts from other applications, disrupting operation during periods when SLAM cannot be restored. Furthermore, recent SLAM functions typically use the so-called VIO (Visual Inertial Odometry) function, which estimates self-position from both the IMU and images. However, the process of analyzing IMU phase shifts using the estimated rotation amount conflicts with the VIO processing, making proper phase synchronization difficult. The Delta Pose estimator 71Aa proposed in this example was designed to address these issues.
[0206] FIG. 52 is a diagram illustrating machine learning of the Delta Pose estimator 71Aa. In the figure, Delta Pose estimator 71Ab represents the Delta Pose estimator 71Aa before learning, and as shown in the figure, Delta Pose estimator 71Ab, like Delta Pose estimator 71b, has a CNN 71-1 and a fully connected layer 71-2 at the subsequent stage. Using the segment approximation feature equation as input data to CNN 71-1 and the IMU quaternion as input data to the fully connected layer 71-2, image-derived quaternions and image-derived translational velocities (three axes) are estimated in the fully connected layer 71-2. In this case, the quaternion and translational velocity (three axes) serving as the teacher are acquired by the posture estimation processing unit 72A in the drawing through SLAM analysis of the left and right images captured by the stereo camera 68. Machine learning is performed as supervised learning for the Delta Pose estimator 71Ab, with the segment approximation feature equation and IMU quaternions as learning input data and the quaternions and translational velocity obtained by the pose estimation processing unit 72 as training data. As a result, an algorithm is generated in the trained Delta Pose estimator 71Aa that estimates (infers) an image-derived quaternion and an image-derived translational velocity using the segment approximation feature equation and IMU quaternions as input data.
[0207] Note that Delta Pose estimator 71Aa uses a segment approximation feature equation based on thinned images as input data, which tends to result in lower measurement accuracy compared to general SLAM, but the impact is small here because it is treated as a self-positioning function aimed at IMU image synchronization and SLAM stabilization. Furthermore, by using it in conjunction with a Global Navigation Satellite System (GNSS) such as a Global Positioning System (GPS), it is possible to ensure a certain degree of self-positioning accuracy, and in use cases that do not necessarily require the retention of absolute coordinates, such as outdoor AR navigation, it is expected to be extremely lightweight and ensure stable operation.
[0208] FIG. 53 is a diagram illustrating the experimental results regarding the estimation accuracy of the image-derived quaternion and image-derived translational velocity by the Delta Pose estimator 71Aa. As a comparative example, Fig. 53A shows the results of estimating translation amounts (three axes) and quaternions (rotation) using SLAM (PTAM), and Fig. 53B shows the results of estimating image-derived translation amounts (three axes) and image-derived quaternions using the Delta Pose estimator 71Aa. The image-derived translation amounts are obtained by integrating the estimated image-derived translational velocity. In the experiment, translation amounts and quaternions were estimated for a movement of approximately 50 m.
[0209] 53A and 53B, it can be seen that the estimation results obtained by SLAM are generally accurate. As for the image-derived quaternions, because it is the so-called VIO method in which IMU quaternions are input, it has enough accuracy to analyze the phase shift between the IMU and the image. However, the Delta Pose estimator 71Aa estimates the amount of movement between frames using DNN from extremely rough image information with 1 / 16 thinning, and calculates the amount of translation by integration. Therefore, due to two factors, the accuracy of the input data and the accuracy of machine learning, errors accumulate over a long period of time, and in the results shown in this figure, an error of about 5 m (about 10%) is observed after 60 seconds. This results in lower accuracy than typical feature-point-based SLAM. However, the Delta Pose estimator 71Aa is designed with consideration for synchronization between the IMU and images by the image synchronization processor 5 that follows. It is intended to constantly measure movement from rough contrast changes in the image to avoid tracking stopping due to feature point loss, as occurs with conventional SLAM. Therefore, the long-term cumulative error described above is acceptable. Essentially, the Delta Pose estimator 71Aa is intended to provide a fail-safe response when feature points are lost, which conventional SLAM cannot handle, and to provide robustness to ensure AR navigation operation in scenes with intense movement, such as outdoor sports. Therefore, the cumulative error is corrected by calibration using GNSS when the use case is outdoors, and by feedback from the relocalization function using SLAM map data when used indoors, thereby contributing to the stabilization of the AR system.
[0210] Fig. 54 is an explanatory diagram of experimental results related to tracking performance using the Delta Pose estimator 71Aa. Here, the results of comparing the number of successful feature point tracking among SLAM using depth images from an iToF sensor, SLAM using extended depth images obtained by the aforementioned Mono (single-eye) depth analysis, SLAM using extended depth images obtained by iToF depth analysis, and SLAM using image-derived quaternions and image-derived translational velocities from the Delta Pose estimator 71Aa are shown for a scene assuming AR navigation (Fig. 54A) and a scene assuming an AR game in a home environment (Fig. 54B). It can be seen that the method using augmented depth images significantly improved tracking performance compared to situations where SLAM using depth images from an iToF sensor stopped tracking. As mentioned above, the method using the Delta Pose estimator 71Aa does not track feature points, but has a simple configuration that estimates the amount of change in translation and rotation (quaternion) between frames from changes in image contrast, so tracking never stopped in the first place.
[0211] FIG. 55 is a diagram illustrating the evaluation result of the RMSE (Root Mean Squared Error) for the translation amount estimated using the Delta Pose estimator 71Aa. Specifically, the evaluation results are shown for SLAM using depth images from an iToF sensor, SLAM using extended depth images obtained by Mono (monocular) depth analysis, SLAM using extended depth images obtained by iToF depth analysis, and Delta Pose estimator 71Aa, using the RMSE of the estimated X-axis translation amount (Figure 55A), Y-axis translation amount (Figure 55B), and Z-axis translation amount (Figure 55C). It should be noted that if tracking is lost, it is difficult to measure RMSE correctly, but in this case it is calculated as a state hold.
[0212] The experimental results confirmed that with the iToF sensor, there were many cases outside the measurement range and translational motion could not be estimated correctly in most scenes, whereas the other methods were generally able to estimate translational motion. In principle, the method using extended depth images compromises measurement accuracy in order to ensure robustness from the perspective of IMU synchronization using thinned images as input, but in reality it achieved a relatively high score. This shows that it is possible to achieve lighter translational motion estimation without requiring high-resolution image input such as HD or 4k.
[0213] <6. Image synchronization processing section (Delta Phase estimation)> FIG. 56 is a block diagram showing an example of the internal configuration of the image synchronization processing unit 5. As shown in FIG. The image synchronization processing unit 5 adjusts the phase of the IMU based on the image-derived quaternion obtained by the image recognition processing unit 10 so that the IMU and the image are synchronized.
[0214] As shown in the figure, the image synchronization processing unit 5 includes a delta phase estimator 81, a phase adjustment unit 88, and a bias removal unit 89. The delta phase estimator 81 includes a quaternion calculation unit 82, an average calculation unit 83, a difference calculation unit 84, a delay score analysis learning unit 85a, LPFs (low pass filters) 86, 86, and a determiner 87. The delay score analysis learning device 85a is configured with a DNN learning device such as a CNN learning device, and has an average-side SAE 85-1, a difference-side SAE 85-2, and a delay analysis layer 85-3a and a score analysis layer 85-4a as fully connected layers subsequent to these SAEs.
[0215] In the image synchronization processing unit 5, the basic processing flow is as follows: first, a quaternion calculation unit 82 calculates a quaternion from the IMU information (three-axis acceleration, three-axis angular velocity) obtained by the IMU 2, and an average calculation unit 83 and a difference calculation unit 84 calculate the average value and difference value, respectively, between this quaternion and the image-derived quaternion from the image recognition processing unit 10. These average value and difference value are provided as input data to a delay score analysis learning device 85a.
[0216] In the delay score analysis learning device 85a, the average-side SAE85-1 uses the above average value as an input, and the difference-side SAE85-2 uses the above difference value as an input, and pre-training and control line association learning are performed. Therefore, the average-side SAE85-1 compresses the dimension of the above average value to obtain an approximate feature equation of the average value in the intermediate layer, and the difference-side SAE85-2 compresses the dimension of the above difference value to obtain an approximate feature equation of the difference value in the intermediate layer.
[0217] The delay analysis layer 85-3a estimates the amount of phase shift with the IMU image based on an approximate feature equation of the average value and an approximate feature equation of the difference value by performing machine learning, which will be described later. The score analysis layer 85-4a estimates a score value indicating the likelihood (reliability of the estimation result) of the amount of phase shift estimated by the delay analysis layer 85-3a based on an approximate feature equation of the average value and an approximate feature equation of the difference value by performing machine learning, which will be described later.
[0218] The phase shift amount estimated in the delay analysis layer 85-3a is input to the decision unit 87 via one LPF 86, and the score value estimated in the score analysis layer 85-4a is input to the decision unit 87 via the other LPF 86. The determiner 87 determines one of three states, delay, neutral, or advance, for the phase shift with the IMU image side, based on the amount of phase shift and the score value input via the LPF 86. Specifically, the determiner 87 basically determines whether the phase shift is delay when the amount of phase shift exceeds a negative (or positive) threshold, whether it is advance when the amount of phase shift exceeds a positive (or negative) threshold, or whether it is neutral if it is within the range of both thresholds, but if the score value is equal to or less than a predetermined value, it continues to use the previous determination result.
[0219] The phase adjustment unit 88 adjusts the phase of each sample value of the 3-axis acceleration and 3-axis angular velocity as IMU information based on the determination result (delay / neutral / advance) by the determiner 87. The phase adjustment process here advances the phase by a predetermined amount if the determination result of the determiner 87 is delay, delays the phase by a predetermined amount if the determination result is advance, and does not adjust the phase if the determination result is neutral. Note that the phase adjustment process in the phase adjustment unit 88 can be a simple process of shifting samples by an integer amount, or sample interpolation using floating-point amounts is also possible, although this increases the calculation cost.
[0220] Based on the image-derived quaternion, the bias removal unit 89 removes bias from the IMU information after the phase adjustment process by the phase adjustment unit 88. Since the bias on the IMU side can be estimated from the difference with the image-derived quaternion, this is removed from the IMU information. The IMU information from which the bias has been removed by the bias removal unit 89 is supplied as image-synchronized IMU information to the attitude control processing unit 6 at the subsequent stage (see FIG. 1).
[0221] Here, the data input to the Delta Phase estimator 81 may be either a sliding window in which the sampling region is sequentially slid, or a fixed window in which the region is divided into fixed block units. In order to recognize the flow of time at fixed intervals, an approximate feature equation is generated from data of, for example, 100 samples or more. Furthermore, the phase adjustment unit 88 performs a phase adjustment process on the IMU samples according to the delay / neutral / advance determination result by the determiner 87. This process determines the direction of the IMU phase shift and shifts the phase of the samples in the corresponding direction, aiming to achieve neutral. The reason for this time-consuming adjustment as a direction determination rather than estimating a specific delay amount is that, while phase shift estimation typically requires an approach similar to pattern matching, pattern matching often leads to false peaks and significant false detections. If an error occurs in the IMU phase adjustment, abnormal distortions will occur when correction processes such as rolling shutter distortion correction are performed later, and similar distortions may also cause false detections in subsequent SLAM processing. To avoid this, an approach that determines delay / neutral / advance by also referencing the score value is adopted. Although this approach cannot be executed immediately upon device startup and causes some delay, it is intended to perform AR operation as stably as possible. Regarding bias anomalies in the IMU, the bias on the IMU side can be removed based on the image, eliminating the IMU bias caused by temperature and improving the performance of the subsequent SLAM.If there is a large discrepancy between the quaternion data estimated from the IMU and the quaternion data estimated from the image, the system can be equipped with a diagnostic function that determines whether there is an abnormality on the IMU side, thereby preventing the subsequent SLAM from operating abnormally.
[0222] The machine learning technique of the Delta Phase estimator 81 will be described with reference to FIGS. FIG. 57 is an explanatory diagram of the learning method of the delay analysis layer 85-3a. In this case, a synchronized set of image-derived quaternions and IMU information is used to prepare the training dataset. The quaternions calculated from the IMU information by the quaternion calculation unit 82 are delayed / advanced by the phase adjustment unit 80 according to random delay / advance instruction information. The quaternions adjusted by the phase adjustment unit 80 are provided to the average calculation unit 83 and the difference calculation unit 84. The average calculation unit 83 calculates the average value with respect to the image-derived quaternion, and the difference calculation unit 84 calculates the difference value with respect to the image-derived quaternion. The average value is input to the average side SAE 85-1, and the difference value is input to the difference side SAE 85-2, and converted into approximate feature equations. These approximate feature equations of the average value and the difference value are then provided as training input data for the delay analysis layer 85-3b before training. The random delay / advance instruction information described above is provided to the training data, and supervised machine learning is performed.
[0223] By performing the above-described machine learning, an algorithm for estimating the phase shift (amount and direction of shift) of the IMU is generated in the trained delay analysis layer 85-3a using as input data the approximate feature equation of the average value obtained in the average-side SAE 85-1 and the approximate feature equation of the difference value obtained in the difference-side SAE 85-2.
[0224] Although not illustrated, in this case, the control line association learning for the average side SAE85-1 and the difference side SAE85-2 involves inputting a sample identification value indicating which sample in the above-mentioned window the data is (for example, if the window consists of 100 samples, a value that identifies the 1st sample to the 100th sample) as the control line value.
[0225] FIG. 58 is an explanatory diagram of the learning method of the score analysis layer 85-4a. First, in this case as well, the method of obtaining the mean value approximation feature equation and the difference value approximation feature equation as learning input data using the quaternion calculation unit 82, the phase adjustment unit 80, the average calculation unit 83, the difference calculation unit 84, the average side SAE 85-1, and the difference side SAE 85-2 based on the delay / advance random instruction information is the same as in the case of FIG. 57. As shown in the figure, a trained delay analysis layer 85-3a is used to obtain an estimation result for the IMU phase shift based on the mean value approximation feature equation and the difference value approximation feature equation, and the estimation result is input to a score calculation unit 91. The score calculation unit 91 inputs true values of the delay / advance (e.g., true values of the phase shift amount and shift direction), compares these true values with the phase shift value estimated by the delay analysis layer 85-3a, calculates a true value of the score value, and provides the true value as training data to the score analysis layer 85-4b before training. As shown in the figure, the learning input data for the score analysis layer 85-4b are the mean value approximation feature equation and the difference value approximation feature equation.
[0226] By performing machine learning in the score analysis layer 85-4b using the learning input data and teacher data as described above, an algorithm is generated in the trained score analysis layer 85-4a that uses the mean value approximation feature equation and the difference value approximation feature equation as input data to estimate a score value for IMU phase shift estimation.
[0227] Here, since the correct answer is unknown when analyzing the amount of IMU delay in a mobile device such as a smartphone, it is extremely difficult to confirm the validity of the estimation results.Therefore, using a dedicated device that synchronizes the image and IMU, we artificially applied a +1 frame delay and a -1 frame delay to conduct an experiment to check whether the delay / advance judgment by the judge 87 was performed correctly. The results are shown in Figures 59 and 60. Figure 59 shows the results of a +1 frame delay, and Figure 60 shows the results of a -1 frame delay. In each figure, "Delay" indicates the result of the IMU phase shift estimation by the delay analysis layer 85-3a, "Score" indicates the result of the score value estimation by the score analysis layer 85-4a, "LPF" indicates the output of the LPF 86, and "Direction" indicates the result of the delay / neutral / advance determination by the determiner 87.
[0228] In principle, the estimation is performed using LPF86 over a period of about 10 seconds, so although the phase shift cannot be estimated correctly immediately after startup, it has been shown that the system gradually becomes able to correctly determine the delay / advance. Although it is difficult to publish all the experimental data due to space limitations, we confirmed that the direction could be correctly determined for approximately 90% of the scenes.
[0229] <7. Attitude control processing unit> FIG. 61 is a block diagram showing an example of the internal configuration of the attitude control processing unit 6. As described above, the attitude control processing unit 6 inputs the image-synchronized IMU information from the image synchronization processing unit 5 and obtains attitude-controlled IMU information, which is IMU information that has been subjected to various attitude control processes, such as centrifugal force removal. Here, examples of IMU information correction processing for attitude control include correction processing for eliminating centrifugal force, correction processing for state machine control, correction processing for adding effects, and correction processing for stabilizer braking.
[0230] In the correction process for removing centrifugal force, a process is performed to remove an offset caused by centrifugal force that occurs in the acceleration quaternion.
[0231] In the correction process for state machine control, correction process of IMU quaternions for state machine control is performed. The state machine control here means control to stop the gimbal function for the display image of the signal processing device 1. The signal processing device 1 in this example has a gimbal function that corrects the Z axis of the CG space in AR so that it coincides with the vertical direction in real space, but the gimbal function should be stopped to ensure stability of attitude control when the camera direction of the signal processing device 1 becomes directly upward or downward, etc. Control that stops the gimbal function in this way is called state machine control here. Actual field tests revealed that user movements are extremely complex, and that a wide variety of AR usage scenarios are expected, making rule-based processing difficult to address. For example, the horizontal correction function implemented using an acceleration sensor requires exception handling when there is no defined horizontal reference when the camera is facing the zenith, as well as various exception handling procedures depending on the camera's intended use, such as cartwheels (roll rotation) and backflips (pitch rotation). While it is conceivable to use threshold-based control to account for the effects of acceleration noise, this is also difficult in practice. For this reason, in this example, AI is used to automatically estimate when the gimbal function should be disabled. This estimation process is realized using a mechanism that concatenates (cascades) mathematical transformations of approximate feature equations.
[0232] Here, to clarify the differences from Patent Document 1, Patent Document 1 employs a method of adjusting the blending ratio of attitude obtained from gyro and acceleration according to the sensor observation conditions. This is a necessary process for horizontal correction to deal with scenes where horizontal correction is not possible due to the principles described above. Furthermore, there are frame restrictions on stabilizer rotation, so processes such as clipping the rotation amount before hitting the edge require limiting processing based on acceleration and gyro information. In contrast, this patent employs a method of transforming gyro and acceleration attitude information into a feature formula that expresses the time flow, and then using DNN formula processing to determine the current scene and transform the control feature formula accordingly, finally decoding the feature formula into a quantized control value.
[0233] In comparison with Patent Document 1, the former employs a numerical calculation method, whereas the present invention is novel in principle in that it encodes quantized sensor information into a feature formula and repeatedly transforms the feature formula using a formula manipulation method, and because the feature formula is developed using a machine learning mechanism, it also possesses the inventive step of being able to acquire a higher learning ability than a data set.Numerical calculation is realized by a combination of arithmetic operations, while formula manipulation is realized by a combination of arithmetic operations and symbolic processing, and numerical calculation and formula manipulation are also academically different.
[0234] Regarding the application of effects, the term "effects" refers to vibration effects used to create a more realistic image. The correction process for applying effects involves correcting the IMU quaternions so that vibrations acting as effects are not removed during stabilization. Specifically, the system learns vibrations that correspond to a sense of realism, and removes the vibration components from the IMU quaternions. This prevents the vibration components from being removed during stabilization, improving the sense of realism.
[0235] The correction process for stabilized braking involves correction to improve camera work. When converting from world coordinates to camera local coordinates, if control is performed using a simple proportional component, the tracking speed will be slow when actively moving the viewpoint in use cases such as when the camera is attached to the user's head. Therefore, IMU quaternion correction is performed to more actively control the posture in a manner similar to the user's viewpoint movement.
[0236] Furthermore, in the posture control processor 6 of this example, in addition to the correction process for posture control regarding the IMU quaternions as described above, it also infers rotation prediction information for one frame later. The inferred rotation prediction information can be used for buffering the input image in the stabilization process, as will be exemplified later.
[0237] Based on the above assumptions, the configuration of the attitude control processor 6 will be described. The posture control processing unit 6 includes an angular velocity quaternion calculation unit 101, an acceleration quaternion calculation unit 102, an SAE 103, a centrifugal force removal correction learner 104a, a state machine correction learner 105a, an effect correction learner 106a, a stabilizer braking correction learner 107a, and a rotation prediction learner 108a.
[0238] The angular velocity quaternion calculation unit 101 calculates angular velocity quaternions based on three-axis angular velocity information from the image-synchronized IMU information obtained by the image synchronization processing unit 5, and the acceleration quaternion calculation unit 102 calculates acceleration quaternions based on three-axis acceleration information from the image-synchronized IMU information obtained by the image synchronization processing unit 5.
[0239] The SAE 103 is an SAE that has undergone pre-training and control line association learning using the angular velocity quaternion calculated by the angular velocity quaternion calculation unit 101 and the acceleration quaternion calculated by the acceleration quaternion calculation unit 102 as inputs, and obtains approximate feature equations indicative of the features of the angular velocity quaternion and the acceleration quaternion in the intermediate layer.
[0240] The centrifugal force elimination correction learner 104a, the state machine correction learner 105a, the effect correction learner 106a, the stabilizer braking correction learner 107a, and the rotation prediction learner 108a are configured with DNN learners such as CNN learners, and each performs a predetermined inference using the approximate feature equation obtained in the previous SAE as input data.
[0241] Specifically, centrifugal force removal correction learning device 104a infers the result of removing the offset caused by centrifugal force from the acceleration quaternion, using as input data the approximate feature equation obtained in the intermediate layer of SAE 103. At this time, the intermediate layer of SAE in centrifugal force removal correction learning device 104a obtains an approximate feature equation that indicates the features of the acceleration quaternion from which the offset caused by centrifugal force has been removed and the input angular velocity quaternion.
[0242] The state machine correction learning device 105a uses the approximate feature equation obtained by the centrifugal force elimination correction learning device 104a as input data to infer a correction result of the IMU quaternion for stopping the gimbal function in accordance with a scene where the above-mentioned state machine control should be executed. At this time, an approximate feature equation indicating the characteristics of the corrected IMU quaternion (acceleration and angular velocity) is obtained in the intermediate layer of the SAE in the state machine correction learning device 105a.
[0243] Here, in this example, the state machine correction learning device 105a performs control line association learning so as to decode and output the inferred angular velocity quaternion. The angular velocity quaternion decoded by the state machine correction learning device 105a is fed back to the angular velocity quaternion calculation unit 101. Based on the angular velocity quaternion thus fed back, the angular velocity quaternion calculation unit 101 performs processing to remove bias that occurs in the angular velocity quaternion calculated from the IMU information, as a measure to avoid the so-called butterfly effect.
[0244] In this example, centrifugal force removal correction learning device 104a and state machine correction learning device 105a are relative value learning modules that learn difference values before and after. On the other hand, effect correction learning device 106a and stabilizer braking correction learning device 107a, which are located after these, correspond to absolute value learning modules that learn absolute values. In this example, a feedback loop that feeds back angular velocity quaternions to angular velocity quaternion calculation unit 101 is formed within a range that includes only the relative value learning modules. This allows angular velocity quaternion calculation unit 101 to appropriately remove bias that occurs in the angular velocity quaternions, thereby avoiding the butterfly effect.
[0245] The effect correction learning device 106a uses the approximate feature equation obtained by the state machine correction learning device 105a as input data to infer the correction result of the IMU quaternion for applying vibration as the above-mentioned effect. At this time, an approximate feature equation indicating the characteristics of the corrected IMU quaternion is obtained in the intermediate layer of SAE in the effect correction learning device 106a.
[0246] The stabilizer braking correction learner 107a uses the approximate feature equation obtained by the effect correction learner 106a as input data to infer the correction result of the IMU quaternion for achieving the stabilizer braking described above. At this time, the intermediate layer of the SAE in the stabilizer braking correction learner 107a obtains an approximate feature equation that indicates the characteristics of the IMU quaternion after correction. The stabilizer braking correction learner 107a is trained on control line association so as to decode and output the corrected IMU quaternion, and as shown in the figure, the decoded output of the stabilizer braking correction learner 107a is output to the outside of the attitude control processing unit 6 as attitude-controlled IMU information.
[0247] The rotation prediction learner 108a uses the approximate feature equation obtained by the stabilizer braking correction learner 107a as input data to infer rotation prediction information as prediction information for the rotation amount one frame later. The rotation prediction learner 108a is trained using control line association so as to decode and output the inferred rotation prediction information. As mentioned above, the decoded rotation prediction information can be used for buffering the input image in the stabilization process.
[0248] With reference to Figures 62 to 66, the machine learning techniques of the centrifugal force elimination correction learner 104a, the state machine correction learner 105a, the effect correction learner 106a, the stabilizer braking correction learner 107a, and the rotation prediction learner 108a will be described.
[0249] FIG. 62 is a diagram illustrating the learning method of the centrifugal force removal correction learning device 104a. As shown in the figure, centrifugal force elimination unit 110 receives the angular velocity quaternion calculated by angular velocity quaternion calculation unit 101 and the acceleration quaternion calculated by acceleration quaternion calculation unit 102, and removes offset caused by centrifugal force from the acceleration quaternion based on the true value of centrifugal force. The acceleration quaternion and angular velocity quaternion after the offset removal are used as training data for centrifugal force elimination correction learner 104b before learning, and the approximate feature equation obtained in the intermediate layer of SAE 103 is used as learning input data to perform machine learning in centrifugal force elimination correction learner 104b. As a result, in centrifugal force removal correction learning device 104a after learning, an algorithm is generated that uses the approximate feature equation obtained in the intermediate layer of SAE 103 as input data to infer the result of removing the offset caused by centrifugal force from the acceleration quaternion, and an approximate feature equation that shows the characteristics of the IMU quaternion from which the offset has been removed is obtained in the intermediate layer of SAE in centrifugal force removal correction learning device 104a.
[0250] Note that angular velocity quaternions are less susceptible to the influence of centrifugal force, while acceleration quaternions are more susceptible to the influence of centrifugal force. Therefore, it is conceivable to use the difference between the angular velocity quaternion and the acceleration quaternion as the true value of the centrifugal force.
[0251] FIG. 63 is a diagram illustrating the learning method of the state machine correction learning device 105a. As shown in the figure, state machine correction unit 111 inputs the angular velocity quaternions calculated by angular velocity quaternion calculation unit 101 and the acceleration quaternions calculated by acceleration quaternion calculation unit 102, and performs correction processing for the above-mentioned state machine control on these input quaternions based on the corresponding scene true value. The acceleration quaternions and angular velocity quaternions after this correction processing are used as teacher data for state machine correction learner 105b before learning, and machine learning is performed by state machine correction learner 105b using the approximate feature equation obtained by centrifugal force removal correction learner 104a as learning input data as shown in the figure. As a result, in the state machine correction learner 105a after learning, an algorithm is generated using the approximate feature equation obtained by the centrifugal force elimination correction learner 104a as input data to infer the correction result of the IMU quaternion for stopping the gimbal function in accordance with the scene where state machine control should be executed, and an approximate feature equation showing the characteristics of the IMU quaternion as the correction result is obtained in the intermediate layer of SAE in the state machine correction learner 105a.
[0252] FIG. 64 is a diagram illustrating the learning method of the effect correction learning device 106a. As shown in the figure, effect correction unit 112 inputs the angular velocity quaternion calculated by angular velocity quaternion calculation unit 101 and the acceleration quaternion calculated by acceleration quaternion calculation unit 102, and performs correction processing on these input quaternions based on the true effect value. Specifically, correction processing is performed to remove the vibration component indicated by the true effect value from the input quaternion. The result of this correction processing is used as training data for effect correction learner 106b before learning, and machine learning is performed in effect correction learner 106b using the approximate feature equation obtained by state machine correction learner 105a as input data for learning as shown in the figure. As a result, in the effect correction learning device 106a after learning, an algorithm is generated that uses the approximate feature equation obtained by the state machine correction learning device 105a as input data to infer the correction result of the IMU quaternion for imparting the above-mentioned effect vibration, and an approximate feature equation that indicates the characteristics of the IMU quaternion as the correction result is obtained in the intermediate layer of SAE in the effect correction learning device 106a.
[0253] FIG. 65 is an explanatory diagram of the learning method of the stabilizer braking correction learning device 107a. As shown in the figure, stabilizer braking correction unit 113 inputs the angular velocity quaternion calculated by angular velocity quaternion calculation unit 101 and the acceleration quaternion calculated by acceleration quaternion calculation unit 102, performs correction processing on these input quaternions based on the stabilizer braking true value, and uses the result of this correction processing as training data for stabilizer braking correction learner 107b before learning.Also, as shown in the figure, the approximate feature equation obtained by effect correction learner 106a is used as learning input data to perform machine learning in stabilizer braking correction learner 107b. As a result, in the stabilizer braking correction learner 107a after learning, an algorithm is generated that uses the approximate feature equation obtained by the effect correction learner 106a as input data to infer the correction result of the IMU quaternion to achieve the stabilizer braking described above, and an approximate feature equation that shows the characteristics of the IMU quaternion as the correction result is obtained in the intermediate layer of SAE in the stabilizer braking correction learner 107a. As described above, the stabilizer braking correction learner 107a performs control line association learning so that it can decode and output an IMU quaternion as a correction result.
[0254] FIG. 66 is a diagram illustrating the learning method of rotation prediction learning device 108a. As shown in the figure, the true value of rotation information one frame later is given as training data for the rotation prediction learner 108b before learning, and the approximate feature equation obtained by the stabilizer braking correction learner 107a is given as input data for learning, and machine learning is performed on the rotation prediction learner 108b. As a result, the learned rotation prediction learner 108a generates an algorithm for inferring rotation prediction information indicating the amount of rotation one frame later, using the approximate feature equation obtained by the stabilizer braking correction learner 107a as input data. As described above, control line association learning is also performed on rotation prediction learner 108a so that rotation prediction information can be decoded and output.
[0255] FIG. 67 shows the results of an experiment related to the performance evaluation of rotation prediction learning device 108a. Specifically, Figure 67 shows the inferred values (solid lines) of rotation prediction information by the rotation prediction learner 108a for each direction of pitch (Figure 67A), yaw (Figure 67B), and roll (Figure 67C) compared with the correct values (dotted lines). From these results, it can be confirmed that rotation predictions are generally equivalent to the correct values, regardless of the rotation direction of pitch, yaw, or roll.
[0256] <8. Variations> Although the signal processing device 1 has been described above as an embodiment, the present technology is not limited to the specific examples described so far, and various modified configurations can be adopted. Various modifications will be described below.
[0257] [8-1. Fusion SLAM using Delta Pose and Delta Phase estimators] FIG. 68 is an explanatory diagram of a configuration example for realizing fusion SLAM to which the above-described Delta Pose estimator 71Aa and Delta Phase estimator 81 are applied. Here, we will explain an example of a configuration for SLAM self-position estimation using a fusion method, which uses the image-derived translational velocity and image-derived quaternion obtained by the Delta Pose estimator 71Aa, the self-position coordinates detected by GNSS (GNSS coordinates in the figure), and the feature point detection results based on the captured image.
[0258] As shown in the figure, in this case, a Delta Pose estimator 71Aa, a GNSS absolute coordinate correction Kalman filter 116, a feature point registration and deviation adjustment unit 117, a Delta Phase estimator 81, and a phase adjustment unit 88 are provided. As described above, the Delta Pose estimator 71Aa (see FIG. 49) receives the segment approximation feature equation and the IMU quaternion as input data, and infers the image-derived translational velocity and the image-derived quaternion. The Delta Phase estimator 81 receives the image-derived quaternion estimated by the Delta Pose estimator 71Aa and IMU information (three axes of angular velocity and three axes of acceleration) as input data, and outputs a state determination result of delay / neutral / advance regarding the phase shift on the IMU side relative to the image. A phase adjustment unit 88 adjusts the phase of the IMU information based on the state determination result, and obtains image-synchronized IMU information. This image-synchronized IMU information is used for AR image processing.
[0259] The image-derived translational velocities (three axes) estimated by the Delta Pose estimator 71Aa are integrated by a translational integration processor 115 and input to a GNSS absolute coordinate correction Kalman filter 116 as image-derived translational coordinates. The GNSS absolute coordinate correction Kalman filter 116 calibrates the input image-derived translation coordinates based on GNSS coordinates. Because image-derived translation information has some accuracy deviation, Kalman filter processing using GNSS coordinates is used to calibrate the translation coordinates to obtain higher accuracy. This calibration processing also functions as error correction for GNSS coordinates. While GNSS can obtain absolute coordinate information, its sample rate is slower than the image frame rate, at around 10 Hz, and errors of 1 m to a maximum of 100 m can occur depending on the surrounding environment. Processing using a Kalman filter can correct such GNSS errors using the image-derived translation coordinate values.
[0260] The feature point registration / shift adjustment unit 117 inputs a captured image as an original resolution image and performs a process of extracting feature points from the captured image. Furthermore, based on the translation coordinates obtained by the GNSS absolute coordinate correction Kalman filter 116 and the image-derived quaternions from the Delta Pose estimator 71Aa, the feature point registration / shift adjustment unit 117 performs a process of registering, in map data, feature points extracted from the previous frame whose positional change in the current frame matches the positional change determined from the translation coordinates and the image-derived quaternions. Furthermore, the feature point registration / shift adjustment unit 117 compares the map data with the feature points extracted in the current frame, and if a positional shift is detected, performs a correction to fill in the absolute value there.
[0261] Conventional SLAM extracts feature points from high-resolution images, analyzes the amount of movement of the feature points between frames to estimate the vehicle's position and orientation, and simultaneously registers the feature points in map data. On the other hand, with the configuration illustrated in Figure 68, the vehicle's position is estimated by fusing thinned images with IMU and GNSS, and in the subsequent processing, the estimated vehicle's position is input, and feature points extracted from high-resolution image frames that match the input vehicle's position are searched for and registered in map data, thereby reducing the number of falsely detected feature points.
[0262] FIG. 69 is an explanatory diagram of an example of a signal processing device that uses an integrated sensor specialized for AR. As mentioned above, we have achieved excellent self-localization by analyzing thinned images (e.g., 32 x 24 pixels) using DNNs as the information source. Therefore, by incorporating a lightweight 32 x 24 pixel image sensor and stacked logic for calculations into IMU2, we can propose a sensor device specifically designed for AR sensing in smartphones. In particular, with conventional SLAM, the image input / output interface consumes a large amount of power due to the increasing resolution of image sensors today. This makes it unrealistic to develop a dedicated SLAM LSI, especially for smartphones. The typical approach has been to use a downstream GPU in the AP layer for SLAM processing. However, GPU resources are shared by multiple applications, which hinders stable AR operation. Furthermore, architectural constraints make it difficult to reduce power consumption by reducing the frame buffer and synchronize with the IMU. Here, by installing an extremely lightweight image sensor inside the IMU, data bandwidth can be kept to a minimum, and because it is very small, multiple sensors can be used even in products with limited volume, such as smartphones, and more accurate AR surveying can be expected.
[0263] Figure 69A shows an overview of a lightweight AR sensor packed with a MEMS (Micro Electro Mechanical Systems) that detects acceleration and angular velocity and a 32 x 24 pixel monochrome sensor. Although it has a very low pixel count compared to recent high-resolution sensors, as in the experiment shown in this example, even with 32 x 24 pixels, it is possible to estimate the self-position to a certain extent, and by providing a function for absolute coordinate correction using GNSS or map data in the subsequent stage, even such a lightweight AR sensor can be fully operational.
[0264] Figure 69B shows an example of a smartphone with compact AR sensors, as shown in Figure 69A, positioned at each of its four corners. Using lenses with as wide an angle as possible reduces the number of devices. While narrow-angle lenses are often used with iToF sensors, tracking is often lost with narrow-angle lenses during panning, so it is desirable to use wide-angle lenses for AR sensing. This method can capture objects from multiple viewpoints, enabling expansion to stereo and even four-viewpoint surveying. The adoption of the universal interface mentioned above allows for analysis of a wide variety of variations. Furthermore, while AR processing using high-resolution image sensors requires the reuse of frame memory, particularly in LSIs, which consumes a large amount of power due to the interface, this method is extremely lightweight (approximately 32 x 24 pixels), allowing it to transmit acceleration and angular velocity data at a transmission bandwidth of less than 1 Mbps. This lightweight processing is possible on the CPU and LSI, reducing power consumption during AR use. AR sensing with distributed compact AR sensors can thus be considered an optimal solution specialized for AR surveying.
[0265] FIG. 70 is a diagram for explaining a configuration example of a stacked logic unit and an AP layer to be provided in the subsequent stage of the small AR sensor illustrated in FIG. 69A. The stacked logic unit here refers to the logic unit built into the compact AR sensor, and the AP layer refers to the application layer operated by a GPU or CPU outside the compact AR sensor. As shown in the figure, the angular velocity, acceleration, and image data (low resolution) obtained by the compact AR sensor are input, and the angular velocity and acceleration are synchronized with the image at the hardware level. The image is encoded into an approximate feature equation by a segmentation analysis learner, and this approximate feature equation and angular velocity are used as input to the Delta Pose estimator to estimate the image-derived translational velocity (IMG) and quaternion (IMG). The quaternion (IMG) estimated from the image is compared with the quaternion (IMU) calculated from the IMU to remove bias, and the bias-calibrated rotational quaternion is sent to the subsequent AP layer. On the other hand, the image-derived translational velocity (IMG) is differentiated to obtain the acceleration estimate (IMG), which is then compared with the acceleration (IMU) obtained from the acceleration sensor to separate it into a gravity component and other components. This takes advantage of the fact that the image-derived acceleration (IMG) does not include gravitational acceleration. As a result, the gravity acceleration component and translational coordinate information are sent to the AP layer. In the AP layer, the self-position is estimated with high accuracy using a Kalman filter of the translation data (translation coordinates) sent from the stacked logic unit and the translation data (translation coordinates) derived from GMSS, and the feature points at this position are registered in the map data.
[0266] In conventional SLAM, the rotation and translation of feature points are calculated and then the feature points are registered in map data, but this method simply estimates the self-position using the Delta Pose estimator on the sensor edge, IMU, and GNSS, and then registers selected feature points that match these conditions in map data.As this approach selects only good feature points and registers them in map data, it is expected to be more robust than conventional SLAM.As errors accumulate with the Delta Pose estimator alone, the amount of deviation is appropriately corrected using GNSS and map data.
[0267] Here, the absolute value coordinate operation and relative value coordinate operation of AR will be described. Generally, most SLAM applications assume absolute coordinates for their own position from the start of the application. This is because most applications, such as AR games and 3D scanning (three-dimensional surveying), require absolute coordinates. On the other hand, AR navigation systems intended for outdoor navigation do not necessarily require absolute coordinates at the start of the AR experience. AR navigation systems require the camera image and CG image to fit naturally without delay or misalignment in response to movement. Furthermore, maintaining absolute coordinates for the AR start position is not required in most cases. One approach is to update absolute values from GNSS when the self-position is lost due to accumulated errors. Furthermore, in AR navigation applications, the operation of returning to a location with map data when tracking is lost is extremely inconvenient, and even more so in AR navigation systems for heavy machinery and construction applications. In this example, we assume a two-step process: absolute position calibration from map data for applications that use absolute coordinates, and absolute value correction from GNSS for applications that use relative coordinates.
[0268] [8-2. Memory control variations in stabilization processing] As explained above, in stabilization processing, the input image is buffered in the buffer memory 42, and then the output image is rendered based on the reference coordinates CR. The buffer memory 42 can be, for example, a frame memory or a line memory. Line memory is generally used in lightweight LSI designs, but due to its limited memory capacity, it can only perform small-scale image stabilization. Frame memory is a common memory in PC and smartphone environments, but because it requires a large number of resources, its use in LSIs is generally limited. Even in smartphone applications, it is often used for GPU processing. Taking these points into consideration, in this example, we consider using a memory that buffers an image of a circular partial area within the input image (hereinafter referred to as a "rotary memory"), rather than a frame memory that buffers the entire input image.
[0269] Rotary memory requires very special memory access, and requires a unique hardware design, including the cache controller. This memory control does not have a garbage collection mechanism, and if stabilization processing continues using the conventional method, the memory will gradually become fragmented. To address this, a memory reference scheduler and a memory release scheduler were designed to minimize memory misses and immediately release memory that is no longer needed.
[0270] Therefore, in this example, by setting a practical upper limit on roll rotation and predetermining the writable area in addition to the memory release schedule method, a garbage collection mechanism is provided in the buffering structure, which makes it possible for both the rendering unit and the buffering unit to refer to and overwrite memory in approximately raster scan order.
[0271] FIG. 71 is an explanatory diagram of memory control of a rotary memory to which elements of garbage collection are applied. First, as described above, this memory control is based on the premise that an upper limit is set for the amount of roll angle rotation between one frame as a specification of the stabilization process. 71A shows the relationship between the buffer area (matt-finished portion) for rendering the Nth frame and the image portion ("Ar-n" in the figure) used for rendering the Nth frame when the input image is 4k x 4k (approximately 4000 x 4000 pixels). The buffer area for the input image is set based on the reference coordinate CR. In the rendering phase of the Nth frame, the first line to be read (output) in the buffer area is determined from the reference coordinate CR of the pixel on the first line of the output coordinate system ("Lt" in the diagram). Once this first line Lt is determined, it is determined that the area above it will not be used for rendering due to the constraint of the upper roll angle limit value mentioned above. Therefore, when rendering of the Nth frame begins, it is possible to allow overwriting of the memory stored above it for the N+1th frame (area "Ao" in the diagram). Conventional methods did not take into account memory areas that are obviously unused like this, and it was necessary to wait for a notification from the memory release scheduler. As shown in Figures 71B and 71C, the input image for the N+1th frame is stored in the memory area that is allowed to be overwritten as described above in approximately raster order. As shown by the scan direction Sn+1 in Figure 71B, the scan direction of the input image for the N+1th frame is shifted from the scan direction of the input image for the Nth frame by the amount of rotation indicated by the rotation prediction information (rotation prediction information for the Nth frame). By controlling memory as described above, although fragments of edge conditions occur due to roll angle rotation, the data is arranged in memory space in a state close to raster scan order, and this series of behaviors results in memory buffering with functionality equivalent to garbage collection.
[0272] In this example, a method is adopted in which AI processing is used to generate a memory lookup table and a memory release table for the buffer memory 42 to achieve the above-described memory control. The memory lookup table here refers to a table that is looked up to identify the buffer area for the input image, i.e., the image area to be buffered in the buffer memory 42. The memory release table refers to a table that indicates the overwritable area in the memory space of the buffer memory 42 when buffering the N+1th frame.
[0273] Referring to Figure 72, an example of a machine learning method for realizing the generation process of these memory lookup tables and memory release tables as AI processing will be described. As shown in the figure, in this case, the machine learning uses the lattice point mesh generating / shaping unit 21′, segment matrix generating unit 22, segment searching unit 23, remesh data generating unit 24, and remesh learning device 27a described above, as well as a memory release table generating unit 121p, a memory lookup table generating unit 122p, a memory release table generating learning device 121b, and a memory lookup table generating learning device 122b.
[0274] As shown in the figure, a memory release table generation unit 121p and a memory reference table generation unit 122p receive input of remesh data (reference coordinates CR at segment granularity) from a remesh data generation unit 24. The memory reference table generation unit 122p generates a memory reference table by rule-based processing based on the remesh data.
[0275] Rotation prediction information is input to the memory release table generation unit 121p together with the remesh data. This rotation prediction information is obtained by the rotation prediction learning device 108a described above (see FIG. 61). The memory release table generating unit 121p generates a memory release table indicating the overwritable areas of the buffer memory 42 on the basis of the remesh data and the rotation prediction information, using rule-based processing, assuming the above-mentioned memory control.
[0276] The memory lookup table generation learning device 122b and the memory release table generation learning device 121b are configured to include a DNN learning device such as a CNN learning device. The memory lookup table generation learner 122b performs machine learning using the approximate feature equation indicating the features of the remesh data obtained by the remesh learner 27a as learning input data and the memory lookup table obtained by the memory lookup table generator 122p as training data. This machine learning involves control line association learning, in which the x and y coordinate values of the output image of the frame to be rendered are used as control line inputs. By performing the above-described machine learning, the trained memory lookup table generation learner 122b (hereinafter referred to as "122a") generates an algorithm for inferring the address of the buffer area in the input image using the approximate feature equation obtained by the remeshing learner 27a as input data. Then, by performing the above-described control line association learning, the memory lookup table generation learner 122a is enabled to decode and output information indicating the address of the buffer area.
[0277] The memory release table generation learner 121b performs machine learning using the approximate feature equation obtained by the remesh learner 27a and the rotation prediction information as learning input data and the memory release table obtained by the memory release table generator 121p as training data. This machine learning also involves control line association learning, for example, in which the x and y coordinate values of the output image of the frame to be rendered are used as control line inputs. By performing such machine learning, the learned memory release table generation learner 121b (hereinafter referred to as "121a") generates an algorithm that infers the address of the overwritable area in the buffer memory 42 using the approximate feature equation obtained by the remeshing learner 27a and rotation prediction information as input data, and by performing the above-mentioned control line association learning, the memory release table generation learner 121a is able to decode and output information indicating the address of the overwritable area.
[0278] FIG. 73 shows an example of a configuration for realizing the above-mentioned memory control using a trained memory lookup table generation learning device 122a and a memory release table generation learning device 121a. As shown in the figure, the memory lookup table generation learner 122a is given the approximate feature equation from the remeshing learner 27a as input data, and the memory release table generation learner 121a is given the approximate feature equation and rotation prediction information from the remeshing learner 27a as input data.
[0279] In this case, the memory control unit 41 controls the buffer memory 42 to buffer the image in the buffer area of the input image based on the address information of the buffer area that is decoded and output in the memory lookup table generation learning device 122a in response to the input on the control line. In this case, the memory control unit 41 controls the buffering of the N+1th frame of the input image in the overwritable area in the buffer memory 42 based on the address information of the overwritable area that is decoded and output in the memory release table generation learning device 121a in response to the input on the control line.
[0280] [8-3. Approximate feature equation analysis method] For AI processing using approximate feature formulas, it is possible to analyze the internal structure of the developed algorithm to determine whether it behaves as hypothesized. Such analysis functions are necessary as debugging functions for AI processing systems that use approximate feature formulas, and in principle involve factorizing the feature formula. Furthermore, in a broad sense, even conventional DNN technology has not been able to directly and effectively utilize the internal features of the network, and in most cases only the learning results of the fully connected layer are utilized. However, in face recognition, for example, internal features are used to create image structural information, such as constructing the eyes, nose, and mouth from image edge information, and then constructing the face from these features. The analysis function proposed in this study is a promising means of effectively utilizing these intermediate products.
[0281] A method for analyzing an approximate feature formula will be described with reference to Figures 74 and 75. Here, an example will be described in which the analysis target is the extended depth analysis learning device 66a described above with reference to Figure 36 and the like. First, as shown in FIG. 74, the extended depth analysis learning device 66a obtains approximate feature equations for the extended depth image in a mode that uses IMU information and a mode that does not use IMU information (a mode in which zero data is input as IMU information). Then, the difference between the approximate feature equations obtained in each mode is calculated by the difference calculation unit 131, and the calculated difference is used as input data for the SAE 132 to perform pre-training of the SAE 132. The approximate feature equation obtained in the intermediate layer of the pre-trained SAE 132 in this manner is an approximate feature equation that mainly indicates the features of the depth data of the component estimated from the IMU information. Hereinafter, this approximate feature equation will be referred to as approximate feature equation De.
[0282] As shown in Fig. 75, the extended depth analysis learning device 66a performs re-learning using this approximate feature equation De as training data. Note that the re-learning is performed in a mode that uses IMU information. Through this re-learning, the extended depth analysis learning device 66a generates an algorithm that uses the segment approximate feature equation, the optical flow approximate feature equation, the thinned depth image approximate feature equation, and the IMU information as input data to infer an approximate feature equation (De) that indicates the features of the depth data of the component estimated from the IMU information. By decoding and analyzing the approximate feature equation De inferred by the extended depth analysis learner 66a after re-learning, it is possible to debug the extended depth analysis learner 66a to determine the extent to which the IMU information contributes to the inference of the extended depth image.
[0283] The above example shows the analysis of the contribution of IMU information. However, by using the same method, the contribution of other input elements (segment approximation feature equation, optical flow approximation feature equation, thinned depth image approximation feature equation) to the inference of the augmented depth image can also be analyzed.
[0284] [8-4. Game Mining] Here, by analyzing the approximate feature formula as described above, it is possible to individually extract only the information required for inference and reduce the weight of the inference-related configuration. Furthermore, it is expected to have effects such as improving inference performance by isolating specific events.
[0285] Here, in order to improve inference accuracy, it is important not only to consider which input elements to use but also to consider the blend ratio at which multiple input elements should be used. In this case, finding a solution for determining which input elements and the blend ratio that are optimal for improving inference accuracy requires analyzing various combinations of input elements and blend ratios, which results in an enormous workload.
[0286] Therefore, here we will explain an example of adopting a technique called "game mining" in this case to find the optimal combination solution. Game mining converts propositions into puzzle models and finds optimal combinations by solving puzzle games. For example, video game players from around the world cooperate to solve puzzle games and try out various combinations. For each combination, a corresponding score is calculated, and this can be confirmed as the score in the puzzle game. If the calculated score is equal to or exceeds a predetermined value, the game is solved, and the combination obtained by solving the game is considered the optimal solution.
[0287] Combinatorial optimization here refers to factoring a trained DNN network as a feature equation. No efficient method for factorizing a trained network using machine learning has been discovered to date. Instead, this study employs a solution approach that leverages human intuition by converting the game into a visually comprehensible puzzle game. This approach can be thought of as converting the potential energy of game players into productive energy. According to estimates, if even 0.001% of the productive energy of game players worldwide could be harnessed, it could potentially exceed the total algorithmic productivity of a large corporation. Game players can be considered "virtual employees" who collaborate with algorithm development. In game mining, virtual employees tackle business problems converted into puzzle game models, and the solvers receive rewards for their work.
[0288] For example, it is possible to anticipate collaboration between IT companies and game players by utilizing recent gaming events with prizes, such as e-sports. The more difficult the business problem, the more collaborators you can get, and the higher the chance of solving it. Furthermore, even if you don't have knowledge of the problem, you can simply rely on your intuition to solve the puzzle, so people of all ages and genders can participate in the gameplay. This is not limited to the stabilizer system, AR, and VR (Virtual Reality) systems envisioned in this project, but can also be applied to fields such as medicine.
[0289] The ultimate goal is to convert the puzzle proposition into a Hamiltonian equation of the quantum Ising model and automatically factorize it using a quantum computer. However, current quantum computers do not have sufficient specifications, and converting events into a puzzle model is currently being attempted as a preparatory step.
[0290] Furthermore, as shown in Figure 1, this project has realized all pipelines in a seamless manner by linking the mathematical processing of feature equations, from IMU adjustment (image synchronization processing unit 5) to attitude control (attitude control processing unit 6), image reference (image reference processing unit 7), image processing (image processing unit 8), image analysis (image analysis processing unit 9), and image recognition (image recognition processing unit 10), with the aim of creating new added value from this in the future through a factorization approach of trained networks.We believe that the approach of using internal features by factorizing trained networks will become a major theme in the academic field of AI algorithm development in the future.
[0291] First, referring to Figure 76, we will explain the combination elements including the blend ratio in the DNN estimator. When control line association learning is performed, 4*n pieces of difference information (hereafter referred to as "extended Wavelet transform information") are obtained for n control lines. It is also necessary to determine 4*(4*n) blend ratio variations for the obtained extended Wavelet transform algorithm. The more complex this network becomes, the greater the number of variations becomes, resulting in an astronomical combinatorial optimization problem to solve. In a trained network, a massive search is required to factorize the feature formulas formed at each layer, from low to high levels of abstraction.
[0292] Based on the above premise, a method for searching for optimal combinatorial solutions using game mining will be explained with reference to Figure 77. Here, it is assumed that the AI processing involves providing the approximate feature equation obtained in the intermediate layer of the extended depth analysis learning device 66a as input data to the fully connected layer 71-2 of the Delta Pose estimator 71A (see FIG. 51) to infer image-derived translational velocity and image-derived quaternion. In this example, IMU information is not used in the inference of the extended depth image.
[0293] For the extended depth analysis learner 66a, the factorization control lines in the figure allow you to specify which of the inputs of the segment approximation feature equation, optical flow approximation feature equation, and depth image approximation feature equation to use (in other words, which input to set as zero information). You can also specify 4*(4*n) blend ratio variations for the extended Wavelet transform algorithm in the network. These input and blend ratio variations are specified according to the gamer's gameplay operations, which are shown in the figure as gamer actions.
[0294] The extended depth analysis learning device 66a sets various combinations of inputs and blend ratios through the work of the game player, indicated as "gamer work" in the figure. Here, the score analysis unit 137 calculates a score value for the set combination. Specifically, the score analysis unit 137 calculates a score value based on the decoded value (decoded value of the extended depth image) of the extended depth analysis learning device 66a and the correct depth value. Here, score analysis is not performed based on the final image-derived translational velocity or image-derived quaternion because the network configuration requires time for these calculations. In puzzle games, in order to reduce the stress caused by image latency, it is required that the screen be refreshed at most 0.1 seconds after the user's controller operation. If the image-derived translational velocity or image-derived quaternion were calculated and the score value were calculated based on these values, the time from the operation to the refresh of the game image generation unit 138 in accordance with the score value would be long, resulting in a decrease in the game's playability. When attempting such a time-consuming parameter search, a penalty process that requires time in the game is imposed. For example, by setting combination conditions in advance and inserting a few seconds of effect processing in the game, the calculation time required to update the puzzle data can be secured, reducing the stress of waiting time for game players. In this way, if puzzle search is performed under different conditions while incurring penalty time, it may be possible to achieve better combination optimization.
[0295] When the score value obtained by the score analysis unit 137 is equal to or greater than a predetermined value, the combination of input and blend ratio when the score value is obtained is identified as the combination that is the optimal solution.
[0296] Here, in the network of the extended depth analysis learning device 66a, in response to the input of approximate feature equations for each of the segments, optical flow, and depth image, an approximate feature equation showing the response of the factorized components related to the segments (hereinafter referred to as the "first factorized approximate feature equation"), an approximate feature equation showing the response of the factorized components related to the optical flow (hereinafter referred to as the "second factorized approximate feature equation"), and an approximate feature equation showing the response of the factorized components related to the depth image (hereinafter referred to as the "third factorized approximate feature equation") are obtained.
[0297] After searching for the optimal solution, the combination of inputs and blend ratios of the extended depth analysis learner 66a is set as the combination of the optimal solution, and the first, second, and third factorization approximate feature equations obtained in the extended depth analysis learner 66a in this state are given as input data to the fully connected layer 71-2 of the Delta Pose estimator 71Aa, as shown in FIG. 78, to infer the image-derived translational velocity and the image-derived quaternion. This makes it possible to improve the accuracy of inferring image-derived translational velocity and image-derived quaternion.
[0298] Figures 79 and 80 are explanatory diagrams of the effects of game mining described above, with Figure 79 showing the inference results for image-derived translation coordinates and image-derived quaternions in the first scene, and Figure 80 showing the inference results for image-derived translation coordinates and image-derived quaternions in the second scene (in this case, the translation coordinates were also obtained by integrating the translation velocity). Figures 79A and 80A show the comparison results between the true and inferred values for the translation coordinates and quaternions before the game was completed, and Figures 79B and 80B show the comparison results between the true and inferred values for the translation coordinates and quaternions after the game was completed.
[0299] As shown in FIG. 79, a reduction in rotation noise was confirmed in the first scene, and as shown in FIG. 80, an improvement in translation accuracy was confirmed in the second scene.
[0300] [8-5.Other] The technology explained above uses a mathematical processing architecture for the majority of embedded system architectures, and its algorithms are an extension of so-called deep learning technology using autoencoders, making it possible to transform mathematical formulas by preparing a learning set of pairs of inputs and expected values. This is a very significant design approach in algorithm design, as it makes it possible to assemble, apply, and develop mathematical formulas without requiring specialized mathematical knowledge. It will be a powerful tool to support the design and development of many program engineers around the world, and will contribute to the advancement of science.
[0301] This technology is not limited to applications in sensing technology, as exemplified above. If observed data and predicted transformation events can be obtained as data, the mathematical relationship between them can be expressed in the form of a mathematical transformation of an approximate feature formula, which can be used in industrial applications. This technology is promising and can contribute not only to sensing technology, but also to a wide range of academic fields, including chemistry, medicine, physics, and astronomy.
[0302] Unlike typical CNN techniques, this technology primarily uses machine learning with autoencoders, and unfortunately, the resulting AI capabilities are inferior to those of CNNs. However, unlike CNNs, which synthesize algorithms from training sets without understanding the underlying logic, this technology has value in that it requires a certain level of computer vision and image processing knowledge, allowing the designer to carefully consider the pipeline structure and build the architecture. This approach is not as computationally intensive as typical CNNs, and it also makes it possible to incorporate multiple functions simultaneously into a single large network by utilizing intermediate products. As a result, this technology makes it possible to build a DNN pipeline with a consistent mathematical processing structure that covers everything from IMU synchronization adjustment to attitude control, image reference, image processing, image analysis, and image recognition. Furthermore, this technology has succeeded in integrating everything from inertial odometry, which estimates self-position from the IMU, to visual odometry, which estimates self-position from images, in a seamless DNN network. By deploying this mathematical processing architecture in this way, it will be possible to provide the market with a highly versatile development framework that can be adapted to a variety of needs, simply by updating the dictionary data and the firmware of the DSP (Digital Signal Processor) and CPU.
[0303] <9. Summary of embodiments> As described above, the signal processing device (same 1) as an embodiment has a first learner including a stacked autoencoder, and after pre-training of the stacked autoencoder using a predetermined type of data as learning input data, control line association learning is performed on the first learner, thereby obtaining approximate feature equations in the intermediate layer of the stacked autoencoder that can be regarded as equations that indicate the characteristics of the predetermined type of data (such as a lattice point mesh approximation curved surface generation unit 39, a 1 / 8-dimensional compression encoder 52a, a Delta Pose estimator 71a, a depth image mathematical formula encoder 63, an average-side SAE85-1, a difference-side SAE85-2, and an SAE103), and an inference processing unit (such as a lens distortion correction learner 32a, a high frequency correction learner 45a, an extended depth analysis learner 66a, a delay analysis layer 85-3a, a score analysis layer 85-4a, and a centrifugal force removal correction learner 104a) that perform predetermined inference using the feature approximation equations as input data. According to the above configuration, inference processing is performed on approximate feature expressions obtained by dimensionally compressing data of a predetermined type using a layered autoencoder. Therefore, the amount of data input to the inference processing unit can be reduced, and the calculation cost for inference can be reduced.
[0304] In the signal processing device according to the embodiment, the predetermined type of data is image-based data. As mentioned above, in this specification, the term "image" broadly refers to information in which a value is indicated for each of a plurality of pixels, each of which has a fixed position on a two-dimensional plane. This concept includes, for example, a gradation image in which the brightness value for each pixel is indicated, a distance image (depth image) in which the value of the distance to the subject is indicated for each pixel, a polarization image in which a value representing the polarization state of incident light is indicated for each pixel, and a thermal image in which a value representing the temperature is indicated for each pixel. According to the above configuration, when performing inference related to an image, the amount of data input to the inference processing unit can be reduced, and the calculation cost for inference can be reduced.
[0305] Furthermore, in the signal processing device according to the embodiment, the inference processing unit (high frequency correction learning device 45a) performs inference on the image processing result image, which is an image that has been subjected to predetermined image processing. This reduces the amount of data input to the inference processing unit when performing inference on an image resulting from image processing such as noise removal and flicker correction. Therefore, it is possible to reduce the calculation cost when performing inference on the image resulting from image processing.
[0306] Furthermore, the signal processing device as an embodiment includes a thinning unit (thinning stabilizer processing unit 47) that thins out an image, and a high-frequency inference unit (high-frequency estimation learner 48a) that infers high-frequency components from the thinned image obtained by the thinning unit, and an approximate feature equation generation unit (1 / 8-dimensional compression encoder 52a) obtains a high-frequency approximate feature equation, which is an approximate feature equation that indicates the features of the high-frequency components in an intermediate layer, by pre-training a stacked autoencoder using the high-frequency components as learning input data, and an inference processing unit (high-frequency correction learner 45a) infers an image resulting from image processing using the high-frequency approximate feature equation as input data. This reduces the amount of data input to the inference processing unit when inferring an image resulting from image processing related to image processing that mainly targets high frequency components, such as noise removal or flicker correction. Therefore, it is possible to reduce the calculation cost when performing image inference related to the correction of high frequency components.
[0307] In addition, in the signal processing device as an embodiment, the thinning unit performs thinning stabilization processing as stabilization processing on the input image to obtain a thinned image of the stabilized image, and is equipped with a synthesis unit (identical 49) that synthesizes the high-frequency components inferred by the high-frequency inference unit with the thinned image. As mentioned above, thinning stabilization processing refers to processing that refers to corresponding pixel values in the input coordinate system in units of pixel groups consisting of a predetermined number of pixels in the output coordinate system, rather than referring to corresponding pixel values in the input coordinate system for each pixel position in the output coordinate system. By performing the stabilization process as a thinning-out stabilization process, the amount of calculation required for the stabilization process can be reduced, and by combining high-frequency components with the thinned image obtained by the thinning-out stabilization process, it is possible to increase the resolution of the thinned image obtained by the thinning-out stabilization process. Therefore, it is possible to reduce the processing load of the stabilization process while suppressing a decrease in resolution of the stabilized image.
[0308] Furthermore, in the signal processing device as an embodiment, the inference processing unit performs inference on the image processing result image using, as input data, a synthesized image of the thinned image and high-frequency components obtained by the synthesis unit and the high-frequency approximate feature equation generated by the approximate feature equation generation unit for the image one frame before the synthesized image. This configuration is particularly suitable for inferring an image resulting from image processing, such as noise removal or flicker correction, based on the difference between high frequency components between adjacent frames.
[0309] Furthermore, the signal processing device as an embodiment is provided with a low-frequency image processing unit (low-frequency image processing unit 57) that performs predetermined image processing on the thinned image, and the synthesis unit synthesizes the high-frequency components inferred by the high-frequency inference unit with the thinned image that has been image-processed by the low-frequency image processing unit (see Figures 33 and 34). This makes it possible to reduce the processing load of image processing, such as shading correction, when performing image processing that mainly targets low-frequency components of an image.
[0310] In the signal processing device according to the embodiment, the inference processing unit (extended depth analysis learning device 66a) performs inference on an extended image, which is an image with more extended information than the input image to the approximate feature equation generating unit. This reduces the amount of data input to the inference processing unit, for example, when inferring an extended image as an image with a dynamic range extended from the input image. Therefore, it is possible to reduce the computational cost when performing inference on an extended image.
[0311] Furthermore, in the signal processing device according to the embodiment, the image is a distance image, and the extended image is a distance image in which the dynamic range of distance is extended more than that of the input image. As mentioned above, the dynamic range of distance corresponds to the range of distances that can be measured by a distance measuring sensor. For example, changing the distance image of a distance measuring sensor whose measurable distance range is 50 cm to 5 m to that of a distance measuring sensor whose measurable distance range is 50 cm to 10 m corresponds to the term "extension" here. In this case, the dynamic range of distance can be expressed as being extended from "50 cm to 5 m range" to "50 cm to 10 m range." According to the above configuration, it is possible to reduce the computational cost when inferring an extended image from a distance image.
[0312] Furthermore, in the signal processing device as an embodiment, the approximate feature equation generation unit (depth image mathematical equation encoder 63) obtains an approximate feature equation for the distance image obtained by the distance measurement sensor, and the inference processing unit (extended depth analysis learner 66a) infers the extended image using as input data the approximate feature equation for the distance image and acceleration information generated based on a detection signal from an acceleration sensor (same 2b) provided in a device equipped with the distance measurement sensor. By using the acceleration information, it is possible to estimate the approximate distance between each position in the image, using the direction of gravity as a reference. Therefore, by inferring an extended image using not only the approximate feature equation of the distance image but also the acceleration information as input data as described above, it is possible to improve the inference accuracy of the extended image.
[0313] In addition, in the signal processing device as an embodiment, the inference processing unit infers an extended image using as input data an approximate feature equation for the distance image, acceleration information, and segment information indicating the results of semantic segmentation performed on an image captured by an image sensor whose sensing range overlaps with that of the ranging sensor. By using segment information, it becomes possible to infer an augmented image based on region information for each subject, such as which region of the image various subjects, such as the ground, the sky, objects on the ground, etc., are present in. For example, for the ground region, the distance to each position within the region can be estimated based on the direction of gravity, and for regions of objects with relatively little thickness, such as a person on the ground, it is possible to estimate that the distance to each position within the region is approximately the same. Therefore, by inferring an extended image using not only the approximate feature equation of the distance image and acceleration information but also segment information as described above, it is possible to further improve the accuracy of inferring an extended image of a distance image.
[0314] Furthermore, in the signal processing device according to the embodiment, the inference processing unit (Delta Pose estimators 71a, 71Aa) performs inference related to image recognition. This reduces the amount of input data to the inference processing unit when performing inferences related to image recognition, such as inferring image-derived motion information (device motion information based on changes in the content of the captured image) or recognizing the area where a specific object exists in an image. Therefore, it is possible to reduce the calculation cost when performing inference related to image recognition.
[0315] Furthermore, in the signal processing device according to the embodiment, the inference processing unit infers motion information derived from an image. This reduces the amount of data input to the inference processing unit when inferring image-derived motion information, that is, device motion information based on changes in the contents of captured images. Therefore, it is possible to reduce the computational cost required to infer image-derived motion information.
[0316] Furthermore, in the signal processing device as an embodiment, the predetermined type of data is motion data indicating motion information of an object obtained from a detection signal of a motion sensor attached to the object to detect the motion of the object, the approximate feature equation generation unit (SAE103) has a stacked autoencoder that is pre-trained using the motion data as learning input data, and using the motion data as input data, obtains an approximate feature equation indicating the features of the motion data in an intermediate layer of the stacked autoencoder, and the inference processing unit () uses the feature approximation equation as input data to infer motion data that has been corrected so as to reduce the influence of a specific event on the motion data. This allows for the inference of corrected motion data that reduces the influence of certain events, such as centrifugal force, on the motion data. Therefore, the accuracy of the motion data can be improved.
[0317] Furthermore, in the signal processing device as an embodiment, one or more other inference processing units are provided at a subsequent stage of the inference processing unit, each performing inference using an approximate feature equation obtained in an intermediate layer of the previous-stage inference processing unit as input data, and the multiple inference processing units are connected in cascade, the motion sensor includes an angular velocity sensor and is equipped with an angular velocity data generation unit (angular velocity quaternion calculation unit 101) that generates angular velocity data as motion data based on a detection signal from the angular velocity sensor, and the angular velocity data inferred by any one of the inference processing units connected in cascade is fed back to the angular velocity data generation unit, and the angular velocity data generation unit performs processing to remove bias occurring in the angular velocity data based on the fed-back angular velocity data (see FIG. 61). This makes it possible to remove the gyro bias and improve the accuracy of the angular velocity data.
[0318] Furthermore, in the signal processing device as an embodiment, as an inference processing system having an inference processing unit, a plurality of inference processing systems each performing different inferences are provided, and an inference processing unit in one inference processing system performs inference using, as input data, an approximate feature formula generated by an approximate feature formula generation unit in another inference processing system (see FIGS. 1, 36, and 49). This makes it possible to share the approximate feature formula generation unit between different inference processing systems. Therefore, when a plurality of inference processing systems are provided and each inference processing system performs a different inference, it is possible to improve the efficiency of the process related to the generation of approximate feature expressions.
[0319] Moreover, a signal processing method as an embodiment is a signal processing method including: an approximate feature equation generation procedure in which the approximate feature equation generation unit has a first learning device including a stacked autoencoder, and the approximate feature equation generation unit has performed control line association learning for the first learning device after pre-training of the stacked autoencoder using a predetermined type of data as learning input data, and the predetermined type of data is provided as input data to obtain an approximate feature equation that can be regarded as an equation that shows the features of the predetermined type of data in an intermediate layer of the stacked autoencoder; and an inference processing procedure in which an inference processing unit has an inference unit that has performed machine learning as supervised learning using the approximate feature equation as learning input data, and performs a predetermined inference using the feature approximation equation as input data. This signal processing method can also provide the same functions and effects as the signal processing device described above.
[0320] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0321] <10. This Technology> The present technology can also be configured as follows. (1) an approximate feature equation generation unit including a first learning device including a layered autoencoder, the approximate feature equation generation unit performing control line association learning on the first learning device after pre-training of the layered autoencoder using a predetermined type of data as learning input data, thereby obtaining an approximate feature equation in an intermediate layer of the layered autoencoder that can be regarded as an equation indicating the features of the predetermined type of data; an inference processing unit that has a second learning device that has undergone machine learning as supervised learning using the approximate feature formula as learning input data, and that performs predetermined inference using the feature approximation formula as input data Signal processing device. (2) The predetermined type of data is image-based data. The signal processing device according to (1) above. (3) The inference processing unit performs inference on an image processing result image as an image that has been subjected to predetermined image processing. The signal processing device according to (2) above. (4) a thinning unit that thins out the image; a high frequency inference unit that infers high frequency components for the thinned image obtained by the thinning unit, the approximate feature equation generation unit obtains a high-frequency approximate feature equation, which is the approximate feature equation indicating a feature of the high-frequency component, in the intermediate layer by pre-training the layered autoencoder using the high-frequency component as learning input data; The inference processing unit performs inference on the image processing result image using the high-frequency approximation feature equation as input data. The signal processing device according to (3) above. (5) the thinning unit performs thinning stabilization processing as stabilization processing on the input image to obtain the thinned image of the stabilized image; a synthesis unit that synthesizes the high-frequency components inferred by the high-frequency inference unit with the thinned image; The signal processing device according to (4) above. (6) The inference processing unit performs inference on the image processing result image using, as input data, a synthesized image of the thinned image and the high-frequency component obtained by the synthesis unit and the high-frequency approximate feature equation generated by the approximate feature equation generation unit for the image one frame before the synthesized image. The signal processing device according to (5) above. (7) a low-frequency image processing unit that performs predetermined image processing on the thinned image, The synthesis unit synthesizes the high-frequency component inferred by the high-frequency inference unit with the thinned image that has been image-processed by the low-frequency image processing unit. The signal processing device according to (5) or (6). (8) The inference processing unit performs inference on an extended image, which is an image in which information is extended more than that of an input image to the approximate feature equation generating unit. The signal processing device according to (2) above. (9) the image is a range image, The extended image is a distance image in which the dynamic range of distance is extended more than that of the input image. The signal processing device according to (8) above. (10) the approximate feature equation generation unit obtains the approximate feature equation for a range image obtained by a range measurement sensor; The inference processing unit infers the augmented image using the approximate feature equation for the distance image and acceleration information generated based on a detection signal from an acceleration sensor provided in a device equipped with the distance measuring sensor as input data. The signal processing device according to (9) above. (11) The inference processing unit infers the augmented image using, as input data, the approximate feature equation for the distance image, the acceleration information, and segment information indicating a result of performing semantic segmentation on an image captured by an image sensor whose sensing range overlaps with that of the distance measuring sensor. The signal processing device according to (10) above. (12) The inference processing unit performs inference related to image recognition. The signal processing device according to (2) above. (13) The inference processing unit infers motion information derived from an image. The signal processing device according to (12) above. (14) the predetermined type of data is motion data indicating motion information of the object obtained from a detection signal of a motion sensor attached to the object and detecting the motion of the object, The approximate feature equation generation unit a layered autoencoder that has been pre-trained using the motion data as learning input data, and the layered autoencoder uses the motion data as input data to obtain the approximate feature equation that indicates features of the motion data in an intermediate layer of the layered autoencoder; The inference processing unit uses the feature approximation formula as input data to infer motion data corrected so as to reduce the influence of a specific event on the motion data. The signal processing device according to (1) above. (15) one or more other inference processing units are provided at a subsequent stage of the inference processing unit, each performing inference using the approximate feature formula obtained in an intermediate layer of the inference processing unit at the previous stage as input data, and the plurality of inference processing units are connected in cascade; the motion sensor includes an angular velocity sensor; an angular velocity data generation unit that generates angular velocity data as the motion data based on the detection signal of the angular velocity sensor; Angular velocity data inferred by any one of the cascade-connected inference processing units is fed back to the angular velocity data generation unit, The angular velocity data generating unit performs a process to remove bias occurring in the angular velocity data based on the fed back angular velocity data. The signal processing device according to (14) above. (16) The inference processing system includes a plurality of inference processing systems each performing a different inference, The inference processing unit in one of the inference processing systems performs inference using the approximate feature expression generated by the approximate feature expression generating unit in another of the inference processing systems as input data. The signal processing device according to any one of (1) to (15). (17) an approximate feature equation generation step of providing the predetermined type of data as input data to the approximate feature equation generation unit having a first learning device including a layered autoencoder, the approximate feature equation generation unit having performed control line association learning for the first learning device after pre-training of the layered autoencoder using the predetermined type of data as learning input data, and obtaining an approximate feature equation that can be regarded as an equation indicating the features of the predetermined type of data in an intermediate layer of the layered autoencoder; an inference processing procedure for performing a predetermined inference using the feature approximation formula as input data by an inference processing unit having a second learner that has undergone machine learning as supervised learning using the approximate feature formula as learning input data; Signal processing methods. [Explanation of symbols]
[0322] 1. Signal Processing Device 2 IMU 2a Acceleration sensor 2b Angular rate sensor 3. Image Sensor 4 ToF sensors 5 Image synchronization processing section 6 Attitude control processing unit 7,7' Image reference processing section 8,8' Image processing section 9. Image analysis processing section 10 Image recognition processing section 11 AR processing section 21,21' Grid point mesh generation and shaping section 22 Segment matrix generation unit 23 Segment Search Unit 24 Remesh data generation unit 25 Pixel coordinate interpolation section CR Reference Coordinates 31 Grid Point Mesh Generator 32 Lens Distortion Corrector 33 Projector 34 Rotator 35 Free curvature perspective projector 36 Scan Controller 37 Mesh mask 38 Grid point reference coordinate calculator 39 Grid point mesh approximation curved surface generation unit 32b, 32a Lens distortion correction learning machine 33b, 33a Projection learner 34b,34a Rotational Learner 35b, 35a Free curvature perspective projection learning device 36b, 36a Scanning control learner 37b,37a Mesh mask learner 26b, 26a Grid point mesh segment matrix transformation learner 27b,27a Remesh learner 28b,28a Remesh extended decoding learner 41' Memory control unit 42 buffer memory 43 Cache Memory 44 Stabilizer interpolation processing unit 45,45' Correction processing section 41 Memory control unit 46 Four corner extraction part 47 Thinning stabilizer processing unit 48a, 48b High frequency estimation learner 49 Synthesis Section 50,51 8x8 buffer 52a, 52b 1 / 8-dimensional compression encoder 53 Delay 55 High frequency correct value calculation unit 45b,45b' High frequency correction learning unit 48a',48a' High frequency estimation learner 56,57 Low-frequency image processing unit 48Aa, 48Ab High frequency estimation learner 58 Packing Processing Section 61a, 61b Segmentation analysis learner 62a, 62b Optical flow analysis learning machine 63 Depth Image Mathematical Encoder 64 Delay Unit 65 Thinning processing unit 66a, 66b Extended depth analysis learner 61r Segmentation Recognizer 62r Optical flow detection unit 62-1 SAE 62-2 Fully connected autoencoder 68 Stereo Camera 69 Stereo Depth Analysis 70 ToF pseudo-deterioration processing unit 66aA Extended Depth Analysis Learning Machine 71a,71b Delta Pose Estimator 71-1 CNN 71-2 Fully connected layer 72 Attitude estimation processing unit 71Aa,71Ab Delta Pose Estimator 72A Attitude and translational velocity estimation processing unit 81 Delta Phase Estimator 82 Quaternion Calculation Unit 82 83 Average calculation section 84 Difference calculation part 85-1 Average side SAE 85-2 Differential side SAE 85-3a, 85-3b Delay analysis layer 85-4a, 85-4b Score analysis layer 86 LPF 87 Judgment device 88 Phase adjustment section 89 Bias removal section 90 Phase adjustment section 91 Score Calculation Section 101 Angular velocity quaternion calculation unit 102 Acceleration Quaternion Calculation Unit 103 SAE 104a, 104b Centrifugal force removal correction learning device 105a, 105b State machine correction learner 106a, 106b Effect correction learner 107a, 107b Stabilizer braking correction learning device 108a, 108b Rotation prediction learner 110 Centrifugal force removal section 111 State machine correction unit 112 Effect Correction Section 113 Stabilizer braking correction unit
Claims
1. an approximate feature equation generation unit including a first learning device including a layered autoencoder, the approximate feature equation generation unit performing control line association learning on the first learning device after pre-training of the layered autoencoder using a predetermined type of data as learning input data, thereby obtaining an approximate feature equation in an intermediate layer of the layered autoencoder that can be regarded as an equation indicating the features of the predetermined type of data; an inference processing unit having a second learning device that has undergone machine learning as supervised learning using the approximate feature formula as learning input data, and that performs predetermined inference using the feature approximation formula as input data Signal processing device.
2. The predetermined type of data is image-based data. The signal processing device according to claim 1 .
3. The inference processing unit performs inference on an image processing result image as an image that has been subjected to predetermined image processing. The signal processing device according to claim 2 .
4. a thinning unit that thins out the image; a high frequency inference unit that infers high frequency components for the thinned image obtained by the thinning unit, the approximate feature equation generation unit pre-trains the layered autoencoder using the high-frequency components as learning input data to obtain a high-frequency approximate feature equation in the intermediate layer, the high-frequency approximate feature equation being the approximate feature equation indicating features of the high-frequency components; The inference processing unit performs inference on the image processing result image using the high-frequency approximation feature equation as input data. The signal processing device according to claim 3 .
5. the thinning unit performs thinning stabilization processing as stabilization processing on the input image to obtain the thinned image of the stabilized image; a synthesis unit that synthesizes the high-frequency components inferred by the high-frequency inference unit with the thinned image; The signal processing device according to claim 4 .
6. The inference processing unit performs inference on the image processing result image using, as input data, a synthesized image of the thinned image and the high-frequency component obtained by the synthesis unit and the high-frequency approximate feature equation generated by the approximate feature equation generation unit for the image one frame before the synthesized image. The signal processing device according to claim 5 .
7. a low-frequency image processing unit that performs predetermined image processing on the thinned image, The synthesis unit synthesizes the high-frequency component inferred by the high-frequency inference unit with the thinned image that has been image-processed by the low-frequency image processing unit. The signal processing device according to claim 5 .
8. The inference processing unit performs inference on an extended image, which is an image in which information is extended more than that of an input image to the approximate feature equation generating unit. The signal processing device according to claim 2 .
9. the image is a range image, The extended image is a distance image in which the dynamic range of distance is extended more than that of the input image. The signal processing device according to claim 8 .
10. the approximate feature equation generation unit obtains the approximate feature equation for a range image obtained by a range measurement sensor; The inference processing unit infers the augmented image using the approximate feature equation for the distance image and acceleration information generated based on a detection signal from an acceleration sensor provided in a device equipped with the distance measuring sensor as input data. The signal processing device according to claim 9 .
11. The inference processing unit infers the augmented image using, as input data, the approximate feature equation for the distance image, the acceleration information, and segment information indicating a result of performing semantic segmentation on an image captured by an image sensor whose sensing range overlaps with that of the distance measuring sensor. The signal processing device according to claim 10.
12. The inference processing unit performs inference related to image recognition. The signal processing device according to claim 2 .
13. The inference processing unit infers motion information derived from an image. The signal processing device according to claim 12.
14. the predetermined type of data is motion data indicating motion information of the object obtained from a detection signal of a motion sensor attached to the object and detecting the motion of the object, The approximate feature equation generation unit a layered autoencoder that has been pre-trained using the motion data as learning input data, and the layered autoencoder uses the motion data as input data to obtain the approximate feature equation that indicates features of the motion data in an intermediate layer of the layered autoencoder; The inference processing unit uses the feature approximation formula as input data to infer motion data corrected so as to reduce the influence of a specific event on the motion data. The signal processing device according to claim 1 .
15. one or more other inference processing units are provided at a subsequent stage of the inference processing unit, each performing inference using the approximate feature formula obtained in an intermediate layer of the inference processing unit at the previous stage as input data, and the plurality of inference processing units are connected in tandem; the motion sensor includes an angular velocity sensor; an angular velocity data generation unit that generates angular velocity data as the motion data based on the detection signal of the angular velocity sensor; Angular velocity data inferred by any one of the cascade-connected inference processing units is fed back to the angular velocity data generation unit, The angular velocity data generating unit performs a process to remove bias occurring in the angular velocity data based on the fed back angular velocity data. The signal processing device according to claim 14.
16. The inference processing system includes a plurality of inference processing systems each performing a different inference, The inference processing unit in one of the inference processing systems performs inference using the approximate feature expression generated by the approximate feature expression generating unit in another of the inference processing systems as input data. The signal processing device according to claim 1 .
17. an approximate feature equation generation step of providing the predetermined type of data as input data to the approximate feature equation generation unit having a first learning device including a layered autoencoder, the approximate feature equation generation unit having performed control line association learning for the first learning device after pre-training of the layered autoencoder using the predetermined type of data as learning input data, and obtaining an approximate feature equation that can be regarded as an equation indicating the features of the predetermined type of data in an intermediate layer of the layered autoencoder; an inference processing procedure for performing a predetermined inference using the feature approximation formula as input data by an inference processing unit having a second learner that has undergone machine learning as supervised learning using the approximate feature formula as learning input data; Signal processing methods.
Citation Information
Patent Citations
Imaging apparatus, image processing apparatus, and image processing method
JP2014066995A