An autonomous driving device based on a photonic convolutional reservoir of multi-channel optically injected lasers

Through the photon convolution reserve pool computing device of the multi-channel light injection laser, the emission laser of the delayed feedback loop is used to replace the fully connected layer, combining ridge regression and winner-take-all strategies, it solves the difficulty of training and gradient problems in autonomous driving, and improves the information processing rate and accuracy.

CN116861963BActive Publication Date: 2025-07-08XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310580135.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-22
Publication Date
2025-07-08
Estimated Expiration
2043-05-22

AI Technical Summary

Technical Problem

In the existing autonomous driving technology, the training parameters of convolutional neural networks are large and the loss function is non-convex, which leads to high training difficulties and difficult hardware implementation, and there are problems of gradient vanishing and gradient explosion, affecting efficiency and accuracy.

Method used

The photon convolution reserve pool computing device based on a multi-channel light injection laser is adopted, and the fully connected layer is replaced by a transmitting laser of a delayed feedback loop, combining ridge regression and winner-take-all strategies to avoid gradient disappearance and explosion and improve data processing rate.

Benefits of technology

It effectively solves the problems of gradient disappearance and explosion, improves the information processing rate and accuracy of autonomous driving, and shows good voice command recognition and environmental perception performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116861963B_ABST
    Figure CN116861963B_ABST
Patent Text Reader

Abstract

The present invention provides an autonomous driving device based on a multi-channel optically injected laser photon convolution reservoir. Driving information is obtained from the image acquisition device of the vehicle through an information acquisition device, and an identification device identifies the driving information through a pre-trained target recognition and classification network to obtain a category. The information to be identified in the present invention obtains a feature vector through an input layer, and the non-linear transient response of the reservoir is output through the reservoir. Then, the non-linear transient response is combined through an output layer and the output weight is calculated to obtain a target category. Since the present invention uses an emitting laser with a delay feedback loop to replace the traditional fully connected layer, the defects of gradient disappearance and gradient explosion are avoided. In addition, the multi-injection emitting laser used in the present invention can improve the data processing rate, and the post-processing method combining ridge regression and the winner-takes-all strategy is used to obtain the category to which the information to be identified belongs. Therefore, the present invention exhibits good performance for speech command recognition and autonomous driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to an autonomous driving device based on a multi-channel optically injected laser photon convolution reservoir. Background Art

[0002] With the rapid development of technology, autonomous driving provides a new service mode and a new experience for social development. However, the most core and difficult problem that autonomous driving needs to solve currently is perception. At present, the industry mainly focuses on multi-sensor fusion solutions dominated by vision and technical solutions dominated by lidar with other sensors as auxiliary.

[0003] At present, in autonomous driving, object detection and tracking based on images are mainly achieved through convolutional neural networks, and there are mainly algorithms such as two-stage detection, single-stage detection, and Transformer detection.

[0004] The two-stage detection specifically includes two steps: extracting the object region and classification recognition. In the first stage, a region proposal network is used to generate candidate boxes based on the feature map. In the second stage, a fully connected layer is used to achieve refined classification and regression. Compared with the two-stage algorithm, the single-stage detection only needs to perform feature extraction once to achieve object detection, and has a faster detection speed. The Transformer detection introduces the attention mechanism into the field of object detection, models the relationship between different objects, and integrates relationship information into the features to achieve the purpose of feature enhancement.

[0005] Due to the problems of large number of training parameters and non-convex loss function in the convolutional neural network widely used in the existing environmental perception tasks in the field of autonomous driving, there are problems of large training difficulty and difficult hardware implementation. For example, to process more complex image content, the number of layers of the neural network needs to be continuously deepened. However, as the number of layers of the neural network deepens, the optimization function is prone to fall into local optimal solutions, and the problems of gradient dispersion and gradient explosion are more prominent during the training process, resulting in low efficiency and accuracy of autonomous driving. Summary of the Invention

[0006] In order to solve the above problems existing in the prior art, the present invention provides a photon convolution reservoir computing device with a multi-channel optically injected laser for autonomous driving. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0007] The present invention provides an autonomous driving device based on a multi-channel optically injected laser photon convolution reservoir, including:

[0008] An information acquisition device, configured to acquire driving information from an image acquisition device of a vehicle, and process the driving information to obtain information to be recognized;

[0009] Among them, the driving information includes environmental images and / or voice information around the vehicle;

[0010] An identification device for identifying the category of the information to be identified through a trained target recognition and classification network of its own;

[0011] Among them, the target recognition and classification network includes an input layer, a reservoir layer, and an output layer. The information to be identified obtains a feature vector through the input layer, outputs the non-linear transient response of the reservoir layer through the reservoir layer, and then combines the non-linear transient response and calculates the output weight through the output layer to obtain the category of the information to be identified.

[0012] The present invention provides an autonomous driving device based on a multi-channel optically injected laser photon convolution reservoir. The driving information is obtained from the image acquisition device of the vehicle through an information acquisition device, and the driving information is processed to obtain the information to be identified; the identification device identifies the category of the information to be identified through a trained target recognition and classification network of its own. The target recognition and classification network of the present invention includes an input layer, a reservoir layer, and an output layer. The information to be identified obtains a feature vector through the input layer, outputs the non-linear transient response of the reservoir layer through the reservoir layer, and then combines the non-linear transient response and calculates the output weight through the output layer to obtain the category of the information to be identified. Since the present invention uses a transmitting laser with a delayed feedback loop to replace the traditional fully connected layer, the defects of gradient disappearance and gradient explosion are avoided. In addition, the multiple-injection transmitting laser used in the present invention can improve the data processing rate, and the category to which the information to be identified belongs is obtained by combining the post-processing method of ridge regression and the winner-takes-all strategy. Therefore, the present invention shows good performance for voice command recognition and autonomous driving.

[0013] The following will further describe the present invention in detail with reference to the drawings and embodiments. Description of the Drawings

[0014] Figure 1 is a schematic diagram of an autonomous driving device based on a multi-channel optically injected laser photon convolution reservoir provided by the present invention;

[0015] Figure 2 is a schematic diagram of the structure of a target recognition and classification network provided by the present invention;

[0016] Figure 3 is a schematic diagram of the process of changing the data dimension during the process of the convolution preprocessing module extracting the eigenvalue of the original data;

[0017] Figure 4 is a schematic diagram of the process of training the target recognition and classification network with environmental images provided by the present invention;

[0018] Figure 5It is a schematic diagram of the process of training the target recognition and classification network for voice information provided by the present invention. Detailed implementation manners

[0019] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0020] As Figure 1 shown, the present invention provides an autonomous driving device based on a multi-channel optically injected laser photon convolution reservoir, including:

[0021] An information acquisition device, configured to acquire driving information from an image acquisition device of a vehicle, and process the driving information to obtain information to be recognized;

[0022] Wherein, the driving information includes environmental images and / or voice information around the vehicle;

[0023] The information acquisition device of the present invention is specifically configured to: acquire driving information from an image acquisition device of a vehicle; when the driving information includes an environmental image, convert the environmental image into a picture with a size of 28×28 pixels; when the driving information includes a voice command, convert the voice command into a two-dimensional matrix with a size of 86×P, where P is the length of voice data, and convert the 86×P matrix into a two-dimensional matrix with a size of 28×28, and normalize the values to between [0, 1]; determine the 8×28 pixel picture and / or the normalized result as the information to be recognized.

[0024] A recognition device, configured to recognize the information to be recognized through a trained target recognition and classification network of its own to obtain the category of the information to be recognized;

[0025] Referring to Figure 2 , the target recognition and classification network includes an input layer, a reservoir, and an output layer. The information to be recognized obtains a feature vector through the input layer, outputs the non-linear transient response of the reservoir through the reservoir, and then combines the non-linear transient response and calculates the output weight through the output layer to obtain the category of the information to be recognized.

[0026] The input layer includes a convolution preprocessing module and a masking processing module;

[0027] Wherein, the convolution preprocessing module is configured to first perform convolution on the driving information, then perform downsampling of data through average pooling, and then perform non-linear activation to obtain a one-dimensional feature vector; the masking processing module is configured to multiply the one-dimensional feature vector by a masking matrix to obtain a multi-channel input signal, and inject it into the reservoir through a Mach-Zehnder modulator.

[0028] The convolution preprocessing module includes five parts: Cov1, Cov2, Pool1, Pool2, and Sigmod;

[0029] Among them, convolution is performed using Cov1 and Cov2, Pool1 and Pool2 implement downsampling of data through average pooling, and Sigmod is used for non-linear activation;

[0030] Cov1 and Cov2 apply 6 and 12 5×5 convolutional kernels respectively. Each convolutional kernel slides one pixel each time, and Sigmod uses the simgod function to process the output results of Cov1 and Cov2; Pool1 and Pool2 use a 2×2 pooling kernel as the minimum unit.

[0031] Pool1 and Pool2 significantly improve the test accuracy of the system while further reducing the data size and processing time. A one-dimensional vector u(t) with 192 elements is generated through convolutional preprocessing. Figure 3 Details show the dimensional changes of data during the convolutional preprocessing process. The original image 2(a) in the figure is a 28×28 matrix, which is converted into Figure 3 the six 24×24 matrices shown in figure (b) through Cov1 and Sigmod. Pool1 is for downsampling, which converts the data into 6 12×12 matrices, as shown in figure (c) of figure 3. Similarly, after Cov2 and Sigmod, 12 8×8 matrices are obtained ( Figure 3 figure (d) in the figure). Finally, 12 4×4 matrices as shown in Figure 3 figure (e) in the figure are obtained by Pool2. After flattening the Figure 3 matrix shown in figure (e) in the figure, a one-dimensional vector u(t) containing 192 elements shown in figure Figure 3 (f) in the figure is obtained. At the same time, the weights of the convolutional layers used in the convolutional preprocessing module are obtained by training Cov1 and Cov2 using the backpropagation algorithm.

[0032] The present invention uses convolutional preprocessing to extract the features of the original data, reduces the amount of data required to be processed by the system, effectively improves the system information processing rate and reduces the system power consumption.

[0033] The mask preprocessing module of the present invention multiplies the one-dimensional eigenvalue vector u(t) by the mask matrix m(t) to generate the input signal S(t). Therefore, the input signal S(t) is a random linear combination of the one-dimensional eigenvalue vector u(t) and the mask matrix. m(t) is a 192×virtual node number matrix, as Figure 3As shown in (a). The mask matrix m(t) is of size 192×the number of virtual nodes, and the number of virtual nodes is equal to the ratio of the total delay time of the delay feedback loop used in the reservoir to the sampling interval. In the present invention, the input layer divides the input data into F paths and injects it into the reservoir composed of the VCSEL and the delay loop through a Mach-Zehnder modulator. The present invention adopts the method of multi-path injection into the VCSEL, which improves the information processing rate of the system.

[0034] The reservoir includes a vertical cavity surface emitting laser and a delay feedback loop;

[0035] Among them, the degree of freedom of the emitting laser is increased through the delay feedback loop; multiple input signals generate a non-linear transient response under the action of the degree of freedom and polarization components through the emitting laser and are transmitted to the output layer.

[0036] The reservoir of the present invention is as shown in Figure 2 Figure (b) therein, and the VCSEL with a feedback loop is used as a non-linear node in the reservoir. In addition, under appropriate operating conditions, two orthogonal polarization components (referred to as X polarization (X-PC) and Y polarization (Y-PC)) can coexist in the VCSEL simultaneously, thereby generating a richer non-linear dynamic state. Here we collect the transient X1, X2...X of the virtual nodes in the X-PC F .

[0037] The non-linear transient response is combined in sequence through the output layer, and the combined result is sampled at a certain sampling interval to obtain a state matrix; and the output layer determines the category with the largest output weight as the category to which the information to be recognized belongs according to the state matrix and the corresponding output weights, using the winner-takes-all strategy.

[0038] The output layer in this example is as shown in Figure 2 Figure (c) therein, and a post-processing method combining ridge regression and the winner-takes-all strategy is specifically adopted to obtain the final result from the non-linear transient response result state(n).

[0039] The training process of the target recognition and classification network of the present invention includes:

[0040] a. Obtain autonomous driving images and voices with prior information and form a training set;

[0041] Among them, the prior information indicates the true category to which the samples in the training set belong;

[0042] b. Perform convolution processing on each sample in the training set through a convolution preprocessing module to convert the sample into a one-dimensional feature vector;

[0043] c. Multiply the one-dimensional feature vector by the corresponding mask matrix through the mask preprocessing module to obtain multiple input signals, and inject them into the vertical-cavity surface-emitting laser

[0044] d. Through the vertical-cavity surface-emitting laser, generate a non-linear transient response under the action of degrees of freedom and polarization components, and transmit it to the output layer; where the degrees of freedom are generated by the delay feedback loop for the emitting laser

[0045] e. Sample the non-linear transient response at a predetermined time interval through the output layer to obtain a state matrix, and use the ridge regression algorithm to calculate the output weight between the state matrix and the output vector with the true category to which the sample belongs as the target

[0046] The relationship among the state matrix, output vector, and output weight in e is expressed as

[0047] y(i) = W(i)state(i);

[0048] Among them, the output weight is expressed as W(i), the output vector is expressed as y(i), y(n) is a one-dimensional column vector including N elements, N is the number of categories, and i represents the category serial number

[0049] Among them, the output vector represents the predicted category to which the sample belongs

[0050] d. Repeat b to e for each sample, and compare the predicted category represented by the output vector with the true category. If they are inconsistent, adjust and return to e to recalculate the output weight

[0051] Obtain the environmental image around the vehicle from the image acquisition device on the vehicle body, and adjust the environmental image to a picture of 28×28 pixels and use it as the image to be recognized

[0052] Reference Figure 4 , the specific implementation of the present invention to obtain a trained target recognition and classification network using the environmental image includes: recognizing the image to be recognized through the output weight matrix obtained through the training process to obtain the category of the image to be recognized

[0053] The specific training process is as follows: Use several pictures of all target categories included in the prior information as the training set, and perform convolution preprocessing on all pictures used for training through the input layer of the reservoir to convert the training pictures into one-dimensional feature vectors u(t) with 192 elements. Further, in the mask preprocessing task, divide the one-dimensional feature vector into F paths of signals, and multiply them by mask1, mask2... mask F , to obtain input signals S1(t), S2(t)... S F (t). Further, the input signal S 1(t), S2(t)...S F (t) is injected into the VCSEL and the nonlinear transient responses X1, X2...X are collected F , and after combination and sampling, state(n) is obtained. Finally, the output weight matrix W obtained by the ridge regression algorithm is multiplied by state(n) to obtain the output vector y(n), where y(n) is a 1×N vector and the winner-takes-all strategy is used to obtain the final recognition result, and N is the number of categories.

[0054] Reference Figure 5 , the implementation of the present invention to obtain a trained target recognition and classification network using voice commands specifically includes:

[0055] The voice information is acquired by using the voice command acquisition module in the vehicle, and the Lyon model is used to convert the voice command into a two-dimensional matrix of 86×P, where P is the length of the voice. Further, the 86×P matrix is converted into a two-dimensional matrix of size 28×28, and the values are normalized to between [0,1]. And this 28×28 two-dimensional matrix is used as the voice command to be recognized.

[0056] The voice command to be recognized is recognized through the output weight matrix obtained in the training process to obtain the category of the voice command to be recognized.

[0057] In the training process, several signals of all voice commands included in the prior information are used to form a training set, and the remaining training process is the same as the training process in the environmental perception scheme of the self-driving image based on multi-channel optical injection VCSEL provided by the implementation of the present invention.

[0058] The training scheme of the present invention avoids the problems of a large number of training parameters, high training difficulty, and easy occurrence of the loss function falling into a local optimal solution, gradient dispersion, and gradient explosion in the neural network during the training process. It can show good performance in voice command recognition and self-driving.

[0059] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0060] Although the present application has been described in connection with various embodiments, those skilled in the art will recognize other variations of the disclosed embodiments while practicing the claimed present application by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the singular "a" or "an" does not exclude a plurality.

[0061] The above is a further detailed description of the present invention in connection with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited only to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as falling within the protection scope of the present invention.

Claims

1. An autonomous driving device based on a photonic convolutional reservoir of a multi-channel optically injected laser, characterized in that, Including: An information acquisition device, configured to acquire driving information from an image acquisition device of a vehicle, and process the driving information to obtain information to be recognized; wherein the driving information includes environmental images around the vehicle and / or voice information; An identification device, configured to identify the information to be recognized through a trained target recognition and classification network of its own to obtain the category of the information to be recognized; wherein the target recognition and classification network includes an input layer, a reservoir layer, and an output layer. The information to be recognized obtains a feature vector through the input layer, outputs the non-linear transient response of the reservoir layer through the reservoir layer, and then combines the non-linear transient response and calculates the output weight through the output layer to obtain the category of the information to be recognized; The input layer includes a convolution preprocessing module and a masking processing module; wherein the convolution preprocessing module is configured to first perform convolution on the driving information, then perform downsampling of data through average pooling, and then perform non-linear activation to obtain a one-dimensional feature vector; the masking processing module is configured to multiply the one-dimensional feature vector by a masking matrix to obtain a multi-channel input signal, and inject it into the reservoir layer through a Mach-Zehnder modulator; The reservoir layer includes a vertical cavity surface emitting laser and a delay feedback loop; wherein the delay feedback loop increases the degree of freedom for the emitting laser; the multi-channel input signal generates a non-linear transient response under the action of the degree of freedom and polarization components through the emitting laser, and is transmitted to the output layer; The non-linear transient response is combined according to time series through the output layer, and the combined result is sampled at a certain sampling interval to obtain a state matrix; and the output layer determines the category with the largest output weight as the category to which the information to be recognized belongs by using the winner-takes-all strategy based on the state matrix and the corresponding output weight.

2. The autonomous driving device based on a photonic convolutional reservoir of a multi-channel optically injected laser according to claim 1, characterized in that, The convolution preprocessing module includes five parts: Cov1, Cov2, Pool1, Pool2, and Sigmod; wherein Cov1 and Cov2 are used for convolution, Pool1 and Pool2 perform downsampling of data through average pooling, and Sigmod is used for non-linear activation; Cov1 and Cov2 respectively apply 6 and 12 5×5 convolution kernels, each convolution kernel slides one pixel each time, and Sigmod uses the simgod function to process the output results of Cov1 and Cov2; Pool1 and Pool2 use a 2×2 pooling kernel as the minimum unit.

3. The autonomous driving device of the photonic convolutional reservoir based on a multi-channel optically injected laser according to claim 1, wherein The training process of the target recognition and classification network includes: a. Acquire self-driving images and voices with prior information, and form a training set; wherein the prior information represents the true category to which the samples in the training set belong; b. Perform convolution processing on each sample in the training set through the convolution preprocessing module to convert the sample into a one-dimensional feature vector; c. Multiply the one-dimensional feature vector by the corresponding masking matrix through the masking preprocessing module to obtain a multi-channel input signal, and inject it into the vertical cavity surface emitting laser d. Generate a non-linear transient response through the vertical cavity surface emitting laser under the action of the degree of freedom and polarization components, and transmit it to the output layer; wherein the degree of freedom is generated by the delay feedback loop for the emitting laser; e. Sampling the non-linear transient response by the output layer at a predetermined time interval to obtain a state matrix, and using the ridge regression algorithm to calculate the output weight between the state matrix and the output vector with the true category to which the sample belongs as the target; wherein, the output vector represents the predicted category to which the sample belongs; d. Repeat steps b to e for each sample, and compare the predicted category represented by the output vector with the true category. If they are inconsistent, adjust and return to step e to recalculate the output weight.

4. The autonomous driving device of a photonic convolutional reservoir based on a multi-channel optically injected laser according to claim 3, wherein The relationship among the state matrix, the output vector, and the output weight in step e is expressed as: y(i) = W(i)state(i); wherein, the output weight is expressed as W(i), the output vector is expressed as y(i), y(n) is a one-dimensional column vector including n elements, n is the number of categories, and i represents the category serial number.

5. The autonomous driving device of the photonic convolutional reservoir based on a multi-channel optically injected laser according to claim 1, wherein The information acquisition device is specifically configured to: Obtain driving information from the image acquisition device of the vehicle; When the driving information includes an environmental image, convert the environmental image into a picture with a size of 28×28 pixels; When the driving information includes a voice command, convert the voice command into a two-dimensional matrix with a size of 86×P, and then convert the 86×P matrix into a two-dimensional matrix with a size of 28×28, and normalize the values to between [0, 1]; Determine the 28×28 pixel picture and / or the normalized result as the information to be recognized.

Citation Information

Patent Citations

  • Action video recognition method based on single-node photon reserve pool calculation

    CN113343813A

  • Depth reserve pool calculation system and method based on spinning vertical cavity surface emitting laser

    CN115145535A