Imaging recognition system, training method and training device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]由于识别系统的硬件开发和软件开发一般由不同的部门完成,硬件工程师追求成像的较高质量,软件工程师追求图像的准确分类,不同部门分别在各自的领域追求局部最优解,但这不一定是整个识别系统的全局最优解,因而现有的识别系统的设计效率不一定是最优的,削弱了企业的产品竞争力,结合识别系统在各行业的大量使用,造成了社会资源的大量浪费
[0041]与现有技术相比,本发明具有以下有益效果:该成像识别系统以优化识别对象的识别准确率为目标,统筹成像模块和识别模块的参数设计,追求全局效率最优的参数,所以相对于现有技术的分别追求局部最优,一方面设计效率更高,另一方面由于成像模块不需要追求局部最优,所以成像模块的成本相对于现有技术可以更低,即该识别系统的成本可以更低,所以该识别系统通过软硬一体的设计方式,实现了设计效率更高、成本更低的效果,增强了产品的竞争力,也有利于节约社会的资源。
Smart Images

Figure CN117935227B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of optical system design, and more particularly to an imaging recognition system, its training method, and training device. Background Technology
[0002] There is a significant demand for image recognition and classification across various industries, such as healthcare, security, and public transportation. Recognition systems that can accurately, quickly, efficiently, and cost-effectively identify data in diverse scenarios are crucial for the development of these industries.
[0003] In the process of developing this invention, the inventors discovered the following problems with existing identification systems:
[0004] Because the hardware and software development of recognition systems are generally completed by different departments, hardware engineers pursue high image quality, while software engineers pursue accurate image classification. Different departments pursue local optimal solutions in their respective fields, but this is not necessarily the global optimal solution for the entire recognition system. Therefore, the design efficiency of existing recognition systems may not be optimal, which weakens the product competitiveness of enterprises. Combined with the large-scale use of recognition systems in various industries, this has resulted in a large waste of social resources. Summary of the Invention
[0005] In order to solve at least one of the above-mentioned problems in the prior art, the present invention aims to provide an imaging recognition system, training method and training device that can identify and classify more efficiently and at a lower cost.
[0006] To achieve the above-mentioned objective, one embodiment of the present invention provides an imaging recognition system, comprising:
[0007] An imaging module, which includes hardware for imaging;
[0008] The recognition module includes a recognition neural network, which acquires the image on the imaging surface and outputs the recognition result. The hardware parameters of the imaging module are set according to the recognition accuracy of the recognition neural network.
[0009] As a further improvement of the present invention, the imaging module includes an aperture, a metasurface lens and an imaging surface arranged in sequence. Light passes through the aperture and is imaged on the imaging surface by the metasurface lens. The metasurface lens includes a substrate and a plurality of primitives on the substrate.
[0010] The number of metasurface lenses is set to one.
[0011] As a further improvement of the present invention, the hardware parameters include volume parameters, wherein the volume parameter is a minimum volume, and the minimum volume is the volume that satisfies the recognition accuracy of the recognition neural network within the target range of recognition accuracy.
[0012] As a further improvement of the present invention, the volume parameters include lens length parameters and metasurface area parameters, wherein the lens length parameter is the shortest length, and the metasurface area parameter is the minimum area, and the shortest length and the minimum area are the length and area that satisfy the recognition accuracy of the recognition neural network to be within the target range of recognition accuracy.
[0013] As a further improvement of the present invention, the hardware parameters include performance parameters, wherein the performance parameters are minimum performance parameters, wherein the minimum performance parameters are the performance parameters that satisfy the recognition accuracy of the recognition neural network to be within the target recognition accuracy range, and the performance parameters include transfer function parameters, field of view parameters, and focal number parameters.
[0014] As a further improvement of the present invention, the transfer function parameter of the imaging module is not less than 20% in the spatial frequency domain at a full field of view of 10 lp / mm.
[0015] As a further improvement of the present invention, the imaging module is used to acquire an image of the driver inside the car cockpit, and the recognition neural network is used to recognize the driver's driving behavior.
[0016] To achieve one of the above-mentioned objectives, an embodiment of the present invention provides a training method for an imaging recognition system, comprising the following steps:
[0017] Step S10: Obtain a set of real images, wherein each image in the set of real images has its corresponding real classification information;
[0018] Step S20: Input the real image set into the imaging simulation model to generate a simulated image set, wherein the imaging simulation model is used to simulate the image generated by the imaging module from the real image;
[0019] Step S30: Use the simulated image set as the training set of the recognition neural network, use the difference between the predicted classification information obtained by the recognition neural network from the simulated image set and the real classification information as the loss function, use the parameters of the imaging simulation model as trainable parameters, train the recognition neural network, update the parameters of the imaging simulation model during the training process, update the simulated image set after the imaging simulation model parameters are updated, and continue to train the recognition neural network.
[0020] Step S40: After training is completed, the parameters of the recognition neural network and the parameters of the imaging module are obtained.
[0021] As a further improvement of the present invention, the step of updating the imaging simulation model parameters during training includes:
[0022] Determine the recognition accuracy of the recognition neural network;
[0023] If the recognition accuracy is within the target range, reduce the setting parameters of the imaging simulation model to obtain an updated imaging simulation model, and repeat steps S20 to S30.
[0024] As a further improvement of the present invention, the step of training the recognition neural network further includes:
[0025] If the recognition accuracy of the recognition neural network cannot exceed the minimum recognition rate threshold, the setting parameters of the imaging simulation model are increased to obtain an updated imaging simulation model, and steps S20 to S30 are repeated.
[0026] If the setting parameters of the simulation model are further reduced, the corresponding recognition accuracy cannot converge to the target recognition accuracy range. Then, the lowest parameter of the setting parameters of the imaging simulation model within the target recognition accuracy range is determined as the required parameter, and the recognition neural network training is completed.
[0027] As a further improvement of the present invention, the imaging module includes a metasurface lens, and the setting parameters include volume parameters and / or performance parameters. The volume parameters include lens length parameters and / or metasurface area parameters, and the performance parameters include transfer function parameters and / or field of view parameters and / or focal number parameters.
[0028] As a further improvement of the present invention, the real image set is an image of the driver inside the car cockpit, the real classification information is a classification of the driver's driving behavior, and the recognition neural network is used to identify the driver's driving behavior.
[0029] As a further improvement of the present invention, the images in the real image set correspond to multiple real classification information;
[0030] The step of training the recognition neural network includes:
[0031] Multiple images are randomly selected from the simulated image set as a training set to train the recognition neural network.
[0032] To achieve one of the above-mentioned objectives, an embodiment of the present invention provides a training device for an imaging recognition system, comprising:
[0033] The acquisition module is used to acquire a set of real images, wherein each image in the set of real images has its corresponding real classification information;
[0034] The simulation module is used to input the real image set into the imaging simulation model to generate a simulated image set, wherein the imaging simulation model is used to simulate the image generated by the imaging module from the real images;
[0035] The training module is used to use the simulated image set as the training set of the recognition neural network, use the difference between the predicted classification information obtained by the recognition neural network from the simulated image set and the real classification information as the loss function, use the parameters of the imaging simulation model as trainable parameters, train the recognition neural network, update the parameters of the imaging simulation model during the training process, update the simulated image set after the imaging simulation model parameters are updated, and continue to train the recognition neural network.
[0036] The generation module is used to obtain the parameters of the recognition neural network and the parameters of the imaging module after training is completed.
[0037] To achieve one of the above-mentioned objectives, one embodiment of the present invention provides an electronic device, comprising:
[0038] Storage module, used to store computer programs;
[0039] The processing module, when executing the computer program, can implement the steps in the training method of the imaging recognition system described above.
[0040] To achieve one of the above-mentioned objectives, one embodiment of the present invention provides a readable storage medium storing a computer program that, when executed by a processing module, can implement the steps in the training method of the above-mentioned imaging recognition system.
[0041] Compared with existing technologies, the present invention has the following beneficial effects: This imaging recognition system aims to optimize the recognition accuracy of the object, coordinates the parameter design of the imaging module and the recognition module, and pursues the parameters with the best global efficiency. Therefore, compared with the existing technologies that pursue local optima separately, the design efficiency is higher. On the other hand, since the imaging module does not need to pursue local optima, the cost of the imaging module can be lower than that of the existing technologies. That is, the cost of the recognition system can be lower. Therefore, the recognition system achieves higher design efficiency and lower cost through the integrated hardware and software design, which enhances the competitiveness of the product and is also conducive to saving social resources. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of an imaging recognition system according to an embodiment of the present invention;
[0043] Figure 2 This is a schematic diagram of the imaging module according to an embodiment of the present invention;
[0044] Figure 3 This is a schematic diagram of the structure of a metasurface lens according to an embodiment of the present invention;
[0045] Figure 4 This is a flowchart of a training method for an imaging recognition system according to an embodiment of the present invention;
[0046] Figure 5 This is a flowchart of the training process of a recognition neural network according to an embodiment of the present invention;
[0047] Figure 6 This is a schematic diagram of the training device of an imaging recognition system according to an embodiment of the present invention;
[0048] Among them, 100 is the imaging recognition system; 10 is the imaging module; 11 is the aperture; 12 is the metasurface lens; 121 is the substrate; 122 is the element; 13 is the filter; 14 is the imaging surface; and 20 is the recognition module. Detailed Implementation
[0049] The present invention will now be described in detail with reference to the specific embodiments shown in the accompanying drawings. However, these embodiments do not limit the present invention, and any structural, methodological, or functional modifications made by those skilled in the art based on these embodiments are included within the scope of protection of the present invention.
[0050] One embodiment of the present invention provides an imaging recognition system, training method, and training device that can identify and classify objects more efficiently and at a lower cost.
[0051] Image recognition system
[0052] An imaging recognition system according to this embodiment, such as Figure 1 As shown, it includes an imaging module 10 and a recognition module 20. The imaging module 10 is used to acquire images of the current scene, and the recognition module 20 is used to recognize the images of the imaging module 10 and obtain the recognition results corresponding to the images.
[0053] Taking the application of the recognition system in the transportation field, especially for detecting abnormal driving behavior in the car cockpit, as an example, the imaging module 10 is used to acquire images of the driver inside the car cockpit, and the recognition neural network is used to identify the driver's driving behavior. This embodiment can address scenarios where, with the continuous improvement of the level of vehicle driving automation, drivers no longer need to focus on driving operations for extended periods, giving them more free time. Some drivers may fail to react promptly to dangerous situations due to a lack of focus. Through this imaging recognition system 100, the driver's behavior inside the car cockpit can be identified, and timely warnings can be issued for abnormal behaviors, such as reminding the driver of fatigued driving, using a mobile phone, drinking water, fixing hair and makeup, reaching into the back seat, or turning their head to talk to other passengers. This allows the driver to quickly return to a focused driving state, thereby improving driving safety and reducing traffic accidents.
[0054] Imaging module 10 includes hardware for imaging, which can be a conventional imaging module, including an aperture, a lens module composed of several conventional lenses, and an imaging surface, or... Figure 2 As shown, the imaging module 10 includes an aperture 11, a metasurface lens 12, and an imaging surface 14 arranged sequentially. Light passes through the aperture 11 and is imaged on the imaging surface 14 by the metasurface lens 12. In this embodiment, the imaging module 10 using a metasurface lens 12 is used as an example for explanation. The aperture 11 is used to collect light field information, the metasurface lens 12 is used to transmit the light field onto the imaging surface 14, and the imaging surface 14 is used to complete photoelectric conversion and record the distribution of light intensity. The imaging surface 14 can be a CMOS (Complementary Metal Oxide Semiconductor).
[0055] In this embodiment, the imaging module 10 can be equipped with only one metasurface lens 12. On the one hand, since the imaging module 10 does not use a conventional lens, i.e., a lens that relies on factors such as thickness, shape, and refractive index of the material to control the focal point, the imaging module 10 using the metasurface lens 12 can be made thinner and smaller than existing imaging modules. On the other hand, based on the training method described below, the imaging module 10 using only one metasurface lens 12 can meet the recognition rate requirements with a minimal volume.
[0056] Specific structural parameters of metasurface lens 12 Figure 3 As shown, the metasurface lens 12 includes a substrate 121 and a plurality of primitives 122 on the substrate 121. These primitives 122 have a structure composed of subwavelength scatterers arranged periodically or quasi-periodically. By selecting the material of the metasurface lens 12 and designing the parameters of the substrate 121 and the primitives 122, the amplitude and phase of light can be controlled by the metasurface lens 12 to achieve the expected optical response.
[0057] Furthermore, in this embodiment, the number of metasurface lenses 12 can be set to only one, which minimizes the volume of the entire imaging module 10. The specific reasons for setting it to one without affecting the recognition result are explained below.
[0058] Another example Figure 2 As shown, the imaging module 10 may also include a filter 13 to filter out unwanted wavelengths in the light.
[0059] The recognition module 20 includes a recognition neural network, which acquires the image on the imaging surface 14 and outputs the recognition result. The hardware parameters of the imaging module 10 are set according to the recognition accuracy of the recognition neural network.
[0060] The recognition neural network can be set as a Convolutional Neural Network (CNN) for accurate image classification. Hardware parameters can include volume parameters and performance parameters, and different types of parameters can be selected depending on the application requirements. The hardware parameters of the imaging module 10 are set according to the recognition accuracy of the recognition neural network. That is to say, the hardware parameters of the imaging module 10 do not depend on the imaging quality of the images it captures, but on the final recognition accuracy of the recognition neural network. This achieves a collaborative design between software and hardware, pursuing the global optimal solution of the entire recognition system rather than the optimal solution of the hardware or software design.
[0061] Furthermore, the performance parameters of the imaging module 10 are generally positively correlated with its price, imaging quality, and volume. In other words, higher cost and larger size generally result in better imaging quality, while lower cost and smaller size generally result in lower imaging quality. Since the imaging module 10 does not need to pursue local optima, meaning its imaging quality can be relatively lowered to meet only the target accuracy range, the hardware parameters can be further set based on the recognition accuracy of the neural network, and further based on the requirement that the recognition accuracy falls within the target accuracy range. The target accuracy range can be determined based on actual needs, for example, 99.5% to 100%. This embodiment aims to maintain a relatively high accuracy while pursuing lower values for the performance and volume parameters of the imaging module 10. Alternatively, the target accuracy range can be any range that meets the recognition requirements but does not need to be extremely high.
[0062] Therefore, in general, this embodiment minimizes the cost and size of the recognition system by ensuring that the hardware parameters of the imaging module 10 meet the requirement that the recognition accuracy of the recognition neural network is within the target accuracy range. Considering the scenario described above where the recognition system is applied to detect abnormal driving behavior inside a car's cockpit, reducing the size of the recognition system, given the limited space inside a car, facilitates its installation and use within the vehicle, lowers its cost, promotes its widespread adoption in more cars, and enhances the product's competitiveness.
[0063] Furthermore, the volume parameter is the minimum volume, which is the volume that satisfies the recognition accuracy of the recognition neural network within the target range of recognition accuracy.
[0064] The volume parameters include lens length parameters and metasurface area parameters. The lens length parameter is the shortest length, and the metasurface area parameter is the minimum area. The shortest length and the minimum area are the length and area that satisfy the recognition accuracy of the recognition neural network within the target range of recognition accuracy.
[0065] The lens length parameter can be the length between aperture 11 and imaging plane 14. The smaller this length, the smaller the volume of imaging module 10 can be. A smaller metasurface area results in lower performance and lower cost, and consequently, a smaller volume of imaging module 10. In pursuing a recognition accuracy rate within the target range for the recognition neural network, both the lens length parameter and the metasurface area parameter can be adjusted towards smaller values, thereby reducing the overall volume of imaging module 10. Specific adjustment methods are described in the training method section below.
[0066] Furthermore, the performance parameter is the minimum performance, wherein the minimum performance is the performance that satisfies the recognition accuracy of the recognition neural network within the target recognition accuracy range.
[0067] Based on the above, lower performance generally means lower cost. Adjusting the performance parameters to a smaller value can reduce the overall cost of the imaging module 10.
[0068] The performance parameters include transfer function parameters, field of view parameters, and focal number parameters. These parameters are key indicators in the design of the imaging module 10, affecting its performance and application. Different parameter settings can be used for different application scenarios to achieve the desired imaging effect.
[0069] The transfer function parameter is used to measure the ability of an optical imaging system to capture image details at different spatial frequencies. The transfer function represents the proportion of details of different frequencies in the image that are transmitted by the system. A high transfer function value indicates that the system can transmit high-frequency details well, thus obtaining a clear image. Conversely, a low transfer function value corresponds to a relatively blurry image.
[0070] In this implementation, the transfer function parameters are not less than 20% in the spatial frequency domain at a full field of view of 10 lp / mm.
[0071] The field of view parameter is the angular range of objects or scenes that the imaging system can capture. It determines how wide an area the imaging module 10 can see. A larger field of view means that the camera can capture a wider scene, while a smaller field of view means that the imaging module 10 captures a smaller area.
[0072] The focal number parameter is the ratio of the lens diameter to the focal length in an imaging system. A smaller focal number means a larger aperture, allowing more light to pass through, but also a shallower depth of field, meaning only a few objects are in focus. A larger focal number means a smaller aperture, allowing less light to pass through, but also a relatively larger depth of field, where objects remain relatively sharp both in front of and behind the focal point.
[0073] In conjunction with the above, although larger values for the transfer function parameter, field of view parameter, and focal number parameter generally correspond to better performance, this embodiment does not aim to maximize these parameters. Instead, during the training of the recognition neural network, these parameters are designed in a direction that meets the requirements for recognition accuracy. Furthermore, these parameters can be their respective performance parameters when the recognition accuracy is within the target range.
[0074] Training methods for imaging recognition systems
[0075] The following is a training method for the imaging recognition system described above. Specifically, it is tailored to the required application scenario. During the training of the recognition neural network, the parameters of the recognition neural network and the imaging module are optimized to obtain a recognition system that meets the requirements of the desired scenario.
[0076] The following is combined Figures 4-5 This invention provides a training method for an imaging recognition system according to an embodiment of the present invention. Although the present application provides method operation steps as shown in the following embodiments or flowcharts, the execution order of these steps is not limited to the execution order provided in the embodiments of the present application, based on conventional or non-creative labor, where there is no necessary causal relationship in the logical process.
[0077] Specifically, the training method for the imaging recognition system provided in this embodiment includes the following steps S10 to S40:
[0078] Step S10: Obtain a set of real images, wherein each image in the set of real images has its corresponding real classification information.
[0079] Taking the above-described recognition system as an example of detecting abnormal driving behavior inside a car's cockpit, the images in the real image set can be images of the driver inside the car's cockpit. The real classification information is a classification of the driver's driving behavior. The images in the real image set correspond to multiple real classification information, which may include categories such as: driver fatigue, using a mobile phone, drinking water, fixing hair and makeup, reaching into the back seat, turning head to talk to other passengers, etc. Correspondingly, the recognition neural network is used to identify the driver's driving behavior.
[0080] Step S20: Input the real image set into the imaging simulation model to generate a simulated image set, wherein the imaging simulation model is used to simulate the image generated by the imaging module from the real image.
[0081] In this embodiment, the imaging module including a metasurface lens is used as an example for illustration. The imaging simulation model can be a model based on the imaging module with the metasurface lens, established in optical product design and simulation software. This model can simulate the corresponding image generated by the imaging module based on the input real image. The imaging module can be simulated based on this imaging simulation model, without the need to produce different physical imaging modules to iterate the imaging module parameters.
[0082] The volume and performance parameters of the imaging module mentioned above can more specifically include lens length parameters, metasurface area parameters, transfer function parameters, field of view parameters, and focal number parameters.
[0083] In addition, parameters such as light source parameters, aperture parameters, CMOS parameters, and substrate parameters of metasurface lenses can be set in optical product design and simulation software. More specific parameters include the waveform and wavelength of the light source, the material and size of the substrate, the distance between the metasurface lens and the light source and the imaging surface, etc. Other parameters can also be set according to the application scenario, such as the required pixel size and resolution.
[0084] The final imaging simulation model can simulate the imaging effect of the imaging module in the scene in which it is used. Through its image simulation function, it can convert the real image set into a simulated image set obtained by the imaging module.
[0085] In addition, the images in the simulated image set can be converted into tensor form and normalized to facilitate their import into the recognition neural network for training in subsequent steps.
[0086] Step S30: Use the simulated image set as the training set of the recognition neural network, use the difference between the predicted classification information obtained by the recognition neural network from the simulated image set and the real classification information as the loss function, use the parameters of the imaging simulation model as trainable parameters, train the recognition neural network, update the parameters of the imaging simulation model during the training process, update the simulated image set after the imaging simulation model parameters are updated, and continue to train the recognition neural network.
[0087] Before training the recognition neural network, the preliminary preparations may also include steps S301 to S304:
[0088] Step S301: Construction of the recognition neural network. The recognition neural network can be set as a convolutional neural network with a total of 8 layers: Layer 1 is a convolutional layer; Layers 2 and 3 contain convolutional blocks (kernel function 3, stride 1) + batch normalization + convolutional blocks (kernel function 3, stride 1) + batch normalization + short-circuit connections; Layers 4-7 contain convolutional blocks (kernel function 3, stride 2) + batch normalization + convolutional blocks (kernel function 3, stride 2) + batch normalization + short-circuit connections; Layer 8 is a linear layer. All short-circuit connections are convolutional blocks (kernel function 1, stride 1) + batch normalization.
[0089] Step S302: Initialization of the recognition neural network. This specifically includes random orthogonal initialization of the weight parameters in the recognition neural network. The randomly generated weight matrix is orthogonalized to ensure that the initial weights are independent and orthogonal. This initial recognition neural network can, to some extent, avoid redundancy and overfitting problems in the subsequent training process.
[0090] Step S303: Define the loss function. The loss function calculates the loss value, which guides the learning and adjustment of the model. In this embodiment, the loss function can be the cross-entropy loss function, the expression of which is:
[0091] H(p,q)=-∑ x p(x)logq(x),
[0092] Where H(p,q) is the cross-entropy between the distribution p of the true classification information and the distribution q of the predicted classification information. P(x) is the true probability of class x in the data category, and q(x) is the probability of class x in the predicted classification information.
[0093] Step S303: Optimizer selection. In this embodiment, the optimizer is selected as Stochastic Gradient Descent (SGD), which randomly selects a small batch of labeled data during each training process, calculates the gradient of the loss function, and uses it to update the weight parameters.
[0094] Step S304: Setting parameters for dynamic adjustment of the learning rate. Reducing the model's learning rate after each training period can prevent overfitting after multiple training sessions. In this embodiment, a learning rate scheduler (lrscheduler) can be used to adjust the learning rate.
[0095] Then, the training of the recognition neural network begins. During training, the weight parameters of the recognition neural network and the parameters of the imaging module are trained simultaneously. Figure 5 As shown, it includes the following steps:
[0096] Step S31: Input the simulated image set into the recognition neural network.
[0097] Step S32: Train the recognition neural network using a loss function.
[0098] Step S33: Determine if the loss value meets a preset condition. The preset condition may be that the loss value converges to a set threshold.
[0099] If the loss value does not meet the preset conditions, you can continue to step S331: update the weights in the recognition neural network, and then return to step S31 to repeat the above loop until the loss value meets the preset conditions.
[0100] After the initial training of the recognition neural network described above is completed, the following steps can be continued:
[0101] Step S34: Update the imaging simulation model parameters. The parameter update can be based on the judgment of the recognition accuracy of the recognition neural network. Specifically, it can be based on whether the recognition accuracy of the recognition neural network is within the target recognition accuracy range.
[0102] Step S34 specifically includes:
[0103] Step S341: If the recognition accuracy cannot be greater than the minimum recognition rate threshold, increase the setting parameters of the imaging simulation model to obtain an updated imaging simulation model, and then repeat steps S20 to S30.
[0104] Specifically, the minimum recognition rate threshold must be less than or equal to the minimum value of the target recognition accuracy range. For example, if the target recognition accuracy range is 99.5% to 100% as mentioned above, the minimum recognition rate threshold could be 95% or 99.5%. If the minimum recognition rate threshold is equal to the minimum value of the target recognition accuracy range, the adjustment goal is to directly bring the recognition accuracy within the target range. If the minimum recognition rate threshold is less than the minimum value of the target recognition accuracy range, other optimization methods can be used to adjust the parameters between the minimum recognition rate threshold and the minimum value of the target recognition accuracy range, such as optimizing the scene or adjusting the model's hyperparameters.
[0105] For example, in step S341, if the current recognition accuracy is 93%, it can be determined that the recognition accuracy is too low. The parameters of the imaging module can be further increased to generate a new set of simulated images and retrain the recognition neural network so that the recognition accuracy can be greater than the minimum recognition rate threshold.
[0106] Step S342: If the recognition accuracy is within the target range, reduce the setting parameters of the imaging simulation model to obtain an updated imaging simulation model, and repeat steps S20 to S30.
[0107] Here, the target range for recognition accuracy, as mentioned above, is the range of 99.5% to 100%. For example, if the current recognition accuracy is 99.8%, the parameters of the imaging module can be further reduced to generate a new set of simulated images and retrain the recognition neural network.
[0108] Step S343: If the setting parameters of the simulation model are further reduced, the corresponding recognition accuracy cannot converge to the target recognition accuracy range. Then, the lowest parameter of the setting parameters of the imaging simulation model within the target recognition accuracy range is determined as the required parameter, and the recognition neural network training is completed.
[0109] For example, if the settings of the current imaging simulation model are further reduced, the corresponding recognition accuracy cannot be maintained in the range of 99.5% to 100%. Therefore, the current settings are the lowest parameters that can be reduced to meet the target range of recognition accuracy.
[0110] As described above, the setting parameters include volume parameters and / or performance parameters. The volume parameters include lens length parameters and / or metasurface area parameters. The performance parameters include transfer function parameters and / or field of view parameters and / or focal number parameters. One or more of these five parameters—volume parameters, performance parameters, transfer function parameters, field of view parameters, and focal number parameters—can be adjusted, and the adjustment logic is as described above regarding the specific meanings of each parameter. Through the above steps, after the recognition neural network is trained, each parameter can pursue a suitable value that meets the recognition accuracy requirements; this value is the globally optimal value.
[0111] Furthermore, during each training session, multiple images are randomly selected from the simulated image set as the training set to train the recognition neural network.
[0112] Multiple randomly selected images may correspond to different real classification information. Recognizing different real classification information each time can simulate the non-ideal image acquisition process in the usage scenario and improve the robustness of the recognition neural network.
[0113] The training process of the recognition neural network enables the reverse adjustment of the parameters of the imaging module, allowing the hardware and software parameters to be optimized in the same training process under the same recognition accuracy, thus achieving the design effect of integrated hardware and software.
[0114] Step S40: After training is completed, the parameters of the recognition neural network and the parameters of the imaging module are obtained.
[0115] The parameters of the imaging module obtained from training can be used to design the specific parameters of the hardware. The parameters of the recognition neural network obtained from training can enable the recognition neural network to better classify and recognize the current scene and the image obtained with the current hardware parameters.
[0116] Compared with the prior art, this embodiment has the following beneficial effects:
[0117] This imaging recognition system aims to optimize the recognition accuracy of the object. It coordinates the parameter design of the imaging module and the recognition module, and pursues parameters with the best global efficiency. Therefore, compared with the existing technology that pursues local optima separately, it has higher design efficiency. On the other hand, since the imaging module does not need to pursue local optima, the cost of the imaging module can be lower than that of the existing technology. That is, the cost of this recognition system can be lower. Therefore, this recognition system achieves higher design efficiency and lower cost through the integrated hardware and software design, which enhances the competitiveness of the product and also helps to save social resources.
[0118] Training device for image recognition system
[0119] In one embodiment, a training device for an imaging recognition system is provided, such as... Figure 6 As shown. The training device includes the following modules, and the specific functions of each module are as follows:
[0120] The acquisition module is used to acquire a set of real images, wherein each image in the set of real images has its corresponding real classification information;
[0121] The simulation module is used to input the real image set into the imaging simulation model to generate a simulated image set, wherein the imaging simulation model is used to simulate the image generated by the imaging module from the real images;
[0122] The training module is used to use the simulated image set as the training set of the recognition neural network, use the difference between the predicted classification information obtained by the recognition neural network from the simulated image set and the real classification information as the loss function, use the parameters of the imaging simulation model as trainable parameters, and update the simulated image set after the parameters of the imaging simulation model are updated, thereby training the recognition neural network.
[0123] The generation module is used to obtain the parameters of the recognition neural network and the parameters of the imaging module after training is completed.
[0124] It should be noted that for details not disclosed in the training device of this embodiment, please refer to the details disclosed in the training method of this embodiment.
[0125] Those skilled in the art will understand that the schematic diagram of the module is merely an example of the training device and does not constitute a limitation on the terminal device of the training device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the training device may also include input / output devices, network access devices, buses, etc.
[0126] The training device may also include computing devices such as computers, laptops, handheld computers, and cloud servers, as well as, but not limited to, processing modules, storage modules, and computer programs stored in the storage modules and executable on the processing modules, such as the training method programs described above. When the processing module executes the computer program, it implements the steps in the various training method embodiments described above, for example... Figure 4 and 5 The steps are shown.
[0127] In addition, the present invention also proposes an electronic device, which includes a storage module and a processing module. When the processing module executes the computer program, it can implement the steps in the above-mentioned training method, that is, implement the steps in any of the technical solutions of the above-mentioned training method.
[0128] The electronic device can be integrated into the training device, a local terminal device, or part of a cloud server.
[0129] The processing module can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. The processing module is the control center of the training device, connecting all parts of the entire training device through various interfaces and lines.
[0130] The storage module can be used to store the computer programs and / or modules. The processing module implements various functions of the training device by running or executing the computer programs and / or modules stored in the storage module and by calling the data stored in the storage module. The storage module may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc. In addition, the storage module may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0131] For example, the computer program can be divided into one or more modules / units, which are stored in a storage module and executed by a processing module to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the training device.
[0132] Furthermore, one embodiment of the present invention provides a readable storage medium storing a computer program that, when executed by a processing module, can implement the steps in the above-described training method, that is, implement the steps in any of the technical solutions of the above-described training method.
[0133] If the modules integrated in the training method are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by the processing module, it can implement the steps of the various method embodiments described above.
[0134] The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include any entity or device capable of carrying the computer program code, recording media, U disks, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0135] It should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This way of describing the specification is only for clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
[0136] The detailed descriptions listed above are merely specific descriptions of feasible embodiments of the present invention, and are not intended to limit the scope of protection of the present invention. All equivalent embodiments or modifications made without departing from the spirit of the present invention should be included within the scope of protection of the present invention.
Claims
1. An imaging recognition system, characterized in that, include: An imaging module, which includes hardware for imaging; The recognition module includes a recognition neural network that acquires an image on an imaging surface and outputs a recognition result. The hardware parameters of the imaging module are set according to the recognition accuracy of the recognition neural network. The imaging module includes an aperture, a metasurface lens, and an imaging surface arranged sequentially. Light passes through the aperture and is imaged on the imaging surface by the metasurface lens. The metasurface lens includes a substrate and multiple primitives on the substrate. The number of metasurface lenses is set to one. The hardware parameters include volume parameters and / or performance parameters. The volume parameter is the minimum volume required for the recognition neural network to achieve a recognition accuracy within the target accuracy range. The volume parameter includes lens length and metasurface area parameters. The lens length parameter is the shortest length required for the recognition neural network to achieve a recognition accuracy within the target accuracy range. The metasurface area parameter is the minimum area required for the recognition neural network to achieve a recognition accuracy within the target accuracy range. The performance parameters are the lowest performance requirements required for the recognition neural network to achieve a recognition accuracy within the target accuracy range. The performance parameters include transfer function parameters, field of view parameters, and focal number parameters. The imaging recognition system is trained according to the following steps: Step S10: Obtain a set of real images, wherein each image in the set of real images has its corresponding real classification information; Step S20: Input the real image set into the imaging simulation model to generate a simulated image set, wherein the imaging simulation model is used to simulate the image generated by the imaging module from the real image; Step S30: Use the simulated image set as the training set of the recognition neural network, use the difference between the predicted classification information obtained by the recognition neural network from the simulated image set and the real classification information as the loss function, use the parameters of the imaging simulation model as trainable parameters, train the recognition neural network, update the parameters of the imaging simulation model during the training process, update the simulated image set after the imaging simulation model parameters are updated, and continue to train the recognition neural network. Step S40: After training is completed, the parameters of the recognition neural network and the hardware parameters of the imaging module are obtained.
2. The imaging recognition system according to claim 1, characterized in that, The transfer function parameter of the imaging module is not less than 20% in the spatial frequency domain at a full field of view of 10 lp / mm.
3. The imaging recognition system according to claim 1, characterized in that, The imaging module is used to acquire images of the driver inside the car cockpit, and the recognition neural network is used to recognize the driver's driving behavior.
4. The imaging recognition system according to claim 1, characterized in that, Updating the parameters of the imaging simulation model during training includes: Determine the recognition accuracy of the recognition neural network; If the recognition accuracy is within the target range, reduce the setting parameters of the imaging simulation model to obtain an updated imaging simulation model, and repeat steps S20 to S30.
5. The imaging recognition system according to claim 1, characterized in that, Training the recognition neural network further includes: If the recognition accuracy of the recognition neural network cannot exceed the minimum recognition rate threshold, the setting parameters of the imaging simulation model are increased to obtain an updated imaging simulation model, and steps S20 to S30 are repeated. If the setting parameters of the simulation model are further reduced, the corresponding recognition accuracy cannot converge to the target recognition accuracy range. Then, the lowest parameter of the setting parameters of the imaging simulation model within the target recognition accuracy range is determined as the required parameter, and the recognition neural network training is completed.
6. The imaging recognition system according to claim 1, characterized in that, The real image set consists of images of the driver inside the car cockpit, and the real classification information is a classification of the driver's driving behavior.
7. The imaging recognition system according to claim 1, characterized in that, The images in the real image set correspond to multiple real classification information; The step of training the recognition neural network includes: Multiple images are randomly selected from the simulated image set as a training set to train the recognition neural network.
8. A training device for an imaging recognition system, characterized in that, For an imaging recognition system as described in any one of claims 1-7, the apparatus comprises: The acquisition module is used to acquire a set of real images, wherein each image in the set of real images has its corresponding real classification information; The simulation module is used to input the real image set into the imaging simulation model to generate a simulated image set, wherein the imaging simulation model is used to simulate the image generated by the imaging module from the real images; The training module is used to use the simulated image set as the training set of the recognition neural network, use the difference between the predicted classification information obtained by the recognition neural network from the simulated image set and the real classification information as the loss function, use the parameters of the imaging simulation model as trainable parameters, train the recognition neural network, update the parameters of the imaging simulation model during the training process, update the simulated image set after the imaging simulation model parameters are updated, and continue to train the recognition neural network. The generation module is used to obtain the parameters of the recognition neural network and the hardware parameters of the imaging module after training is completed.
Citation Information
Patent Citations
Neural network model training method and device, image processing method and device and terminal equipment
CN111950723A
Silicon-based photoelectronic integrated imaging system
CN112637525A