Machine vision system and method for new energy vehicles
By deploying multiple cameras and artificial intelligence-based image processing algorithms on new energy vehicles, we can collect and analyze vehicle perspective environmental images in real time, solving the problems of high cost, limited detection range and lack of semantic information of traditional perception solutions, achieving more efficient surrounding environment perception and a safer driving experience.
Patent Information
- Application Number
- CN202410684328.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-05-30
AI Technical Summary
The vehicle surrounding environment perception scheme of traditional new energy vehicles relies on expensive lidar sensors, has a limited detection range and cannot provide semantic information of the target object, limiting perception ability and driving experience.
By deploying multiple cameras to acquire vehicle viewing environment images in real time, and introducing artificial intelligence-based image processing and analysis algorithms on the back end, performing coordination and correlation analysis, identifying target objects in the vehicle's surrounding environment, and generating a detailed surrounding environment map.
It overcomes the high cost of traditional perception solutions, limited detection range and lack of semantic information, improves the surrounding environment perception capabilities of new energy vehicles, and enhances driving experience and safety.
Smart Images

Figure CN118506323B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine vision, and more specifically, to a machine vision system and method for new energy vehicles. Background Art
[0002] New energy vehicles are developing and gaining popularity due to their environmental benefits and lower operating costs. Compared with traditional internal combustion engine vehicles, new energy vehicles have the advantage of lower emissions, which helps reduce air pollution and greenhouse gas emissions. As new energy vehicles become more popular, the demand for improving their safety, reliability and driving experience is also growing.
[0003] The vehicle surroundings perception technology of new energy vehicles is crucial to the safety and driving experience of new energy vehicles. By sensing the vehicle's surroundings, new energy vehicles can make more intelligent decisions, for example, by sensing and monitoring other vehicles, pedestrians and objects in the vehicle's surroundings to avoid collisions.
[0004] However, traditional vehicle surrounding environment perception solutions usually rely on sensors such as LiDAR to detect target objects. LiDAR sensors are expensive, which increases the overall cost of the car. In addition, LiDAR sensors have a limited detection range, which may limit their effectiveness in some cases. In addition, LiDAR sensors can only provide basic information such as the location and speed of the target object, but cannot provide semantic information about the type of the target object.
[0005] Therefore, a machine vision system for new energy vehicles is desired. Summary of the invention
[0006] In order to solve the above technical problems, the present invention is proposed. The embodiment of the present invention provides a machine vision system and method for new energy vehicles, which collects multiple vehicle-perspective environment images in real time through multiple cameras deployed in new energy vehicles, and introduces an artificial intelligence-based image processing and analysis algorithm at the back end to perform coordination and correlation analysis of these vehicle-perspective environment images, so as to identify other vehicles, pedestrians, objects and other target objects in the vehicle's surrounding environment, and more accurately create a detailed map of the vehicle's surrounding environment based on the importance and correlation relationship of different target objects in the environment. In this way, it is possible to overcome the problems of high sensor prices, limited detection range, and inability to provide semantic information about the type of target object in traditional perception solutions, thereby improving the surrounding environment perception capabilities of new energy vehicles based on machine vision systems.
[0007] According to one aspect of the present invention, there is provided a machine vision system for a new energy vehicle, comprising:
[0008] A vehicle perspective environment image acquisition module, used to acquire multiple vehicle perspective environment images acquired by multiple cameras deployed on new energy vehicles;
[0009] A target object detection module, used for passing the multiple vehicle perspective environment images through a multi-target object detector based on a YOLO network to obtain multiple target object interest area images;
[0010] A target region of interest division module, used for dividing each of the plurality of target object region of interest images according to the region where the target object is located to obtain a sequence of a plurality of target object local region of interest images;
[0011] A target region of interest global semantic feature extraction module is used to extract features from the sequence of the plurality of target object local region of interest images and to perform a global semantic fusion representation of the region of interest to obtain a plurality of target object region of interest global semantic feature vectors;
[0012] A vehicle multi-directional perspective environment semantic fusion representation module, used for transmitting the global semantic feature vectors of the regions of interest of the multiple target objects through a target object semantic feature information transmission reinforcement module based on an information transmission network to obtain a multi-directional vehicle environment semantic fusion representation feature vector as a multi-directional vehicle environment semantic fusion representation feature;
[0013] The surrounding environment map generation module is used to generate a surrounding environment map of the new energy vehicle based on the multi-directional vehicle environment semantic fusion representation features.
[0014] According to another aspect of the present invention, there is provided a machine vision method for a new energy vehicle, comprising:
[0015] Acquire multiple vehicle perspective environment images collected by multiple cameras deployed on new energy vehicles;
[0016] Passing the multiple vehicle perspective environment images through a multi-target object detector based on a YOLO network to obtain multiple target object region of interest images;
[0017] Dividing each of the plurality of target object region of interest images according to the region where the target object is located to obtain a sequence of a plurality of target object local region of interest images;
[0018] Performing feature extraction and global semantic fusion representation of the regions of interest on the sequences of the multiple target object local region images to obtain global semantic feature vectors of the regions of interest of the multiple target objects;
[0019] The global semantic feature vectors of the regions of interest of the multiple target objects are passed through a target object semantic feature information transfer enhancement module based on an information transfer network to obtain a multi-directional vehicle environment semantic fusion representation feature vector as a multi-directional vehicle environment semantic fusion representation feature;
[0020] Based on the multi-directional vehicle environment semantic fusion representation features, a surrounding environment map of the new energy vehicle is generated.
[0021] Compared with the prior art, the present invention provides a machine vision system and method for new energy vehicles, which collects multiple vehicle-view environmental images in real time through multiple cameras deployed on new energy vehicles, and introduces an image processing and analysis algorithm based on artificial intelligence at the back end to perform coordination and correlation analysis of these vehicle-view environmental images, so as to identify other vehicles, pedestrians, objects and other target objects in the vehicle's surrounding environment, and more accurately create a detailed map of the vehicle's surrounding environment based on the importance and correlation relationship of different target objects in the environment. In this way, it is possible to overcome the problems of high sensor prices, limited detection range, and inability to provide semantic information about the type of target object in traditional perception solutions, thereby improving the surrounding environment perception capabilities of new energy vehicles based on machine vision systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other purposes, features and advantages of the present invention will become more apparent by describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0023] Figure 1 A block diagram of a machine vision system for a new energy vehicle according to an embodiment of the present invention;
[0024] Figure 2 A system architecture diagram of a machine vision system for a new energy vehicle according to an embodiment of the present invention;
[0025] Figure 3 A block diagram of a training phase of a machine vision system for new energy vehicles according to an embodiment of the present invention;
[0026] Figure 4 A block diagram of a global semantic feature extraction module for a target region of interest in a machine vision system for a new energy vehicle according to an embodiment of the present invention;
[0027] Figure 5 The figure is a flow chart of a machine vision method for new energy vehicles according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described here.
[0029] As shown in the present invention and claims, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular, but also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list, and the method or device may also include other steps or elements.
[0030] Although the present invention has made various references to certain modules in the system according to an embodiment of the present invention, any number of different modules can be used and run on a user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0031] The present invention uses a flow chart to illustrate the operations performed by the system according to an embodiment of the present invention. It should be understood that the preceding or following operations are not necessarily performed precisely in order. On the contrary, various steps may be processed in reverse order or simultaneously as required. At the same time, other operations may also be added to these processes, or one or more operations may be removed from these processes.
[0032] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described here.
[0033] The machine vision system can enhance the environmental perception function of the vehicle by providing images and semantic information of the vehicle's surrounding environment, effectively reducing the probability of vehicle collision accidents. In the machine vision system of new energy vehicles, the environmental images collected by multiple cameras can provide more comprehensive and accurate environmental perception information, helping the system to better understand the surrounding environment.
[0034] In the technical solution of the present invention, a machine vision system for new energy vehicles is proposed. Figure 1 4 is a block diagram of a machine vision system for new energy vehicles according to an embodiment of the present invention. Figure 2 FIG. 1 is a system architecture diagram of a machine vision system for new energy vehicles according to an embodiment of the present invention. Figure 1 and Figure 2As shown, the machine vision system for new energy vehicles according to an embodiment of the present invention includes: a vehicle perspective environment image acquisition module 310, which is used to obtain multiple vehicle perspective environment images acquired by multiple cameras deployed in the new energy vehicle; a target object detection module 320, which is used to obtain multiple target object interest area images by passing the multiple vehicle perspective environment images through a multi-target object detector based on a YOLO network; a target interest area division module 330, which is used to divide each target object interest area image in the multiple target object interest area images according to the area where the target object is located to obtain a sequence of multiple target object interest local area images; a target interest area global language The semantic feature extraction module 340 is used to perform feature extraction and global semantic fusion representation of the regions of interest on the sequence of the multiple target object local area images to obtain global semantic feature vectors of the regions of interest of the multiple target objects; the vehicle multi-perspective environment semantic fusion representation module 350 is used to obtain a multi-directional vehicle environment semantic fusion representation feature vector as a multi-directional vehicle environment semantic fusion representation feature through the target object semantic feature information transmission enhancement module based on the information transmission network; the surrounding environment map generation module 360 is used to generate a surrounding environment map of the new energy vehicle based on the multi-directional vehicle environment semantic fusion representation feature.
[0035] In particular, the vehicle perspective environment image acquisition module 310 and the target object detection module 320 are used to obtain multiple vehicle perspective environment images collected by multiple cameras deployed on new energy vehicles; and, the multiple vehicle perspective environment images are passed through a multi-target object detector based on a YOLO network to obtain multiple target object interest area images. Considering that the vehicle perspective environment image collected by the camera not only contains semantic information about the vehicle surrounding environment, but also contains a large amount of background interference noise. Therefore, when actually perceiving the vehicle surrounding environment, it is necessary to detect different target objects in the vehicle perspective environment image to provide more accurate target interest area semantic information, which is helpful for subsequent target object type recognition and semantic analysis. Based on this, in the technical solution of the present invention, the multiple vehicle perspective environment images are passed through a multi-target object detector based on a YOLO network to obtain multiple target object interest area images. In particular, the multi-target object detector based on the YOLO network can simultaneously detect multiple target objects in the image and generate an interest area surrounding these target objects. This is crucial for the perception of the vehicle surrounding environment, because multiple target objects, such as other vehicles, pedestrians and objects, need to be detected and tracked simultaneously during the vehicle driving process.
[0036] In particular, the target region of interest division module 330 is used to divide each target object region of interest image in the multiple target object region of interest images according to the region where the target object is located to obtain a sequence of multiple target object local region of interest images. That is, after using the YOLO network to detect and identify the target object in each vehicle perspective environment image, it is necessary to divide and accurately analyze these target objects, so as to use the local region of interest and semantic feature information where each target object is located to more accurately and meticulously perceive the environment. Based on this, each target object region of interest image in the multiple target object region of interest images is further divided according to the region where the target object is located to obtain a sequence of multiple target object local region of interest images.
[0037] In particular, the target region of interest global semantic feature extraction module 340 is used to extract features from the sequence of the plurality of target object local region of interest images and perform region of interest global semantic fusion representation to obtain a plurality of target object region of interest global semantic feature vectors. In particular, in a specific example of the present invention, Figure 4 As shown, the target region of interest global semantic feature extraction module 340 includes: a target object local region of interest feature extraction unit 341, used to pass the sequence of the multiple target object local region of interest images through a target object feature extractor based on a convolutional neural network model to obtain a sequence of multiple target object local region of interest feature vectors; a region of interest global semantic association unit 342, used to pass the sequence of the multiple target object local region of interest feature vectors through a region of interest multi-target feature fuser based on an autocorrelation saliency network to obtain the multiple target object region of interest global semantic feature vectors.
[0038] Specifically, the target object local area of interest feature extraction unit 341 is used to pass the sequence of the multiple target object local area of interest images through a target object feature extractor based on a convolutional neural network model to obtain a sequence of multiple target object local area of interest feature vectors. That is, the sequence of the multiple target object local area of interest images is subjected to feature mining in a target object feature extractor based on a convolutional neural network model to respectively extract the implicit semantic feature information about the target object contained in each target object local area of interest image in each of the vehicle perspective environment images, thereby obtaining a sequence of multiple target object local area of interest feature vectors. Among them, the target object feature extractor based on a convolutional neural network model uses an AlexNet network as a feature extractor.
[0039] Specifically, the global semantic association unit 342 of the region of interest is used to pass the sequence of the feature vectors of the local region of interest of the multiple target objects through the multi-target feature fuser of the region of interest based on the autocorrelation saliency network to obtain the global semantic feature vectors of the region of interest of the multiple target objects. It should be understood that since each target object local region of interest feature vector in the sequence of the feature vectors of the local region of interest of the multiple target objects represents the semantic feature information of the target object contained in the local region of interest of a specific target object in a vehicle perspective environment image, and since different target object semantic features in a vehicle perspective environment image have different correlations and importance semantics, for example, different vehicles in front may have different distances from new energy vehicles, at this time different vehicles should receive different attention, and there is also a driving semantic association relationship between each vehicle. Therefore, in order to be able to perform an overall target object semantic analysis for each vehicle perspective environment image, in the technical solution of the present invention, the sequence of the feature vectors of the local region of interest of the multiple target objects is further passed through the multi-target feature fuser of the region of interest based on the autocorrelation saliency network to obtain the global semantic feature vectors of the region of interest of the multiple target objects. In particular, the region of interest multi-target feature fuser based on the autocorrelation saliency network can learn and capture the global features of the correlation between the semantic features of the local regions of interest of each target object contained in each vehicle-view environment image, which is conducive to the global semantic feature expression of the target object in each vehicle-view environment image. In addition, in the process of global semantic expression of the image, the autocorrelation saliency network can also pay attention to the categories, attributes, and relationships of different target objects to judge the importance of different target objects, so as to give different circles to different target objects, so as to more accurately perceive the vehicle's surrounding environment and generate a more accurate surrounding environment map. More specifically, the sequence of feature vectors of the local region of interest of each target object is processed by the region of interest multi-target feature fuser based on the autocorrelation saliency network with the following autocorrelation saliency fusion formula to obtain the global semantic feature vector of the region of interest of the target object; wherein the autocorrelation saliency fusion formula is:
[0040]
[0041] Among them, h i is the i-th target object interesting local area feature vector in the sequence of target object interesting local area feature vectors, and W i Represent the weight coefficient vector and weight coefficient matrix respectively, B i is the offset vector, Selu(·) represents the Selu function, e iis the attention score of the i-th feature vector of the local area of interest of the target object, λ and α are both hyperparameters, softmax(·) represents the softmax function, t is the number of vectors in the sequence of feature vectors of the local area of interest of the target object, and V is the global semantic feature vector of the target object area of interest.
[0042] It is worth mentioning that in other specific examples of the present invention, feature extraction and global semantic fusion representation of the regions of interest of the sequences of the multiple target object local region images can be performed in other ways to obtain global semantic feature vectors of the regions of interest of the multiple target objects, for example: input the sequence of the multiple target object local region images of interest; feature extraction is performed on each local region image, and a pre-trained convolutional neural network (CNN) model (such as VGG, ResNet, or EfficientNet) can be used to extract the feature representation of the local region; features are extracted from each local region image to obtain a sequence of local feature vectors; global semantic fusion is performed on the sequence of local feature vectors to capture the semantic information of the entire sequence and fuse the correlation between the local regions; attention mechanism, recurrent neural network (RNN) or Transformer and other methods are used to achieve global semantic fusion; representation learning is performed on the feature sequence after global semantic fusion to learn the global semantic features of the regions of interest of the multiple target objects; a deep learning model such as a recurrent neural network (RNN), a convolutional neural network (CNN), or an attention mechanism is used to learn the global semantic feature vector; the global semantic feature vectors of the regions of interest of the multiple target objects are integrated to obtain a set of global semantic feature vectors of the regions of interest of the multiple target objects.
[0043] In particular, the vehicle multi-directional perspective environment semantic fusion representation module 350 is used to pass the global semantic feature vectors of the regions of interest of the multiple target objects through the target object semantic feature information transmission reinforcement module based on the information transmission network to obtain the multi-directional vehicle environment semantic fusion representation feature vector as the multi-directional vehicle environment semantic fusion representation feature. Considering that in the process of vehicle surrounding environment perception and surrounding environment map generation, it is necessary to perform association analysis on the global semantics of the environmental image of each perspective, so as to generate a more comprehensive and accurate surrounding environment map of the new energy vehicle. Based on this, in the technical solution of the present invention, the global semantic feature vectors of the regions of interest of the multiple target objects are further passed through the target object semantic feature information transmission reinforcement module based on the information transmission network to obtain the multi-directional vehicle environment semantic fusion representation feature vector. It should be understood that through the processing of the target object semantic feature information transmission reinforcement module based on the information transmission network, the global association semantics of each target object contained in the environmental image of different vehicle perspectives can be transmitted and fused based on the semantic information of the global perspective. Moreover, in the process of information transmission and fusion, different weights can be applied to each perspective based on the correlation relationship between the global semantics of multiple target objects under different perspectives, so as to focus on the target objects that need to be paid attention to during the driving of the vehicle, so as to further improve the perception ability of the vehicle's surrounding environment and enhance the comprehensive understanding of the vehicle's surrounding environment, so as to generate a surrounding environment map of new energy vehicles that meets user needs. Specifically, the global semantic feature vectors of the regions of interest of the multiple target objects are processed by the target object semantic feature information transmission enhancement module based on the information transmission network with the following information transmission enhancement formula to obtain the multi-directional vehicle environment semantic fusion representation feature vector; wherein, the information transmission enhancement formula is:
[0044]
[0045] Among them, x i and x j denote the i-th and j-th target object region of interest global semantic feature vectors in the multiple target object region of interest global semantic feature vectors, respectively, g(x j ) represents x j With x i The number of eigenvectors between them, f(x i ,x j ) represents x j With x i similarity between them, M represents the total number of feature vectors in the global semantic feature vectors of the regions of interest of the multiple target objects, n is equal to the numerical value of M, and V represents the multi-directional vehicle environment semantic fusion representation feature vector.
[0046] In particular, the surrounding environment map generation module 360 is used to generate a surrounding environment map of the new energy vehicle based on the multi-directional vehicle environment semantic fusion representation feature. In particular, in a specific example of the present invention, the multi-directional vehicle environment semantic fusion representation feature vector is passed through a vehicle surrounding environment map generator based on AI GC to obtain a generation result, and the generation result is a surrounding environment map of the new energy vehicle. In other words, the vehicle's multi-directional environmental semantic information transmission fusion representation feature is used to generate the vehicle surrounding environment map. In this way, a detailed map of the vehicle's surrounding environment can be created more accurately based on the importance and correlation relationship of different target objects in the vehicle's surrounding environment. In this way, the problems of high sensor prices, limited detection range, and inability to provide semantic information about the target object type in traditional perception solutions can be overcome, thereby improving the surrounding environment perception capabilities of new energy vehicles based on machine vision systems.
[0047] It should be understood that before using the above-mentioned neural network model for inference, it is necessary to train the target object feature extractor based on the convolutional neural network model, the multi-target feature fuser of the region of interest based on the autocorrelation saliency network, the target object semantic feature information transfer enhancement module based on the information transfer network, and the vehicle surrounding environment map generator based on AIGC. That is, according to the machine vision system 300 for new energy vehicles of the present invention, it also includes a training module 400 for training the target object feature extractor based on the convolutional neural network model, the multi-target feature fuser of the region of interest based on the autocorrelation saliency network, the target object semantic feature information transfer enhancement module based on the information transfer network, and the vehicle surrounding environment map generator based on AIGC.
[0048] Figure 3 FIG. 4 is a block diagram of a training phase of a machine vision system for a new energy vehicle according to an embodiment of the present invention. Figure 3As shown, according to an embodiment of the present invention, a machine vision system 300 for new energy vehicles includes: a training module 400, including: a training data acquisition unit 410, used to acquire training data, wherein the training data includes a plurality of training vehicle perspective environment images collected by a plurality of cameras deployed on the new energy vehicle; a training target object detection unit 420, used to pass the plurality of training vehicle perspective environment images through a multi-target object detector based on a YOLO network to obtain a plurality of training target object interest region images; a training target interest region division unit 430, used to divide each of the plurality of training target object interest region images according to the region where the target object is located to obtain a sequence of a plurality of training target object interest local region images; a training target object interest local region feature extraction unit 440, used to pass the plurality of training target object interest local region images through a target object feature extractor based on a convolutional neural network model to obtain a sequence of a plurality of training target object interest local region feature vectors; a training interest region global semantic association unit 450, used to associate the plurality of training target objects with the target object; The sequence of feature vectors of local regions of interest of objects is passed through a multi-target feature fuser of regions of interest based on an autocorrelation saliency network to obtain a plurality of global semantic feature vectors of regions of interest of training target objects; a training semantic fusion representation unit 460 is used to pass the global semantic feature vectors of regions of interest of the plurality of training target objects through a target object semantic feature information transfer enhancement module based on an information transfer network to obtain a training multi-directional vehicle environment semantic fusion representation feature vector; a loss function calculation unit 470 is used to pass the training multi-directional vehicle environment semantic fusion representation feature vector through a vehicle surrounding environment map generator based on AIGC to obtain a mean square error loss function value; a training unit 480 is used to train the target object feature extractor based on a convolutional neural network model, the multi-target feature fuser of regions of interest based on an autocorrelation saliency network, the target object semantic feature information transfer enhancement module based on an information transfer network and the vehicle surrounding environment map generator based on AIGC based on the mean square error loss function value, wherein the training multi-directional vehicle environment semantic fusion representation feature vector is optimized at each model iteration.
[0049] In particular, in the technical solution described above, the sequence of the multiple training target object local region of interest feature vectors represents the image semantic features of the sequence of the training target object local region of interest images. Thus, after the sequence of the multiple training target object local region of interest feature vectors passes through the region of interest multi-target feature fuser based on the autocorrelation saliency network, the multiple training target object region of interest global semantic feature vectors have image semantic feature expression differences based on the global distribution of source image semantics of different target object region of interest images.
[0050] In this way, considering that the global semantic feature vectors of the regions of interest of the multiple training target objects are further obtained through the target object semantic feature information transmission reinforcement module based on the information transmission network to obtain the training multi-directional vehicle environment semantic fusion representation feature vector, if under the cascade model serialization feature extraction structure of the image semantic feature, based on the feature image semantic sequence correspondence between the global semantic feature vectors of the regions of interest of the multiple training target objects and the training multi-directional vehicle environment semantic fusion representation feature vector, the overflow features can be suppressed on the basis of maintaining the image semantic regression discrimination under the cascade dimension, which is beneficial to improving the image quality of the surrounding environment map of the monitored new energy vehicle generated during model inference.
[0051] Based on this, the present application optimizes the training multi-aspect vehicle environment semantic fusion representation feature vector at each model iteration, for example, when back propagation is performed based on the loss function between the real image and the inferred image, including: calculating the training multi-aspect vehicle environment semantic fusion representation mean matrix and the training multi-aspect vehicle environment semantic fusion representation variance matrix of the training multi-aspect vehicle environment semantic fusion representation feature vector, wherein the value of each position of the training multi-aspect vehicle environment semantic fusion representation mean matrix is the mean of a pair of eigenvalues of the two positions of the training multi-aspect vehicle environment semantic fusion representation feature vector corresponding to the position coordinates respectively, and the value of each position of the training multi-aspect vehicle environment semantic fusion representation variance matrix is the variance of a pair of eigenvalues of the two positions of the training multi-aspect vehicle environment semantic fusion representation feature vector corresponding to the position coordinates respectively; matrix multiplying the transposed vector of the training multi-aspect vehicle environment semantic fusion representation feature vector with the training multi-aspect vehicle environment semantic fusion representation mean matrix to obtain a training first multi-aspect vehicle environment semantic fusion representation intermediate vector and transforming the training multi-aspect vehicle environment semantic fusion representation feature vector into the training multi-aspect vehicle environment semantic fusion representation feature vector; The method comprises the following steps: performing matrix multiplication of the variance matrix of the training multi-directional vehicle environment semantic fusion representation and the feature vector of the training multi-directional vehicle environment semantic fusion representation to obtain an intermediate vector of the training second multi-directional vehicle environment semantic fusion representation, wherein the feature vector of the training multi-directional vehicle environment semantic fusion representation is a column vector; calculating the point sum of the transposed vector of the intermediate vector of the training first multi-directional vehicle environment semantic fusion representation and the intermediate vector of the training second multi-directional vehicle environment semantic fusion representation to obtain an intermediate vector of the training third multi-directional vehicle environment semantic fusion representation; calculating the matrix product of the mean matrix of the training multi-directional vehicle environment semantic fusion representation and the variance matrix of the training multi-directional vehicle environment semantic fusion representation, and performing matrix multiplication of the transposed vector of the feature vector of the training multi-directional vehicle environment semantic fusion representation and the matrix product to obtain an intermediate vector of the training fourth multi-directional vehicle environment semantic fusion representation; calculating the point sum of the transposed vector of the intermediate vector of the training third multi-directional vehicle environment semantic fusion representation and the intermediate vector of the training fourth multi-directional vehicle environment semantic fusion representation to obtain an optimized training multi-directional vehicle environment semantic fusion representation vector.
[0052] That is, by using the mean matrix of the training multi-directional vehicle environment semantic fusion representation vector and the variance matrix of the training multi-directional vehicle environment semantic fusion representation vector for the group aggregation statistical evaluation of the local distribution of the feature value granularity of the training multi-directional vehicle environment semantic fusion representation vector as the retrieval and response distribution enhancement of the training multi-directional vehicle environment semantic fusion representation vector, and constructing a reference-free distribution retrieval response framework under the open domain of the feature distribution based on group aggregation of the training multi-directional vehicle environment semantic fusion representation vector, the distribution response redundancy caused by the local overflow feature of the training multi-directional vehicle environment semantic fusion representation vector is avoided through response superposition, so as to realize the faithfulness constraint of the self-aggregation statistics correlation of the response of the training multi-directional vehicle environment semantic fusion representation vector to the semantic space domain of the target generated image, and improve the image quality of the surrounding environment map of the monitored new energy vehicle generated during model inference. In this way, a detailed map of the vehicle surrounding environment can be created more accurately according to the importance and correlation relationship of different target objects in the environment, the perception ability of the new energy vehicle to the surrounding environment is improved, and the driving safety of the new energy vehicle is ensured.
[0053] As described above, the machine vision system 300 for new energy vehicles according to an embodiment of the present invention can be implemented in various wireless terminals, such as a server with a machine vision algorithm for new energy vehicles. In one possible implementation, the machine vision system 300 for new energy vehicles according to an embodiment of the present invention can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the machine vision system 300 for new energy vehicles can be a software module in the operating system of the wireless terminal, or can be an application developed for the wireless terminal; of course, the machine vision system 300 for new energy vehicles can also be one of the many hardware modules of the wireless terminal.
[0054] Alternatively, in another example, the machine vision system 300 for new energy vehicles and the wireless terminal may also be separate devices, and the machine vision system 300 for new energy vehicles may be connected to the wireless terminal via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.
[0055] Furthermore, a machine vision method for new energy vehicles is also provided.
[0056] Figure 5 FIG. 1 is a flow chart of a machine vision method for new energy vehicles according to an embodiment of the present invention. Figure 5As shown, according to an embodiment of the present invention, a machine vision method for new energy vehicles includes the following steps: S1, acquiring multiple vehicle perspective environment images collected by multiple cameras deployed on the new energy vehicle; S2, passing the multiple vehicle perspective environment images through a multi-target object detector based on a YOLO network to obtain multiple target object interest area images; S3, dividing each target object interest area image in the multiple target object interest area images according to the area where the target object is located to obtain a sequence of multiple target object interest local area images; S4, performing feature extraction and global semantic fusion representation of the interest area on the sequence of the multiple target object interest local area images to obtain multiple target object interest area global semantic feature vectors; S5, passing the multiple target object interest area global semantic feature vectors through a target object semantic feature information transfer enhancement module based on an information transfer network to obtain a multi-directional vehicle environment semantic fusion representation feature vector as a multi-directional vehicle environment semantic fusion representation feature; S6, generating a surrounding environment map of the new energy vehicle based on the multi-directional vehicle environment semantic fusion representation feature.
[0057] In summary, the machine vision method for new energy vehicles according to the embodiment of the present invention is explained, which collects multiple vehicle-perspective environment images in real time through multiple cameras deployed on new energy vehicles, and introduces an artificial intelligence-based image processing and analysis algorithm at the back end to perform coordination and correlation analysis of these vehicle-perspective environment images, so as to identify other vehicles, pedestrians, objects and other target objects in the vehicle's surrounding environment, and more accurately create a detailed map of the vehicle's surrounding environment based on the importance and correlation relationship of different target objects in the environment. In this way, it is possible to overcome the problems of high sensor prices, limited detection range, and inability to provide semantic information about the type of target object in traditional perception solutions, thereby improving the surrounding environment perception capabilities of new energy vehicles based on machine vision systems.
[0058] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A machine vision system for new energy vehicles, characterized in that: include: A vehicle perspective environment image acquisition module, used to acquire multiple vehicle perspective environment images acquired by multiple cameras deployed on new energy vehicles; A target object detection module, used for passing the multiple vehicle perspective environment images through a multi-target object detector based on a YOLO network to obtain multiple target object interest area images; A target region of interest division module, used for dividing each of the plurality of target object region of interest images according to the region where the target object is located to obtain a sequence of a plurality of target object local region of interest images; A target region of interest global semantic feature extraction module is used to extract features from the sequence of the plurality of target object local region of interest images and to perform a global semantic fusion representation of the region of interest to obtain a plurality of target object region of interest global semantic feature vectors; A vehicle multi-directional perspective environment semantic fusion representation module, used for transmitting the global semantic feature vectors of the regions of interest of the multiple target objects through a target object semantic feature information transmission reinforcement module based on an information transmission network to obtain a multi-directional vehicle environment semantic fusion representation feature vector as a multi-directional vehicle environment semantic fusion representation feature; A surrounding environment map generation module, used to generate a surrounding environment map of the new energy vehicle based on the multi-directional vehicle environment semantic fusion representation features; The target region of interest global semantic feature extraction module includes: A target object local region of interest feature extraction unit, used for passing the sequence of the plurality of target object local region of interest images through a target object feature extractor based on a convolutional neural network model to obtain a sequence of a plurality of target object local region of interest feature vectors; An area of interest global semantic association unit, used for passing the sequences of feature vectors of the local areas of interest of the multiple target objects through an area of interest multi-target feature fuser based on an autocorrelation saliency network to obtain global semantic feature vectors of the areas of interest of the multiple target objects; The region of interest global semantic association unit is used to: process the sequence of feature vectors of each local region of interest of the target object through the region of interest multi-target feature fuser based on the autocorrelation saliency network according to the following autocorrelation saliency fusion formula to obtain the global semantic feature vector of the region of interest of the target object; The autocorrelation saliency fusion formula is: Among them, h i is the i-th target object interesting local area feature vector in the sequence of target object interesting local area feature vectors, and W i Represent the weight coefficient vector and weight coefficient matrix respectively, B i is the offset vector, Selu(·) represents the Selu function, e i is the attention score of the i-th feature vector of the local region of interest of the target object, λ and α are both hyperparameters, softmax(·) represents the softmax function, t is the number of vectors in the sequence of feature vectors of the local region of interest of the target object, and V is the global semantic feature vector of the region of interest of the target object; The vehicle multi-directional perspective environment semantic fusion representation module is used to: process the global semantic feature vectors of the regions of interest of the multiple target objects through the target object semantic feature information transfer enhancement module based on the information transfer network with the following information transfer enhancement formula to obtain the multi-directional vehicle environment semantic fusion representation feature vector; Wherein, the information transmission enhancement formula is: Among them, x i and x j denote the i-th and j-th target object region of interest global semantic feature vectors in the multiple target object region of interest global semantic feature vectors, respectively, g(x j ) represents x j With x i The number of eigenvectors between them, f(x i ,x j ) represents x j With x i similarity between them, M represents the total number of feature vectors in the global semantic feature vectors of the regions of interest of the multiple target objects, n is equal to the numerical value of M, and V represents the multi-directional vehicle environment semantic fusion representation feature vector.
2. The machine vision system for new energy vehicles according to claim 1, characterized in that: The target object feature extractor based on the convolutional neural network model adopts the AlexNet network as the feature extractor.
3. The machine vision system for new energy vehicles according to claim 2, characterized in that: The surrounding environment map generation module is used to: pass the multi-directional vehicle environment semantic fusion representation feature vector through the vehicle surrounding environment map generator based on AIGC to obtain a generation result, and the generation result is the surrounding environment map of the new energy vehicle.
4. The machine vision system for new energy vehicles according to claim 3, characterized in that: It also includes a training module for training the target object feature extractor based on the convolutional neural network model, the region of interest multi-target feature fuser based on the autocorrelation saliency network, the target object semantic feature information transfer enhancement module based on the information transfer network, and the vehicle surrounding environment map generator based on AIGC.
5. The machine vision system for new energy vehicles according to claim 4, characterized in that: The training module comprises: A training data acquisition unit, used to acquire training data, wherein the training data includes a plurality of training vehicle perspective environment images collected by a plurality of cameras deployed on the new energy vehicle; A training target object detection unit, used for passing the plurality of training vehicle perspective environment images through a multi-target object detector based on a YOLO network to obtain a plurality of training target object region of interest images; A training target region of interest division unit, configured to divide each of the plurality of training target object region of interest images according to the region where the target object is located to obtain a sequence of a plurality of training target object local region of interest images; A training target object local region of interest feature extraction unit, used for passing the sequence of the plurality of training target object local region of interest images through a target object feature extractor based on a convolutional neural network model to obtain a sequence of a plurality of training target object local region of interest feature vectors; A training region of interest global semantic association unit is used to obtain a plurality of global semantic feature vectors of regions of interest of training target objects by passing the sequences of feature vectors of local regions of interest of the training target objects through a region of interest multi-target feature fuser based on an autocorrelation saliency network; A training semantic fusion representation unit, used for transmitting the global semantic feature vectors of the regions of interest of the multiple training target objects through a target object semantic feature information transmission reinforcement module based on an information transmission network to obtain a training multi-directional vehicle environment semantic fusion representation feature vector; A loss function calculation unit, used for passing the training multi-directional vehicle environment semantic fusion representation feature vector through an AIGC-based vehicle surrounding environment map generator to obtain a mean square error loss function value; A training unit is used to train the target object feature extractor based on the convolutional neural network model, the multi-target feature fuser of the region of interest based on the autocorrelation saliency network, the target object semantic feature information transfer enhancement module based on the information transfer network, and the vehicle surrounding environment map generator based on AIGC based on the mean square error loss function value, wherein the training multi-directional vehicle environment semantic fusion representation feature vector is optimized at each model iteration.
6. A machine vision system method for a new energy vehicle using a machine vision system according to any one of claims 1 to 5, characterized in that: include: Acquire multiple vehicle perspective environment images collected by multiple cameras deployed on new energy vehicles; Passing the multiple vehicle perspective environment images through a multi-target object detector based on a YOLO network to obtain multiple target object region of interest images; Dividing each of the plurality of target object region of interest images according to the region where the target object is located to obtain a sequence of a plurality of target object local region of interest images; Performing feature extraction and global semantic fusion representation of the regions of interest on the sequences of the multiple target object local region images to obtain global semantic feature vectors of the regions of interest of the multiple target objects; The global semantic feature vectors of the regions of interest of the multiple target objects are passed through a target object semantic feature information transfer enhancement module based on an information transfer network to obtain a multi-directional vehicle environment semantic fusion representation feature vector as a multi-directional vehicle environment semantic fusion representation feature; Based on the multi-directional vehicle environment semantic fusion representation features, a surrounding environment map of the new energy vehicle is generated.
Citation Information
Patent Citations
Door and window opening prompting method and system after vehicle collision
CN116386005A
Real estate surveying and mapping information management system and method based on GIS technology
CN118096479A