Target detection method, device, equipment, system and storage medium

By setting the interface module on the vehicle to detect and configure the newly inserted image sensor, the problem of high difficulty in image sensor integration in the prior art is solved, and the rapid start-up and high efficiency of object detection are achieved.

CN120107926APending Publication Date: 2025-06-06HEFEI IFLY DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510030816.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Among the existing object detection methods, the integration of image sensors is difficult, and manufacturers or professionals need to modify the wiring and add new image sensors.

Method used

By setting up an interface module on the vehicle, a newly inserted image sensor is detected and configured. The interface module receives image data input from the image sensor and performs object detection based on these data.

Benefits of technology

Reduces the difficulty of integrating image sensors, allowing users to flexibly add new image sensors without the need to modify wiring, and achieves fast start-up and high efficiency of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107926A_ABST
    Figure CN120107926A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method, device, equipment and system and a storage medium, and the method comprises the steps: detecting that an interface module is newly inserted into an image sensor, the interface module is disposed on a vehicle, and the image sensor can carry out the image collection of the external environment of the vehicle after being inserted; the newly inserted image sensor is configured; receiving image data input by the image sensor through the interface module; and performing target detection based on the image data to obtain a target detection result representing whether the target exists in the external environment of the vehicle or not. Through the method, the integration difficulty of the image sensor can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision and artificial intelligence, and in particular to a target detection method, device, equipment, system and storage medium. Background Art

[0002] In the existing computer vision and artificial intelligence fields, vehicles need to use image sensors to obtain corresponding image data for target detection. However, in existing target detection methods, image sensors used for target detection are generally pre-integrated. If new image sensors need to be added for target detection in the future, manufacturers or professionals are required to modify the vehicle wiring to integrate the new image sensors into the vehicle. In this process, the integration of image sensors is difficult. Summary of the invention

[0003] The main technical problem solved by the present application is to provide a target detection method, device, equipment, system and storage medium, which can reduce the difficulty of integrating image sensors.

[0004] In order to solve the above technical problems, a technical solution adopted by the present application is: to provide a target detection method, the method comprising: detecting that an interface module is newly inserted into an image sensor, the interface module is arranged on a vehicle, and the image sensor can collect images of the external environment of the vehicle after insertion. The newly inserted image sensor is configured. Image data input by the image sensor is received through the interface module. Target detection is performed based on the image data to obtain a target detection result characterizing whether there is a target in the external environment of the vehicle.

[0005] In order to solve the above technical problems, another technical solution adopted by the present application is: to provide a target detection device, the device comprising: an image acquisition module, used to detect that an interface module is newly inserted into an image sensor, the interface module is arranged on the vehicle, and the image sensor can capture images of the external environment of the vehicle after insertion. A configuration module, used to configure the newly inserted image sensor. An image data acquisition module, used to receive image data input by the image sensor through the interface module. A target detection module, used to perform target detection based on image data, and obtain a target detection result characterizing whether there is a target in the external environment of the vehicle.

[0006] In order to solve the above technical problems, another technical solution adopted by the present application is: to provide an electronic device, the electronic device includes a memory and a processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the above target detection method.

[0007] To solve the above technical problems, another technical solution adopted in the present application is: to provide a target detection system, including an electronic device and an interface module, the electronic device and the interface module can communicate with each other; wherein the interface module includes multiple interfaces for inserting an image sensor; the electronic device includes a memory and a processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the above target detection method.

[0008] In order to solve the above technical problems, another technical solution adopted by the present application is: providing a computer-readable storage medium, which is used to store program instructions, and the program instructions can be executed to implement the above target detection method.

[0009] The above scheme detects that an interface module is newly inserted into an image sensor, and the interface module is arranged on the vehicle. After the image sensor is inserted, it can capture images of the external environment of the vehicle. The newly inserted image sensor is configured. The image data input by the image sensor is received through the interface module. Target detection is performed based on the image data to obtain a target detection result that characterizes whether there is a target in the external environment of the vehicle. This process configures the newly added sensor by setting an interface module, so that the user can flexibly add new image sensors as needed without the manufacturer or professionals having to modify the wiring of the vehicle, thereby reducing the difficulty of integrating the image sensor. Moreover, the image sensor of the present application can be plug-and-play, so the target detection can be started quickly, thereby improving the efficiency of target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a flowchart of an embodiment of a target detection method provided by the present application;

[0011] Figure 2 is a flowchart of an embodiment of step S14 in the target detection method provided by the present application;

[0012] Figure 3 It is a flow chart of an embodiment of a target detection device of the present application;

[0013] Figure 4 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0014] Figure 5 It is a schematic diagram of the framework of an embodiment of the target detection system of the present application;

[0015] Figure 6 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and effect of the present application clearer and more specific, the present application is further described in detail below with reference to the accompanying drawings and examples.

[0017] It should be noted that the term "several" in this article means at least one, and the terms "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. The term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of types. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0018] See also Figure 1 , Figure 1 is a flow chart of an embodiment of the target detection method provided by the present application. It should be noted that if there are substantially the same results, this embodiment is not based on Figure 1 The process sequence shown is limited. Figure 1 As shown, this embodiment includes:

[0019] Step S11: It is detected that the interface module is newly inserted into the image sensor, the interface module is arranged on the vehicle, and the image sensor can collect images of the external environment of the vehicle after the insertion.

[0020] The image sensor in this article can capture light in the corresponding spectral range, convert it into an electrical signal, and generate a corresponding spectral image. For example, a visible light sensor, an infrared sensor, a thermal imaging sensor, etc. Among them, by using the above-mentioned image sensors, it is possible to capture light in different spectral ranges and generate different spectral images. For example, a visible light spectral image is generated using a visible light sensor, and an infrared spectral image is generated using an infrared sensor. And the above different spectral images can represent the diverse information in the external environment of the current vehicle. For example, the visible light spectral image can clearly show the surrounding scene in the external environment of the vehicle, including key information such as road conditions, pedestrian dynamics, and traffic signs. Infrared spectral images are good at capturing the thermal radiation information of objects at night or in low light conditions, and can indirectly obtain the image information of pedestrians or vehicles hidden behind shadows or obstacles in the external environment of the vehicle.

[0021] In one embodiment, the image sensor can be connected to the interface module through an interface method such as a parallel interface, a serial interface, or a custom interface. For example, the image sensor can be directly connected to the interface module through a parallel interface, or the image sensor can be indirectly connected to the interface module through a serial bus. There is no limitation on the method of connecting the image sensor to the interface module, and the specific method can be set according to the type of image sensor, system requirements, and interface standards.

[0022] In this embodiment, the interface module can detect that the image sensor is newly inserted by electrical connection detection, physical structure detection, communication protocol detection, etc. For example, the interface module can be designed with an electrical connection detection mechanism, such as using specific pins to detect the insertion status of the image sensor. When the image sensor is inserted into the interface module, the level state on these pins will change, thereby triggering the detection mechanism. Alternatively, a specific communication protocol is usually followed between the image sensor and the interface module. When a new image sensor is inserted, the interface module will attempt to establish communication with it and confirm the presence of the new image sensor through the device identification process in the protocol.

[0023] The interface module can be directly or indirectly arranged on the vehicle. For example, the interface module can be directly embedded in the electrical system of the vehicle, such as installed in a fixed position inside the vehicle, and connected to external devices such as image sensors through the power cord and data cable of the vehicle. Alternatively, the interface module is connected to the vehicle system through some intermediate device, such as an independent device box, adapter or device. This intermediate device can be set at a certain position inside the vehicle, such as under the dashboard or in the trunk, and communicate with the vehicle system through a dedicated connection line or wireless communication technology. In one embodiment, the interface module can be set on a fusion device, which can be set on the vehicle, wherein the fusion device can be used to process different spectral images generated by different image sensors.

[0024] Step S12: configuring the newly inserted image sensor.

[0025] The newly inserted image sensor may be configured manually, semi-automatically or automatically. For example, when it is detected that the interface module is newly inserted into the image sensor, the interface module or the fusion device automatically adjusts the configuration of the new image sensor.

[0026] In one embodiment, configuring the newly inserted image sensor includes at least one of the following steps: configuring the working parameters and communication mode of the newly inserted image sensor. In the case where the interface module includes multiple interfaces, the association information between the newly inserted image sensor and the inserted image sensor is saved, wherein the newly inserted image sensor and the inserted image sensor are respectively inserted into different interfaces, and the inserted image sensor is an image sensor that has been inserted into the interface and has completed the working parameters and communication mode, and the association information is used to perform at least one of the following instructions when indicating that both interfaces are inserted with image sensors: selecting at least one image sensor adapted to the current environment for target detection, and determining information used for target detection in image data of each image sensor when multiple image sensors are used for target detection.

[0027] The correlation information in this embodiment refers to the relationship and complementarity between different image sensor data under specific environments or conditions. For example, the correlation information includes data redundancy, environmental adaptability, time consistency, noise characteristics and feature complementarity. The above correlation information is conducive to the use of image sensors or the fusion between image sensors. For example, due to different environmental conditions, the performance of different image sensors is different. In haze weather, lidar sensors may be affected, while millimeter wave radar sensors can still provide reliable data. Alternatively, when there is noise in the surrounding environment, different image sensors have different sensitivities to noise, and some image sensors perform poorly in high-noise environments. The quality of image data can be improved by fusing low-noise data from other image sensors. Alternatively, the RGB sensor and the infrared sensor capture target image information in the same environment, and the two may provide complementary information. Among them, the RGB sensor provides color and shape information, while the infrared sensor provides heat source information under low light conditions.

[0028] In a specific embodiment, the correlation information includes at least one of data redundancy, time consistency, environmental adaptability, and noise characteristics.

[0029] In yet another specific embodiment, the operating parameter includes at least one of an operating frequency, a data transmission rate, and a data format.

[0030] In another specific embodiment, configuring the working parameters and communication mode of the newly inserted image sensor includes: identifying and obtaining sensor information of the newly inserted image sensor, and acquiring a communication protocol matching the newly inserted sensor, wherein the sensor information includes at least one of a sensor type and a sensor characteristic. Configuring the working parameters matching the sensor information for the newly inserted image sensor. And, configuring the communication mode of the newly inserted image sensor according to the communication interface protocol matching the newly inserted sensor.

[0031] In another embodiment, the relevant parameters of the image sensor inserted into the interface module can be adaptively adjusted according to the external environment of the vehicle, so as to select an image sensor suitable for target detection. In a specific embodiment, before receiving the image data collected by the image sensor through the interface module, it also includes: adjusting at least one of the operating frequency and exposure time of the image sensor based on the current environment. In the case where the interface module includes multiple interfaces and at least two interfaces are respectively inserted with image sensors, the image sensor used for target detection is determined.

[0032] In another embodiment, when the detection effect of the image sensor currently used for target detection is not good, the efficiency and accuracy of target detection can be improved by reselecting a new image sensor for target detection, adding a new image sensor to cooperate with the current image sensor, or using multiple identical image sensors to work simultaneously and realizing the cooperation of multiple image sensors through a data fusion algorithm. In a specific embodiment, in response to the image sensor currently used for target detection being blocked or failing, the image sensor used for target detection is re-determined or a new image sensor is added for target detection.

[0033] For example, the vehicle's external environmental characteristics (such as light, temperature, humidity, etc.) are analyzed in real time, and the operating frequency, exposure time and other parameters of different image sensors are dynamically adjusted according to the current external environmental characteristics to optimize the fusion effect. In strong light or foggy conditions, the working modes of visible light sensors and infrared sensors are adaptively adjusted. The collaborative mechanism between multiple image sensors allows data to compensate for each other. During pedestrian detection, if an image sensor is blocked or fails, the image data of other image sensors will be filled in in time, and the integrity of the data will be ensured through real-time multimodal fusion.

[0034] In another embodiment, the newly inserted image sensor can be configured through the interface module. For example, each image sensor can be connected through the interface in the interface module, wherein the interface can not only be used for data transmission of the image sensor, but also can set the configuration channel, that is, it allows the image sensor to be dynamically adjusted according to different task requirements. When a new image sensor is connected, the image sensor type and image sensor characteristics are automatically identified, such as the manufacturer, model, supported communication protocol, etc. On this basis, the appropriate communication protocol is automatically adjusted and configured. For example, according to the supported multiple communication protocols (such as USB, Ethernet, CAN, etc.), the appropriate communication protocol is automatically switched or configured, that is, different image sensors can dynamically select an adapted data transmission mode. According to the acquired image sensor characteristics, relevant parameters such as operating frequency, data transmission rate, data format, etc. are automatically configured. At the same time, when a new image sensor is detected to be inserted into the interface module, the characteristics of the image sensor and the correlation information with other image sensor data can be learned through the self-learning module to further enhance the coordination between multiple image sensors.

[0035] Step S13: receiving image data input by the image sensor through the interface module.

[0036] In this embodiment, the image data input by the image sensor can be a modal image directly acquired by the image sensor without any processing, or it can be an initial image feature obtained after the image sensor performs preliminary feature extraction on the modal image locally. For example, each image sensor has a preprocessing capability, that is, preliminary feature extraction or denoising operations can be performed locally on the image sensor to reduce the burden on transmission and central processors. After the modal image is preprocessed locally on the image sensor, the initial image features are obtained and output.

[0037] In one embodiment, the image data is a modal image acquired by an image sensor, or initial image features extracted by the image sensor from a modal image.

[0038] Step S14: performing target detection based on the image data to obtain a target detection result indicating whether there is a target in the external environment of the vehicle.

[0039] It should be noted that when the image sensor is always running, the acquired image data may be video image data. The video image data includes at least one frame of image data. By performing feature extraction on multiple frames of image data in the video image data, corresponding multiple target image features may be obtained. One of the target image features may be used for target detection. Contextual information (upper and lower timing information) between multiple target image features may also be used, that is, multiple target image features may be fused for target detection.

[0040] When there are at least two image sensors for target detection and the image sensors are always running, the image data obtained can be at least two video image data. Since each image sensor is independent of each other, the video image data obtained by each of them is also independent of each other. When using video image data for target detection, it is necessary to align each video image data in the time series dimension to achieve time consistency, that is, to obtain multiple frames of image data with consistent time series. Perform feature extraction on each frame of image data with consistent time series, and obtain corresponding fused image features to perform target detection. Among them, before feature extraction, each frame of image data with consistent time series corresponding to each image sensor can be subjected to image data fusion, and feature extraction can be performed on the fused image data to obtain corresponding fused image features. Alternatively, feature extraction can be performed on each frame of image data corresponding to each image sensor, and image data feature fusion can be performed on each target image feature corresponding to each image sensor with consistent time series to obtain corresponding fused image features.

[0041] In one embodiment, the interface module includes a plurality of interfaces and at least two interfaces are respectively inserted into image sensors, the modalities of the image sensors inserted into different interfaces are the same or different, at least one of the image sensors inserted into the at least two interfaces is used for target detection, and the image sensor used for target detection is a target image sensor. Target detection is performed on the image data to obtain a target detection result representing whether there is a target in the external environment of the vehicle, including: extracting target image features of the image data corresponding to each target image sensor respectively. Performing spatiotemporal alignment on each target image feature. Fusing each spatiotemporally aligned target image features to obtain a fused image feature. Performing target detection based on the fused image feature to obtain a target detection result.

[0042] Among them, target detection can be performed in the vehicle-mounted device or in the cloud. For example, the vehicle-mounted device includes a fusion device or an in-vehicle image sensor. The vehicle-mounted device is used to perform target detection on the image data of each image sensor, or the vehicle-mounted device is used to perform feature extraction on the image data of each image sensor, and the corresponding fused image features are obtained, and the fused image features are sent to the cloud, and the cloud performs target detection on the fused image features. Or the vehicle-mounted device sends the image data of each image sensor to the cloud, and the cloud obtains the corresponding fused image features to perform target detection. The specific implementation method can be adjusted according to actual needs and is not limited here. Among them, 5G or Vehicle-to-Everything communication technology can be used to ensure fast and stable data transmission between the cloud and the vehicle-mounted device.

[0043] Before target detection, task allocation can be dynamically adjusted according to real-time computing power. In a specific embodiment, before performing target detection on image data and obtaining a target detection result that characterizes whether there is a target in the external environment of the vehicle, it also includes: obtaining resource occupancy assessments of each task in target detection, wherein each task includes a feature extraction task, a spatiotemporal alignment task, a fusion task, and a target detection task, and the resource occupancy assessment characterizes the occupancy amount or occupancy ratio of the corresponding task to the processing resources in the processing unit. According to the resource occupancy assessment corresponding to each task, the processing unit to execute each task is determined. In another specific embodiment, according to the resource occupancy assessment corresponding to each task, it is determined whether a heterogeneous computing unit needs to be added to the task, wherein the task of adding a heterogeneous computing unit is collaboratively executed by the processing unit corresponding to the task and the added heterogeneous computing unit. Among them, the resource share assessment of each task in target detection can be obtained by calculation or table lookup mapping.

[0044] For example, in target detection tasks, if computing resources are tight, image data from key areas (such as the center of the road) will be processed first to ensure that information that is critical to the task can be analyzed in a timely manner. Correspondingly, sensor data from other non-critical areas may be processed later or in a more streamlined manner. Therefore, in terms of resource occupancy assessment, image data from key areas will naturally be given a higher priority than image data from other areas. In addition, the urgency and complexity of the task can also be important considerations when evaluating resource occupancy. For those urgent and complex tasks, the system will give a higher resource occupancy assessment to ensure that these tasks can receive timely and sufficient processing resources. For different types of tasks, they can be assigned to the most suitable heterogeneous computing units for execution. For example, complex neural network inference tasks, due to their computational intensity and the need for high-performance computing units, have a higher resource occupancy assessment and can be handled by a dedicated neural processing unit (NPU). For some relatively simple preprocessing tasks, such as data format conversion or basic data cleaning, the resource occupancy is low, and they can be chosen to be completed by the central processing unit (CPU).

[0045] In another specific implementation, in scenarios where the efficiency of the fusion device needs to be further improved, FPGA (field programmable gate array) can be used to accelerate some key algorithms. FPGA has efficient parallel computing capabilities and can accelerate the fusion and reasoning process of multimodal data while maintaining low power consumption. Alternatively, a dedicated AI acceleration chip (such as Tesla's FSD chip or Mobileye) is integrated to optimize the operating efficiency of the deep learning model and support real-time target detection and feature fusion from the hardware level.

[0046] In yet another embodiment, the target is a pedestrian.

[0047] This embodiment detects that an interface module is newly inserted into an image sensor. The interface module is arranged on the vehicle. After being inserted, the image sensor can capture images of the external environment of the vehicle. The newly inserted image sensor is configured. The image data input by the image sensor is received through the interface module. Target detection is performed based on the image data to obtain a target detection result that characterizes whether there is a target in the external environment of the vehicle. This process configures the newly added sensor by setting up an interface module, so that the user can flexibly add new image sensors as needed without the need for manufacturers or professionals to modify the wiring of the vehicle, thereby reducing the difficulty of integrating the image sensor. Moreover, the image sensor of the present application can be plug-and-play, so that target detection can be started quickly, thereby improving the efficiency of target detection.

[0048] See also Figure 2 , Figure 2 is a flow chart of an embodiment of step S14 in the target detection method provided in the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to step S14. Figure 2 The process sequence shown is limited. Figure 2 As shown, this embodiment includes:

[0049] Step S21: extracting target image features of image data corresponding to each target image sensor respectively.

[0050] The target image features of the image data corresponding to each target image sensor include multi-level and multi-resolution image features. Among them, the multi-level and multi-resolution image features corresponding to each target image sensor can be obtained and feature fused to obtain the corresponding target image features. Among them, the feature extraction of image data can be performed through a feature extraction network, a focus detection algorithm, an edge detection algorithm, etc. The multi-level and multi-resolution image features can be fused through a feature pyramid, a wavelet transform, a convolutional neural network, an attention mechanism, etc. to obtain the corresponding target image features.

[0051] For example, feature extraction may be performed layer by layer on the image data corresponding to each target image sensor, and the image features extracted layer by layer may be merged to obtain the corresponding target image features. Alternatively, feature extraction of different resolutions may be performed on the image data corresponding to each target image sensor based on different levels, and the extracted features may be merged to obtain the corresponding target image features.

[0052] In one embodiment, target image features of image data corresponding to each target image sensor are extracted respectively, including: for image data corresponding to each target image sensor, a plurality of image features are extracted from the image data using a feature extraction network, wherein the plurality of image features include image features corresponding to multiple levels, and the image features corresponding to each level include image features corresponding to multiple resolutions. Based on the plurality of image features, the target image features of the image data corresponding to the target image sensor are obtained. The target image sensors corresponding to different modalities use different feature extraction networks.

[0053] For example, feature extraction is performed using the feature extraction network corresponding to the target image sensor, including: extracting features independently from image data at each level at different resolutions. And further extracting features at each resolution by level (shallow, middle, deep). At a specific level (such as shallow, middle or deep), the features extracted from each resolution are merged. The importance of image features at different resolutions is dynamically adjusted using an attention mechanism or weighted merging. Through a multi-resolution feature pyramid structure, image features at different resolutions and levels are fused in the same framework to obtain target image features corresponding to each target image sensor. Among them, low-resolution images can capture a wide range of overall information for preliminary detection of global background and large-scale objects. High-resolution images focus on details and are used to capture small-scale objects and local information. The information captured at different levels gradually changes from simple features such as low-level edges and textures to complex features at high levels (such as object relationships and semantics).

[0054] In a specific embodiment, the feature extraction network corresponding to the target image sensor can be combined with multi-scale spatiotemporal convolution fusion to obtain target image features with spatiotemporal characteristics. The image features extracted by the feature extraction network are further extracted, and multi-scale convolution kernels are applied in the spatiotemporal dimension to extract image features at different time periods and different spatial resolutions as target image features, which can better capture the local spatiotemporal dependencies of the data. For example, a larger convolution kernel can capture a wide range of environmental changes, while a smaller convolution kernel can focus on detecting local features of the target. For target detection scenarios, the algorithm will extract features of multiple scales from different spectra and time periods to ensure that the target can still be accurately identified in complex scenarios such as distance and lighting changes.

[0055] Furthermore, multi-level and multi-resolution image features in image data can be obtained by integrating different types of processing units and algorithms. In one embodiment, the extraction of image features at each level and resolution is performed by the first processing unit, or multiple image features are divided into a first type of image features and a second type of image features, the first type of image features are image features with a level higher than a preset level and / or a resolution higher than a preset resolution, the first type of image features are executed by the first processing unit, the second type of image features are executed by the second processing unit, and the first processing unit and the second processing unit meet at least one of the following conditions: the performance of the first processing unit is higher than that of the second processing unit, the first processing unit is a cloud processing unit, and the second processing unit is a local processing unit.

[0056] For example, use local devices to perform preliminary and rapid feature extraction and preprocessing of image data to meet real-time requirements under low-latency conditions. Use low-performance devices to process low-level, low-resolution feature extraction, which is suitable for preprocessing and fast response tasks and has a smaller data processing load. Use high-performance devices to perform complex high-resolution feature extraction and deep processing or to fuse the final multi-level and multi-resolution image features.

[0057] Step S22: performing spatiotemporal alignment on the features of each target image.

[0058] Since each target image feature corresponds to a target image feature corresponding to a certain frame, before fusing the target image features, it is necessary to align the target image features in time and space to ensure the consistency of the target image features in time sequence. That is, when the target image features corresponding to each target image sensor are obtained, the target image features can be aligned in time and space.

[0059] In this paper, spatiotemporal alignment refers to converting the image features captured by different image sensors at different times and different spatial positions into a unified spatiotemporal framework for feature fusion. Before target detection, each target image feature can be preprocessed to reduce the heterogeneity between different target image features. For example, each target image feature can be normalized, spatiotemporally transformed, etc.

[0060] In one embodiment, performing spatiotemporal alignment on each target image feature includes: preprocessing each target image feature, performing spatiotemporal alignment on each preprocessed target image feature using an alignment model, and obtaining spatiotemporal aligned target image features, wherein the alignment model is generated using a generative adversarial network.

[0061] In a specific embodiment, the target image features are preprocessed before target detection to ensure the standardization and consistency of the target image features. The target image features at different time or spatial positions are transformed using spatiotemporal transformation to adapt them to a unified reference frame, or the target image features corresponding to different target image sensors are standardized using normalization, such as by linear transformation or Z-score normalization to compare different target image features on the same scale, or the target image features are pre-aligned before the alignment model performs spatiotemporal alignment on the target image features to map them to the same spatiotemporal feature space.

[0062] In another specific implementation, each target image feature may be preprocessed, each preprocessed target image feature may be encoded to obtain a corresponding target encoded image feature, and each target encoded image feature may be spatiotemporally aligned using an alignment model to obtain each spatiotemporally aligned target image feature.

[0063] Among them, modality-specific encoders can be introduced at the input end to target the heterogeneity of image data corresponding to different image sensors. Each image sensor corresponds to an independent Transformer encoder, which processes features from images of different modalities respectively. For example, the visible light encoder: mainly extracts visually clear details, such as the contour and color information of the target. Infrared light encoder: focuses on extracting heat source data, especially at night or in bad weather, infrared sensors can capture human body heat signals. Thermal imaging encoder: extracts heat signals based on temperature differences to identify targets within a specific temperature range. Each encoder learns the high-order semantic features of its own modality through an independent Transformer module, and exchanges information through a shared attention mechanism in the subsequent fusion stage.

[0064] In another embodiment, the training process of the alignment model includes: after encoding the features of different target images, the features of different spectra are mapped to similar distributions in the latent space through adversarial training of a generative adversarial network (GAN). By introducing spatiotemporal constraints in the discriminator of the generative adversarial network, it is ensured that the fused features are not only aligned in the spatial dimension, but also consistent in the temporal dimension. By continuously adjusting the corresponding parameters, the alignment model is trained so that the alignment model has the ability to perform spatiotemporal alignment on the features of each target image. For example, in target detection, infrared spectrum is used at night, and visible spectrum is used during the day. These two types of data are automatically aligned in the latent space through adversarial learning to ensure the spatiotemporal consistency after feature fusion. Among them, the spatiotemporal constraints are mainly used to ensure the alignment consistency between different spectral features in the time sequence and space dimensions. The spatiotemporal constraints generally include: Time step: control the interval between different time nodes to balance real-time and information integrity. For example, spectral features are acquired every 0.5 seconds to ensure that the target is not lost due to a long sampling interval. Spatial resolution: The resolution of different image sensors is usually different and needs to be adjusted to the same benchmark. Taking the camera as an example, high-resolution images can be downsampled or low-resolution images can be super-resolved to align the features of different image sensors in the spatial dimension. Sampling Frequency: Control the data acquisition frequency to keep all image data synchronized. For example, the sampling frequencies of infrared cameras and visible light cameras are different, which will cause data timing misalignment. Adjusting the sampling frequency can synchronize sampling. Frame Alignment: Ensure that each frame corresponds to the same physical moment in time to avoid misalignment due to sensor delay or system delay. For example, align the images of pedestrians taken by infrared and visible light cameras by frame sequence numbering. Spectral Calibration: Adjust the performance differences of different image sensors for the same target to keep the encoding of image features consistent in the latent space.

[0065] In yet another embodiment, before performing spatiotemporal alignment on each target image feature, a first attention weight corresponding to each target image feature is determined, and the spatiotemporal alignment is performed based on the first attention weight.

[0066] In a specific implementation, before using the alignment model to perform spatiotemporal alignment on each target image feature, the weight of each target image feature may be dynamically adjusted to determine the corresponding first attention weight.

[0067] For example, the key area in the target image feature can be determined by the corresponding first attention weight, and the target image features corresponding to the key area are focused on spatial and temporal alignment. Alternatively, the target image features corresponding to the image sensors corresponding to the key frame are determined by the corresponding first attention weight, and the target image features corresponding to the key frame are focused on spatial and temporal alignment.

[0068] Step S23: Fusing the spatiotemporally aligned target image features to obtain fused image features.

[0069] The target image features aligned in time and space can be directly fused, weighted fused or context fused to obtain the corresponding fused image features. For example, the target image features aligned in time and space are directly fused to obtain the corresponding fused image features. Alternatively, different attention weights are assigned to the target image features aligned in time and space, and feature fusion is performed to obtain the corresponding fused image features. Alternatively, the target image features of the target image sensors aligned in time and space corresponding to different time sequences are fused (such as the target image features of the previous frame are fused with the target image features of the current frame), and the target image features of the target image sensors after fusion are fused to obtain the corresponding fused image features. Alternatively, the target image features corresponding to the target image sensors corresponding to the current frame and the previous frame are fused to obtain the current frame fused image data and the previous frame fused image data, and the current frame fused image data and the previous frame fused image data are fused to obtain the corresponding fused image features. The specific implementation method can be adjusted according to the actual situation and is not limited here.

[0070] In one embodiment, the target image features aligned in time and space are fused to obtain the fused image features, including: determining the second attention weight corresponding to each target image feature based on a weight reference factor, wherein the weight reference factor includes at least one of the current environment, the image acquisition time, and the sensor characteristics. The target image features aligned in time and space are fused using the second attention weights to obtain the fused image features.

[0071] For example, in a daytime environment, the second attention weight of the visible light sensor corresponding to the target image feature can be adaptively increased. At night or in low light conditions, the second attention weight of the infrared sensor corresponding to the target image feature or the thermal imaging sensor corresponding to the target image feature can be adaptively increased. In the case of high noise in the current environment, the second attention weight of the thermal noise sensor corresponding to the target image feature can be adaptively increased. In the case of high sensitivity of the target image sensor, its corresponding second attention weight can be adaptively increased.

[0072] Specifically, the second attention weight corresponding to each target image feature may be determined by means of a covariance matrix, an adaptive weighting algorithm, or an adaptive weight matrix.

[0073] In a specific implementation, the target image features corresponding to each target image sensor include target image features corresponding to the historical frame and target image features corresponding to the current frame. Using each second attention weight, fusing each spatiotemporally aligned target image features to obtain a fused image feature, including: using each second attention weight, fusing the target image features corresponding to the current frame of each target image sensor to obtain a fused image feature corresponding to the current frame.

[0074] In another specific embodiment, after using each second attention weight to fuse each target image feature that has been aligned in time and space to obtain a fused image feature, it also includes: using the third attention weights corresponding to the historical frame and the current frame respectively, to fuse the fused image features corresponding to the historical frame and the fused image features corresponding to the current frame to obtain a new fused image feature, and the new fused image feature is used for target detection.

[0075] Step S24: Perform target detection based on the fused image features to obtain target detection results.

[0076] In one embodiment, target detection is performed based on fused image features to obtain target detection results, including: using a target detection model to perform local target detection and global target detection based on the fused image features, respectively, to obtain local detection results and global detection results. The local detection results and the global detection results are combined to obtain the target detection results. The local target detection and the global target detection are performed based on the local attention mechanism and the global attention mechanism, respectively.

[0077] Among them, a multi-resolution feature pyramid can be used to integrate local detection results and global detection results. For example, by using target image features of different resolutions, the target detection results from global target detection and local target detection are fused. This process can effectively process target information of different scales and ensure that the target detection model can detect targets of various scales. The target detection results of global target detection and local target detection are integrated through a fusion layer (multi-resolution pyramid) to improve the overall detection accuracy.

[0078] The attention distribution mechanism can be used to generate local target detection results, global target detection results, and visualization images corresponding to the target detection results, respectively, to show the weights of different modal images in target detection (such as the first attention weight). In a specific embodiment, the attention distribution mechanism of Transformer is used to generate an attention heat map for each target detection result to show the weights of different modalities in the target detection process. For example, the system can show whether a certain target detection is based on infrared thermal image or visible light image detection.

[0079] In yet another embodiment, the object detection model is implemented using a bidirectional spatiotemporal transformer network.

[0080] The bidirectional space-time transformation network in this embodiment refers to the effective fusion of temporal context features through a bidirectional time stream structure and a recurrent network, and the use of a multi-stage feature fusion mechanism and a prediction method based on inter-frame features for target detection. Among them, the bidirectional time stream in the bidirectional space-time transformation network generally refers to the use of a bidirectional recurrent neural network (Bi-RNN) architecture, which can simultaneously consider forward and backward time series information. A bidirectional recurrent network refers to a special recurrent neural network (RNN), which includes a forward RNN and a backward RNN. The corresponding fused image features are obtained through the above-mentioned bidirectional space-time transformation network, and target detection is performed based on the key features in the fused image features. Among them, the key features in the fused image features can be fused image features corresponding to key areas or key frames (a frame in the historical frame or the current frame corresponding to the third attention weight being larger).

[0081] For example, when the historical frame includes the previous frame, the second attention weight is used to obtain the fused image features corresponding to the current frame and the fused image features corresponding to the previous frame. The fused image features corresponding to the current frame are preliminarily fused with the fused image features of the previous frame. And through the attention mechanism (third attention weight) or feature mapping, the fused image features corresponding to the previous frame or the current frame are fused to obtain the corresponding fused image features. Combine the features of the current frame and the previous frame to obtain key features, and perform target prediction for the next frame, that is, perform target detection. Alternatively, combine the image features of the current frame and the previous frame to obtain key features, and perform target detection on the current frame, or combine the image features of the historical frame and the current frame to obtain key features and perform target detection on the current frame.

[0082] The alignment model, target detection model, etc. in this article can be made lightweight by model pruning or knowledge distillation. For example, model pruning is used to remove redundant parameters or nodes in the alignment model and target detection model that are not important for the target detection task. Among them, redundant parameters or nodes are parts that contribute little to the target detection task and have little impact on model performance but consume a lot of computing resources. Alternatively, knowledge distillation is a method that transfers knowledge from a large and complex model (teacher model) to a smaller and lightweight model (student model). First, a teacher model with good performance is trained to perform high-precision pedestrian detection. Knowledge is extracted from the teacher model, such as intermediate layer features, output probability distribution, etc. The extracted knowledge is used to train a smaller student model so that it can imitate the behavior of the teacher model and achieve similar detection performance.

[0083] In another embodiment, the vehicle is further provided with a reference sensor, and the reference sensor collects local environment information of the vehicle. Before performing target detection based on the fused image features to obtain the target detection result, it also includes: extracting features from the local environment information collected by the reference sensor to obtain local features. Performing target detection based on the fused image features to obtain the target detection result includes: fusing the local features with the fused image features to obtain the target fused features. Performing target detection based on the target fused features to obtain the target detection result.

[0084] In a specific embodiment, the reference sensor includes at least one of an image sensor, a temperature sensor, and a humidity sensor.

[0085] In another specific embodiment, the local features and the fused image features are fused to obtain the target fused features, including: adjusting the importance weights corresponding to the local features and the fused image features, respectively. Based on the importance weights, the local features and the fused image features are fused to obtain the target fused features.

[0086] In another specific embodiment, before fusing the local features with the fused image features to obtain the target fused features, the method further includes: performing local context analysis on the local features to obtain context information, wherein the context information includes at least one of temporal context information and spatial context information. The local features are enhanced using the context information.

[0087] For example, local environmental information from reference sensors is collected in real time under a specific environment. This information may include visual information (such as camera images), temperature, humidity and other data. Use a deep learning model (such as CNN or other feature extraction networks) to extract features from the collected local environmental information and obtain local features. Perform local context analysis on the extracted local features and enhance the local features using temporal context or spatial context information. Use importance weights to fuse the extracted local features with the fused image features to obtain target fusion features. The edge sensor can send target fusion features or local features to the fusion device or cloud for target detection. Among them, through self-learning or online learning mechanisms, the reference sensor can continuously adjust the parameters of its fusion algorithm according to changes in the external environment to improve the accuracy of detection or recognition.

[0088] In yet another embodiment, the target is a pedestrian.

[0089] For example, in intelligent driving, vehicle-based target detection can be pedestrian detection. A fusion device can be set on the vehicle for pedestrian detection. Among them, the fusion device is provided with an interface module, a low-performance device, a high-performance device, etc., and the interface module is provided with at least one interface. The image sensor can be connected through the interface in the interface module. When a newly inserted image sensor is detected in the interface module, the type and characteristics of the image sensor are obtained, the correlation information between the image sensor and the inserted image sensor is obtained, and the current environment information of the vehicle is obtained. Based on the type, characteristics, correlation information, and current environment information of the above-mentioned image sensors, the newly inserted image sensor is configured. At the same time, the inserted image sensor can also be adaptively adjusted. When the inserted image sensor is blocked or invalid, other image sensors will be adjusted in time. Through the above steps, plug-and-play of the image sensor is realized, showing the automation and reliability of the fusion device. The image features corresponding to each image sensor are obtained by using the convolutional neural network corresponding to each image sensor. The image features are aligned in time and space by using the cross-spectral spatiotemporal transformation and normalization algorithm, spatiotemporal alignment and feature adversarial learning. The cross-modal attention mechanism is used to perform weighted fusion of image features from different modalities (such as different image sensors or different spectra). The attention weight is dynamically adjusted according to the environmental signal to determine the importance of different image features in the fusion process and obtain the corresponding fused image features. The real-time performance of pedestrian detection is guaranteed by the efficient spatiotemporal alignment of the fusion device and the deep fusion capability of the Transformer model. The CPU is used for preliminary image feature extraction or simple feature fusion on low-performance devices, and the NPU is used for complex image feature extraction on high-performance devices (such as the cloud). Through 5G or Vehicle-to-Everything communication technology, the data transmission between the cloud and low-performance devices is ensured to be fast and stable. The fused image features of the historical frame are fused with the fused image features of the current frame using a bidirectional spatiotemporal transformation network. By processing the fused image features of consecutive frames in the autonomous driving scenario, the Transformer can capture features such as pedestrian movement trajectory and behavioral changes before and after occlusion for subsequent target detection. In the pedestrian detection task, the key areas are often concentrated in the area where the pedestrians are located, and there is a large amount of irrelevant background in the entire image. The local attention mechanism can be used to guide the Transformer to focus on areas where pedestrians may appear (such as the road surface in front of the vehicle, crosswalks, etc.) for pedestrian detection. At the same time, the global attention mechanism is used to perform pedestrian detection on the whole. By using feature maps of different resolutions, the pedestrian detection results from global pedestrian detection and local pedestrian detection are fused to obtain the final pedestrian detection result. Pedestrian detection is performed based on the temporal context fusion features (fused image features of historical frames and current frames) extracted in the feature fusion step.Through temporal context fusion, feature information from different time frames can be integrated. This context information helps to improve the accuracy and robustness of pedestrian detection. When performing any of the above steps, some algorithms can be accelerated by using FPGA (field programmable gate array) or AI acceleration chips. The flexible plug-and-play architecture and multimodal fusion capabilities of the fusion device enable it to perform better pedestrian detection in autonomous driving technology, especially in complex and harsh driving environments.

[0090] See also Figure 3 , Figure 3 It is a flow chart of an embodiment of the target detection device of the present application. The target detection device 300 includes an image acquisition module 310, a configuration module 320, an image data acquisition module 330, and a target detection module 340. The image acquisition module 310 is used to detect that the interface module is newly inserted into the image sensor, and the interface module is arranged on the vehicle. After the insertion, the image sensor can perform image acquisition on the external environment of the vehicle. The configuration module 320 is used to configure the newly inserted image sensor. The image data acquisition module 330 is used to receive image data input by the image sensor through the interface module. The target detection module 340 is used to perform target detection based on the image data to obtain a target detection result that characterizes whether there is a target in the external environment of the vehicle.

[0091] In some embodiments, the configuration module 320, when executing the configuration of the newly inserted image sensor, includes at least one of the following steps: configuring the working parameters and communication mode of the newly inserted image sensor. In the case where the interface module includes multiple interfaces, the association information between the newly inserted image sensor and the inserted image sensor is saved, wherein the newly inserted image sensor and the inserted image sensor are respectively inserted into different interfaces, and the inserted image sensor is an image sensor that has been inserted into the interface and has completed the working parameters and communication mode, and the association information is used to perform at least one of the following instructions when indicating that both interfaces are inserted with image sensors: selecting at least one image sensor adapted to the current environment for target detection, and determining the information used for target detection in the image data of each image sensor when multiple image sensors are used for target detection.

[0092] In some embodiments, the relevance information includes at least one of data redundancy, time consistency, environmental adaptability, and noise characteristics.

[0093] In some embodiments, the operating parameter includes at least one of an operating frequency, a data transmission rate, and a data format.

[0094] In some embodiments, the configuration module 320 configures the working parameters and communication mode of the newly inserted image sensor, including: identifying and obtaining sensor information of the newly inserted image sensor, and obtaining a communication protocol matching the newly inserted sensor, wherein the sensor information includes at least one of a sensor type and a sensor characteristic. The working parameters matching the sensor information are configured for the newly inserted image sensor. Furthermore, the communication mode of the newly inserted image sensor is configured according to the communication interface protocol matching the newly inserted sensor.

[0095] In some embodiments, before the image data acquisition module 330 receives the image data collected by the image sensor through the interface module, it also includes: adjusting at least one of the operating frequency and exposure time of the image sensor based on the current environment. In the case where the interface module includes multiple interfaces and at least two interfaces are respectively inserted into the image sensor, determining the image sensor used for target detection, and / or, in response to the image sensor currently used for target detection being blocked or invalid, re-determining the image sensor used for target detection or adding a new image sensor used for target detection.

[0096] In some embodiments, the interface module includes multiple interfaces and at least two interfaces are respectively inserted into image sensors, the modalities of the image sensors inserted into different interfaces are the same or different, at least one of the image sensors inserted into at least two interfaces is used for target detection, and the image sensor used for target detection is a target image sensor. The target detection module 340 performs target detection on the image data to obtain a target detection result that characterizes whether there is a target in the external environment of the vehicle, including: extracting target image features of the image data corresponding to each target image sensor respectively. Performing spatiotemporal alignment on each target image feature. Fusing the spatiotemporally aligned target image features to obtain a fused image feature. Performing target detection based on the fused image feature to obtain a target detection result.

[0097] In some embodiments, before the target detection module 340 performs target detection on the image data and obtains the target detection result characterizing whether there is a target in the external environment of the vehicle, it also includes: obtaining the resource occupancy evaluation of each task in the target detection, wherein each task includes a feature extraction task, a spatiotemporal alignment task, a fusion task, and a target detection task, and the resource occupancy evaluation characterizes the occupancy amount or occupancy ratio of the corresponding task to the processing resources in the processing unit. According to the resource occupancy evaluation corresponding to each task, the processing unit to execute each task is determined, and / or, according to the resource occupancy evaluation corresponding to each task, it is determined whether a heterogeneous computing unit needs to be added for the task, wherein the task of adding a heterogeneous computing unit is executed collaboratively by the processing unit corresponding to the task and the added heterogeneous computing unit.

[0098] In some embodiments, the target detection module 340 performs the steps of extracting target image features of image data corresponding to each target image sensor, including: for image data corresponding to each target image sensor, extracting multiple image features from the image data using a feature extraction network, wherein the multiple image features include image features corresponding to multiple levels, and the image features corresponding to each level include image features corresponding to multiple resolutions. Based on the multiple image features, the target image features of the image data corresponding to the target image sensor are obtained. Among them, the target image sensors corresponding to different modalities use different feature extraction networks.

[0099] In some embodiments, the extraction of image features at each level and each resolution is performed by the first processing unit, or multiple image features are divided into a first category of image features and a second category of image features, the first category of image features are image features with a level higher than a preset level and / or a resolution higher than a preset resolution, the first category of image features are performed by the first processing unit, the second category of image features are performed by the second processing unit, and the first processing unit and the second processing unit meet at least one of the following conditions: the performance of the first processing unit is higher than that of the second processing unit, the first processing unit is a cloud-based processing unit, and the second processing unit is a local processing unit.

[0100] In some embodiments, the target detection module 340 performs the spatiotemporal alignment of each target image feature, including: preprocessing each target image feature. Using the alignment model to perform spatiotemporal alignment on each preprocessed target image feature to obtain each spatiotemporal aligned target image feature. The alignment model is generated using a generative adversarial network. And / or, before performing the spatiotemporal alignment on each target image feature, determining a first attention weight corresponding to each target image feature, and performing the spatiotemporal alignment based on the first attention weight.

[0101] In some embodiments, the target detection module 340 performs the fusion of each target image feature that is aligned in time and space to obtain the fused image feature, including: determining the second attention weight corresponding to each target image feature based on the weight reference factor, wherein the weight reference factor includes at least one of the current environment, the image acquisition time, and the sensor characteristics. Using each second attention weight, the target image features that are aligned in time and space are fused to obtain the fused image feature.

[0102] In some embodiments, the target image features corresponding to each target image sensor include target image features corresponding to the historical frame and target image features corresponding to the current frame. The target detection module 340 uses the second attention weights to fuse the target image features aligned in time and space to obtain the fused image features, including: using the second attention weights to fuse the target image features corresponding to the current frame of each target image sensor to obtain the fused image features corresponding to the current frame. After the target detection module 340 uses the second attention weights to fuse the target image features aligned in time and space to obtain the fused image features, it also includes: using the third attention weights corresponding to the historical frame and the current frame respectively to fuse the fused image features corresponding to the historical frame and the fused image features corresponding to the current frame to obtain a new fused image feature, and the new fused image feature is used for target detection.

[0103] In some embodiments, the target detection module 340 performs target detection based on the fused image features to obtain the target detection result, including: using the target detection model to perform local target detection and global target detection based on the fused image features, respectively, to obtain local detection results and global detection results. The local detection results and the global detection results are combined to obtain the target detection result. Among them, the local target detection and the global target detection are performed based on the local attention mechanism and the global attention mechanism, respectively, and / or the target detection model is implemented using a bidirectional spatiotemporal transformation network.

[0104] In some embodiments, the vehicle is further provided with a reference sensor, which collects local environmental information of the vehicle. Before the target detection module 340 performs target detection based on the fused image features to obtain the target detection result, it also includes: extracting features from the local environmental information collected by the reference sensor to obtain local features. When the target detection module 340 performs target detection based on the fused image features to obtain the target detection result, it includes: fusing the local features with the fused image features to obtain the target fused features. Target detection is performed based on the target fused features to obtain the target detection result.

[0105] In some embodiments, the reference sensor includes at least one of an image sensor, a temperature sensor, and a humidity sensor.

[0106] In some embodiments, the target detection module 340 performs the fusion of the local features and the fused image features to obtain the target fused features, including: adjusting the importance weights corresponding to the local features and the fused image features, respectively. Based on the importance weights, the local features and the fused image features are fused to obtain the target fused features.

[0107] In some embodiments, before the target detection module 340 performs the fusion of the local features and the fused image features to obtain the target fused features, it further includes: performing local context analysis on the local features to obtain context information, where the context information includes at least one of temporal context information and spatial context information. The local features are enhanced using the context information.

[0108] In some embodiments, the target is a pedestrian.

[0109] In some embodiments, the image data is a modal image acquired by an image sensor, or initial image features extracted by the image sensor from a modal image.

[0110] See also Figure 4 , Figure 4 4 is a schematic diagram of a framework of an embodiment of an electronic device of the present application. The electronic device 40 includes a memory 41 and a processor 42 coupled to each other, and the processor 42 is used to execute program instructions stored in the memory 41 to implement the steps in any of the above target detection method embodiments. In a specific implementation scenario, the electronic device 40 may include but is not limited to: a microcomputer, a server, and in addition, the electronic device 40 may also include a mobile device such as a laptop computer and a tablet computer, which is not limited here.

[0111] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned target detection method embodiments. The processor 42 can also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip having signal processing capabilities. The processor 42 may also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. In addition, the processor 42 may be implemented by an integrated circuit chip.

[0112] See also Figure 5 , Figure 5 The configuration system 50 of the application service includes an electronic device 51 and an interface module 52. The electronic device 51 and the interface module 52 can communicate with each other. The interface module 52 includes multiple interfaces for inserting an image sensor. The electronic device 51 is Figure 4 The electronic device 40 described in.

[0113] See also Figure 6 , Figure 6 The computer-readable storage medium 60 stores program instructions 61 that can be executed by a processor, and the program instructions 61 are used to implement the steps in any of the above target detection method embodiments.

[0114] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0115] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0116] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0117] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0118] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

Claims

1. A target detection method, characterized in that: The method comprises: Detecting that an interface module is newly inserted into an image sensor, the interface module being provided on a vehicle, and the image sensor being able to capture images of an external environment of the vehicle after being inserted; configuring the newly inserted image sensor; receiving image data input by the image sensor through the interface module; Target detection is performed based on the image data to obtain a target detection result indicating whether the target exists in the external environment of the vehicle.

2. The method according to claim 1, characterized in that: The configuring of the newly inserted image sensor comprises at least one of the following steps: Configuring the operating parameters and communication mode of the newly inserted image sensor; In the case where the interface module includes multiple interfaces, the association information between the newly inserted image sensor and the already inserted image sensor is saved, wherein the newly inserted image sensor and the already inserted image sensor are respectively inserted into different interfaces, and the already inserted image sensor is an image sensor that has been inserted into the interface and completed the working parameters and communication methods, and the association information is used to provide at least one of the following guidance when indicating that image sensors are inserted into two of the interfaces: selecting at least one image sensor adapted to the current environment for target detection, and determining information used for target detection in image data of each image sensor when multiple image sensors are used for target detection.

3. The method according to claim 2, characterized in that The association information includes at least one of data redundancy, time consistency, environmental adaptability and noise characteristics; And / or, the operating parameters include at least one of an operating frequency, a data transmission rate, and a data format; And / or, configuring the working parameters and communication mode of the newly inserted image sensor includes: Identify and obtain sensor information of the newly inserted image sensor, and acquire a communication protocol matching the newly inserted sensor, wherein the sensor information includes at least one of a sensor type and a sensor characteristic; configuring operating parameters matching the sensor information for the newly inserted image sensor; and, The communication mode of the newly inserted image sensor is configured according to the communication interface protocol matching the newly inserted sensor.

4. The method according to claim 1, characterized in that Before the image data collected by the image sensor is received through the interface module, the method further includes: adjusting at least one of an operating frequency and an exposure time of the image sensor based on a current environment; In the case where the interface module includes multiple interfaces and at least two of the interfaces are respectively inserted with image sensors, the image sensor used for target detection is determined, and / or, in response to the image sensor currently used for target detection being blocked or failing, the image sensor used for target detection is re-determined or a new image sensor used for target detection is added.

5. The method according to claim 1, characterized in that: The interface module comprises a plurality of interfaces and at least two interfaces are respectively inserted into image sensors, the image sensors inserted into different interfaces have the same or different modes, at least one of the image sensors inserted into the at least two interfaces is used for target detection, and the image sensor used for target detection is a target image sensor; The performing target detection on the image data to obtain a target detection result indicating whether the target exists in the external environment of the vehicle includes: Respectively extracting target image features of the image data corresponding to each of the target image sensors; Performing spatiotemporal alignment on each of the target image features; Fusing the target image features that have been aligned in time and space to obtain fused image features; Target detection is performed based on the fused image features to obtain the target detection result.

6. The method according to claim 5, characterized in that Before performing target detection on the image data to obtain a target detection result indicating whether the target exists in the external environment of the vehicle, the method further includes: Respectively obtain resource occupancy evaluations of the tasks in the target detection, wherein the tasks include feature extraction tasks, spatiotemporal alignment tasks, fusion tasks, and target detection tasks, and the resource occupancy evaluations represent the occupancy amount or occupancy ratio of the corresponding tasks to the processing resources in the processing unit; Determine the processing unit that executes each of the tasks according to the resource occupancy assessment corresponding to each of the tasks, and / or determine whether it is necessary to add a heterogeneous computing unit for the task according to the resource occupancy assessment corresponding to each of the tasks, wherein the task of adding the heterogeneous computing unit is collaboratively executed by the processing unit corresponding to the task and the added heterogeneous computing unit.

7. The method according to claim 5, characterized in that The step of respectively extracting target image features of image data corresponding to each target image sensor comprises: For the image data corresponding to each of the target image sensors, a plurality of image features are extracted from the image data using a feature extraction network, wherein the plurality of image features include image features corresponding to multiple levels, and the image features corresponding to each level include image features corresponding to multiple resolutions; Based on the multiple image features, obtaining target image features of the image data corresponding to the target image sensor; Wherein, the target image sensors corresponding to different modalities use different feature extraction networks; And / or, the extraction of image features at each of the levels and each of the resolutions is performed by the first processing unit, or the multiple image features are divided into a first category of image features and a second category of image features, the first category of image features are image features whose levels are higher than a preset level and / or whose resolutions are higher than a preset resolution, the first category of image features are performed by the first processing unit, and the second category of image features are performed by the second processing unit, and the first processing unit and the second processing unit meet at least one of the following conditions: the performance of the first processing unit is higher than that of the second processing unit, the first processing unit is a cloud-based processing unit, and the second processing unit is a local processing unit.

8. The method according to claim 5, characterized in that The performing spatiotemporal alignment on each of the target image features comprises: Preprocessing each of the target image features; Using an alignment model, performing spatiotemporal alignment on the preprocessed target image features to obtain spatiotemporal aligned target image features; Wherein, the alignment model is generated by using a generative adversarial network; and / or, before performing spatiotemporal alignment on each of the target image features, a first attention weight corresponding to each of the target image features is determined, and the spatiotemporal alignment is performed based on the first attention weight.

9. The method according to claim 5, characterized in that The step of fusing the spatiotemporally aligned target image features to obtain fused image features includes: Determining a second attention weight corresponding to each of the target image features based on a weight reference factor, wherein the weight reference factor includes at least one of a current environment, an image acquisition moment, and a sensor characteristic; The second attention weights are used to fuse the spatiotemporally aligned target image features to obtain fused image features.

10. The method according to claim 9, characterized in that The target image features corresponding to each of the target image sensors include target image features corresponding to the historical frame and target image features corresponding to the current frame; The step of fusing the spatiotemporally aligned target image features using the second attention weights to obtain fused image features includes: Using each of the second attention weights, fusing the target image features corresponding to the current frame of each of the target image sensors to obtain the fused image features corresponding to the current frame; After fusing the spatiotemporally aligned target image features using the second attention weights to obtain fused image features, the method further includes: The fused image features corresponding to the historical frame and the fused image features corresponding to the current frame are fused using the third attention weights respectively corresponding to the historical frame to obtain new fused image features, and the new fused image features are used for the target detection.

11. The method according to claim 5, characterized in that The performing target detection based on the fused image features to obtain the target detection result includes: Using the target detection model to perform local target detection and global target detection based on the fused image features, respectively, to obtain local detection results and global detection results accordingly; Combining the local detection result and the global detection result to obtain the target detection result; Wherein, the local target detection and the global target detection are performed based on the local attention mechanism and the global attention mechanism respectively, and / or, the target detection model is implemented using a bidirectional spatiotemporal transformation network.

12. The method according to claim 5, characterized in that The vehicle is also provided with a reference sensor, and the reference sensor collects local environment information of the vehicle; Before performing target detection based on the fused image features to obtain the target detection result, the method further includes: Extracting features from the local environment information collected by the reference sensor to obtain local features; The performing target detection based on the fused image features to obtain the target detection result includes: Fusing the local features with the fused image features to obtain target fused features; Target detection is performed based on the target fusion features to obtain the target detection result.

13. The method according to claim 12, characterized in that The reference sensor includes at least one of an image sensor, a temperature sensor, and a humidity sensor; And / or, fusing the local feature with the fused image feature to obtain a target fused feature, including: Adjusting the importance weights corresponding to the local features and the fused image features respectively; Based on the importance weight, the local feature and the fused image feature are fused to obtain a target fused feature; And / or, before fusing the local feature with the fused image feature to obtain a target fused feature, the method further includes: Performing local context analysis on the local features to obtain context information, wherein the context information includes at least one of temporal context information and spatial context information; The local features are enhanced using the context information.

14. The method according to claim 1, characterized in that The target is a pedestrian; and / or the image data is a modal image acquired by the image sensor, or initial image features extracted by the image sensor from the modal image.

15. A target detection device, characterized in that: The device comprises: An image acquisition module, used to detect that an interface module is newly inserted into an image sensor, the interface module being provided on a vehicle, and the image sensor being able to acquire images of an external environment of the vehicle after being inserted; A configuration module, used for configuring the newly inserted image sensor; An image data acquisition module, used for receiving image data input by the image sensor through the interface module; The target detection module is used to perform target detection based on the image data to obtain a target detection result indicating whether the target exists in the external environment of the vehicle.

16. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores program instructions, and the processor is used to execute the program instructions to implement the target detection method according to any one of claims 1-14.

17. A target detection system, characterized in that: It includes an electronic device and an interface module, wherein the electronic device and the interface module can communicate with each other; Wherein, the interface module comprises a plurality of interfaces for inserting an image sensor; The electronic device is the electronic device according to claim 16.

18. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program instructions, and the program instructions can be executed to implement the target detection method according to any one of claims 1-14.