Target identification method and device, electronic device, computer readable medium

CN115731486BActive Publication Date: 2026-09-25LYNXI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110996812.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-27
Publication Date
2026-09-25
Estimated Expiration
2041-08-27

AI Technical Summary

Benefits of technology

[0080]本公开实施例提供的目标识别方法中,对事件信息采集装置获取的原始数据进行数据增强变换,得到至少一路增强数据,并基于至少一路增强数据分别进行目标识别,得到多路初始识别结果,最终根据多路初始识别结果得到目标识别结果,通过数据增强变换能够由原始数据得到不同情况下的增强数据,从而使得对目标物体进行识别具有更好的泛化性能,将多路初始识别结果融合得到目标识别结果,能够针对事件信息采集装置采集的任意原始数据都具有良好的识别精度,即提升了基于事件信息采集装置采集的数据进行目标识别的精度

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731486B_ABST
    Figure CN115731486B_ABST
Patent Text Reader

Abstract

The present disclosure provides a target recognition method, comprising: performing enhancement processing on original data collected by an event information collection device according to at least one data enhancement transformation to obtain at least one piece of enhanced data, wherein each piece of enhanced data corresponds to at least one data enhancement transformation; inputting multiple pieces of to-be-recognized data into a target recognition network to obtain multiple initial recognition results of recognizing target objects in the original data, wherein the multiple initial recognition results correspond to the multiple pieces of to-be-recognized data one by one; the multiple pieces of to-be-recognized data comprise one piece of original data and at least one piece of enhanced data; and obtaining a target recognition result of recognizing the target objects in the original data according to the multiple initial recognition results. The present disclosure also provides a target recognition device, an electronic device and a computer readable medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer vision technology, and in particular to a target recognition method, a target recognition device, an electronic device, and a computer-readable medium. Background Technology

[0002] An event information acquisition device is an event-driven sensor that can output a data stream reflecting changes in the scene when an event occurs.

[0003] In some related technologies, the accuracy of target identification based on data collected by event information acquisition devices is not yet ideal. Summary of the Invention

[0004] This disclosure provides a target recognition method, a target recognition device, an electronic device, and a computer-readable medium.

[0005] In a first aspect, embodiments of this disclosure provide a target recognition method, including:

[0006] The raw data collected by the event information acquisition device is enhanced according to at least one data enhancement transformation to obtain at least one enhanced data channel, wherein each enhanced data channel corresponds to at least one data enhancement transformation.

[0007] Multiple channels of data to be identified are input into a target recognition network to obtain multiple initial recognition results for identifying target objects in the original data. The multiple initial recognition results correspond one-to-one with the multiple channels of data to be identified. The multiple channels of data to be identified include one channel of the original data and at least one channel of enhanced data.

[0008] The target recognition result is obtained by recognizing the target object in the original data based on the multi-path initial recognition result.

[0009] In some embodiments, the raw data includes multiple raw event data, each raw event data corresponding to an event; the raw event data includes coordinate components, time components, and polarity components, wherein the coordinate components represent the coordinates of the corresponding event, the time components represent the occurrence time of the corresponding event, and the polarity components represent the polarity of the corresponding event; the step of enhancing the raw data collected by the event information acquisition device according to at least one data enhancement transformation to obtain at least one channel of enhanced data includes:

[0010] Based on at least one combination of various data augmentation transformations, including at least one data augmentation transformation that transforms the coordinate components of the event data and at least one data augmentation transformation that transforms the polar components of the event data, each of the original event data in the original data is processed to obtain the at least one channel of augmented data, wherein each combination of augmentation transformations includes at least one data augmentation transformation.

[0011] In some embodiments, processing each of the original event data in the original data according to a data augmentation transformation that transforms the polarity components of the event data includes:

[0012] The polarity components of each of the original event data are transformed to obtain each of the enhanced event data constituting the enhanced data, wherein the polarity of the events corresponding to the enhanced event data is opposite to the polarity of the events corresponding to the original event data.

[0013] In some embodiments, processing each of the original event data in the original data according to a data augmentation transformation that transforms the coordinate components of the event data includes:

[0014] The coordinate components of each original event data in the original data are scaled and transformed according to the enhancement amplitude vector; or

[0015] The coordinate components of each original event data in the original data are shifted and transformed according to the enhancement amplitude vector; or

[0016] The coordinate components of each original event data in the original data are rotated and transformed according to the enhancement amplitude vector;

[0017] The enhancement amplitude vector represents the range of amplitude of the data augmentation transformation performed on each of the original event data in the original data, based on the transformation of the coordinate components of the event data.

[0018] In some embodiments, for any one of the original event data in the original data, scaling the coordinate components of each of the original event data in the original data according to the enhancement magnitude vector includes:

[0019] The scaling ratio is determined based on the enhancement amplitude vector;

[0020] The coordinate components of the original event data are scaled and transformed according to the scaling ratio to obtain an enhanced event data that constitutes the enhanced data.

[0021] If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

[0022] In some embodiments, for any one of the original event data in the original data, performing a shift transformation on the coordinate components of each of the original event data in the original data according to the enhancement amplitude vector includes:

[0023] The shift vector is determined based on the enhancement amplitude vector;

[0024] The coordinate components of the original event data are shifted and transformed according to the shift vector to obtain an enhanced event data that constitutes the enhanced data.

[0025] If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

[0026] In some embodiments, for any one of the original event data in the original data, performing a rotation transformation on the coordinate components of each of the original event data in the original data according to the enhancement amplitude vector includes:

[0027] The rotation angle is determined based on the enhancement amplitude vector;

[0028] The coordinate components of the original event data are rotated and transformed relative to the center coordinates according to the rotation angle to obtain an enhanced event data that constitutes the enhanced data.

[0029] If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

[0030] In some embodiments, the step of inputting multiple streams of data to be identified into a target recognition network to obtain multiple initial recognition results for identifying target objects in the original data includes:

[0031] Generate the corresponding identification frame data for each channel of the data to be identified based on each channel of the data to be identified;

[0032] Each of the target recognition frame data is input into the target recognition network to obtain the initial recognition result for each channel.

[0033] In some embodiments, the step of generating identification frame data corresponding to each channel of identification data based on each channel of identification data includes:

[0034] The event stream composed of each stream of data to be identified is aggregated to generate event frame data;

[0035] The frame data to be identified is generated by superimposing multiple sequentially adjacent event frame data.

[0036] In some embodiments, the step of obtaining a target recognition result for recognizing a target object in the original data based on the multi-path initial recognition result includes:

[0037] The initial recognition results from each path are weighted and summed to obtain the target recognition result.

[0038] In some embodiments, before the step of enhancing the raw data collected by the event information acquisition device according to at least one data augmentation transformation to obtain at least one enhanced data stream, the target recognition method further includes:

[0039] The enhancement amplitude vector is determined based on the original data.

[0040] In some embodiments, determining the enhancement magnitude vector based on the original data includes:

[0041] The original data is input into the enhancement amplitude control network to generate the enhancement amplitude vector, wherein the enhancement amplitude control network is trained based on training sample data collected by the event information acquisition device.

[0042] Secondly, embodiments of this disclosure provide a target recognition device, comprising:

[0043] A data augmentation module is used to augment the raw data collected by the event information acquisition device according to at least one data augmentation transformation to obtain at least one channel of augmented data, wherein each channel of augmented data corresponds to at least one of the data augmentation transformations.

[0044] The recognition module is used to input multiple channels of data to be recognized into the target recognition network to obtain multiple initial recognition results for recognizing the target objects in the original data. The multiple initial recognition results correspond one-to-one with the multiple channels of data to be recognized. The multiple channels of data to be recognized include one channel of the original data and at least one channel of enhanced data.

[0045] The fusion module is used to obtain the target recognition result for recognizing the target object in the original data based on the multi-path initial recognition results.

[0046] In some embodiments, the raw data includes multiple raw event data, each of which corresponds to an event; the raw event data includes coordinate components, time components, and polarity components, wherein the coordinate components represent the coordinates of the corresponding event, the time components represent the occurrence time of the corresponding event, and the polarity components represent the polarity of the corresponding event.

[0047] The data augmentation module is used to process each of the original event data in the original data according to at least one combination of various data augmentation transformations, including at least one data augmentation transformation that transforms the coordinate components of the event data and at least one combination of data augmentation transformations that transforms the polar components of the event data, to obtain the at least one channel of augmented data. Each combination of augmentation transformations includes at least one data augmentation transformation.

[0048] In some embodiments, in the data augmentation module, processing each of the original event data in the original data according to the data augmentation transformation that transforms the polarity components of the event data includes:

[0049] The polarity components of each of the original event data are transformed to obtain each of the enhanced event data constituting the enhanced data, wherein the polarity of the events corresponding to the enhanced event data is opposite to the polarity of the events corresponding to the original event data.

[0050] In some embodiments, in the data augmentation module, processing each of the original event data in the original data according to the data augmentation transformation that transforms the coordinate components of the event data includes:

[0051] The coordinate components of each original event data in the original data are scaled and transformed according to the enhancement amplitude vector; or

[0052] The coordinate components of each original event data in the original data are shifted and transformed according to the enhancement amplitude vector; or

[0053] The coordinate components of each original event data in the original data are rotated and transformed according to the enhancement amplitude vector;

[0054] The enhancement amplitude vector represents the range of amplitude of the data augmentation transformation performed on each of the original event data in the original data, based on the transformation of the coordinate components of the event data.

[0055] In some embodiments, in the data augmentation module, for any one of the original event data in the original data, scaling the coordinate components of each of the original event data in the original data according to the augmentation magnitude vector includes:

[0056] The scaling ratio is determined based on the enhancement amplitude vector;

[0057] The coordinate components of the original event data are scaled and transformed according to the scaling ratio to obtain an enhanced event data that constitutes the enhanced data.

[0058] If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

[0059] In some embodiments, in the data enhancement module, for any one of the original event data in the original data, performing a shift transformation on the coordinate components of each of the original event data in the original data according to the enhancement magnitude vector includes:

[0060] The shift vector is determined based on the enhancement amplitude vector;

[0061] The coordinate components of the original event data are shifted and transformed according to the shift vector to obtain an enhanced event data that constitutes the enhanced data.

[0062] If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

[0063] In some embodiments, in the data augmentation module, for any one of the original event data in the original data, performing a rotation transformation on the coordinate components of each of the original event data in the original data according to the augmentation magnitude vector includes:

[0064] The rotation angle is determined based on the enhancement amplitude vector;

[0065] The coordinate components of the original event data are rotated and transformed relative to the center coordinates according to the rotation angle to obtain an enhanced event data that constitutes the enhanced data.

[0066] If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

[0067] In some embodiments, the identification module includes:

[0068] A framing unit is used to generate frame data corresponding to each channel of data to be identified based on each channel of data to be identified.

[0069] The identification unit is used to input the frame data to be identified from each channel into the target identification network to obtain the initial identification results from each channel.

[0070] In some embodiments, the framing unit is used to aggregate the event stream composed of each of the data to be identified to generate event frame data; and to superimpose multiple sequentially adjacent event frame data to generate the data to be identified.

[0071] In some embodiments, the fusion module is used to perform a weighted summation of the initial recognition results from each path to obtain the target recognition result.

[0072] In some embodiments, the target identification device further includes:

[0073] An amplitude control module is used to determine the enhanced amplitude vector based on the original data.

[0074] In some embodiments, the amplitude control module is used to input the original data into the enhanced amplitude control network to generate the enhanced amplitude vector, wherein the enhanced amplitude control network is trained based on training sample data collected by the event information acquisition device.

[0075] Thirdly, embodiments of this disclosure provide an electronic device, including:

[0076] One or more processors;

[0077] A memory having stored one or more programs, which, when executed by one or more processors, enable the one or more processors to implement any of the target recognition methods described in the first aspect of the present disclosure.

[0078] One or more I / O interfaces are connected between the processor and the memory and configured to enable information interaction between the processor and the memory.

[0079] Fourthly, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements any of the target recognition methods described in the first aspect of this disclosure.

[0080] In the target recognition method provided in this disclosure, the raw data acquired by the event information acquisition device is subjected to data augmentation transformation to obtain at least one enhanced data stream. Target recognition is then performed based on this enhanced data stream to obtain multiple initial recognition results. Finally, the target recognition result is obtained based on these multiple initial recognition results. Through data augmentation transformation, enhanced data under different conditions can be obtained from the raw data, thereby improving the generalization performance of target object recognition. By fusing the multiple initial recognition results to obtain the target recognition result, good recognition accuracy can be achieved for any raw data acquired by the event information acquisition device, thus improving the accuracy of target recognition based on data acquired by the event information acquisition device.

[0081] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0082] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:

[0083] Figure 1 This is a flowchart of a target recognition method according to an embodiment of this disclosure;

[0084] Figure 2 This is a schematic diagram of a target recognition architecture according to an embodiment of this disclosure;

[0085] Figure 3 This is a flowchart of some steps in another target recognition method according to an embodiment of this disclosure;

[0086] Figure 4 This is a flowchart of some steps in another target recognition method according to an embodiment of this disclosure;

[0087] Figure 5 This is a flowchart of some steps in another target recognition method according to an embodiment of this disclosure;

[0088] Figure 6 This is a flowchart of some steps in another target recognition method according to an embodiment of this disclosure;

[0089] Figure 7 This is a flowchart of some steps in another target recognition method according to an embodiment of this disclosure;

[0090] Figure 8 This is a flowchart of some steps in another target recognition method according to an embodiment of this disclosure;

[0091] Figure 9 This is a block diagram of a target recognition device according to an embodiment of the present disclosure;

[0092] Figure 10 This is a block diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0093] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0094] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.

[0095] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0096] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0097] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.

[0098] Firstly, referring to Figure 1 This disclosure provides a target recognition method, including:

[0099] In step S100, the raw data collected by the event information acquisition device is enhanced according to at least one data enhancement transformation to obtain at least one enhanced data channel, wherein each enhanced data channel corresponds to at least one data enhancement transformation.

[0100] In step S200, multiple channels of data to be identified are input into the target recognition network to obtain multiple initial recognition results for identifying the target objects in the original data. The multiple initial recognition results correspond one-to-one with the multiple channels of data to be identified. The multiple channels of data to be identified include one channel of the original data and at least one channel of enhanced data.

[0101] In step S300, a target recognition result is obtained based on the multi-path initial recognition result to identify the target object in the original data.

[0102] In some embodiments, the event information acquisition device includes a dynamic vision sensor (DVS). A DVS is an event-driven sensor that outputs a data stream reflecting changes in pixel brightness at a specific location within a scene when an event occurs, and has become an important component of computer vision. In some embodiments, the dynamic vision sensor is an event-driven camera.

[0103] Figure 2 This is a schematic diagram of a target recognition architecture according to an embodiment of this disclosure. Figure 2 The target recognition architecture includes multiple branches, each corresponding to at least one data augmentation transformation. For example, such as... Figure 2 As shown, the target recognition architecture includes branches corresponding to data augmentation transformation 1, data augmentation transformation 2, ..., data augmentation transformation n, and also includes branches that concatenate at least two of data augmentation transformation 1, data augmentation transformation 2, ..., data augmentation transformation n. Concatenating at least two of data augmentation transformation 1, data augmentation transformation 2, ..., data augmentation transformation n means that at least two of data augmentation transformation 1, data augmentation transformation 2, ..., data augmentation transformation n are sequentially performed on the original data in that branch to obtain one path of augmented data. When the original data collected by the event information acquisition device passes through each branch, the corresponding data augmentation transformation of that branch is performed to obtain each path of augmented data. In this embodiment, the network layer is a target recognition network.

[0104] In this embodiment of the disclosure, the original data is enhanced by data augmentation transformation, which can yield enhanced data under different conditions. The enhanced data obtained can more comprehensively reflect multiple aspects and situations of the event acquired by the event information acquisition device compared to the original data. For example, through data augmentation transformation, at least one of the following enhanced data can be obtained: enhanced data showing different changes in light intensity corresponding to the target object; enhanced data showing different sizes of the target object within the visual range; enhanced data showing different positions of the target object within the visual range; and enhanced data showing the target object after rotation and / or mirror flipping within the visual range. This embodiment of the disclosure does not impose any special limitations on this.

[0105] It should be noted that, in the embodiments of this disclosure, identifying target objects in the raw data means recognizing and distinguishing target objects. For example, target recognition determines what the target object is and classifies the target object.

[0106] like Figure 2As shown, in step S200, target recognition is performed on each path of data to be recognized through a target recognition network at the network layer to obtain initial recognition results for each path. This embodiment of the disclosure does not impose any special limitations on the target recognition network. For example, the target recognition network can be a convolutional neural network or a spiking neural network.

[0107] In some embodiments, the data to be identified from each input channel is simultaneously input into the target identification network. This disclosure does not impose any special limitations on the structure of the target identification network. In some embodiments, the target identification network has multiple input terminals, each corresponding to one input channel of data to be identified.

[0108] It should be noted that, in the embodiments of this disclosure, the target recognition network can be trained using any training method, and this disclosure does not impose any special limitations on it. For example, it can be trained using enhanced data obtained by enhancing the original training data collected by the event information acquisition device. The enhancement processing can include any enhancement transformation method in the embodiments of this disclosure, and this disclosure does not impose any limitations on it. Since the original training data has been enhanced, the resulting training data has high complexity. Training the neural network with high-complexity training data can solve the problem of overfitting during training. Furthermore, the trained neural network can correspond to more possible scenarios during inference, thereby enabling the trained neural network to have better processing accuracy.

[0109] In the target recognition method provided in this embodiment, the raw data acquired by the event information acquisition device is subjected to data augmentation transformation to obtain at least one enhanced data stream. Target recognition is then performed based on the at least one enhanced data stream to obtain multiple initial recognition results. Finally, the target recognition result is obtained based on the multiple initial recognition results. Through data augmentation transformation, enhanced data under different conditions can be obtained from the raw data, thereby enabling better generalization performance in target object recognition. By fusing the multiple initial recognition results to obtain the target recognition result, good recognition accuracy can be achieved for any raw data acquired by the event information acquisition device, thus improving the accuracy of target recognition based on data acquired by the event information acquisition device.

[0110] The following description uses an event information acquisition device as an example of a dynamic visual sensor to illustrate the embodiments of this disclosure.

[0111] In some embodiments, the raw data acquired by the dynamic vision sensor consists of multiple event data in the form of (x, y, t, p) quadruples, with each event data corresponding to an event. Here, the coordinate component (x, y) represents the coordinates of the event; the time component t represents the time of the event's occurrence; and the polarity component p represents the polarity of the event, i.e., whether the light intensity increases or decreases; for example, 0 indicates decreased light intensity, and 1 indicates increased light intensity.

[0112] Accordingly, in some embodiments, the raw data includes multiple raw event data, each raw event data corresponding to one event; the raw event data includes coordinate components, time components, and polarity components, wherein the coordinate components represent the coordinates of the corresponding event, the time components represent the occurrence time of the corresponding event, and the polarity components represent the polarity of the corresponding event; refer to Figure 3 Step S100 includes:

[0113] In step S110, each of the original event data in the original data is processed according to at least one combination of various data augmentation transformations, including at least one data augmentation transformation that transforms the coordinate components of the event data and at least one combination of data augmentation transformations that transforms the polar components of the event data, to obtain the at least one channel of augmented data. Each combination of augmentation transformations includes at least one data augmentation transformation.

[0114] It should be noted that the data augmentation transformations that make up different combinations of augmentation transformations are not all the same.

[0115] In this embodiment of the disclosure, processing each of the original event data in the original data according to a data augmentation transformation that transforms the polarity component of the event data can yield augmented data with different changes in illumination intensity corresponding to the target object; processing each of the original event data in the original data according to at least one data augmentation transformation that transforms the coordinate component of the event data can yield at least one of a variety of augmented data, such as augmented data with different sizes of the target object within the visual range, augmented data with different positions of the target object within the visual range, and augmented data after the target object is rotated and / or mirrored within the visual range.

[0116] This disclosure does not impose any special limitations on how the data augmentation transformation, which transforms the polarity components of the event data, is applied to process each of the original event data in the original data.

[0117] In some embodiments, refer to Figure 4 In step S110, the data augmentation transformation that transforms the polarity components of the event data to process each of the original event data in the original data includes:

[0118] In step S111, the polarity components of each of the original event data are transformed to obtain each of the enhanced event data constituting the enhanced data, wherein the polarity of the events corresponding to the enhanced event data is opposite to the polarity of the events corresponding to the original event data.

[0119] It should be noted that, in this embodiment of the disclosure, step S111 is performed on each original event data in the original data to obtain multiple enhanced event data.

[0120] In this embodiment of the disclosure, enhanced data with different changes in light intensity corresponding to the target object can be obtained through step S111.

[0121] In some embodiments, step S111 can be performed as follows:

[0122] i. Read raw event data (x, y, t, p) from the raw data;

[0123] ii. If the polar component p of the original event data (x, y, t, p) is 0, it is changed to 1 to obtain the enhanced event data corresponding to the original event data; if the polar component p of the original event data (x, y, t, p) is 1, it is changed to 0 to obtain the enhanced event data corresponding to the original event data. The coordinate components of the enhanced event data are the same as those of the original event data, and the time components of the enhanced event data are the same as those of the original event data.

[0124] The above polarity component transformation is performed sequentially on all the original event data of the original data.

[0125] This disclosure does not impose any special limitations on how to process each of the original event data in the original data according to at least one data augmentation transformation that transforms the coordinate components of the event data.

[0126] In this embodiment of the disclosure, when processing each of the original event data in the original data according to at least one data augmentation transformation that transforms the coordinate components of the event data, the range of the augmentation processing on the original data can also be controlled. For example, when adjusting the size of a target object within the visual range through data augmentation transformation, the range of size change of the target object can be controlled; when changing the position of a target object within the visual range through data augmentation transformation, the range of displacement of the target object can be controlled; when rotating a target object within the visual range through data augmentation transformation, the rotation angle of the target object can be controlled. In this embodiment of the disclosure, the range of augmentation processing is controlled using an augmentation amplitude vector.

[0127] In some embodiments, refer to Figure 5 In step S110, processing each of the original event data in the original data according to at least one data augmentation transformation that transforms the coordinate components of the event data includes:

[0128] In step S112, the coordinate components of each of the original event data in the original data are scaled and transformed according to the enhancement amplitude vector; or

[0129] In step S113, the coordinate components of each of the original event data in the original data are shifted according to the enhancement amplitude vector; or

[0130] In step S114, the coordinate components of each of the original event data in the original data are rotated according to the enhancement amplitude vector;

[0131] The enhancement amplitude vector represents the range of amplitude of the data augmentation transformation performed on each of the original event data in the original data, based on the transformation of the coordinate components of the event data.

[0132] In this embodiment of the disclosure, step S112 can obtain enhanced data of different sizes of the target object within the visual range, step S113 can obtain enhanced data of different positions of the target object within the visual range, and step S114 can obtain enhanced data of the target object after rotation and / or mirror flipping within the visual range.

[0133] This disclosure does not impose any special limitations on how to scale and transform the coordinate components of each of the original event data in the original data.

[0134] In some embodiments, for any one of the original event data in the original data, scaling the coordinate components of each of the original event data in the original data according to the enhancement amplitude vector includes: determining a scaling ratio according to the enhancement amplitude vector; scaling the coordinate components of the original event data according to the scaling ratio to obtain an enhanced event data constituting the enhanced data; and deleting the enhanced event data if the coordinate components of the enhanced event data exceed a preset data boundary.

[0135] In some embodiments, scaling transformation of the coordinate components of any one of the original event data in the original data can be performed as follows:

[0136] i. Determine the preset scale α and the preset data boundary (X1, Y1). The preset data boundary represents the size of the raw data acquired by the dynamic vision sensor.

[0137] ii. Read raw event data (x, y, t, p) from the raw data;

[0138] iii. Multiply the coordinate components (x, y) sequentially by a preset ratio α, and round up to obtain the new coordinate components. This results in enhanced event data. The time component of the enhanced event data is the same as the time component of the original event data, and the polarity component of the enhanced event data is the same as the polarity component of the original event data.

[0139] iv. If and / or Delete the enhanced event data.

[0140] The above coordinate scaling transformation is performed sequentially on all the original event data of the original data to obtain enhanced event data with multiple coordinate components that do not exceed the preset data boundary.

[0141] This disclosure does not impose any special limitations on how to perform shift transformations on the coordinate components of each of the original event data in the original data.

[0142] In some embodiments, for any one of the original event data in the original data, shifting and transforming the coordinate components of each of the original event data in the original data according to the enhancement amplitude vector includes: determining a shift vector according to the enhancement amplitude vector; shifting and transforming the coordinate components of the original event data according to the shift vector to obtain an enhanced event data constituting the enhanced data; and deleting the enhanced event data if the coordinate components of the enhanced event data exceed a preset data boundary.

[0143] In some embodiments, shifting and transforming the coordinate components of any one of the original event data in the original data can be performed as follows:

[0144] i. Determine the shift vector (Δx2, Δy2) and the preset data boundary (X2, Y2). The preset data boundary represents the size of the raw data acquired by the dynamic vision sensor.

[0145] ii. Read raw event data (x, y, t, p) from the raw data;

[0146] iii. Add the coordinate components (x, y) of the original event data and the shift vector (Δx2, Δy2) to obtain the coordinate components (x+Δx2, y+Δy2) of the enhanced event data;

[0147] iv. If x + Δx² > X² and / or y + Δy² > Y², delete this enhanced event data.

[0148] The above coordinate shift transformation is performed sequentially on all the original event data of the original data to obtain multiple enhanced event data whose coordinate components do not exceed the preset data boundary.

[0149] This disclosure does not impose any special limitations on how to perform rotation transformation on the coordinate components of each of the original event data in the original data.

[0150] In some embodiments, for any one of the original event data in the original data, rotating the coordinate components of each of the original event data in the original data according to the enhancement amplitude vector includes: determining a rotation angle according to the enhancement amplitude vector; rotating the coordinate components of the original event data according to the rotation angle relative to the center coordinate to obtain an enhanced event data constituting the enhanced data; and deleting the enhanced event data if the coordinate components of the enhanced event data exceed a preset data boundary.

[0151] In some embodiments, performing a rotation transformation on the coordinate components of any one of the original event data in the original data can be performed as follows:

[0152] i. Determine the preset rotation angle Determine the shift vector (Δx3, Δy3) to determine the preset boundary (X3, Y3) and center coordinates.

[0153] ii. Read raw event data (x, y, t, p) from the raw data;

[0154] iii. Calculate the coordinate components (x, y) of the original event data about the center coordinates using angle calculations. The angle θ and the distance R will As the rotation angle, the coordinate components (x1, y1) of the enhanced event data are determined using the rotation angle and the distance R;

[0155] iv. If x1 > X3 and / or y1 > Y3, delete this enhanced event data.

[0156] The above rotation transformation is performed sequentially on all the original event data of the original data to obtain enhanced event data with multiple coordinate components that do not exceed the preset data boundary.

[0157] like Figure 2 As shown, the enhanced data from multiple channels and the original data, after data augmentation processing, are passed through a framing layer to generate frame data, which is then input into the target recognition network in the network layer for target recognition. In this embodiment of the disclosure, generating frame data through the framing layer means combining one channel of data to be recognized within a specific time window into a frame to obtain frame data.

[0158] Accordingly, in some embodiments, reference is made to Figure 6 Step S200 includes:

[0159] In step S210, corresponding identification frame data is generated based on each channel of the identification data.

[0160] In step S220, the initial recognition result of each channel is obtained based on the frame data to be identified in each channel.

[0161] The frame data to be identified may include at least one event frame.

[0162] In some embodiments, refer to Figure 7 Step S210 includes:

[0163] In step S211, the event stream composed of each channel of data to be identified is aggregated to generate event frame data;

[0164] In step S212, multiple sequentially adjacent event frame data are superimposed to generate the frame data to be identified.

[0165] It should be noted that in step S211, the event stream consisting of one stream of data to be identified within a specific time window is aggregated to generate event frame data.

[0166] In some embodiments, step S210 may be performed as follows:

[0167] The input augmentation data for the framing layer consists of multiple augmentation event data in the form of (x, y, t, p) quadruples. The augmentation event data within the time window Δt forms the event stream E. t′ ={e i |e i =[x i y i ,t′,p i The framing layer aggregates the event stream and outputs event frame data X. t =q(E t′ The aggregation function q(·) can be any of the following: nonpolar function aggregation, cumulative aggregation, and logical operation aggregation; this embodiment does not impose any special limitation on this. The event frame data is in the form (c, h, w), where c represents the channel, h represents the height of the frame data, and w represents the width of the frame data. Finally, T adjacent event frame data are superimposed in chronological order to form the frame data to be identified in the form (T, c, h, w).

[0168] In some embodiments, when the frame data to be identified includes multiple event frame data, target identification can be performed using a network layer based on a spiking neural network. A network layer based on a spiking neural network can better understand the temporal information between multiple event frame data in the frame data to be identified, formed by superimposing them in chronological order, thereby improving the accuracy of target identification.

[0169] like Figure 2As shown, the fusion layer obtains the target recognition result based on the multiple initial recognition results. This embodiment of the disclosure does not impose any special limitations on how the target recognition result is obtained from the multiple initial recognition results.

[0170] In some embodiments, the step of obtaining a target recognition result based on the multiple initial recognition results includes: weighting and summing the initial recognition results from each path to obtain the target recognition result.

[0171] In some embodiments, the step of obtaining a target recognition result based on the multiple initial recognition results includes: calculating the average value of each initial recognition result as the target recognition result.

[0172] In this embodiment, the enhancement amplitude vector can be a fixed value or it can change dynamically. This embodiment does not impose any special limitations on this. In this embodiment, when the enhancement amplitude vector changes dynamically, it can be dynamically acquired or automatically generated. For example, a data enhancement amplitude control module can be added to the system to generate and adjust the enhancement amplitude vector.

[0173] Accordingly, refer to Figure 8 In some embodiments, prior to step 100, the target recognition method further includes:

[0174] In step S400, the enhancement amplitude vector is determined based on the original data.

[0175] In some embodiments, an enhancement magnitude vector is determined using a neural network based on the original data. In some embodiments, the neural network used to generate the enhancement magnitude vector is trained using data acquired by a dynamic vision sensor as samples.

[0176] Accordingly, in some embodiments, determining the enhancement amplitude vector based on the raw data includes: inputting the raw data into an enhancement amplitude control network to generate the enhancement amplitude vector, wherein the enhancement amplitude control network is trained based on training sample data collected by the event information acquisition device.

[0177] In some embodiments, the enhancement amplitude control network can be trained based on training sample data collected by the event information acquisition device; the trained amplitude control network can be used to generate a corresponding enhancement amplitude vector based on the original data.

[0178] This disclosure does not impose any special limitations on how the augmentation amplitude control network is trained. For example, the augmentation amplitude control network that generates the augmentation amplitude vector can be jointly trained with the target recognition network so that when the augmented data obtained by augmenting the original data according to the augmentation amplitude vector is input into the target recognition network for recognition, the accuracy of the target recognition network can be improved.

[0179] Secondly, referring to Figure 9 This disclosure provides a target recognition device, including:

[0180] Data augmentation module 110 is used to augment the raw data collected by the event information acquisition device according to at least one data augmentation transformation to obtain at least one channel of augmented data, wherein each channel of augmented data corresponds to at least one of the data augmentation transformations.

[0181] The recognition module 120 is used to input multiple channels of data to be recognized into the target recognition network to obtain multiple initial recognition results for recognizing target objects in the original data. The multiple initial recognition results correspond one-to-one with the multiple channels of data to be recognized. The multiple channels of data to be recognized include one channel of the original data and at least one channel of enhanced data.

[0182] The fusion module 130 is used to obtain a target recognition result for recognizing the target object in the original data based on the multi-path initial recognition result.

[0183] In some embodiments, the raw data includes multiple raw event data, each of which corresponds to an event; the raw event data includes coordinate components, time components, and polarity components, wherein the coordinate components represent the coordinates of the corresponding event, the time components represent the occurrence time of the corresponding event, and the polarity components represent the polarity of the corresponding event.

[0184] The data augmentation module 110 is used to process each of the original event data in the original data according to at least one combination of various data augmentation transformations, including at least one data augmentation transformation that transforms the coordinate components of the event data and at least one combination of data augmentation transformations that transforms the polar components of the event data, to obtain the at least one channel of augmented data. Each combination of augmentation transformations includes at least one data augmentation transformation.

[0185] In some embodiments, in the data augmentation module 110, processing each of the original event data in the original data according to the data augmentation transformation that transforms the polarity components of the event data includes:

[0186] The polarity components of each of the original event data are transformed to obtain each of the enhanced event data constituting the enhanced data, wherein the polarity of the events corresponding to the enhanced event data is opposite to the polarity of the events corresponding to the original event data.

[0187] In some embodiments, in the data augmentation module 110, processing each of the original event data in the original data according to the data augmentation transformation that transforms the coordinate components of the event data includes:

[0188] The coordinate components of each original event data in the original data are scaled and transformed according to the enhancement amplitude vector; or

[0189] The coordinate components of each original event data in the original data are shifted and transformed according to the enhancement amplitude vector; or

[0190] The coordinate components of each original event data in the original data are rotated and transformed according to the enhancement amplitude vector;

[0191] The enhancement amplitude vector represents the range of amplitude of the data augmentation transformation performed on each of the original event data in the original data, based on the transformation of the coordinate components of the event data.

[0192] In some embodiments, in the data enhancement module 110, for any one of the original event data in the original data, scaling the coordinate components of each of the original event data in the original data according to the enhancement magnitude vector includes:

[0193] The scaling ratio is determined based on the enhancement amplitude vector;

[0194] The coordinate components of the original event data are scaled and transformed according to the scaling ratio to obtain an enhanced event data that constitutes the enhanced data.

[0195] If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

[0196] In some embodiments, in the data enhancement module 110, for any one of the original event data in the original data, performing a shift transformation on the coordinate components of each of the original event data in the original data according to the enhancement magnitude vector includes:

[0197] The shift vector is determined based on the enhancement amplitude vector;

[0198] The coordinate components of the original event data are shifted and transformed according to the shift vector to obtain an enhanced event data that constitutes the enhanced data.

[0199] If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

[0200] In some embodiments, in the data enhancement module 110, for any one of the original event data in the original data, performing a rotation transformation on the coordinate components of each of the original event data in the original data according to the enhancement magnitude vector includes:

[0201] The rotation angle is determined based on the enhancement amplitude vector;

[0202] The coordinate components of the original event data are rotated and transformed relative to the center coordinates according to the rotation angle to obtain an enhanced event data that constitutes the enhanced data.

[0203] If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

[0204] In some embodiments, the identification module 120 includes:

[0205] A framing unit is used to generate frame data corresponding to each channel of data to be identified based on each channel of data to be identified.

[0206] The identification unit is used to input the frame data to be identified from each channel into the target identification network to obtain the initial identification results from each channel.

[0207] In some embodiments, the framing unit is used to aggregate the event stream composed of each of the data to be identified to generate event frame data; and to superimpose multiple sequentially adjacent event frame data to generate the data to be identified.

[0208] In some embodiments, the fusion module 130 is used to perform a weighted summation of the initial recognition results from each path to obtain the target recognition result.

[0209] In some embodiments, the target identification device further includes:

[0210] An amplitude control module is used to determine the enhanced amplitude vector based on the original data.

[0211] In some embodiments, the amplitude control module is used to input the original data into the enhanced amplitude control network to generate the enhanced amplitude vector, wherein the enhanced amplitude control network is trained based on training sample data collected by the event information acquisition device.

[0212] Thirdly, referring to Figure 10 This disclosure provides an electronic device, including:

[0213] One or more processors 201;

[0214] The memory 202 stores one or more programs, which, when executed by one or more processors, enable the one or more processors to implement any of the target recognition methods described in the first aspect of the embodiments of this disclosure.

[0215] One or more I / O interfaces 203 are connected between the processor and the memory and configured to enable information exchange between the processor and the memory.

[0216] Among them, processor 201 is a device with data processing capabilities, including but not limited to central processing unit (CPU); memory 202 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH); I / O interface (read-write interface) 203 is connected between processor 201 and memory 202, and can realize information interaction between processor 201 and memory 202, including but not limited to data bus (Bus).

[0217] In some embodiments, the processor 201, memory 202, and I / O interface 203 are interconnected via bus 204, and thus connected to other components of the computing device.

[0218] Fourthly, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements any of the target recognition methods described in the first aspect of this disclosure.

[0219] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0220] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.

Claims

1. A target recognition method, comprising: The raw data collected by the event information acquisition device is enhanced by at least one data augmentation transformation to obtain at least one enhanced data stream, wherein each enhanced data stream corresponds to at least one data augmentation transformation; the event information acquisition device is an event-driven sensor; the raw data includes multiple raw event data, each raw event data corresponding to one event; the raw event data includes coordinate components, time components, and polarity components, wherein the coordinate components represent the coordinates of the corresponding event, the time components represent the occurrence time of the corresponding event, and the polarity components represent the polarity of the corresponding event; Multiple channels of data to be identified are input into a target recognition network to obtain multiple initial recognition results for identifying target objects in the original data. The multiple initial recognition results correspond one-to-one with the multiple channels of data to be identified. The multiple channels of data to be identified include one channel of the original data and at least one channel of enhanced data. The target recognition result is obtained by recognizing the target object in the original data based on the multi-path initial recognition result.

2. The target recognition method according to claim 1, wherein, The steps of enhancing the raw data collected by the event information acquisition device according to at least one data augmentation transformation to obtain at least one enhanced data stream include: Based on at least one combination of various data augmentation transformations, including at least one data augmentation transformation that transforms the coordinate components of the event data and at least one data augmentation transformation that transforms the polar components of the event data, each of the original event data in the original data is processed to obtain the at least one channel of augmented data, wherein each combination of augmentation transformations includes at least one data augmentation transformation.

3. The target recognition method according to claim 2, wherein, The data augmentation transformation, which transforms the polarity components of the event data, processes each of the original event data in the original data, including: The polarity components of each of the original event data are transformed to obtain each of the enhanced event data constituting the enhanced data, wherein the polarity of the events corresponding to the enhanced event data is opposite to the polarity of the events corresponding to the original event data.

4. The target recognition method according to claim 2, wherein, The data augmentation transformation, which transforms the coordinate components of the event data, processes each of the original event data in the original data, including: The coordinate components of each original event data in the original data are scaled and transformed according to the enhancement amplitude vector; or The coordinate components of each original event data in the original data are shifted and transformed according to the enhancement amplitude vector; or The coordinate components of each original event data in the original data are rotated and transformed according to the enhancement amplitude vector; The enhancement amplitude vector represents the range of amplitude of the data augmentation transformation performed on each of the original event data in the original data, based on the transformation of the coordinate components of the event data.

5. The target recognition method according to claim 4, wherein, For any one of the original event data in the original data, scaling transformation of the coordinate components of each of the original event data in the original data according to the enhancement magnitude vector includes: The scaling ratio is determined based on the enhancement amplitude vector; The coordinate components of the original event data are scaled and transformed according to the scaling ratio to obtain an enhanced event data that constitutes the enhanced data. If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

6. The target recognition method according to claim 4, wherein, For any one of the original event data in the original data, the shift transformation of the coordinate components of each of the original event data in the original data according to the enhancement amplitude vector includes: The shift vector is determined based on the enhancement amplitude vector; The coordinate components of the original event data are shifted and transformed according to the shift vector to obtain an enhanced event data that constitutes the enhanced data. If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

7. The target recognition method according to claim 4, wherein, For any one of the original event data in the original data, the rotation transformation of the coordinate components of each of the original event data in the original data according to the enhancement amplitude vector includes: The rotation angle is determined based on the enhancement amplitude vector; The coordinate components of the original event data are rotated and transformed relative to the center coordinates according to the rotation angle to obtain an enhanced event data that constitutes the enhanced data. If the coordinate components of the enhanced event data exceed the preset data boundary, the enhanced event data is deleted.

8. The target recognition method according to any one of claims 1 to 7, wherein, The steps of inputting multiple streams of data to be identified into a target recognition network to obtain multiple initial recognition results for the target objects in the original data include: Generate the corresponding identification frame data for each channel of the data to be identified based on each channel of the data to be identified; Each of the target recognition frame data is input into the target recognition network to obtain the initial recognition result for each channel.

9. The target recognition method according to claim 8, wherein, The steps of generating the corresponding identification frame data for each channel of the identification data include: The event stream composed of each stream of data to be identified is aggregated to generate event frame data; The frame data to be identified is generated by superimposing multiple sequentially adjacent event frame data.

10. The target recognition method according to any one of claims 1 to 7, wherein, The steps for obtaining target recognition results for identifying target objects in the original data based on the multi-path initial recognition results include: The initial recognition results from each path are weighted and summed to obtain the target recognition result.

11. The target recognition method according to any one of claims 4 to 7, wherein, Before the step of enhancing the raw data collected by the event information acquisition device according to at least one data augmentation transformation to obtain at least one channel of enhanced data, the target recognition method further includes: The enhancement amplitude vector is determined based on the original data.

12. The target recognition method according to claim 11, wherein, Determining the enhancement magnitude vector based on the original data includes: The original data is input into the enhancement amplitude control network to generate the enhancement amplitude vector, wherein the enhancement amplitude control network is trained based on training sample data collected by the event information acquisition device.

13. A target recognition device for implementing the target recognition method as described in any one of claims 1-12, comprising: A data augmentation module is used to augment the raw data collected by the event information acquisition device according to at least one data augmentation transformation to obtain at least one channel of augmented data, wherein each channel of augmented data corresponds to at least one of the data augmentation transformations. The recognition module is used to input multiple channels of data to be recognized into the target recognition network to obtain multiple initial recognition results for recognizing the target objects in the original data. The multiple initial recognition results correspond one-to-one with the multiple channels of data to be recognized. The multiple channels of data to be recognized include one channel of the original data and at least one channel of enhanced data. The fusion module is used to obtain the target recognition result for recognizing the target object in the original data based on the multi-path initial recognition results.

14. An electronic device comprising: One or more processors; A memory having stored one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the target recognition method according to any one of claims 1 to 12; One or more I / O interfaces are connected between the processor and the memory and configured to enable information interaction between the processor and the memory.

15. A computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the target recognition method according to any one of claims 1 to 12.