An image processing method and apparatus

By selecting appropriate sensors and adaptive event representation methods in electronic devices, the accuracy and efficiency issues of image processing in different scenarios are solved, achieving more efficient image processing and lower power consumption.

CN116134484BActive Publication Date: 2026-04-21HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2020-12-31
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently acquire accurate motion information in different scenarios, resulting in poor image processing performance.

Method used

By selecting RGB sensors and motion sensors in electronic devices, and choosing to activate different sensors based on scene information, combined with the adaptive event representation switching and decoding methods of the vision sensor chip, data transmission and parsing are optimized.

Benefits of technology

It improves the accuracy and efficiency of image processing, reduces the power consumption of electronic devices, and avoids data loss and transmission bottlenecks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116134484B_ABST
    Figure CN116134484B_ABST
Patent Text Reader

Abstract

An image processing method and apparatus are disclosed for obtaining clearer images based on different duration requirements. The method includes: acquiring motion information, the motion information including information about the motion trajectory of a target object moving within the detection range of a motion sensor; generating at least one event image frame based on the motion information, the at least one event image frame including an image representing the motion trajectory of the target object moving within the detection range; acquiring a target task and acquiring an iteration duration based on the target task; iteratively updating the at least one event image frame to obtain an updated at least one event image frame, wherein the duration of iteratively updating the at least one event image frame does not exceed the iteration duration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more particularly to an image processing method and apparatus. Background Technology

[0002] Typically, information collected by motion sensors can be used for image reconstruction, target detection, photographing moving objects, photographing with moving devices, image deblurring, motion estimation, depth estimation, or target detection and recognition, among other scenarios. Therefore, obtaining more accurate motion information has become an urgent problem to be solved. Summary of the Invention

[0003] This application provides an image processing method and apparatus for obtaining clearer images.

[0004] In a first aspect, this application provides a switching method applied to an electronic device, the electronic device including an RGB sensor and a motion sensor, the RGB (red green blue) sensor being used to acquire images within a shooting range, and the motion sensor being used to acquire information generated when an object moves relative to the motion sensor within the detection range of the motion sensor, the method comprising: selecting at least one from the RGB sensor and the motion sensor based on scene information, and acquiring data through the selected sensor, the scene information including at least one of the state information of the electronic device, the type of the application requesting image acquisition in the electronic device, or environmental information.

[0005] Therefore, in this embodiment, different sensors in the electronic device can be selected and activated according to different scenarios, adapting to more scenarios and demonstrating strong generalization ability. Furthermore, appropriate sensors can be activated based on the actual scenario, eliminating the need to activate all sensors and reducing the power consumption of the electronic device.

[0006] In one possible implementation, the status information includes the remaining battery power and remaining storage capacity of the electronic device; the environmental information includes the change in light intensity within the shooting range of the color RGB sensor and the motion sensor, or information about moving objects within the shooting range.

[0007] Therefore, in the embodiments of this application, the sensor to be activated can be selected according to the status of the electronic device or environmental information, which can adapt to more scenarios and has strong generalization ability.

[0008] Furthermore, in the following different implementations, the activated sensor may be different. When a certain sensor collects data, that sensor is already activated, which will not be described in detail below.

[0009] Secondly, this application provides a vision sensor chip, which may include: a pixel array circuit for generating at least one data signal corresponding to a pixel in the pixel array circuit by measuring changes in light intensity, wherein the at least one data signal indicates a light intensity change event, and the light intensity change event indicates that the light intensity change measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A readout circuit, coupled to the pixel array circuit, is used to read at least one data signal from the pixel array circuit in a first event representation manner. The readout circuit is also used to provide at least one data signal to a control circuit. The readout circuit is further used to switch to reading at least one data signal from the pixel array circuit in a second event representation manner when it receives a conversion signal generated based on the at least one data signal from the control circuit. As can be seen from the first aspect, the vision sensor can adaptively switch between two event representation methods, ensuring that the data reading rate always remains within a predetermined data reading rate threshold, thereby reducing the cost of data transmission, parsing, and storage of the vision sensor and significantly improving the sensor's performance. Furthermore, such a vision sensor can perform data statistics on events generated over a period of time to predict the possible event generation rate in the next time period, thus enabling the selection of a readout mode more suitable for the current external environment, application scenario, and motion state.

[0010] In one possible implementation, the first event is represented by polarity information. The pixel array circuit may include multiple pixels, and each pixel may include a threshold comparison unit. The threshold comparison unit outputs polarity information when the light intensity change exceeds a predetermined threshold. The polarity information indicates whether the light intensity change is increasing or decreasing. A readout circuit is specifically used to read the polarity information output by the threshold comparison unit. In this implementation, the first event is represented by polarity information, which is typically represented by 1-2 bits. This carries less information and avoids the problem of a sudden surge in events for the visual sensor when there is a large amount of data, such as large-area object movement or light intensity fluctuations (e.g., entering or exiting a tunnel, switching lights on or off in a room). Given a fixed preset maximum bandwidth (hereinafter referred to as bandwidth) of the visual sensor, this avoids situations where event data cannot be read, leading to event loss.

[0011] In one possible implementation, the first event is represented by light intensity information. The pixel array may include multiple pixels, and each pixel may include a threshold comparison unit, a readout control unit, a light intensity acquisition unit, and a light intensity detection unit, used to output an electrical signal corresponding to the light signal illuminating it. The electrical signal is used to indicate the light intensity. The threshold comparison unit is used to output a first signal when it determines that the light intensity change exceeds a predetermined threshold based on the electrical signal. The readout control unit is used to instruct the light intensity acquisition unit to acquire and buffer the electrical signal corresponding to the moment the first signal is received, in response to receiving the first signal. The readout circuit is specifically used to read the electrical signal buffered by the light intensity acquisition unit. In this implementation, the first event is represented by light intensity information. When the amount of data transmitted does not exceed the bandwidth limit, light intensity information is used to represent the event. Typically, light intensity information is represented by multiple bits, such as 8 bits to 12 bits. Compared to polarity information, light intensity information can carry more information, which is beneficial for event processing and analysis, such as improving the quality of image reconstruction.

[0012] In one possible implementation, the control circuit is further configured to: determine statistical data based on at least one data signal received from the readout circuit. If the statistical data is determined to meet a predetermined conversion condition, a conversion signal is sent to the readout circuit, the predetermined conversion condition being determined based on a preset bandwidth of the vision sensor chip. This implementation provides a method for switching between two event representation methods, obtaining the conversion condition based on the amount of data to be transmitted. For example, when the amount of data to be transmitted is large, the event is switched to be represented using polarity information to ensure complete data transmission and avoid situations where event data cannot be read, leading to event loss. When the amount of data to be transmitted is small, the event is switched to be represented using light intensity information, allowing the transmitted event to carry more information, which is beneficial for event processing and analysis, such as improving the quality of image reconstruction.

[0013] In one possible implementation, when the first event representation method is to represent the event using light intensity information, and the second event representation method is to represent the event using polarity information, a predetermined conversion condition is that the total amount of data read from the pixel array circuit using the first event representation method is greater than a preset bandwidth, or the predetermined conversion condition is that the number of at least one data signal is greater than the ratio of the preset bandwidth to a first bit, where the first bit is a preset bit of the data format of the data signal. This implementation provides a specific condition for switching from representing the event using light intensity information to representing the event using polarity information. When the amount of data transmitted exceeds the preset bandwidth, the switch to representing the event using polarity information ensures complete data transmission and avoids situations where event data cannot be read, leading to event loss.

[0014] In one possible implementation, when the first event representation method is to represent the event using polarity information, and the second event representation method is to represent the event using light intensity information, a predetermined conversion condition is that if at least one data signal is read from the pixel array circuit using the second event representation method, the total amount of data read is not greater than a preset bandwidth; or the predetermined conversion condition is that the number of at least one data signal is not greater than the ratio of the preset bandwidth to a first bit, where the first bit is a preset bit of the data format of the data signal. This implementation provides a specific condition for switching from representing the event using polarity information to representing the event using light intensity information. When the amount of data transmitted is not greater than the preset bandwidth, the switch to representing the event using light intensity information allows the transmitted event to carry more information, which is beneficial for event processing and analysis, such as improving the quality of image reconstruction.

[0015] Thirdly, this application provides a decoding circuit, which may include: a reading circuit for reading data signals from a vision sensor chip; a decoding circuit for decoding the data signals according to a first decoding method; and a decoding circuit for decoding the data signals according to a second decoding method when a conversion signal is received from a control circuit. The decoding circuit provided in the third aspect corresponds to the vision sensor chip provided in the second aspect, and is used to decode the data signals output by the vision sensor chip provided in the second aspect. The decoding circuit provided in the third aspect can switch between different decoding methods for different event representation methods.

[0016] In one possible implementation, the control circuit is further configured to: determine statistical data based on data signals read from the readout circuit; and if the statistical data is determined to meet predetermined conversion conditions, send a conversion signal to the encoding circuit, the predetermined conversion conditions being determined based on a preset bandwidth of the vision sensor chip.

[0017] In one possible implementation, the first decoding method is to decode the data signal according to the first bit corresponding to the first event representation method, where the first event representation method represents the event through light intensity information. The second decoding method is to decode the data signal according to the second bit corresponding to the second event representation method, where the second event representation method represents the event through polarity information. The polarity information is used to indicate whether the change in light intensity is an increase or a decrease. The conversion condition is that the total amount of data decoded according to the first decoding method is greater than a preset bandwidth, or the predetermined conversion condition is that the number of data signals is greater than the ratio of the preset bandwidth to the first bit. The first bit is a preset bit of the data format of the data signal.

[0018] In one possible implementation, the first decoding method is to decode the data signal according to the first bit corresponding to the first event representation method, where the first event representation method represents the event through polarity information, which is used to indicate whether the change in light intensity is an increase or a decrease. The second decoding method is to decode the data signal according to the second event representation method, where the second event representation method represents the event through light intensity information. The conversion condition is that if the data signal is decoded according to the second decoding method, the total data amount is not greater than a preset bandwidth, or the predetermined conversion condition is that the number of data signals is greater than the ratio of the preset bandwidth to the first bit, where the first bit is a preset bit of the data format of the data signal.

[0019] Fourthly, this application provides a method for operating a vision sensor chip, which may include: measuring a change in light intensity through a pixel array circuit of the vision sensor chip to generate at least one data signal corresponding to a pixel in the pixel array circuit, wherein the at least one data signal indicates a light intensity change event, and the light intensity change event indicates that the light intensity change measured at the corresponding pixel in the pixel array circuit exceeds a predetermined threshold; reading the at least one data signal from the pixel array circuit in a first event representation manner through a readout circuit of the vision sensor chip; providing the at least one data signal to a control circuit of the vision sensor chip through the readout circuit; and, upon receiving a conversion signal generated based on the at least one data signal from the control circuit through the readout circuit, switching to reading the at least one data signal from the pixel array circuit in a second event representation manner.

[0020] In one possible implementation, the first event is represented by polarity information. The pixel array circuit may include multiple pixels, and each pixel may include a threshold comparison unit. The visual sensor chip's readout circuit reads at least one data signal from the pixel array circuit in the first event representation manner. This may include: when the light intensity change exceeds a predetermined threshold, the threshold comparison unit outputs polarity information, which indicates whether the light intensity change is increasing or decreasing. The readout circuit reads the polarity information output by the threshold comparison unit.

[0021] In one possible implementation, the first event is represented by light intensity information. The pixel array may include multiple pixels, and each pixel may include a threshold comparison unit, a readout control unit, and a light intensity acquisition unit. The visual sensor chip's readout circuit reads at least one data signal from the pixel array circuit in the first event representation manner. This may include: the light intensity acquisition unit outputting an electrical signal corresponding to the light signal illuminating it, the electrical signal indicating the light intensity; when the electrical signal determines that the light intensity change exceeds a predetermined threshold, the threshold comparison unit outputs a first signal; in response to receiving the first signal, the readout control unit instructs the light intensity acquisition unit to acquire and buffer the electrical signal corresponding to the moment the first signal is received; and the readout circuit reads the buffered electrical signal from the light intensity acquisition unit.

[0022] In one possible implementation, the method may further include: determining statistical data based on at least one data signal received from the readout circuit; if the statistical data is determined to satisfy a predetermined conversion condition, sending a conversion signal to the readout circuit, the predetermined conversion condition being determined based on a preset bandwidth of the vision sensor chip.

[0023] In one possible implementation, when the first event is represented by light intensity information and the second event is represented by polarity information, a predetermined conversion condition is that the total amount of data read from the pixel array circuit through the first event representation is greater than a preset bandwidth, or the predetermined conversion condition is that the number of at least one data signal is greater than the ratio of the preset bandwidth to the first bit, where the first bit is a preset bit of the data format of the data signal.

[0024] In one possible implementation, when the first event representation method is to represent the event through polarity information and the second event representation method is to represent the event through light intensity information, the predetermined conversion condition is that if at least one data signal is read from the pixel array circuit through the second event representation method, the total amount of data read is not greater than a preset bandwidth, or the predetermined conversion condition is that the number of at least one data signal is not greater than the ratio of the preset bandwidth to the first bit, where the first bit is a preset bit of the data format of the data signal.

[0025] Fifthly, this application provides a decoding method, comprising: reading a data signal from a vision sensor chip via a reading circuit; decoding the data signal according to a first decoding method via a decoding circuit; and decoding the data signal according to a second decoding method via a decoding circuit when a conversion signal is received from a control circuit.

[0026] In one possible implementation, the method further includes: determining statistical data based on data signals read from the readout circuit; and if the statistical data is determined to meet predetermined conversion conditions, sending a conversion signal to the encoding circuit, wherein the predetermined conversion conditions are determined based on a preset bandwidth of the vision sensor chip.

[0027] In one possible implementation, the first decoding method is to decode the data signal according to the first bit corresponding to the first event representation method, where the first event representation method represents the event through light intensity information. The second decoding method is to decode the data signal according to the second bit corresponding to the second event representation method, where the second event representation method represents the event through polarity information. The polarity information is used to indicate whether the change in light intensity is an increase or a decrease. The conversion condition is that the total amount of data decoded according to the first decoding method is greater than a preset bandwidth, or the predetermined conversion condition is that the number of data signals is greater than the ratio of the preset bandwidth to the first bit. The first bit is a preset bit of the data format of the data signal.

[0028] In one possible implementation, the first decoding method is to decode the data signal according to the first bit corresponding to the first event representation method, where the first event representation method represents the event through polarity information, which is used to indicate whether the change in light intensity is an increase or a decrease. The second decoding method is to decode the data signal according to the second event representation method, where the second event representation method represents the event through light intensity information. The conversion condition is that if the data signal is decoded according to the second decoding method, the total data amount is not greater than a preset bandwidth, or the predetermined conversion condition is that the number of data signals is greater than the ratio of the preset bandwidth to the first bit, where the first bit is a preset bit of the data format of the data signal.

[0029] In a sixth aspect, this application provides a vision sensor chip, which may include: a pixel array circuit, used to generate at least one data signal corresponding to a pixel in the pixel array circuit by measuring a change in light intensity, wherein the at least one data signal indicates a light intensity change event, and the light intensity change event indicates that the light intensity change measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A first encoding unit is used to encode the at least one data signal according to a first bit to obtain first encoded data. The first encoding unit is further used to encode the at least one data signal according to a second bit indicated by a first control signal received from a control circuit, wherein the first control signal is determined by the control circuit based on the first encoded data. As can be seen from the solution provided in the sixth aspect, by dynamically adjusting the bit width representing light intensity feature information, when the event generation rate is low and the bandwidth limit has not yet been reached, the event is quantized according to the maximum bit width and encoded. When the event generation rate is high, the bit width representing the light intensity feature information is gradually reduced to meet the bandwidth limit. Subsequently, if the event generation rate decreases again, the bit width representing the light intensity feature information can be increased without exceeding the bandwidth limit. Visual sensors can adaptively switch between multiple event representation methods to better achieve the goal of transmitting all events with greater representational accuracy.

[0030] In one possible implementation, the first control signal is determined by the control circuit based on the first encoded data and the bandwidth preset by the vision sensor chip.

[0031] In one possible implementation, when the amount of data in the first encoded data is not less than the bandwidth, the second bit indicated by the control signal is less than the first bit, so that the total amount of data in at least one data signal encoded by the second bit is not greater than the bandwidth. When the event generation rate is high, the bit width representing light intensity feature information is gradually reduced to meet the bandwidth limit.

[0032] In one possible implementation, when the amount of data in the first encoded data is less than the bandwidth, the second bit indicated by the control signal is greater than the first bit, and the total amount of data in at least one data signal encoded by the second bit is not greater than the bandwidth. If the rate at which events occur decreases, the bit width representing the light intensity feature information can be increased incrementally without exceeding the bandwidth limit, in order to better achieve the goal of transmitting all events with greater representational precision.

[0033] In one possible implementation, the pixel array may include N regions, where at least two regions have different maximum bit values. The maximum bit value represents a preset maximum bit value for encoding at least one data signal generated in a region. A first encoding unit is specifically used to encode at least one data signal generated in a first region based on a first bit to obtain first encoded data. The first bit is not greater than the maximum bit value of the first region, and the first region is any one of the N regions. Specifically, when receiving a first control signal from a control circuit, the first encoding unit is used to encode at least one data signal generated in the first region based on a second bit indicated by the first control signal. The first control signal is determined by the control circuit based on the first encoded data. In this implementation, the pixel array can also be divided into regions, and different weights can be used to set the maximum bit width of different regions to adapt to different regions of interest in the scene. For example, a larger weight can be set in regions that may include target objects, resulting in higher accuracy in the representation of events output in regions including target objects, while a smaller weight can be set in background regions, resulting in lower accuracy in the representation of events output in background regions.

[0034] In one possible implementation, the control circuit is further configured to: send a first control signal to the first encoding unit when it determines that the total data amount of at least one data signal encoded by the third bit is greater than the bandwidth, and the total data amount of at least one data signal encoded by the second bit is not greater than the bandwidth, wherein the third bit and the second bit differ by one bit unit. In this implementation, all events can be transmitted with greater representational precision without exceeding the bandwidth limit.

[0035] In a seventh aspect, this application provides a decoding device, which may include: a reading circuit for reading data signals from a vision sensor chip; a decoding circuit for decoding the data signals according to a first bit; and a decoding circuit further for decoding the data signals according to a second bit indicated by a first control signal received from a control circuit. The decoding circuit provided in the seventh aspect corresponds to the vision sensor chip provided in the sixth aspect, and is used to decode the data signals output by the vision sensor chip provided in the sixth aspect. The decoding circuit provided in the seventh aspect can dynamically adjust the decoding method according to the encoding bits used by the vision sensor.

[0036] In one possible implementation, the first control signal is determined by the control circuit based on the first encoded data and the bandwidth preset by the vision sensor chip.

[0037] In one possible implementation, the second bit is less than the first bit when the total data amount of the data signal decoded based on the first bit is not less than the bandwidth.

[0038] In one possible implementation, when the total amount of data in the data signal decoded based on the first bit is less than the bandwidth, the second bit is greater than the first bit, and the total amount of data in the data signal decoded through the second bit is not greater than the bandwidth.

[0039] In one possible implementation, the readout circuit is specifically used to read the data signal corresponding to a first region from the vision sensor chip. The first region is any one of N regions that the pixel array of the vision sensor may include. At least two of the N regions have different maximum bits, and the maximum bit represents a preset maximum bit for encoding at least one data signal generated in a region. The decoding circuit is specifically used to decode the data signal corresponding to the first region based on the first bit.

[0040] In one possible implementation, the control circuit is further configured to: when it is determined that the total amount of data signal decoded by the third bit is greater than the bandwidth and the total amount of data signal decoded by the second bit is not greater than the bandwidth, send a first control signal to the first encoding unit, wherein the third bit and the second bit differ by 1 bit unit.

[0041] Eighthly, this application provides a method for operating a vision sensor chip, which may include: measuring a change in light intensity through a pixel array circuit of the vision sensor chip to generate at least one data signal corresponding to a pixel in the pixel array circuit, wherein the at least one data signal indicates a light intensity change event, and the light intensity change event indicates that the light intensity change measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. The at least one data signal is encoded according to a first bit by a first encoding unit of the vision sensor chip to obtain first encoded data. When a first control signal is received from a control circuit of the vision sensor chip by the first encoding unit, the at least one data signal is encoded according to a second bit indicated by the first control signal, wherein the first control signal is determined by the control circuit based on the first encoded data.

[0042] In one possible implementation, the first control signal is determined by the control circuit based on the first encoded data and the bandwidth preset by the vision sensor chip.

[0043] In one possible implementation, when the amount of data in the first encoded data is not less than the bandwidth, the second bit indicated by the control signal is less than the first bit, so that the total amount of data in at least one data signal encoded by the second bit is not greater than the bandwidth.

[0044] In one possible implementation, when the amount of data in the first encoded data is less than the bandwidth, the second bit indicated by the control signal is greater than the first bit, and the total amount of data in at least one data signal encoded by the second bit is not greater than the bandwidth.

[0045] In one possible implementation, the pixel array may include N regions, at least two of which have different maximum bit values. The maximum bit represents a preset maximum bit value for encoding at least one data signal generated in a region. Encoding at least one data signal based on the first bit by the first encoding unit of the vision sensor chip may include: encoding at least one data signal generated in a first region based on the first bit by the first encoding unit to obtain first encoded data. The first bit is not greater than the maximum bit value of the first region, and the first region is any one of the N regions. When the first encoding unit receives a first control signal from the control circuit of the vision sensor chip, encoding at least one data signal based on the second bit indicated by the first control signal may include: encoding at least one data signal generated in the first region based on the second bit indicated by the first control signal by the first encoding unit when receiving the first control signal from the control circuit. The first control signal is determined by the control circuit based on the first encoded data.

[0046] In one possible implementation, it may further include: when it is determined that the total data amount of at least one data signal encoded by the third bit is greater than the bandwidth, and the total data amount of at least one data signal encoded by the second bit is not greater than the bandwidth, sending a first control signal to the first encoding unit through a control circuit, wherein the third bit and the second bit differ by 1 bit unit.

[0047] Ninthly, this application provides a decoding method, which may include: reading a data signal from a vision sensor chip via a reading circuit; decoding the data signal according to a first bit via a decoding circuit; and decoding the data signal according to a second bit indicated by the first control signal when the decoding circuit receives a first control signal from a control circuit.

[0048] In one possible implementation, the first control signal is determined by the control circuit based on the first encoded data and the bandwidth preset by the vision sensor chip.

[0049] In one possible implementation, the second bit is less than the first bit when the total data amount of the data signal decoded based on the first bit is not less than the bandwidth.

[0050] In one possible implementation, when the total amount of data in the data signal decoded based on the first bit is less than the bandwidth, the second bit is greater than the first bit, and the total amount of data in the data signal decoded through the second bit is not greater than the bandwidth.

[0051] In one possible implementation, reading data signals from the vision sensor chip via a readout circuit may include: reading data signals corresponding to a first region from the vision sensor chip via the readout circuit. The first region is any one of N regions that the pixel array of the vision sensor may include, where at least two of the N regions have different maximum bits. The maximum bit represents a preset maximum bit value for encoding at least one data signal generated in a region. Decoding the data signal based on the first bit via a decoding circuit may include: decoding the data signal corresponding to the first region based on the first bit via the decoding circuit.

[0052] In one possible implementation, the method may further include: when it is determined that the total amount of data signal decoded by the third bit is greater than the bandwidth, and the total amount of data signal decoded by the second bit is not greater than the bandwidth, sending a first control signal to the first encoding unit, wherein the third bit and the second bit differ by 1 bit unit.

[0053] Tenthly, this application provides a vision sensor chip, which may include: a pixel array circuit for generating multiple data signals corresponding to multiple pixels in the pixel array circuit by measuring changes in light intensity, wherein the multiple data signals indicate at least one light intensity change event, and the at least one light intensity change event indicates that the light intensity change measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A third encoding unit is used to encode a first difference value according to a first preset bit, wherein the first difference value is the difference between the light intensity change and the predetermined threshold. Reducing the precision of event representation, i.e., reducing the bit width representing the event, reduces the information that the event can carry, which is detrimental to event processing and analysis in some scenarios. Therefore, reducing the precision of event representation may not be suitable for all scenarios. In some scenarios, a high bit width is required to represent events, but while a high bit width can carry more data, the data volume is also larger. Given a fixed maximum bandwidth preset by the vision sensor, there may be situations where event data cannot be read, resulting in data loss. The solution provided in the tenth aspect uses differential values ​​to reduce the cost of data transmission, parsing, and storage of visual sensors, while transmitting events with the highest accuracy, significantly improving sensor performance.

[0054] In one possible implementation, the pixel array circuit may include multiple pixels, each pixel may include a threshold comparison unit, which outputs polarity information when the light intensity change exceeds a predetermined threshold. The polarity information indicates whether the light intensity change is increasing or decreasing. A third encoding unit is further configured to encode the polarity information according to a second preset set of bits. In this implementation, the polarity information can also be encoded to indicate whether the light intensity is increasing or decreasing, which helps to obtain the current light intensity information based on the light intensity signal obtained from the previous decoding and the polarity information.

[0055] In one possible implementation, each pixel may include a light intensity detection unit, a readout control unit, and a light intensity acquisition unit. The light intensity detection unit outputs an electrical signal corresponding to the light signal illuminating it, which indicates the light intensity. A threshold comparison unit is specifically used to output polarity information when the light intensity change exceeds a predetermined threshold based on the electrical signal. The readout control unit is used to instruct the light intensity acquisition unit to acquire and buffer the electrical signal at the moment the corresponding polarity information is received, in response to receiving the polarity signal. A third encoding unit is further used to encode the first electrical signal according to a third preset bit. The first electrical signal is the electrical signal at the moment the corresponding polarity information is first received by the light intensity acquisition unit, and the third preset bit is the maximum number of bits preset by the vision sensor to represent the feature information of light intensity. After full encoding of the initial state, subsequent events only need to encode the polarity information and the difference between the light intensity change and the predetermined threshold, which can effectively reduce the amount of encoded data. Here, full encoding refers to encoding an event using the maximum bit width predefined by the vision sensor. Furthermore, by utilizing the light intensity information from the previous event, as well as the decoded polarity information and difference value, the light intensity information at the current moment can be reconstructed losslessly.

[0056] In one possible implementation, the third encoding unit is further configured to: encode the electrical signal acquired by the light intensity acquisition unit according to a third preset number of bits every preset time interval. Full encoding is performed every preset time interval to reduce decoding dependency and prevent bit errors.

[0057] In one possible implementation, the third encoding unit is specifically used to: encode the first difference value according to a first preset bit when the first difference value is less than a predetermined threshold.

[0058] In one possible implementation, the third encoding unit is further configured to: when the first difference value is not less than a predetermined threshold, encode the first remaining difference value and the predetermined threshold according to a first preset bit, wherein the first remaining difference value is the difference between the difference value and the predetermined threshold.

[0059] In one possible implementation, the third encoding unit is specifically used to: when the first residual difference value is not less than a predetermined threshold, encode the second residual difference value according to a first preset bit, wherein the second residual difference value is the difference between the first residual difference value and the predetermined threshold. The predetermined threshold is encoded for the first time according to the first preset bit. The predetermined threshold is encoded for the second time according to the first preset bit. Because the visual sensor may have a certain delay, it may be possible that the light intensity change exceeds the predetermined threshold twice or more before an event occurs. This can lead to problems where the difference value is greater than or equal to the predetermined threshold, and the light intensity change is at least twice the predetermined threshold. For example, the first residual difference value may not be less than the predetermined threshold. In this case, the second residual difference value is encoded. If the second residual difference value is still not less than the predetermined threshold, a third residual difference value can be encoded, wherein the third difference value is the difference between the second residual difference value and the predetermined threshold, and the predetermined threshold is encoded for the third time. This process is repeated until the residual difference value is less than the predetermined threshold.

[0060] Eleventhly, this application provides a decoding device, which may include: an acquisition circuit for reading data signals from a vision sensor chip; and a decoding circuit for decoding the data signal based on a first bit to obtain a differential value, wherein the differential value is less than a predetermined threshold, the differential value being the difference between the light intensity change measured by the vision sensor and the predetermined threshold, and the light intensity change exceeding the predetermined threshold causing the vision sensor to generate at least one light intensity change event. The decoding circuit provided in the eleventh aspect corresponds to the vision sensor chip provided in the tenth aspect, and is used to decode the data signals output by the vision sensor chip provided in the tenth aspect. The decoding circuit provided in the eleventh aspect can employ a corresponding differential decoding method for the differential encoding method used by the vision sensor.

[0061] In one possible implementation, the decoding circuit is further configured to: decode the data signal according to the second bit to obtain polarity information, which is used to indicate whether the change in light intensity is an increase or a decrease.

[0062] In one possible implementation, the decoding circuit is further configured to: decode the data signal received at the first moment according to the third bit to obtain the electrical signal corresponding to the light signal illuminating it output by the vision sensor, wherein the third bit is the maximum bit of feature information preset by the vision sensor to represent the light intensity.

[0063] In one possible implementation, the decoding circuit is further configured to: decode the data signal received at the first moment according to the third bit at preset intervals.

[0064] In one possible implementation, the decoding circuit is specifically configured to: decode the data signal based on the first bit to obtain a differential value and at least one predetermined threshold.

[0065] In a twelfth aspect, this application provides a method for operating a vision sensor chip, which may include: measuring light intensity changes through a pixel array circuit of the vision sensor chip to generate multiple data signals corresponding to multiple pixels in the pixel array circuit, wherein the multiple data signals indicate at least one light intensity change event, and the at least one light intensity change event indicates that the light intensity change measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A third encoding unit of the vision sensor chip encodes a first difference value according to a first preset bit, wherein the first difference value is the difference between the light intensity change and the predetermined threshold.

[0066] In one possible implementation, the pixel array circuit may include multiple pixels, each pixel may include a threshold comparison unit, and the method may further include: when the light intensity change exceeds a predetermined threshold, outputting polarity information through the threshold comparison unit, the polarity information being used to indicate whether the light intensity change is increasing or decreasing; and encoding the polarity information according to a second preset bit by a third encoding unit.

[0067] In one possible implementation, each pixel may include a light intensity detection unit, a readout control unit, and a light intensity acquisition unit. The method may further include: outputting an electrical signal corresponding to the light signal illuminating the pixel through the light intensity detection unit, the electrical signal indicating the light intensity. Outputting polarity information through a threshold comparison unit may include: when the light intensity change exceeds a predetermined threshold based on the electrical signal, outputting polarity information through the threshold comparison unit. The method may further include: in response to receiving a polarity signal, instructing the light intensity acquisition unit to acquire and buffer the electrical signal corresponding to the moment the polarity information is received through the readout control unit. Encoding a first electrical signal according to a third preset bit, the first electrical signal being the electrical signal acquired by the light intensity acquisition unit at the moment of the first reception of the corresponding polarity information, the third preset bit being the maximum number of bits preset by the visual sensor to represent the feature information of light intensity.

[0068] In one possible implementation, the method may further include: encoding the electrical signal acquired by the light intensity acquisition unit according to a third preset bit every preset time interval.

[0069] In one possible implementation, encoding the first difference value according to a first preset bit by the third encoding unit of the vision sensor chip may include: encoding the first difference value according to the first preset bit when the first difference value is less than a predetermined threshold.

[0070] In one possible implementation, the third encoding unit of the vision sensor chip encodes the first difference value according to a first preset bit, and may further include: when the first difference value is not less than a predetermined threshold, encoding the first remaining difference value and the predetermined threshold according to the first preset bit, wherein the first remaining difference value is the difference between the difference value and the predetermined threshold.

[0071] In one possible implementation, when the first difference value is not less than a predetermined threshold, encoding the first remaining difference value and the predetermined threshold according to a first preset bit may include: encoding a second remaining difference value according to the first preset bit when the first remaining difference value is not less than the predetermined threshold, wherein the second remaining difference value is the difference between the first remaining difference value and the predetermined threshold. The predetermined threshold is encoded for the first time according to the first preset bit. The predetermined threshold is encoded for the second time according to the first preset bit, wherein the first remaining difference value may include the second remaining difference value and two predetermined thresholds.

[0072] In a thirteenth aspect, this application provides a decoding method, which may include: reading a data signal from a vision sensor chip via an acquisition circuit; decoding the data signal according to a first bit via a decoding circuit to obtain a difference value, wherein the difference value is less than a predetermined threshold, the difference value being the difference between the light intensity change measured by the vision sensor and the predetermined threshold; if the light intensity change exceeds the predetermined threshold, the vision sensor generates at least one light intensity change event.

[0073] In one possible implementation, it may further include: decoding the data signal according to the second bit to obtain polarity information, which is used to indicate whether the change in light intensity is an increase or a decrease.

[0074] In one possible implementation, it may further include: decoding the data signal received at the first moment according to the third bit to obtain the electrical signal corresponding to the light signal illuminating it output by the vision sensor, wherein the third bit is the maximum bit of feature information preset by the vision sensor to represent the light intensity.

[0075] In one possible implementation, it may further include: decoding the data signal received at the first moment based on the third bit at preset intervals.

[0076] In one possible implementation, decoding the data signal based on the first bit by a decoding circuit to obtain a differential value may include: decoding the data signal based on the first bit to obtain a differential value and at least one predetermined threshold.

[0077] In a fourteenth aspect, this application provides an image processing method, comprising: acquiring motion information, the motion information including information on the motion trajectory of a target object moving within the detection range of a motion sensor; generating at least one event image based on the motion information, the at least one event image being an image representing the motion trajectory of the target object moving within the detection range; acquiring a target task and acquiring an iteration duration based on the target task; iteratively updating the at least one event image to obtain an updated at least one event image, wherein the duration of iteratively updating the at least one event image does not exceed the iteration duration.

[0078] Therefore, in this embodiment, a moving object can be monitored by a motion sensor, and information on the motion trajectory of the object when it moves within the detection range can be collected by the motion sensor. After obtaining the target task, the iteration duration can be determined according to the target task, and the event image can be iteratively updated within the iteration duration to obtain an event image that matches the target task.

[0079] In one possible implementation, any one of the iterative updates in the iterative update of the at least one frame of event image includes: obtaining motion parameters, the motion parameters representing parameters of the relative motion between the motion sensor and the target object; and iteratively updating the target event image in the at least one frame of event image according to the motion parameters to obtain an updated target event image.

[0080] Therefore, in this embodiment of the application, when iteratively updating the event image, the update can be based on the parameters of the relative motion between the object and the motion sensor, thereby compensating for the event image and obtaining a clearer event image.

[0081] In one possible implementation, obtaining the motion parameters includes: obtaining the value of the preset optimization model during the previous iteration update; and calculating the motion parameters based on the value of the optimization model.

[0082] Therefore, in this embodiment, the event image can be updated based on the value of the optimization model, and better motion parameters can be calculated based on the optimization model. Then, the event image can be updated using these motion parameters to obtain a clearer event image.

[0083] In one possible implementation, the step of iteratively updating the target event image in the at least one frame of event images according to the motion parameters includes: compensating for the motion trajectory of the target object in the target event image according to the motion parameters to obtain the target event image obtained in the current iteration update.

[0084] Therefore, in this embodiment, motion parameters can be used to compensate for the motion trajectory of the target object in the event image, making the motion trajectory of the target object in the event image clearer, thereby making the event image clearer.

[0085] In one possible implementation, the motion parameters include one or more of the following: depth, optical flow information, acceleration of the motion sensor or angular velocity of the motion sensor, wherein the depth represents the distance between the motion sensor and the target object, and the optical flow information represents information about the relative motion speed between the motion sensor and the target object.

[0086] Therefore, in the embodiments of this application, motion compensation can be performed on the target object in the event image using various motion parameters to improve the clarity of the event image.

[0087] In one possible implementation, during any one iteration update process, the method further includes: if the result of the current iteration meets a preset condition, then the iteration is terminated, the termination condition including at least one of the following: the number of iterations for updating the at least one frame of event image reaches a preset number or the value change of the optimization model during the update of the at least one frame of event image is less than a preset value.

[0088] Therefore, in the embodiments of this application, in addition to setting the iteration duration, convergence conditions related to the number of iterations or the value of the optimization model can also be set, so as to obtain an event image that meets the convergence conditions under the constraint of the iteration duration.

[0089] In a fifteenth aspect, this application provides an image processing method, comprising: generating at least one event image frame based on motion information, wherein the motion information includes information on the motion trajectory of a target object moving within the detection range of a motion sensor, and the at least one event image frame being an image representing the motion trajectory of the target object moving within the detection range; acquiring motion parameters, wherein the motion parameters represent parameters of the relative motion between the motion sensor and the target object; initializing the value of a preset optimization model based on the motion parameters to obtain the value of the optimization model; and updating the at least one event image frame based on the value of the optimization model to obtain the updated at least one event image frame.

[0090] In this embodiment, the parameters of the relative motion between the motion sensor and the target object can be used to initialize the optimization model, thereby reducing the initial number of iterations of the event image, accelerating the convergence speed of the event image iteration, and obtaining a clearer event image with fewer iterations.

[0091] In one possible implementation, the motion parameters include one or more of the following: depth, optical flow information, acceleration of the motion sensor or angular velocity of the motion sensor, wherein the depth represents the distance between the motion sensor and the target object, and the optical flow information represents information about the relative motion speed between the motion sensor and the target object.

[0092] In one possible implementation, acquiring the motion parameters includes: acquiring data collected by an inertial measurement unit (IMU) sensor; and calculating the motion parameters based on the data collected by the IMU sensor. Therefore, in this embodiment, motion parameters can be calculated using an IMU, thereby obtaining more accurate motion parameters.

[0093] In one possible implementation, after initializing the values ​​of the preset optimization model based on the motion parameters, the method further includes: updating the parameters of the IMU sensor based on the values ​​of the optimization model, wherein the parameters of the IMU sensor are used for data acquisition by the IMU sensor.

[0094] Therefore, in this embodiment, the parameters of the IMU can also be updated according to the value of the optimized model to correct the bias of the IMU and make the data collected by the IMU more accurate.

[0095] In a sixteenth aspect, this application provides an image processing apparatus that has the function of implementing the method of the fourteenth aspect or any possible implementation thereof, or the image processing apparatus has the function of implementing the method of the fifteenth aspect or any possible implementation thereof. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described function.

[0096] In a seventeenth aspect, this application provides an image processing method, comprising: acquiring motion information, the motion information including information on the motion trajectory of a target object moving within the detection range of a motion sensor; generating an event image based on the motion information, the event image being an image representing the motion trajectory of the target object moving within the detection range; and obtaining a first reconstructed image based on at least one event included in the event image, wherein a first pixel and at least one second pixel have different color types, the first pixel is a pixel corresponding to any one of the at least one events in the first reconstructed image, and the at least one second pixel is included among a plurality of pixels adjacent to the first pixel in the first reconstructed image.

[0097] Therefore, in the embodiments of this application, when there is relative motion between the subject being photographed and the motion sensor, image reconstruction can be performed based on the data collected by the motion sensor to obtain the reconstructed image. Even when the image captured by the RGB sensor is unclear, a clear image can still be obtained.

[0098] In one possible implementation, determining the color type corresponding to each pixel in the event image based on at least one event included in the event image to obtain a first reconstructed image includes: scanning each pixel in the event image along a first direction, determining the color type corresponding to each pixel in the event image, and obtaining a first reconstructed image, wherein if the first pixel scanned has an event, the color type of the first pixel is determined to be a first color type, and if a second pixel arranged before the first pixel along the first direction does not have an event, the color type corresponding to the second pixel is a second color type, the first color type and the second color type are different color types, and the pixel with an event represents the pixel in the event image corresponding to the position where the motion sensor detected a change.

[0099] In this embodiment, image reconstruction can be performed based on the events at each pixel in the event image by scanning the event image, thereby obtaining a clearer event image. Therefore, in this embodiment, information collected by a motion sensor can be used for image reconstruction, efficiently and quickly obtaining the reconstructed image, thus improving the efficiency of subsequent image recognition, image classification, etc., on the reconstructed image. Even in scenarios where moving objects are being photographed or where there is camera shake, making it impossible to capture a clear RGB image, image reconstruction can still be performed using information collected by a motion sensor, quickly and accurately reconstructing a clearer image for subsequent recognition or classification tasks.

[0100] In one possible implementation, the first direction is a preset direction, or the first direction is determined based on data collected by the IMU, or the first direction is determined based on an image captured by a color RGB camera. Therefore, in this application embodiment, the direction of scanning the event image can be determined in multiple ways to adapt to more scenarios.

[0101] In one possible implementation, if a plurality of consecutive third pixels arranged in the first direction after the first pixel do not have an event, then the color type corresponding to the plurality of third pixels is the first color type. Therefore, in this embodiment, when there are multiple consecutive pixels that do not have an event, the consecutive pixels correspond to the same color type, avoiding unclear edges caused by the movement of the same object in a real scene.

[0102] In one possible implementation, if a fourth pixel, which is arranged in the first direction after the first pixel and adjacent to the first pixel, has an event, and a fifth pixel, which is arranged in the first direction after the fourth pixel and adjacent to the fourth pixel, does not have an event, then the color type corresponding to the fourth pixel and the fifth pixel is the first color type.

[0103] Therefore, when there are at least two consecutive pixels in the event image that have an event, the reconstructed color type can be maintained when the second event is scanned, thus avoiding unclear edges in the reconstructed image caused by the target object's edges being too wide.

[0104] In one possible implementation, after scanning each pixel in the event image along a first direction to determine the color type corresponding to each pixel in the event image and obtain a first reconstructed image, the method further includes: scanning the event image along a second direction to determine the color type corresponding to each pixel in the event image and obtain a second reconstructed image, wherein the second direction is different from the first direction; and fusing the first reconstructed image and the second reconstructed image to obtain an updated first reconstructed image.

[0105] In this embodiment of the application, the event image can be scanned in different directions to obtain multiple reconstructed images from multiple directions, and then the multiple reconstructed images are fused to obtain a more accurate reconstructed image.

[0106] In one possible implementation, the method further includes: if the first reconstructed image does not meet preset requirements, then updating motion information, updating the event image according to the updated motion information, and obtaining the updated first reconstructed image according to the updated event image.

[0107] In this embodiment, the event image can be updated by combining information collected by the motion sensor, making the updated event image clearer.

[0108] In one possible implementation, before determining the color type corresponding to each pixel in the event image based on at least one event included in the event image to obtain a first reconstructed image, the method further includes: compensating the event image based on motion parameters when the target object and the motion sensor are in relative motion to obtain a compensated event image, wherein the motion parameters include one or more of the following: depth, optical flow information, acceleration of the motion sensor or angular velocity of the motion sensor, wherein the depth represents the distance between the motion sensor and the target object, and the optical flow information represents information on the relative motion speed between the motion sensor and the target object.

[0109] Therefore, in this embodiment of the application, motion compensation can also be performed on the event image by combining motion parameters, so that the event image is clearer, and the reconstructed image obtained is also clearer.

[0110] In one possible implementation, the color type of pixels in the reconstructed image is determined based on the colors captured by an RGB color camera. In this embodiment, the colors in the actual scene can be determined based on the RGB camera, thereby matching the colors of the reconstructed image with those in the actual scene and improving the user experience.

[0111] In one possible implementation, the method further includes: obtaining an RGB image based on data captured by an RGB camera; and fusing the RGB image with the first reconstructed image to obtain an updated first reconstructed image. Therefore, in this embodiment, the RGB image and the reconstructed image can be fused to make the final reconstructed image clearer.

[0112] In an eighteenth aspect, this application also provides an image processing apparatus having the function of implementing the method of the eighteenth aspect or any possible implementation thereof. This function can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described function.

[0113] Nineteenthly, this application provides an image processing method, comprising: acquiring a first event image (event image) and multiple captured first images, wherein the first event image includes information about an object moving within a preset range during a shooting time period of the multiple first images, the multiple first images having different exposure durations, and the preset range being the shooting range of a camera; calculating a first jitter degree corresponding to each of the multiple first images based on the first event image, the first jitter degree representing the degree of camera jitter during the shooting of the multiple first images; determining a fusion weight for each of the multiple first images based on the first jitter degree corresponding to each of the first images, wherein the first jitter degree corresponding to the multiple first images and the fusion weight are negatively correlated; and fusing the multiple first images according to the fusion weight of each of the first images to obtain a target image.

[0114] Therefore, in this embodiment, the degree of jitter during RGB image capture can be quantified using event images, and the fusion weight of each RGB image can be determined based on the jitter level of each RGB image. Generally, RGB images with lower jitter levels correspond to higher fusion weights, resulting in the final target image containing information more aligned with the clearer RGB image, thus yielding a clearer target image. Conversely, RGB images with higher jitter levels typically have lower weight values, while RGB images with lower jitter levels have higher weight values, further ensuring the final target image contains information more aligned with the clearer RGB image, resulting in a clearer target image and improved user experience. Furthermore, if this target image is used for subsequent image recognition or feature extraction, the resulting recognition results or extracted features will be more accurate.

[0115] In one possible implementation, before determining the fusion weight of each of the plurality of first images based on the first jitter level, the method further includes: if the first jitter level is not higher than a first preset value but higher than a second preset value, then performing jitter removal processing on each of the first images to obtain each of the first images after jitter removal.

[0116] Therefore, in the embodiments of this application, the jitter situation can be distinguished based on dynamic data. When there is no jitter, the images can be directly fused. When the jitter is not strong, the RGB images can be adaptively de-jittered. When the jitter is strong, additional RGB images can be captured. This method uses multiple jitter levels and has strong generalization ability.

[0117] In one possible implementation, determining the fusion weight of each of the plurality of first images based on the first degree of jitter includes: if the first degree of jitter is higher than a first preset value, then re-capturing a second image, wherein the second degree of jitter of the second image is not higher than the first preset value; calculating the fusion weight of each first image based on the first degree of jitter of each first image, and calculating the fusion weight of the second image based on the second degree of jitter; fusing the plurality of first images based on the fusion weight of each first image to obtain a target image includes: fusing the plurality of first images and the second image based on the fusion weight of each first image and the fusion weight of the second image to obtain the target image.

[0118] Generally, RGB images with higher jitter levels have lower weight values, while RGB images with lower jitter levels have higher weight values. This results in the final target image containing information more closely resembling that of the clearer RGB image, leading to a sharper final image and improved user experience. Furthermore, if this target image is used for subsequent image recognition or feature extraction, the recognition results or extracted features will be more accurate. For RGB images with high jitter levels, a new RGB image can be captured to obtain a clearer RGB image with lower jitter levels. This allows for the use of the clearer image during subsequent image fusion, resulting in a sharper final target image.

[0119] In one possible implementation, before re-capturing the second image, the method further includes: acquiring a second event image, which was obtained before acquiring the first event image; and calculating exposure parameters based on information included in the second event image, the exposure parameters being used to capture the second image.

[0120] Therefore, in this embodiment, the exposure strategy is adaptively adjusted by using the information collected by the dynamic sensing camera (i.e., motion sensor). That is, by using the high dynamic range sensing characteristics of the texture within the shooting range, the camera adaptively supplements the shooting of images with appropriate exposure time, thereby improving the camera's ability to capture texture information in bright or dark areas.

[0121] In one possible implementation, the re-capturing of the second image further includes: dividing the first event image into multiple regions and dividing the third image into multiple regions, wherein the third image is the first image with the smallest exposure value among the multiple first images, and the multiple regions included in the first event image correspond in position to the multiple regions included in the third image, wherein the exposure value includes at least one of exposure duration, exposure amount, or exposure level; calculating whether each region in the first event image includes first texture information and whether each region in the third image includes second texture information; if the first region in the first event image includes the first texture information and the region in the third image corresponding to the first region does not include the second texture information, then capturing an image according to the exposure parameters to obtain the second image, wherein the first region is any region in the first dynamic region.

[0122] Therefore, in this embodiment, if a region in the first dynamic region includes texture information, and the region in the RGB image with the lowest exposure value that is the same as that region does not include texture information, it indicates that the blurriness of that region in the RGB image is high, and the RGB image can be reshot. However, if no region in the first event image includes texture information, then there is no need to reshot the RGB image.

[0123] In a twentieth aspect, this application provides an image processing method, comprising: first, detecting motion information of a target object, the motion information including information on the motion trajectory of the target object when it moves within a preset range, the preset range being the shooting range of a camera; then, determining focus information based on the motion information, the focus information including parameters for focusing on the target object within the preset range; subsequently, focusing on the target object within the preset range based on the focus information, and capturing an image of the preset range.

[0124] Therefore, in this embodiment, the motion trajectory of the target object within the camera's shooting range can be detected, and then focus information can be determined and focusing completed based on the target object's motion trajectory, thereby capturing a clearer image. Even if the target object is in motion, it can be accurately focused on, capturing a clear image of the motion state and improving the user experience.

[0125] In one possible implementation, the above-mentioned determination of focus information based on motion information may include: predicting the motion trajectory of the target object within a preset time period based on motion information, i.e., information on the motion trajectory of the target object when it moves within a preset range, to obtain a prediction area, wherein the prediction area is the area where the target object is located within the predicted preset time period; determining a focus area based on the prediction area, wherein the focus area includes at least one focus point for focusing on the target object, and the focus information includes the position information of at least one focus point.

[0126] Therefore, in this embodiment, the future trajectory of the target object can be predicted, and the focus area can be determined based on the predicted area, thus accurately focusing on the target object. Even if the target object is moving at high speed, this embodiment can also focus on the target object in advance through prediction, placing the target object in the focus area, thereby capturing a clearer image of the high-speed moving target object.

[0127] In one possible implementation, determining the focus area based on the predicted area may include: if the predicted area meets a preset condition, then the predicted area is determined as the focus area; if the predicted area does not meet the preset condition, then the motion trajectory of the target object within a preset time period is re-predicted based on motion information to obtain a new predicted area, and the focus area is determined based on the new predicted area. The preset condition may be that the predicted area includes the complete target object, or that the area of ​​the predicted area is greater than a preset value, etc.

[0128] Therefore, in this embodiment, the focus area is determined and the camera is triggered to capture images only when the predicted area meets the preset conditions. If the predicted area does not meet the preset conditions, the camera is not triggered to capture images. This avoids incomplete images of the target object or unnecessary captures. Furthermore, the camera can remain inactive when not capturing images, and is only triggered to capture images when the predicted area meets the preset conditions, thus reducing the power consumption of the camera.

[0129] In one possible implementation, the motion information further includes at least one of the target object's motion direction and motion speed; the above-mentioned prediction of the target object's motion trajectory within a preset time period based on the motion information to obtain a prediction area may include: predicting the target object's motion trajectory within a preset time period based on the target object's motion trajectory when moving within a preset range, as well as its motion direction and / or motion speed, to obtain a prediction area.

[0130] Therefore, in this embodiment, the movement trajectory of the target object within a preset time period can be predicted based on the movement trajectory of the target object within a preset range, as well as the movement direction and / or movement speed, thereby accurately predicting the area where the target object will be located within the preset time period, and thus enabling more accurate focusing on the target object, and ultimately capturing clearer images.

[0131] In one possible implementation, the above-mentioned prediction of the target object's trajectory within a preset time period based on the target object's trajectory, direction, and / or speed within a preset range to obtain a prediction region may include: fitting a function of the change of the center point of the target object's region over time based on the target object's trajectory, direction, and / or speed within the preset range; subsequently calculating the predicted center point based on the function, the predicted center point being the center point of the region where the target object is located within the predicted preset time period; and obtaining the prediction region based on the predicted center point.

[0132] Therefore, in this embodiment, the center point of the area where the target object is located can be fitted with a function of time based on the motion trajectory of the target object. Then, the center point of the area where the target object is located at a certain time in the future can be predicted based on the function of time. The predicted area can be determined based on the center point, thereby enabling more accurate focusing on the target object and capturing clearer images.

[0133] In one possible implementation, the image of the prediction range can be captured by an RGB camera, and the above-mentioned focusing on the target object within the preset range based on the focus information can include: focusing on at least one point among the multiple focus points of the RGB camera that has the smallest norm distance to the center point of the focus area.

[0134] Therefore, in this embodiment of the application, at least one point that is closest to the norm distance of the center point of the focusing area can be selected as the focus point and focusing can be performed to complete the focusing of the target object.

[0135] In one possible implementation, the motion information includes the current location of the target object, and the above-mentioned determination of focus information based on the motion information may include: determining the current location of the target object as the focus area, the focus area including at least one focus point for focusing on the target object, and the focus information including the position information of at least one focus point.

[0136] Therefore, in this embodiment, the information of the target object's movement trajectory within a preset range may include the area where the target object is currently located and the area where the target object has historically been located. The area where the target object is currently located can be used as the focus area to complete the focusing of the target object, thereby enabling the capture of a clearer image.

[0137] In one possible implementation, before capturing images within a preset range, the method may further include: acquiring exposure parameters; capturing images within a preset range may include: capturing images within a preset range based on the exposure parameters.

[0138] Therefore, in this embodiment of the application, the exposure parameters can also be adjusted to complete the shooting and obtain a clear image.

[0139] In one possible implementation, the acquisition of exposure parameters described above may include: determining exposure parameters based on motion information, wherein the exposure parameters include exposure duration, the motion information includes the movement speed of the target object, and the exposure duration is negatively correlated with the movement speed of the target object.

[0140] Therefore, in this embodiment, the exposure time can be determined by the movement speed of the target object, matching the exposure time to the movement speed of the target object. For example, the faster the movement speed, the shorter the exposure time, and the slower the movement speed, the longer the exposure time. This avoids overexposure or underexposure, thereby allowing for clearer images to be captured subsequently, improving the user experience.

[0141] In one possible implementation, obtaining the exposure parameters described above may include: determining the exposure parameters based on the light intensity, wherein the exposure parameters include the exposure duration, and the magnitude of the light intensity within a preset range is negatively correlated with the exposure duration.

[0142] Therefore, in this embodiment, the exposure time can be determined based on the detected light intensity. When the light intensity is greater, the exposure time is shorter, and when the light intensity is less, the exposure time is longer, thereby ensuring an appropriate amount of exposure and capturing a clearer image.

[0143] In one possible implementation, after capturing images within a preset range, the method may further include: fusing the images within the preset range based on the motion information of the detected target object and the corresponding image to obtain a target image within the preset range.

[0144] Therefore, in this embodiment of the application, while capturing images, the movement of the target object within a preset range can be monitored to obtain information about the movement of the target object in the image, such as the outline of the target object and the position of the target object within the preset range. This information is then used to enhance the captured image to obtain a clearer target image.

[0145] In one possible implementation, the above-mentioned detection of motion information of target objects within a preset range may include: monitoring the motion of target objects within a preset range using a dynamic vision sensor (DVS) to obtain motion information.

[0146] Therefore, in this embodiment of the application, the DVS can be used to monitor moving objects within the camera's shooting range, thereby obtaining accurate motion information. Even if the target object is in a high-speed motion state, the DVS can capture the target object's motion information in a timely manner.

[0147] In a twentieth aspect, this application also provides an image processing apparatus that has the function of implementing the method of the nineteenth aspect or any possible implementation thereof, or the image processing apparatus has the function of implementing the method of the twentieth aspect or any possible implementation thereof. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described function.

[0148] In a twentieth aspect, embodiments of this application provide a graphical user interface (GUI), characterized in that the GUI is stored in an electronic device, the electronic device including a display screen, a memory, and one or more processors, the one or more processors being used to execute one or more computer programs stored in the memory, the GUI including: responding to a trigger operation for shooting a target object, and shooting an image of a preset range according to focus information, and displaying the image of the preset range, the preset range being the shooting range of a camera, the focus information including parameters for focusing on the target object within the preset range, the focus information being determined based on motion information of the target object, the motion information including information on the motion trajectory of the target object when it moves within the preset range.

[0149] The beneficial effects of the 22nd aspect and any possible implementation thereof can be referred to the description of the 20th aspect and any possible implementation thereof.

[0150] In one possible implementation, the graphical user interface may further include: predicting the motion trajectory of the target object within a preset time period in response to the motion information, obtaining a prediction area, the prediction area being the area where the target object is located within the predicted preset time period, determining the focus area based on the prediction area, displaying the focus area on the display screen, the focus area including at least one focus point for focusing on the target object, and the focus information including the position information of at least one focus point.

[0151] In one possible implementation, the graphical user interface may specifically include: if the predicted area meets preset conditions, then in response to determining the focus area based on the predicted area, displaying the focus area on the display screen; if the predicted area does not meet preset conditions, then in response to re-predicting the motion trajectory of the target object within a preset time period based on the motion information to obtain a new predicted area, and in response to determining the focus area based on the new predicted area, displaying the focus area on the display screen.

[0152] In one possible implementation, the motion information further includes at least one of the motion direction and motion speed of the target object; the graphical user interface may specifically include: in response to predicting the motion trajectory of the target object within a preset time period based on the motion trajectory of the target object moving within a preset range, and the motion direction and / or the motion speed, obtaining the prediction area, and displaying the prediction area on the display screen.

[0153] In one possible implementation, the graphical user interface may specifically include: in response to the motion trajectory of the target object moving within a preset range, and the direction of motion and / or the speed of motion, fitting a function of the change of the center point of the region where the target object is located over time, calculating a predicted center point based on the change function, the predicted center point being the predicted center point of the region where the target object is located, obtaining the predicted region based on the predicted center point, and displaying the predicted region on a display screen.

[0154] In one possible implementation, the image of the prediction range is captured by an RGB camera, and the graphical user interface may specifically include: in response to focusing on at least one point among a plurality of focus points of the RGB camera that has the smallest norm distance to the center point of the focus area, displaying on a display screen the image captured after focusing based on the at least one point as the focus point.

[0155] In one possible implementation, the motion information includes the current location of the target object, and the graphical user interface may specifically include: in response to using the current location of the target object as the focus area, the focus area including at least one focus point for focusing on the target object, and the focus information including the position information of at least one focus point, displaying the focus area on the display screen.

[0156] In one possible implementation, the graphical user interface may further include: in response to information about the motion of the target object and the image corresponding to the detected motion, fusing images within a preset range to obtain a target image within the preset range, and displaying the target image on the display screen.

[0157] In one possible implementation, the motion information is obtained by monitoring the motion of the target object within a preset range using a dynamic vision sensor (DVS).

[0158] In one possible implementation, the graphical user interface may specifically include: in response to acquiring exposure parameters before capturing an image of the preset range, displaying the exposure parameters on a display screen; and in response to capturing an image of the preset range according to the exposure parameters, displaying the image of the preset range captured according to the exposure parameters on a display screen.

[0159] In one possible implementation, the exposure parameters are determined based on the motion information, and the exposure parameters include the exposure duration, which is negatively correlated with the motion speed of the target object.

[0160] In one possible implementation, the exposure parameters are determined based on the light intensity, which can be the light intensity detected by the camera or the light intensity detected by the motion sensor. The exposure parameters include the exposure duration, and the magnitude of the light intensity within the preset range is negatively correlated with the exposure duration.

[0161] In a twentieth aspect, this application provides an image processing method, comprising: firstly, acquiring an event stream and a frame of RGB image (referred to as a first RGB image) using a camera equipped with a motion sensor (e.g., DVS) and an RGB sensor, wherein the acquired event stream includes at least one frame of event image, each frame of which is generated from the motion trajectory information of a target object (i.e., a moving object) moving within the monitoring range of the motion sensor, and the first RGB image is a superposition of the scene captured by the camera at each moment within the exposure time. After acquiring the event stream and the first RGB image, a mask can be constructed based on the event stream. The mask is used to determine the motion region of each frame of event image in the event stream, that is, to determine the position of the moving object in the RGB image. After obtaining the event stream, the first RGB image, and the mask according to the above steps, a second RGB image can be obtained based on the event stream, the first RGB image, and the mask. The second RGB image is an RGB image with the target object removed.

[0162] In the above embodiments of this application, moving objects can be removed based on only one RGB image and event stream, thereby obtaining an RGB image without moving objects. Compared with the prior art, which requires multiple RGB images and event streams to remove moving objects, only one RGB image needs to be captured by the user, resulting in a better user experience.

[0163] In one possible implementation, before constructing the mask based on the event stream, the method may further include: triggering the camera to capture a third RGB image when the motion sensor detects a sudden change in motion within the monitored range at a first moment; obtaining the second RGB image based on the event stream, the first RGB image, and the mask includes: obtaining the second RGB image based on the event stream, the first RGB image, the third RGB image, and the mask. In this case, obtaining the second RGB image based on the event stream, the first RGB image, and the mask can be: obtaining the second RGB image based on the event stream, the first RGB image, the third RGB image, and the mask.

[0164] In the above embodiments of this application, it can be determined whether there is a sudden change in motion in the motion data collected by the motion sensor. When a sudden change in motion exists, the camera is triggered to capture a third RGB image. Then, an event stream and a frame of a first RGB image are acquired in a similar manner, and a mask is constructed based on the event stream. Finally, a second RGB image without moving foreground is obtained based on the event stream, the first RGB image, the third RGB image, and the mask. The obtained third RGB image is highly sensitive because it is automatically captured by the camera under the condition of a sudden change in motion. Thus, an image frame can be obtained as soon as the user notices a change in a moving object. Based on this third RGB image and the first RGB image, a better removal effect for moving objects can be achieved.

[0165] In one possible implementation, the motion sensor detecting a sudden change in motion within the monitoring range at a first moment includes: within the monitoring range, the overlap between the region where the motion sensor acquires a first event stream at the first moment and the region where the motion sensor acquires a second event stream at a second moment is less than a preset value.

[0166] The above embodiments of this application specifically describe the determination conditions for motion mutation, which are feasible.

[0167] In one possible implementation, the mask can be constructed based on the event stream as follows: First, the monitoring range of the motion sensor can be divided into multiple preset neighborhoods (let's call them neighborhoods k). Then, within each neighborhood k, if the number of event images in the event stream within a preset duration Δt exceeds a threshold P, the corresponding neighborhood is determined to be a motion region, which can be marked as 0. If the number of event images in the event stream within the preset duration Δt does not exceed the threshold P, the corresponding neighborhood is determined to be a background region, which can be marked as 1.

[0168] In the above embodiments of this application, a method for constructing a mask is specifically described, which is simple and easy to operate.

[0169] In a twentieth aspect, this application also provides an image processing apparatus having the function of implementing the method of aspect twenty-two or any possible implementation thereof. This function can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described function.

[0170] In a twentieth aspect, this application provides a pose estimation method applied to a simultaneous localization and mapping (SLAM) scenario. The method includes: a terminal acquiring a first event image and a first RGB image, wherein the first event image is temporally aligned with a first target image, and the first target image includes an RGB image or a depth image. The first event image is an image representing the motion trajectory of the target object when it moves within the detection range of a motion sensor. The terminal determines the integration time of the first event image. If the integration time is less than a first threshold, the terminal determines not to perform pose estimation using the first target image. The terminal performs pose estimation based on the first event image.

[0171] In this scheme, when the terminal determines that it is in a scenario where the RGB camera is unable to collect effective environmental information based on the integration time of the event image being less than a threshold, the terminal decides not to perform pose estimation using a low-quality RGB image, in order to improve the accuracy of pose estimation.

[0172] Optionally, in one possible implementation, the method further includes: determining the acquisition time of the first event image and the acquisition time of the first target image; and determining that the first event image and the first target image are time-aligned based on the time difference between the acquisition time of the first target image and the acquisition time of the first event image being less than a second threshold. The second threshold can be determined based on the accuracy of SLAM and the frequency at which the RGB camera acquires RGB images; for example, the second threshold can be 5 milliseconds or 10 milliseconds.

[0173] Optionally, in one possible implementation, acquiring the first event image includes: acquiring N consecutive DVS events; integrating the N consecutive DVS events into a first event image; the method further includes: determining the acquisition time of the first event image based on the acquisition time of the N consecutive DVS events.

[0174] Optionally, in one possible implementation, determining the integration time of the first event image includes: determining N consecutive DVS events for integration into the first event image; and determining the integration time of the first event image based on the acquisition times of the first and last DVS events among the N consecutive DVS events. Since the first event image is obtained by integrating N consecutive DVS events, the terminal can determine the acquisition time of the first event image based on the acquisition times corresponding to the N consecutive DVS events, that is, determine the acquisition time of the first event image as the time period from the acquisition of the first DVS event to the acquisition of the last DVS event among the N consecutive DVS events.

[0175] Optionally, in one possible implementation, the method further includes: acquiring a second event image, the second event image being an image representing the motion trajectory of the target object when it moves within the detection range of a motion sensor. The time period for which the motion sensor detects and acquires the first event image is different from the time period for which the motion sensor detects and acquires the second event image. If no RGB image exists that is time-aligned with the second event image, it is determined that the second event image does not have an RGB image for jointly performing pose estimation; pose estimation is then performed based on the second event image.

[0176] Optionally, in one possible implementation, before determining the pose based on the second event image, the method further includes: if it is determined that the second event image has time-aligned inertial measurement unit (IMU) data, then determining the pose based on the second event image and the corresponding IMU data; if it is determined that the second event image does not have time-aligned IMU data, then determining the pose only based on the second event image.

[0177] Optionally, in one possible implementation, the method further includes: acquiring a second target image, the second target image including an RGB image or a depth image; if there is no event image that is temporally aligned with the second target image, then determining that the second target image does not have an event image for jointly performing pose estimation; and determining a pose based on the second target image.

[0178] Optionally, in one possible implementation, the method further includes: performing loop closure detection based on the first event image and a dictionary, wherein the dictionary is a dictionary constructed based on the event image. That is, before performing loop closure detection, the terminal can pre-construct a dictionary based on the event image so that loop closure detection can be performed based on the dictionary during the loop closure detection process.

[0179] Optionally, in one possible implementation, the method further includes: acquiring multiple event images, which are event images used for training, and which may be event images captured by the terminal in different scenarios; acquiring visual features of the multiple event images, which may include features such as image texture, pattern, or grayscale statistics; clustering the visual features using a clustering algorithm to obtain clustered visual features, each clustered visual feature having a corresponding descriptor; and grouping similar visual features into one category to facilitate subsequent visual feature matching. Finally, constructing the dictionary based on the clustered visual features.

[0180] Optionally, in one possible implementation, performing loop closure detection based on the first event image and the dictionary includes: determining a descriptor for the first event image; determining visual features corresponding to the descriptor of the first event image in the dictionary; determining a bag-of-words vector corresponding to the first event image based on the visual features; and determining the similarity between the bag-of-words vector corresponding to the first event image and the bag-of-words vectors of other event images to determine the event image matched by the first event image.

[0181] In a twentieth aspect, this application provides a keyframe selection method, comprising: acquiring an event image; determining first information of the event image, the first information including events and / or features in the event image; and determining the event image as a keyframe if, based on the first information, the event image at least satisfies a first condition, wherein the first condition is related to the number of events and / or the number of features.

[0182] In this solution, by determining information such as the number of events, event distribution, number of features, and / or feature distribution in the event image, it is possible to determine whether the current event image is a keyframe. This enables rapid selection of keyframes with a small algorithm size, and can meet the needs of rapid keyframe selection in scenarios such as video analysis, video encoding / decoding, or security monitoring.

[0183] Optionally, in one possible implementation, the first condition includes one or more of the following: the number of events in the event image is greater than a first threshold, the number of valid event regions in the event image is greater than a second threshold, the number of features in the event image is greater than a third threshold, and the number of valid feature regions in the event image is greater than a fourth threshold.

[0184] Optionally, in one possible implementation, the method further includes: acquiring a depth image that is time-aligned with the event image; if it is determined based on the first information that the event image at least satisfies a first condition, then the event image and the depth image are determined to be keyframes.

[0185] Optionally, in one possible implementation, the method further includes: acquiring an RGB image that is time-aligned with the event image; acquiring the number of features and / or the effective feature region of the RGB image; if, based on the first information, it is determined that the event image at least satisfies a first condition, and the number of features of the RGB image is greater than a fifth threshold and / or the number of effective feature regions of the RGB image is greater than a sixth threshold, then the event image and the RGB image are determined to be keyframes.

[0186] Optionally, in one possible implementation, determining the event image as a keyframe if the event image at least satisfies a first condition based on the first information includes: determining second information of the event image if the event image at least satisfies the first condition based on the first information, wherein the second information includes motion features and / or pose features in the event image; and determining the event image as a keyframe if the event image at least satisfies a second condition based on the second information, wherein the second condition is related to the amount of motion change and / or pose change.

[0187] Optionally, in one possible implementation, the method further includes: determining the sharpness and / or brightness consistency index of the event image; if, based on the second information, it is determined that the event image at least satisfies the second condition, and the sharpness of the event image is greater than a sharpness threshold and / or the brightness consistency index of the event image is greater than a preset index threshold, then the event image is determined to be a keyframe.

[0188] Optionally, in one possible implementation, determining the brightness consistency index of the event image includes: if the pixels in the event image represent the polarity of light intensity change, then calculating the absolute value of the difference between the number of events in the event image and the number of events in adjacent keyframes, and dividing the absolute value by the number of pixels in the event image to obtain the brightness consistency index of the event image; if the pixels in the event image represent light intensity, then calculating the difference between the event image and adjacent keyframes pixel by pixel, calculating the absolute value of the difference, summing the absolute values ​​corresponding to each group of pixels, and dividing the summation result by the number of pixels to obtain the brightness consistency index of the event image.

[0189] Optionally, in one possible implementation, the method further includes: acquiring an RGB image that is time-aligned with the event image; determining the sharpness and / or brightness consistency index of the RGB image; if, based on the second information, it is determined that the event image at least satisfies the second condition, and the sharpness of the RGB image is greater than a sharpness threshold and / or the brightness consistency index of the RGB image is greater than a preset index threshold, then the event image and the RGB image are determined to be keyframes.

[0190] Optionally, in one possible implementation, the second condition includes one or more of the following: the distance between the event image and the previous keyframe exceeds a preset distance value, the rotation angle between the event image and the previous keyframe exceeds a preset angle value, and the distance between the event image and the previous keyframe exceeds a preset distance value and the rotation angle between the event image and the previous keyframe exceeds a preset angle value.

[0191] In a twentieth aspect, this application provides a pose estimation method, comprising: acquiring a first event image and a target image corresponding to the first event image, wherein the first event image and the target image capture the same environmental information, and the target image includes a depth image or an RGB image; determining a first motion region in the first event image; determining a corresponding second motion region in the image based on the first motion region; and performing pose estimation based on the second motion region in the image.

[0192] In this scheme, dynamic regions in the scene are captured by event images, and pose determination is performed based on these dynamic regions, thereby enabling accurate determination of pose information.

[0193] Optionally, in one possible implementation, determining the first motion region in the first event image includes: if the dynamic vision sensor (DVS) that acquired the first event image is stationary, then acquiring pixels in the first event image that have an event response; and determining the first motion region based on the pixels that have an event response.

[0194] Optionally, in one possible implementation, determining the first motion region based on the event-responsive pixels includes: determining a contour formed by the event-responsive pixels in the first event image; if the area enclosed by the contour is greater than a first threshold, then determining the region enclosed by the contour as the first motion region.

[0195] Optionally, in one possible implementation, determining the first motion region in the first event image includes: if the DVS of the first event image is in motion, then acquiring a second event image, wherein the second event image is the event image of the previous frame of the first event image; calculating the displacement magnitude and displacement direction of a pixel in the first event image relative to the second event image; if the displacement direction of a pixel in the first event image is not the same as the displacement direction of surrounding pixels, or if the difference between the displacement magnitude of a pixel in the first event image and the displacement magnitude of surrounding pixels is greater than a second threshold, then determining that the pixel belongs to the first motion region.

[0196] Optionally, in one possible implementation, the method further includes: determining a corresponding static region in the image based on the first motion region; and determining a pose based on the static region in the image.

[0197] In a twentieth aspect, this application also provides a data processing apparatus that has the function of implementing the method of aspect twenty-five or any possible implementation thereof, or the function of implementing the method of aspect twenty-six or any possible implementation thereof, or the function of implementing the method of aspect twenty-seven or any possible implementation thereof. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described function.

[0198] In a twentieth aspect, embodiments of this application provide an apparatus comprising: a processor and a memory, wherein the processor and the memory are interconnected via a circuit, and the processor invokes program code in the memory to perform processing-related functions in the methods shown in any one of the first to twenty-seventh aspects above. Optionally, the apparatus may be a chip.

[0199] In a thirtieth aspect, this application provides an electronic device comprising: a display module, a processing module, and a storage module.

[0200] The display module is used to display the graphical user interface of the application stored in the storage module, which may be any of the graphical user interfaces described above.

[0201] In a thirty-first aspect, embodiments of this application provide an apparatus, which may also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to perform processing-related functions as described in any of the optional embodiments of the first to twenty-seventh aspects above.

[0202] In a thirty-second aspect, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method in any of the optional embodiments of the first to twenty-seventh aspects described above.

[0203] In a thirty-third aspect, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the method in any of the optional embodiments of the first to twenty-seventh aspects described above. Attached Figure Description

[0204] Figure 1A A schematic diagram of a system architecture is provided for this application;

[0205] Figure 1B A schematic diagram of the structure of an electronic device provided in this application;

[0206] Figure 2 Another system architecture diagram provided for this application;

[0207] Figure 3-a A diagram illustrating the relationship between the amount of data read and the time in an event-stream-based asynchronous read mode;

[0208] Figure 3-b This diagram illustrates the relationship between the amount of data read and the time in a frame-scan-based synchronous read mode.

[0209] Figure 4-a A block diagram of a vision sensor provided in this application;

[0210] Figure 4-b A block diagram of another vision sensor provided in this application;

[0211] Figure 5 This is a schematic diagram illustrating the principles of the frame-scan-based synchronous reading mode and the event-stream-based asynchronous reading mode according to embodiments of this application;

[0212] Figure 6-a This is a schematic diagram illustrating the operation of a visual sensor according to an embodiment of this application in a frame-scan-based readout mode;

[0213] Figure 6-b This is a schematic diagram illustrating the operation of a visual sensor according to an embodiment of this application in an event stream-based reading mode;

[0214] Figure 6-c This is a schematic diagram illustrating the operation of a visual sensor according to an embodiment of this application in an event stream-based reading mode;

[0215] Figure 6-d This is a schematic diagram illustrating the operation of a visual sensor according to an embodiment of this application in a frame-scan-based readout mode;

[0216] Figure 7 A flowchart of a method for operating a vision sensor chip according to a possible embodiment of this application;

[0217] Figure 8 A block diagram of a control circuit provided in this application;

[0218] Figure 9 A block diagram of an electronic device provided in this application;

[0219] Figure 10 A schematic diagram illustrating the change in data volume over time between a single data reading mode and an adaptive switching reading mode according to possible embodiments of this application;

[0220] Figure 11 A schematic diagram of a pixel circuit provided in this application;

[0221] Figure 11-a This is a schematic diagram illustrating the representation of events using light intensity information and the representation of events using polarity information.

[0222] Figure 12-a This is a schematic diagram of a data format control unit in the reading circuit of this application;

[0223] Figure 12-b This is a schematic diagram of another structure of the data format control unit in the reading circuit of this application;

[0224] Figure 13 A block diagram of another control circuit provided in this application;

[0225] Figure 14 A block diagram of another control circuit provided in this application;

[0226] Figure 15 A block diagram of another control circuit provided in this application;

[0227] Figure 16 A block diagram of another control circuit provided in this application;

[0228] Figure 17 A block diagram of another control circuit provided in this application;

[0229] Figure 18 A schematic diagram illustrating the difference between a single event representation and an adaptive event representation provided in this application;

[0230] Figure 19 A block diagram of another electronic device provided in this application;

[0231] Figure 20 A flowchart of a method for operating a vision sensor chip according to a possible embodiment of this application;

[0232] Figure 21 A schematic diagram of another pixel circuit provided in this application;

[0233] Figure 22 A flowchart illustrating an encoding method provided in this application;

[0234] Figure 23 A block diagram of another vision sensor provided in this application;

[0235] Figure 24 This is a schematic diagram illustrating the division of a pixel array into regions.

[0236] Figure 25 A block diagram of another control circuit provided in this application;

[0237] Figure 26 A block diagram of another electronic device provided in this application;

[0238] Figure 27 This is a schematic diagram of a binary data stream;

[0239] Figure 28 A flowchart of a method for operating a vision sensor chip according to a possible embodiment of this application;

[0240] Figure 29-a A block diagram of another vision sensor provided in this application;

[0241] Figure 29-b A block diagram of another vision sensor provided in this application;

[0242] Figure 29-c A block diagram of another vision sensor provided in this application;

[0243] Figure 30 A schematic diagram of another pixel circuit provided in this application;

[0244] Figure 31 A block diagram of a third coding unit provided in this application;

[0245] Figure 32 A flowchart illustrating another encoding method provided in this application;

[0246] Figure 33 A block diagram of another electronic device provided in this application;

[0247] Figure 34 A flowchart of a method for operating a vision sensor chip according to a possible embodiment of this application;

[0248] Figure 35 An event diagram provided for this application;

[0249] Figure 36 An event diagram at a specific moment provided for this application;

[0250] Figure 37 A schematic diagram of a motion area provided in this application;

[0251] Figure 38 A flowchart illustrating an image processing method provided in this application;

[0252] Figure 39 A flowchart illustrating another image processing method provided in this application;

[0253] Figure 40 A flowchart illustrating another image processing method provided in this application;

[0254] Figure 41 An event image diagram provided for this application;

[0255] Figure 42 A flowchart illustrating another image processing method provided in this application;

[0256] Figure 43 A flowchart illustrating another image processing method provided in this application;

[0257] Figure 44 A flowchart illustrating another image processing method provided in this application;

[0258] Figure 45 A flowchart illustrating an image processing method provided in this application;

[0259] Figure 46A Another event image illustration provided for this application;

[0260] Figure 46B Another event image illustration provided for this application;

[0261] Figure 47A Another event image illustration provided for this application;

[0262] Figure 47B Another event image illustration provided for this application;

[0263] Figure 48 A flowchart illustrating another image processing method provided in this application;

[0264] Figure 49 A flowchart illustrating another image processing method provided in this application;

[0265] Figure 50 Another event image illustration provided for this application;

[0266] Figure 51 A schematic diagram of a reconstructed image provided in this application;

[0267] Figure 52 A flowchart illustrating an image processing method provided in this application;

[0268] Figure 53 A schematic diagram illustrating a method for fitting a motion trajectory provided in this application;

[0269] Figure 54 A schematic diagram illustrating a method for determining the focus point provided in this application;

[0270] Figure 55 A schematic diagram illustrating a method for determining a prediction center provided in this application;

[0271] Figure 56 A flowchart illustrating another image processing method provided in this application;

[0272] Figure 57 A schematic diagram of the shooting range provided in this application;

[0273] Figure 58 A schematic diagram of a predicted region provided in this application;

[0274] Figure 59 A schematic diagram of the focusing area provided in this application;

[0275] Figure 60 A flowchart illustrating another image processing method provided in this application;

[0276] Figure 61 A schematic diagram of an image enhancement method provided in this application;

[0277] Figure 62 A flowchart illustrating another image processing method provided in this application;

[0278] Figure 63 A flowchart illustrating another image processing method provided in this application;

[0279] Figure 64 This is a schematic diagram illustrating one application scenario of this application;

[0280] Figure 65 This is a schematic diagram illustrating another application scenario of this application;

[0281] Figure 66 A schematic diagram of a GUI provided in this application;

[0282] Figure 67 A schematic diagram illustrating another GUI provided in this application;

[0283] Figure 68 A schematic diagram illustrating another GUI provided in this application;

[0284] Figure 69A A schematic diagram illustrating another GUI provided in this application;

[0285] Figure 69B A schematic diagram illustrating another GUI provided in this application;

[0286] Figure 69C A schematic diagram illustrating another GUI provided in this application;

[0287] Figure 70 A schematic diagram illustrating another GUI provided in this application;

[0288] Figure 71 A schematic diagram illustrating another GUI provided in this application;

[0289] Figure 72A A schematic diagram illustrating another GUI provided in this application;

[0290] Figure 72B A schematic diagram illustrating another GUI provided in this application;

[0291] Figure 73 A flowchart illustrating another image processing method provided in this application;

[0292] Figure 74 This application provides a schematic diagram of an RGB image with low jitter.

[0293] Figure 75 This application provides a schematic diagram of an RGB image with a high degree of jitter.

[0294] Figure 76 A schematic diagram of an RGB image in a high-contrast scene provided in this application;

[0295] Figure 77 Another event image illustration provided for this application;

[0296] Figure 78 A schematic diagram of an RGB image provided in this application;

[0297] Figure 79 Another RGB image illustration provided for this application;

[0298] Figure 80 Another GUI diagram provided for this application;

[0299] Figure 81 A schematic diagram illustrating the relationship between the photosensitive unit and pixel value provided in this application;

[0300] Figure 82 A flowchart illustrating the image processing method provided in this application;

[0301] Figure 83 A schematic diagram of the event flow provided in this application;

[0302] Figure 84 A schematic diagram showing how multiple shooting scenes provided in this application are exposed and superimposed to obtain a blurred image;

[0303] Figure 85 An illustration of the mask provided in this application;

[0304] Figure 86A schematic diagram of the construction mask provided in this application;

[0305] Figure 87 An image showing the result of removing moving objects from image I to obtain image I′, as provided in this application.

[0306] Figure 88 A schematic diagram of a process for obtaining image I′ from image I by removing moving objects from image I, as provided in this application;

[0307] Figure 89 A schematic diagram illustrating a relatively small movement of an object during the photographing process provided in this application;

[0308] Figure 90 A schematic diagram illustrating the trigger camera capturing a third RGB image provided in this application;

[0309] Figure 91 Image B, captured by a camera based on motion mutation triggering, is provided in this application. k A schematic diagram of an image I actively captured by a user within a certain exposure time;

[0310] Figure 92 This is a schematic diagram of a process for obtaining a second RGB image of a non-moving object based on a first RGB image frame and an event stream E, as provided in this application.

[0311] Figure 93 This is a schematic diagram of a process for obtaining a second RGB image of a non-moving object based on a first RGB image, a third RGB image, and an event stream E, as provided in this application.

[0312] Figure 94A Another GUI diagram provided for this application;

[0313] Figure 94B Another GUI diagram provided for this application;

[0314] Figure 95 A comparative illustration of scenes captured by a traditional camera and a DVS provided for this application.

[0315] Figure 96 A comparative illustration of scenes captured by a conventional camera and a DVS provided in this application;

[0316] Figure 97 A schematic diagram of an outdoor navigation system using DVS provided in this application;

[0317] Figure 98a A station navigation diagram using DVS provided in this application;

[0318] Figure 98bA schematic diagram of a tourist attraction navigation system using DVS provided in this application;

[0319] Figure 99 A shopping mall navigation diagram using DVS provided in this application;

[0320] Figure 100 A schematic diagram of a SLAM execution process provided in this application;

[0321] Figure 101 A flowchart illustrating a pose estimation method 10100 provided in this application;

[0322] Figure 102 A schematic diagram illustrating how to integrate DVS events into an event image, as provided in this application;

[0323] Figure 103 A flowchart illustrating a keyframe selection method 10300 provided in this application;

[0324] Figure 104 This application provides a schematic diagram of region division for an event image;

[0325] Figure 105 A flowchart illustrating a keyframe selection method 10500 provided in this application;

[0326] Figure 106 A flowchart illustrating a pose estimation method 1060 provided in this application;

[0327] Figure 107 A schematic diagram illustrating a process for performing pose estimation based on a static region of an image, as provided in this application;

[0328] Figure 108a A schematic diagram illustrating a process for performing pose estimation based on a motion region of an image, as provided in this application;

[0329] Figure 108b A schematic diagram illustrating a process for performing pose estimation based on the overall region of an image, as provided in this application;

[0330] Figure 109 This application provides a schematic diagram of an AR / VR glasses structure;

[0331] Figure 110 A schematic diagram of a gaze perception structure provided in this application;

[0332] Figure 111 A schematic diagram of a network architecture provided for this application;

[0333] Figure 112 A schematic diagram of the structure of an image processing device provided in this application;

[0334] Figure 113 A schematic diagram of another image processing apparatus provided in this application;

[0335] Figure 114 A schematic diagram of another image processing apparatus provided in this application;

[0336] Figure 115 A schematic diagram of another image processing apparatus provided in this application;

[0337] Figure 116 A schematic diagram of another image processing apparatus provided in this application;

[0338] Figure 117 A schematic diagram of another image processing apparatus provided in this application;

[0339] Figure 118 A schematic diagram of another image processing apparatus provided in this application;

[0340] Figure 119 Another schematic diagram of the data processing apparatus provided in this application;

[0341] Figure 120 Another schematic diagram of the data processing apparatus provided in this application;

[0342] Figure 121 Another schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0343] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0344] The electronic devices, system architecture, and methodologies provided in this application will be described in detail from different perspectives below.

[0345] I. Electronic Equipment

[0346] The method provided in this application can be applied to various electronic devices, or in other words, the method provided in this application can be executed by electronic devices. These electronic devices can be applied to shooting scenarios, such as photography, security, autonomous driving, drone shooting, etc.

[0347] The electronic devices described in this application may include, but are not limited to: smartphones, televisions, tablets, wristbands, head-mounted displays (HMDs), augmented reality (AR) devices, mixed reality (MR) devices, cellular phones, smartphones, personal digital assistants (PDAs), in-vehicle electronic devices, laptop computers, personal computers (PCs), monitoring equipment, robots, in-vehicle terminals, and autonomous vehicles. Of course, the specific form of the electronic device is not limited in the following embodiments.

[0348] For example, the architecture of the electronic device application provided in this application is as follows: Figure 1A As shown.

[0349] Among them, electronic devices, such as Figure 1A The vehicles, mobile phones, AR / VR glasses, security monitoring equipment, cameras, or other smart home terminals described herein can access a cloud platform via wired or wireless networks. The cloud platform contains servers, which can be centralized or distributed. Electronic devices can communicate with the cloud platform's servers via wired or wireless networks to transmit data. For example, after collecting data, electronic devices can save or back it up on the cloud platform to prevent data loss.

[0350] Electronic devices can access access points or base stations to achieve wireless or wired access to the cloud platform. For example, the access point can be a base station, and the electronic device has a SIM card installed. This SIM card is used for network authentication with the operator, thereby accessing the wireless network. Alternatively, the access point can include a router, and the electronic device connects to the router via a 2.4GHz or 5GHz wireless network, thereby accessing the cloud platform through the router.

[0351] Furthermore, electronic devices can process data independently or collaboratively with the cloud, depending on the specific application scenario. For example, a Data View Monitor (DVS) can be installed in the electronic device. The DVS can work in conjunction with a camera or other sensors within the device, or it can work independently. The processor within the DVS or the electronic device can process the data collected by the DVS or other sensors. Alternatively, it can collaborate with cloud devices to process the data collected by the DVS or other sensors.

[0352] The following is an exemplary description of the specific structure of the electronic device.

[0353] For example, see Figure 1B The structure of the electronic device provided in this application will be illustrated below using a specific example.

[0354] It should be noted that the electronic device provided in this application may include, but is not limited to, those that are more advanced than those that are designed for use in electronic devices. Figure 1B More or fewer parts, Figure 1B The electronic device shown is merely an illustrative example. Those skilled in the art can add or remove components in the electronic device as needed, and this application does not limit this.

[0355] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a proximity sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, and an image sensor 180N. The image sensor 180N may include a separate color sensor 1801N and a separate motion sensor 1802N, or it may include a photosensitive unit (which can be called a color sensor pixel) of the color sensor. Figure 1B (not shown in the image) and the photosensitive unit of the motion sensor (which may be called a motion sensor pixel, Figure 1B (Not shown in the image).

[0356] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0357] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0358] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0359] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0360] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0361] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.

[0362] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0363] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0364] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.

[0365] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the electronic device 100 to capture images. The processor 110 and the display screen 194 communicate via the DSI interface to enable the electronic device 100 to display images.

[0366] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0367] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0368] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0369] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0370] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0371] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0372] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0373] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0374] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0375] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0376] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with networks and other devices through wireless communication technology. The wireless communication technologies mentioned may include, but are not limited to: 5th-Generation (5G) systems, Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), Bluetooth, the Global Navigation Satellite System (GNSS), Wireless Fidelity (WiFi), Near Field Communication (NFC), FM (Frequency Modulation Broadcasting), Zigbee, Radio Frequency Identification (RFID), and / or Infrared (IR) technologies. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS), etc.

[0377] In some embodiments, the electronic device 100 may also include a wired communication module ( Figure 1B (not shown in the image), or, the mobile communication module 150 or wireless communication module 160 here can be replaced with a wired communication module (…). Figure 1B (Not shown in the image), this wired communication module enables electronic devices to communicate with other devices via a wired network. This wired network may include, but is not limited to, one or more of the following: optical transport network (OTN), synchronous digital hierarchy (SDH), passive optical network (PON), Ethernet, or flex Ethernet (FlexE), etc.

[0378] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0379] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0380] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0381] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0382] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB camera (or RGB sensor) formats such as O, YUV, etc. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0383] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0384] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0385] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0386] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0387] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.

[0388] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0389] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0390] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.

[0391] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.

[0392] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0393] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0394] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.

[0395] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 about three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.

[0396] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0397] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover using the magnetic sensor 180D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.

[0398] The 180E accelerometer can detect the magnitude of acceleration of electronic device 100 in various directions (typically three axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic devices and applied to applications such as screen orientation switching and pedometers.

[0399] A distance sensor 180F is used to measure distance. Electronic device 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, electronic device 100 can utilize the distance sensor 180F to measure distance for rapid focusing.

[0400] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device 100 emits infrared light outward through the LED. The electronic device 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 may use the proximity sensor 180G to detect when a user holds the electronic device 100 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 180G can also be used in holster mode and pocket mode for automatic unlocking and locking of the screen.

[0401] The ambient light sensor 180L is used to sense the brightness of ambient light. The electronic device 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touches.

[0402] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.

[0403] Temperature sensor 180J is used to detect temperature. In some embodiments, electronic device 100 uses the temperature detected by temperature sensor 180J to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, electronic device 100 performs thermal protection by reducing the performance of a processor located near temperature sensor 180J to reduce power consumption. In other embodiments, when the temperature is below another threshold, electronic device 100 heats battery 142 to prevent abnormal shutdown of electronic device 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, electronic device 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.

[0404] Touch sensor 180K, also known as a "touch device," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touchscreen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.

[0405] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 180M can also be incorporated into headphones to form bone conduction headphones. The audio module 170 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 180M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 180M to realize heart rate detection functionality.

[0406] An image sensor, also known as a photosensitive device or photosensitive element, is a device that converts optical images into electronic signals. It is widely used in digital cameras and other electro-optical devices. An image sensor utilizes the photoelectric conversion function of photoelectric devices to convert the light image on the photosensitive surface into an electrical signal proportional to the light image. Compared to photosensitive elements with "point" light sources such as photodiodes and phototransistors, an image sensor is a functional device that divides the light image on its light-receiving surface into many small units (i.e., pixels) and converts them into usable electrical signals. Each small unit corresponds to a photosensitive unit within the image sensor, also called a sensor pixel. Image sensors are divided into photoconductive image tubes and solid-state image sensors. Compared to photoconductive image tubes, solid-state image sensors have advantages such as small size, light weight, high integration, high resolution, low power consumption, long lifespan, and low price. Based on the different components, they can be divided into two main categories: charge-coupled devices (CCD) and complementary metal-oxide-semiconductor (CMOS). Based on the different types of optical images captured, they can be divided into two main categories: color sensors 1801N and motion sensors 1802N.

[0407] Specifically, the 1801N color sensor, including a traditional RGB image sensor, can be used to detect objects within the camera's field of view. Each photosensitive unit corresponds to a pixel in the image sensor. Since the photosensitive unit can only sense light intensity and cannot capture color information, a color filter must be placed over the photosensitive unit. Different sensor manufacturers have different solutions for how to place the color filter. The most common approach is to use RGB red, green, and blue filters in a 1:2:1 ratio, with four pixels forming one color pixel (i.e., one pixel is covered by a red and one by a blue filter, and the remaining two pixels are covered by a green filter). This ratio is chosen because the human eye is more sensitive to green. After receiving light, the photosensitive unit generates a corresponding current, the magnitude of which corresponds to the light intensity. Therefore, the electrical signal directly output by the photosensitive unit is analog. This analog electrical signal is then converted into a digital signal, and finally, all the digital signals are output as a digital image matrix to a dedicated DSP processing chip for processing. This traditional color sensor outputs a full-frame image of the captured area in frame format.

[0408] Specifically, the motion sensor 1802N can include various types of vision sensors, such as frame-based motion detection vision sensors (MDVS) and event-based motion detection vision sensors. It can be used to detect moving objects within the range captured by the camera, and to acquire the motion contours or trajectories of these objects.

[0409] In one possible scenario, the motion sensor 1802N may include a motion detection (MD) visual sensor, a type of visual sensor that detects motion information derived from the relative motion between the camera and the target. This motion information could be camera motion, target motion, or both of them moving. Motion detection visual sensors include frame-based motion detection and event-based motion detection. Frame-based motion detection visual sensors require exposure integration and obtain motion information through frame differences. Event-based motion detection visual sensors do not require integration and obtain motion information through asynchronous event detection.

[0410] In one possible scenario, the motion sensor 1802N may include a motion detection vision sensor (MDVS), a dynamic vision sensor (DVS), an active pixel sensor (APS), an infrared sensor, a laser sensor, or an inertial measurement unit (IMU). Specifically, the DVS may include sensors such as DAVIS (Dynamic and Active-pixel Vision Sensor), ATIS (Asynchronous Time-based Image Sensor), or a CeleX sensor. The DVS borrows characteristics from biological vision, with each pixel simulating a neuron, independently responding to relative changes in illumination intensity (hereinafter referred to as "light intensity"). For example, if the motion sensor is a DVS, when the relative change in light intensity exceeds a threshold, the pixel will output an event signal, including the pixel's position, timestamp, and characteristic information of the light intensity. It should be understood that in the following embodiments of this application, the motion information, dynamic data, or dynamic images mentioned can all be acquired by the motion sensor.

[0411] For example, the motion sensor 1802N may include an Inertial Measurement Unit (IMU), a device that measures the three-axis angular velocity and acceleration of an object. An IMU typically consists of three single-axis accelerometers and three single-axis gyroscopes, which measure the object's acceleration signal and angular velocity signal relative to the navigation coordinate system, respectively, and use these to calculate the object's attitude. For example, the aforementioned IMU may specifically include the aforementioned gyroscope sensor 180B and accelerometer sensor 180E. The advantage of an IMU is its high data acquisition frequency. IMUs can generally achieve data acquisition frequencies above 100Hz, and consumer-grade IMUs can capture data up to 1600Hz. In a short time, an IMU can provide high-precision measurement results.

[0412] For example, the motion sensor 1802N may include an active pixel sensor (APS). For instance, it captures RGB images at a high frequency >100Hz and subtracts the values ​​from adjacent frames to obtain the change value. If this change value is greater than a threshold (>0), it is set to 1; if it is not greater than the threshold (=0), it is set to 0. The final data is similar to the data obtained by the DVS, thus completing the capture of images of moving objects.

[0413] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0414] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0415] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0416] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0417] II. System Architecture

[0418] In the process of capturing, reading or saving images, electronic devices involve changes between multiple components. The following application provides a detailed description of scenarios such as data acquisition, data encoding and decoding, image enhancement, image reconstruction or application.

[0419] For example, taking an image acquisition and processing scenario as an example, such as... Figure 2 The processing flow of an electronic device is illustrated below.

[0420] Data Acquisition: Data can be acquired using neuromorphic cameras, RGB cameras, or a combination thereof. Neuromorphic cameras may include bio-vision sensors that simulate the biological retina using integrated circuits, with each pixel simulating a biological neuron, expressing changes in light intensity as events. Various types of bio-inspired vision sensors have emerged, all sharing the common feature of independently and asynchronously monitoring changes in light intensity with pixel arrays, outputting these changes as event signals, such as the aforementioned motion sensors DVS or DAVIS. RGB cameras convert analog signals into digital signals, which are then stored in a storage medium. Data can also be acquired by combining neuromorphic cameras and RGB cameras; for example, data acquired by both can be projected onto the same canvas. The value of each pixel can be determined based on the feedback values ​​from the neuromorphic camera and / or the RGB camera, or the value of each pixel can include values ​​from both the neuromorphic camera and RGB as independent channels. Neuromorphic cameras, RGB cameras, or a combination thereof can convert light signals into electrical signals, resulting in a data stream in frames or an event stream in events. In this application, images acquired by an RGB camera are referred to as RGB images, and images acquired by a neuromorphic camera are referred to as event images.

[0421] Data encoding and decoding: This includes data encoding and data decoding. Data encoding can involve encoding the acquired data after data collection and saving the encoded data to a storage medium. Data decoding can involve reading data from the storage medium and decoding it into data that can be used for subsequent identification, detection, etc. Furthermore, the data collection method can be adjusted based on the data encoding and decoding method to achieve more efficient data collection and encoding / decoding. Data encoding and decoding can be categorized in various ways, including encoding and decoding based on neuromorphic cameras, encoding and decoding based on neuromorphic cameras and RGB cameras, or encoding and decoding based solely on RGB cameras. Specifically, during encoding, data acquired by neuromorphic cameras, RGB cameras, or a combination thereof can be encoded and stored in a storage medium according to a specific format. During decoding, the data stored in the storage medium can be decoded into data that can be used subsequently. For example, on the first day, a user can acquire video or image data using a neuromorphic camera, RGB camera, or a combination thereof, encode the video or image data, and store it in a storage medium. The next day, the data can be read from the storage medium and decoded to obtain a playable video or image.

[0422] Image optimization: After the aforementioned neuromorphic camera or RGB camera acquires images, the acquired images are read and then optimized through enhancement or reconstruction to facilitate subsequent processing based on the optimized images. For example, image enhancement and reconstruction may include image reconstruction or motion compensation. Motion compensation, for example, uses motion parameters of moving objects acquired by a DVS to compensate for moving objects in an event image or RGB image, thereby making the obtained event image or RGB image clearer. Image reconstruction, for example, uses images acquired by a neuromorphic vision camera to reconstruct RGB images, thus obtaining clear RGB images even in moving scenes using data acquired by the DVS.

[0423] Application scenarios: After obtaining the optimized RGB image or event image through image optimization, the optimized RGB image or event image can be used for further applications. Of course, it can also be used for further applications of the acquired RGB image or event image. The specific application can be adjusted according to the actual application scenario.

[0424] Specifically, application scenarios can include: motion photography enhancement, DVS and RGB image fusion, detection and recognition, simultaneous localization and mapping (SLAM), eye tracking, keyframe selection, and pose estimation. For example, motion photography enhancement involves enhancing images captured in scenes with moving objects to obtain clearer images of those objects. DVS and RGB image fusion uses moving objects captured by a DVS image to enhance RGB images, compensating for moving objects or objects affected by high contrast ratios to obtain clearer RGB images. Detection and recognition involves performing target detection or recognition based on RGB images or event images. Eye tracking tracks the user's eye movements based on captured RGB images, event images, or optimized RGB images and event images to determine the user's gaze point, gaze direction, and other information. Keyframe selection combines information from a neuromorphic camera to select certain frames as keyframes from video data captured by an RGB camera.

[0425] Furthermore, in the following embodiments of this application, different sensors may need to be activated in different embodiments. For example, when acquiring data, when optimizing event images through motion compensation, a motion sensor may be activated, and optionally an IMU or gyroscope may also be activated. In the image reconstruction embodiment, a motion sensor may be activated to acquire event images, and then the event images may be combined for optimization. Alternatively, in the motion photography enhancement embodiment, a motion sensor and an RGB sensor may be activated. Therefore, in different embodiments, the corresponding sensors may be selected to be activated.

[0426] Specifically, the method provided in this application can be applied to an electronic device, which may include an RGB sensor and a motion sensor, etc. The RGB sensor is used to acquire images within the shooting range, and the motion sensor is used to acquire information generated when an object moves relative to the motion sensor within the detection range of the motion sensor. The method includes: selecting at least one from the RGB sensor and the motion sensor based on scene information, and acquiring data through the selected sensor. The scene information includes at least one of the following: the state information of the electronic device, the type of the application in the electronic device that requests image acquisition, or environmental information.

[0427] In one possible implementation, the aforementioned status information includes information such as the remaining battery power, remaining storage (or available storage) of the electronic device, or CPU load.

[0428] In one possible implementation, the aforementioned environmental information may include changes in light intensity within the shooting range of the color RGB sensor and the motion sensor, or information about moving objects within the shooting range. For example, the environmental information may include changes in light intensity within the shooting range of the RGB sensor or the DVS sensor, or the motion of objects within the shooting range, such as the object's speed and direction of movement, or abnormal motion of objects within the shooting range, such as sudden changes in the object's speed or direction.

[0429] The type of application requesting image acquisition in the aforementioned electronic device can be understood as follows: the electronic device carries an Android, Linux, or HarmonyOS system, and applications can run on this system. The applications running on this system can be divided into various types, such as photo-taking applications or object detection applications.

[0430] Typically, motion sensors are sensitive to changes in motion but not to static scenes. They respond to changes in motion by emitting events. Since static areas emit almost no events, their data only represents the light intensity information of the area where motion changes, not the complete light intensity information of the entire scene. RGB color cameras, on the other hand, excel at recording the complete color of natural scenes and reproducing the texture details within them.

[0431] Taking a mobile phone as an example, the default configuration is that the DVS camera (i.e., the DVS sensor) is off. When using the camera, depending on the type of application being called, such as a photo-taking app calling the camera, if the subject is in a high-speed motion state, both the DVS camera and the RGB camera (i.e., the RGB sensor) need to be turned on. If the app requesting the camera is an app for object detection or motion detection, and does not need to perform object photography and face recognition, then the DVS camera can be turned on, but the RGB camera can be turned off.

[0432] Optionally, the camera activation mode can be selected based on the current status of the device. For example, when the current battery level is below a certain threshold, the user can activate power-saving mode and take photos normally. Only the DVS camera can be turned on because although the DVS camera image is not clear, it has low power consumption and does not require high-definition imaging when detecting moving objects.

[0433] Optionally, the device can sense the surrounding environment to decide whether to switch camera modes. For example, the DVS camera can be activated in night scenes or when the device is in high-speed motion. In static scenes, the DVS camera can be left off.

[0434] Based on the above application types, environmental information, and device status, the camera activation mode is determined, and during operation, a decision can be made on whether to trigger camera mode switching, thereby activating different sensors in different scenarios, demonstrating strong adaptability.

[0435] This can be understood as having three startup modes: RGB camera only, DVS camera only, and both RGB and DVS cameras simultaneously. Furthermore, the reference factors for detection applications and environmental testing may differ depending on the product.

[0436] For example, security cameras with motion detection capabilities only store recordings when they detect moving objects, thus reducing storage space and extending hard drive storage time. Specifically, when DVS and RGB cameras are used in home or security applications, only the DVS camera is enabled by default for motion detection and analysis. When the DVS camera detects abnormal motion or behavior (such as sudden movement or a sudden change in direction), such as when someone approaches or there is a significant change in light intensity, the RGB camera is activated to record a full-scene texture image during that period, serving as monitoring evidence. After the abnormal motion ends, the system switches back to DVS operation, while the RGB camera remains in standby mode, significantly saving data volume and reducing the power consumption of the monitoring equipment.

[0437] The intermittent camera method described above leverages the low power consumption of DVS (Dynamic Sensor Detection). Furthermore, DVS, being event-based motion detection, offers faster response and higher accuracy compared to image-based motion detection, enabling continuous, 24 / 7 detection. This method achieves greater accuracy, lower power consumption, and reduced storage space consumption.

[0438] For example, when DVS and RGB cameras are used in vehicle assistance / autonomous driving, during driving, when encountering oncoming vehicles with high beams, direct sunlight, or entering / exiting tunnels, the RGB camera may not be able to capture effective scene information. In these situations, although the DVS cannot obtain texture information, it can obtain the general outline information of the scene, which is of great assistance to the driver's judgment. In addition, in foggy weather, the outline information captured by the DVS can also help judge road conditions. Therefore, the master-slave working state switch between the DVS and the RGB camera can be triggered in specific scenarios, such as when there are drastic changes in light intensity or in extreme weather conditions.

[0439] The same process applies when the DVS camera is used in AR / VR glasses. When using DVS for SLAM or eye tracking, the camera's activation mode can be determined based on the device status and the surrounding environment.

[0440] In the following embodiments of this application, the sensor is turned on when data collected by a certain sensor is used, and this will not be described in detail below.

[0441] The following section combines the different working modes mentioned above with the aforementioned... Figure 2 Different implementation methods provided in this application are described.

[0442] III. Method and Flow

[0443] The foregoing has provided an exemplary description of the electronic device and system architecture provided in this application. The following description, in conjunction with the foregoing, further details... Figure 1A-Figure 2 This application provides a detailed description of the method provided. Specifically, in conjunction with the foregoing... Figure 2 The architecture is described separately for each module. It should be understood that the methods and steps mentioned below in this application can be implemented individually or in combination in a single device, and can be adjusted according to the actual application scenario.

[0444] 1. Data acquisition and encoding / decoding

[0445] The following example illustrates the combination of data acquisition and data encoding / decoding processes.

[0446] In traditional technologies, visual sensors (i.e., the aforementioned motion sensors) generally employ either an event-stream-based asynchronous readout mode (hereinafter referred to as "event-stream-based readout mode" or "asynchronous readout mode") or a frame-scanning-based synchronous readout mode (hereinafter referred to as "frame-scanning-based readout mode" or "synchronous readout mode"). For a fully manufactured visual sensor, only one of these two modes can be used. Depending on the specific application scenario and motion state, the amount of signal data required to be read per unit time may differ significantly between the two readout modes, and consequently, the cost of outputting the readout data will also differ. Figure 3-a and Figure 3-b The diagrams illustrate the relationship between the amount of data read and the time in the event-stream-based asynchronous read mode and the frame-scan-based synchronous read mode, respectively.

[0447] On the one hand, biomimetic vision sensors, due to their motion sensitivity, and the fact that static areas in the environment typically do not generate light intensity change events (also referred to as "events" in this paper), almost all employ an asynchronous readout mode based on event streams. An event stream refers to events arranged in a certain order. The following uses DVS as an example to illustrate the asynchronous readout mode. According to the sampling principle of DVS, by comparing the current light intensity with the light intensity at the time of the last event, an event is generated and output when the change reaches a predetermined emission threshold C (hereinafter referred to as the predetermined threshold). That is, DVS will typically generate an event when the difference between the current light intensity and the light intensity at the time of the last event exceeds the predetermined threshold C, which can be described by formula 1-1:

[0448] |LL′|≥C(1-1)

[0449] Where L and L' represent the light intensity at the current moment and the light intensity at the time of the last event, respectively.

[0450] In asynchronous read mode, each event can be represented as<x,y,t,m> Let (x, y) represent the pixel location where the event occurred, t represent the time when the event occurred, and m represent the characteristic information of light intensity. Specifically, pixels in the pixel array circuit of the vision sensor measure the change in light intensity in the environment. If the measured change in light intensity exceeds a predetermined threshold, the pixel can output a data signal indicating the event. Therefore, in the asynchronous readout mode based on event stream, the pixels of the vision sensor are further distinguished into pixels that have generated light intensity change events and pixels that have not generated light intensity change events. Light intensity change events can be characterized by the coordinate information (x, y) of the pixel that generated the event, the characteristic information of the light intensity at that pixel, and the time t when the characteristic information of the light intensity was read. The coordinate information (x, y) can be used to uniquely identify pixels in the pixel array circuit. For example, x represents the row index of the pixel in the pixel array circuit, and y represents the column index of the pixel in the pixel array circuit. By identifying the coordinates and timestamp associated with the pixel, the spatiotemporal location of the light intensity change event can be uniquely determined, and all events can be arranged into an event stream in chronological order.

[0451] In some DVS sensors (such as DAVIS and ATIS sensors), m represents the trend of light intensity change, also known as polarity information, typically represented by 1-2 bits. The value can be ON or OFF, where ON indicates light intensity increase and OFF indicates light intensity decrease. Specifically, an ON pulse is generated when the light intensity increases and exceeds a predetermined threshold; an OFF pulse is generated when the light intensity decreases and exceeds a predetermined threshold (in this application, "+1" represents light intensity increase and "-1" represents light intensity decrease). In some DVS sensors, such as the CeleX sensor used for moving object monitoring, m represents absolute light intensity information, also known as light intensity information, typically represented by multiple bits, such as 8-12 bits.

[0452] In asynchronous readout mode, only the data signal at the pixel that generates the light intensity change event is read. Therefore, for bionic vision sensors, the event data to be read has sparse and asynchronous characteristics. For example... Figure 3-a As shown in curve 101, the vision sensor operates in an event-stream-based asynchronous readout mode. As the rate of light intensity change events occurring in the pixel array circuit changes, the amount of data that the vision sensor needs to read also changes over time.

[0453] On the other hand, traditional vision sensors, such as those used in mobile phone cameras and digital camcorders, typically employ a frame-scanning-based synchronous readout mode. This readout mode does not distinguish whether a pixel of the vision sensor experiences a light intensity change event. Regardless of whether a light intensity change event occurs at a particular pixel, the data signal generated by that pixel is read. During data signal reading, the vision sensor scans the pixel array circuit in a predetermined order, synchronously reading the characteristic information m indicating light intensity at each pixel (the characteristic information m of light intensity has been introduced above and will not be repeated here), and outputs it sequentially as the first frame of data, the second frame of data, and so on. Therefore, as... Figure 3-b As shown in curve 102, in synchronous reading mode, the amount of data read by the vision sensor in each frame is the same, and the amount of data remains constant over time. For example, assuming that 8 bits are used to represent the light intensity value of a pixel, and the total number of pixels in the vision sensor is 66, then the amount of data in one frame is 528 bits. Typically, frames are output at equal time intervals, such as at a rate of 30 frames per second, 60 frames per second, or 120 frames per second.

[0454] The applicant found that current vision sensors still have shortcomings, including at least the following aspects:

[0455] First, a single read mode cannot adapt to all scenarios and is not conducive to alleviating the pressure on data transmission and storage.

[0456] like Figure 3-a As shown in curve 101, the vision sensor operates in an event-stream-based asynchronous readout mode. The amount of data the vision sensor needs to read changes over time as the rate of light intensity changes in the pixel array circuitry changes. In static scenes, fewer light intensity change events occur, resulting in a lower total data volume required by the vision sensor. However, in dynamic scenes, such as during vigorous motion, a large number of light intensity change events occur, increasing the total data volume required by the vision sensor. In some scenarios, the sheer volume of light intensity change events can exceed the bandwidth limit, potentially leading to event loss or delayed readout. Figure 3-b As shown in curve 102, the vision sensor operates in a frame-based synchronous readout mode, requiring the representation of the pixel's state or intensity value within a single frame, regardless of whether the pixel has changed. This representation is costly when only a small number of pixels change.

[0457] The output and storage costs of the two modes can differ significantly depending on the application scenario and the state of motion. For example, when shooting a static scene, only a small number of pixels generate light intensity change events over a period of time. For instance, in a single scan, only three pixels in the pixel array circuit generate light intensity change events. In asynchronous readout mode, only the coordinate information (x, y), time information t, and light intensity change amount of these three pixels are needed to characterize the three light intensity change events. Assuming that in asynchronous readout mode, 4, 2, 2, and 2 bits are allocated for the coordinates, read timestamp, and light intensity change amount of a pixel, respectively, the total amount of data required to read in this mode is 30 bits. In synchronous readout mode, although only three pixels generate valid data signals indicating light intensity change events, the data signals output by all pixels in the entire array still need to be read to form a complete frame. Assuming that 8 bits are allocated for each pixel in synchronous readout mode, and the total number of pixels in the pixel array circuit is 66, the total amount of data required to read is 528 bits. Therefore, even with a large number of pixels in the pixel array circuit that have not generated any events, such a large number of bits still need to be allocated in synchronous read mode. This is uneconomical from a representation cost perspective and increases the pressure on data transmission and storage. Therefore, asynchronous read mode is more economical in this situation.

[0458] In another example, when there is intense movement or a sudden change in ambient light intensity, such as a large number of people moving around or lights being suddenly switched on or off, a large number of pixels in the vision sensor measure the light intensity change within a short period and generate data signals indicating the light intensity change event. Since the amount of data representing a single event in asynchronous readout mode is greater than that in synchronous readout mode, using asynchronous readout mode in this situation may require a significant increase in representation cost. Specifically, each row in the pixel array circuit may have multiple consecutive pixels generating light intensity change events. For each event, coordinate information (x, y), time information t, and light intensity characteristic information m need to be transmitted. The coordinate changes between these events often have a deviation of only one unit, and the readout time is also roughly the same. In this case, asynchronous readout mode incurs a high representation cost for coordinate and time information, leading to a surge in data volume. In synchronous readout mode, regardless of the number of light intensity change events generated in the pixel array circuit at any given time, each pixel outputs a data signal indicating only the amount of light intensity change, without needing to allocate bits for coordinate and time information for each pixel. Therefore, synchronous readout mode is more economical for scenarios with dense events.

[0459] Secondly, a single event representation method cannot adapt to all scenarios. Using light intensity information to represent events is not conducive to alleviating the pressure on data transmission and storage, while using polarity information to represent events affects the processing and analysis of events.

[0460] The preceding text introduced synchronous and asynchronous readout modes. All readout events need to be represented by light intensity feature information *m*, which includes polarity information and light intensity information. This paper refers to events represented by polarity information as polarity-formatted events and events represented by light intensity information as light intensity-formatted events. For a fully manufactured vision sensor, only one of these two event formats can be used: either polarity-formatted or light intensity-formatted events. The following section uses asynchronous readout mode as an example to illustrate the advantages and disadvantages of polarity-formatted and light intensity-formatted events.

[0461] In asynchronous reading mode, when polarity information is used to represent events, the polarity information p is typically represented by 1-2 bits, carrying relatively little information. It can only indicate whether the light intensity is increasing or decreasing. Therefore, using polarity information to represent events affects the processing and analysis of events. For example, events represented by polarity information are more difficult to reconstruct in images, and the accuracy for object recognition is also poor. When light intensity information is used to represent events, it is usually represented by multiple bits, such as 8-12 bits. Compared to polarity information, light intensity information can carry more information, which is beneficial for event processing and analysis, such as improving the quality of image reconstruction. However, due to the larger data volume, it takes longer to obtain the light intensity information representing the event. According to the DVS sampling principle, an event will be generated when the light intensity change of a pixel exceeds a predetermined threshold. Therefore, when there is large-area object movement or light intensity fluctuation in the scene (such as entering or exiting a tunnel, turning lights on or off in a room), the visual sensor will face the problem of a sudden increase in events. With a fixed preset maximum bandwidth of the visual sensor (hereinafter referred to as bandwidth), there may be situations where event data cannot be read. Currently, random discarding is usually used to handle this. While random data discarding ensures that the transmitted data volume does not exceed the bandwidth, it results in data loss. In certain special application scenarios (such as autonomous driving), randomly discarded data may be of high importance. In other words, when a large number of events are triggered, the data volume exceeds the bandwidth, and the light intensity format data cannot be completely output to the DVS, leading to the loss of some events. These lost events may hinder event processing and analysis; for example, they may cause ghosting and incomplete outlines during brightness reconstruction.

[0462] To address the aforementioned issues, this application provides a visual sensor that, based on statistical results of light intensity change events generated by a pixel array circuit, compares the data volume under two reading modes, thereby enabling switching to a reading mode suitable for the current application scenario and motion state. Furthermore, based on the statistical results of light intensity change events generated by the pixel array circuit, it compares the relationship between the data volume of events represented by light intensity information and bandwidth, thereby adjusting the event representation precision. Under the premise of meeting bandwidth limitations, it transmits all events in a suitable representation method, transmitting all events with the highest possible representation precision.

[0463] The following describes a vision sensor provided in an embodiment of this application.

[0464] Figure 4-a A block diagram of a vision sensor provided in this application is shown. The vision sensor can be implemented as a vision sensor chip and is capable of reading data signals indicating events in at least one of a frame-scan-based readout mode and an event-stream-based readout mode. Figure 4-aAs shown, the vision sensor 200 includes a pixel array circuit 210 and a readout circuit 220. The vision sensor is coupled to a control circuit 230. It should be understood that... Figure 4-a The visual sensor shown is for illustrative purposes only and does not imply any limitation on the scope of this application. Embodiments of this application can also be embodied in different sensor architectures. Furthermore, it should be understood that the visual sensor may also include other elements or entities for achieving image acquisition, image processing, image transmission, etc., which are not shown for ease of description, but do not imply that embodiments of this application do not include these elements or entities.

[0465] The pixel array circuit 210 may include one or more pixel arrays, and each pixel array includes multiple pixels, each pixel having location information for unique identification, such as coordinates (x, y). The pixel array circuit 210 can be used to measure changes in light intensity and generate multiple data signals corresponding to the multiple pixels. In some possible embodiments, each pixel is configured to respond independently to changes in light intensity in the environment. In some possible embodiments, the pixel compares the measured change in light intensity with a predetermined threshold. If the measured change in light intensity exceeds the predetermined threshold, the pixel generates a first data signal indicating a light intensity change event. For example, the first data signal includes polarity information, such as +1 or -1, or the first data signal may also be absolute light intensity information. In this example, the first data signal may indicate a trend of light intensity change or an absolute light intensity value at the corresponding pixel. In some possible embodiments, if the measured change in light intensity does not exceed the predetermined threshold, the pixel generates a second data signal different from the first data signal, such as 0. In embodiments of this application, the data signal may indicate, but is not limited to, light intensity polarity, absolute light intensity value, changes in light intensity, etc. Light intensity polarity can represent the trend of light intensity change, such as increasing or decreasing, and is usually represented by +1 and -1. Absolute light intensity value can represent the light intensity value measured at the current moment. Depending on the structure, application, and type of sensor, light intensity or the amount of light intensity change can have different physical meanings. The scope of this application is not limited in this respect.

[0466] The readout circuit 220 is coupled to and can communicate with the pixel array circuit 210 and the control circuit 230. The readout circuit 220 is configured to read the data signal output by the pixel array circuit 210, which can be understood as the readout circuit 220 reading the data signal output by the pixel array 210 and sending it to the control circuit 230. The control circuit 230 is configured to control the mode in which the readout circuit 220 reads the data signal. The control circuit 230 can also be configured to control the representation of the output data signal; in other words, to control the representation precision of the data signal. For example, the control circuit can control the visual sensor output of events represented by polarity information, or by light intensity information, or by a fixed number of bits, etc., which will be described below in conjunction with specific embodiments.

[0467] According to possible embodiments of this application, the control circuit 230 can be as follows: Figure 4-a The circuitry or chip shown is an external component of the vision sensor 200 and is connected to the vision sensor 200 via a bus interface. In some other possible embodiments, the control circuitry 230 may also be an internal component of the vision sensor, integrated with the pixel array circuitry and readout circuitry therein. Figure 4-b A block diagram of another vision sensor 300 according to a possible embodiment of this application is shown. Vision sensor 300 can be implemented as an example of vision sensor 200. Figure 4-b The diagram shown is another block diagram of a vision sensor provided in this application. The vision sensor includes a pixel array circuit 310, a readout circuit 320, and a control circuit 330. The pixel array circuit 310, the readout circuit 320, and the control circuit 330 are functionally similar to... Figure 4-a The pixel array circuit 210, readout circuit 220, and control circuit 230 shown are identical and will not be described again here. It should be understood that the vision sensor is for illustrative purposes only and does not imply any limitation on the scope of this application. Embodiments of this application can also be embodied in different vision sensors. Furthermore, it should be understood that the vision sensor may also include other elements, modules, or entities not shown for clarity, but this does not mean that embodiments of this application do not include these elements or entities.

[0468] Based on the aforementioned architecture of the visual sensor, the visual sensor provided in this application will be described in detail below.

[0469] The readout circuit 220 can be configured to scan pixels in the pixel array circuit 210 in a predetermined order to read data signals generated by the corresponding pixels. In embodiments of this application, the readout circuit 220 is configured to read data signals output by the pixel array circuit 210 in more than one signal readout mode. For example, the readout circuit 220 can read in one of a first readout mode and a second readout mode. In the context of this document, the first readout mode and the second readout mode correspond to one of a frame-based readout mode and an event-stream-based readout mode, respectively. Further, the first readout mode can refer to the current readout mode of the readout circuit 220, and the second readout mode can refer to a switchable alternative readout mode.

[0470] refer to Figure 5 It illustrates the principles of a frame-scan-based synchronous read mode and an event-stream-based asynchronous read mode according to embodiments of this application. Figure 5 The upper half of the diagram shows pixels with black dots representing those that generate light intensity change events and pixels without such events representing white dots. The dashed box on the left represents the synchronous reading mode based on frame scanning, where all pixels generate voltage signals based on the received light signals, which are then converted from analog to digital and output as data signals. In this mode, the reading circuit 220 constructs a frame of data by reading the data signals generated by all pixels. The dashed box on the right represents the asynchronous reading mode based on event streams. In this mode, when the reading circuit 220 scans a pixel that generates a light intensity change event, it can acquire the pixel's coordinate information (x, y). Then, it reads only the data signals generated by the pixels that generate light intensity change events and records the reading time t. When there are multiple pixels generating light intensity change events in the pixel array circuit, the reading circuit 220 reads the data signals generated by the multiple pixels sequentially according to the scanning order and constructs an event stream as the output.

[0471] Figure 5 The lower half describes the two reading modes from the perspective of representation cost (e.g., the amount of data required to be read). For example... Figure 5 As shown, in synchronous read mode, the read circuit 220 reads the same amount of data each time, for example, one frame of data. Figure 5 The data shown is frame 1 data 401-1 and frame 2 data 401-2. This is based on the amount of data representing a single pixel (e.g., the number of bits B). p The amount of data to be read in one frame can be determined by the total number of pixels M in the pixel array circuit. pIn asynchronous reading mode, the reading circuit 220 reads the data signal indicating a change in light intensity, and then constructs an event stream 402 by arranging all the events in chronological order. In this case, the amount of data read by the reading circuit 220 each time is equal to the amount of event data B used to represent a single event. ev (For example, the sum of the coordinates (x, y) of the pixel that generated the event, the read timestamp t, and the number of bits representing the light intensity feature information) and the number N of light intensity change events. ev related.

[0472] In some implementations, the reading circuit 220 may be configured to provide at least one read data signal to the control circuit 230. For example, the reading circuit 220 may provide the control circuit 230 with data signals read over a period of time for the control circuit 230 to perform historical data statistics and analysis.

[0473] In some possible embodiments, where the currently employed first reading mode is an event-stream-based reading mode, the reading circuit 220 reads data signals generated by pixels in the pixel array circuit 210 that produce light intensity change events. For ease of description, these data signals are hereinafter also referred to as first data signals. Specifically, the reading circuit 220 determines the position information (x, y) of pixels related to light intensity change events by scanning the pixel array circuit 210. Based on the pixel position information (x, y), the reading circuit 220 reads the first data signal generated by the pixels from among multiple data signals to obtain characteristic information of the light intensity indicated by the first data signal and reading time information t. By way of example, in the event-stream-based reading mode, the amount of event data read per second by the reading circuit 220 can be represented as B. ev ·N ev One bit, that is, the data reading rate of the reading circuit 220 is B. ev ·N ev Bits per second (bps), where B ev In event-stream based read mode, the amount of event data (e.g., number of bits) allocated for each light intensity change event, where the first b x and b y One bit is used to represent the pixel coordinates (x, y), and the next bit... t One bit is used to represent the timestamp t when the data signal was read, and finally b f Each bit is used to represent the characteristic information of the light intensity indicated by the data signal, i.e., B ev =b x +b y +b t +b f N evThe average number of events per second is calculated by the readout circuit 220 based on historical statistics of the number of light intensity change events generated in the pixel array circuit 210 over a period of time. Since the readout mode is based on frame scanning, the amount of data in each frame read by the readout circuit 220 can be expressed as M·B. p Each bit, the amount of data read per second is M·B p •f bits, that is, the data reading rate of the read circuit 220 is M·B p ·f bps, where the total number of pixels in a given visual sensor 200 is M, B p To determine the amount of pixel data (e.g., number of bits) allocated to each pixel in the frame-scanning readout mode, f is a predetermined frame rate of the readout circuit 220 in the frame-scanning readout mode, i.e., the readout circuit 220 scans the pixel array circuit 210 at a predetermined frame rate f Hz in this mode to read the data signals generated by all pixels in the pixel array circuit 210. Therefore, M, B p Both f and f are known quantities, and the data reading rate of the reading circuit 220 in the frame scanning reading mode can be directly obtained.

[0474] In some possible embodiments, where the currently employed first reading mode is a frame-scan-based reading mode, the reading circuit 220 can derive the average number of events N generated per second based on historical statistics of the number of light intensity change events generated in the pixel array circuit 210 over a period of time. ev N is obtained according to the frame scan reading mode. ev It can be calculated that in the event stream-based reading mode, the amount of event data read per second by the reading circuit 220 is B. ev ·N ev In the event-stream-based read mode, the read data rate of the read circuit 220 is B bits. ev ·N ev bps.

[0475] As can be seen from the above two implementation methods, the data reading rate of the reading circuit 220 in the frame scanning-based reading mode can be directly calculated based on predefined parameters, while the data reading rate of the reading circuit 220 in the event stream-based reading mode can be obtained from N obtained in either of the two modes. ev Calculated.

[0476] Control circuit 230 is coupled to read circuit 220 and configured to control read circuit 220 to read data signals generated by pixel array circuit 210 in a specific read mode. In some possible embodiments, control circuit 230 may acquire at least one data signal from read circuit 220 and, based at least on the at least one data signal, determine which of the current read mode and alternative read modes is more suitable for the current application scenario and motion state. Furthermore, in some embodiments, control circuit 230 may, based on this determination, instruct read circuit 220 to switch from the current data read mode to another data read mode.

[0477] In some possible embodiments, control circuit 230 may send an instruction to readout circuit 220 regarding switching readout modes based on historical statistics of light intensity change events. For example, control circuit 230 may determine statistical data related to at least one light intensity change event based on at least one data signal received from readout circuit 220. If the statistical data is determined to meet predetermined switching conditions, control circuit 230 sends a mode switching signal to readout circuit 220 to switch readout circuit 220 to a second readout mode. For ease of comparison, the statistical data may be used to measure the readout data rate of the first readout mode and the second readout mode, respectively.

[0478] In some embodiments, the statistical data may include the total amount of data representing the number of events measured by the pixel array circuit 210 per unit time. If the total amount of data representing light intensity change events read by the readout circuit 220 in the first readout mode is greater than or equal to the total amount of data representing light intensity change events in the second readout mode, then the readout circuit 220 should switch from the first readout mode to the second readout mode. In some embodiments, the first readout mode is a frame-scan-based readout mode and the second readout mode is an event-stream-based readout mode. The control circuit 230 may be based on the number of pixels M, frame rate f, and pixel data amount B of the pixel array circuit. p To determine the total data volume M·B of light intensity change events read in the first reading mode. p • f. The control circuit 230 can be based on the number N of light intensity change events. ev And the amount of event data B associated with the event stream-based reading pattern ev To determine the total amount of data B for light intensity change events. ev ·N ev That is, the total amount of data B of light intensity change events read in the second reading mode. ev ·N ev In some embodiments, a switching parameter can be used to adjust the relationship between the total data volume in the two reading modes, as shown in the following formula (1), where the total data volume M·B of the light intensity change event read in the first reading mode is... p• f is greater than or equal to the total data volume B of light intensity change events in the second reading mode ev ·N ev The reading circuit 220 should switch to the second reading mode:

[0479] η·M·B P ·f≥B ev ·N ev (1)

[0480] Where η is the switching parameter used for adjustment. From the above formula (1), it can be further derived that the first threshold data volume d1 = M·B p ·f·η. That is, if the total data volume B of the light intensity change event... ev ·N ev If the data amount is less than or equal to the threshold data amount d1, it indicates that the total data amount of light intensity change events read in the first reading mode is greater than or equal to the total data amount of light intensity change events read in the second reading mode. The control circuit 230 can then determine that the statistical data of the light intensity change events meets the predetermined switching conditions. In this embodiment, the switching conditions can be determined at least based on the number of pixels M of the pixel array circuit, the frame rate f associated with the frame-scan-based reading mode, and the pixel data amount B. p To determine the threshold data volume d1.

[0481] As an alternative embodiment of the above embodiments, the total data volume M·B of the light intensity change events read in the first reading mode is... p • f is greater than or equal to the total data volume B of light intensity change events in the second reading mode ev ·N ev It can be shown in the following formula (2):

[0482] M·B P ·fB ev ·N ev ≥θ(2)

[0483] Where θ is the switching parameter used for adjustment. From the above formula (2), we can further derive the second threshold data volume.

[0484] d2=M·B p ·f-θ

[0485] That is, if the total amount of data B for light intensity change events... ev ·N evIf the data amount is less than or equal to the second threshold data amount d2, it indicates that the total data amount of light intensity change events read in the first reading mode is greater than or equal to the total data amount of light intensity change events read in the second reading mode. The control circuit 230 can then determine that the statistical data of the light intensity change events meets the predetermined switching conditions. In this embodiment, this can be based at least on the number of pixels M of the pixel array circuit, the frame rate f associated with the frame-scan-based reading mode, and the pixel data amount B. p To determine the threshold data volume d2.

[0486] In some embodiments, the first reading mode is an event stream-based reading mode and the second reading mode is a frame scan-based reading mode. In the event stream-based reading mode, the reading circuit 220 only reads the data signals generated by the pixels that produce the events. Therefore, the control circuit 230 can directly determine the number N of light intensity change events generated in the pixel array circuit 210 based on the number of data signals provided by the reading circuit 220. ev The control circuit 230 can be based on the number of events N. ev And the amount of event data B associated with the event stream-based reading pattern ev Determine the total amount of data for light intensity change events, that is, the total amount of data B of events read in the first reading mode. ev ·N ev Similarly, the control circuit 230 can also be based on the number of pixels M, frame rate f, and pixel data volume B of the pixel array circuit. p To determine the total data volume M·B of light intensity change events read in the second reading mode. p ·f. As shown in the following formula (3), the total amount of data B of light intensity change events read in the first reading mode is ev ·N ev The total data volume M·B of light intensity change events greater than or equal to that of the second reading mode p • f, The reading circuit 220 should switch to the second reading mode:

[0487] B ev ·N ev ≥η·M·B P ·f(3)

[0488] Where η is the switching parameter used for adjustment. From the above formula (3), it can be further derived that the first threshold data volume d1 = η·M·B P ·f. If the total data volume B of the light intensity change event ev ·N ev If the data volume d1 is greater than or equal to the threshold data volume, the control circuit 230 determines that the statistical data of the light intensity change event meets the predetermined switching conditions. In this embodiment, this can be based at least on the number of pixels M, frame rate f, and pixel data volume B of the pixel array circuit. pTo determine the threshold data volume d1.

[0489] As an alternative embodiment of the above embodiments, the total data volume B of the light intensity change event read in the first reading mode is... ev ·N ev The total data volume M·B of light intensity change events greater than or equal to that of the second reading mode p • f can be represented by the following formula (4):

[0490] M·B P ·fB ev ·N ev ≤θ(4)

[0491] Where θ is the switching parameter used for adjustment. From the above formula (4), it can be further derived that the second threshold data volume d2 = M·B P ·f-θ, if the total data volume B of the light intensity change event ev ·N ev If the data volume d2 is greater than or equal to the threshold data volume, the control circuit 230 determines that the statistical data of the light intensity change event meets the predetermined switching conditions. In this embodiment, this can be based at least on the number of pixels M, frame rate f, and pixel data volume B of the pixel array circuit. p To determine the threshold data volume d2.

[0492] In other embodiments, the statistical data may include the number N of events measured by the pixel array circuit 210 per unit time. ev If the first reading mode is a frame-scan-based reading mode and the second reading mode is an event-stream-based reading mode, the control circuit 230 determines the number N of light intensity change events based on the number of the first data signals among the multiple data signals provided by the reading circuit 220. ev If the statistical data indicates the number N of light intensity change events... ev If the number of light intensity change events is less than the first threshold n1, then the control circuit 230 determines that the statistical data of the light intensity change event meets the predetermined switching conditions, which can be based at least on the number of pixels M of the pixel array circuit, the frame rate f associated with the frame scan-based readout mode, and the pixel data volume B. p And the amount of event data B associated with the event stream-based reading pattern. ev To determine the number of first thresholds n1. For example, in the aforementioned embodiment, based on formula (1), the following formula (5) can be further obtained:

[0493]

[0494] That is, the number of the first thresholds n1 can be determined as

[0495] As an alternative embodiment of the above embodiments, based on formula (2), the following formula (6) can be further obtained:

[0496]

[0497] Accordingly, the number of second thresholds n2 can be determined as

[0498] In some other embodiments, if the first reading mode is an event stream-based reading mode and the second reading mode is a frame scan-based reading mode, the control circuit 230 can directly determine the number N of light intensity change events based on the number of at least one data signal provided by the reading circuit 220. ev If the statistical data indicates the number N of light intensity change events... ev If the number of light intensity changes is greater than or equal to the first threshold number n1, then the control circuit 230 determines that the statistical data of the light intensity change event meets the predetermined switching conditions. This can be based at least on the number of pixels M of the pixel array circuit 210, the frame rate f associated with the frame scan-based readout mode, and the pixel data volume B. p And the amount of event data B associated with the event stream-based reading pattern. ev Determine the number of the first threshold n1 = M·B p ·f / (η·B ev For example, in the foregoing embodiments, based on formula (3), the following formula (7) can be further obtained:

[0499]

[0500] That is, the number of the first thresholds n1 can be determined as

[0501] As an alternative embodiment of the above embodiments, based on formula (4), the following formula (8) can be further obtained:

[0502]

[0503] Accordingly, the number of second thresholds n2 can be determined as

[0504] It should be understood that the formulas, switching conditions and related calculation methods given above are merely an example implementation of the embodiments of this application. Other suitable mode switching conditions, switching strategies and calculation methods may also be adopted, and the scope of this application is not limited in this respect.

[0505] Figure 6-a A schematic diagram of a vision sensor according to an embodiment of the present application operating in a frame-scan-based readout mode is shown. Figure 6-bA schematic diagram illustrating the operation of a vision sensor according to an embodiment of this application in an event-stream-based readout mode is shown. Figure 6-a As shown, the read circuit 220 or 320 is currently operating in a first read mode, namely a frame-scan-based read mode. Because the control circuit 230 or 330 determines, based on historical statistics, that the number of events generated in the current pixel array circuit 210 or 310 is relatively small—for example, only four valid data points in a frame—it predicts a low event generation rate in the next time period. If the read circuit 220 or 320 continues to use the frame-scan-based read mode, it will need to repeatedly allocate bits for pixels generating events, resulting in a large amount of redundant data. In this case, the control circuit 230 or 330 sends a mode switching signal to the read circuit 220 or 320 to switch from the first read mode to the second read mode. After the switch, Figure 6-b As shown, the reading circuit 220 or 320 operates in the second reading mode, reading only valid data signals, thereby avoiding the transmission bandwidth and storage resources occupied by a large number of invalid data signals.

[0506] Figure 6-c A schematic diagram of a vision sensor according to an embodiment of the present application operating in an event stream-based readout mode is shown. Figure 6-d A schematic diagram illustrating the operation of a visual sensor according to an embodiment of this application in a frame-scan-based readout mode is shown. Figure 6-c As shown, the reading circuit 220 or 320 is currently operating in a first reading mode, namely, an event-stream-based reading mode. Because the control circuit 230 or 330 determines, based on historical statistics, that the number of events generated in the pixel array circuit 210 or 310 is currently high—for example, nearly all pixels in the pixel array circuit 210 or 310 generate data signals indicating light intensity changes exceeding a predetermined threshold within a short period—the reading circuit 220 or 320 can predict a potentially high event generation rate in the next time period. Since the read data signals contain a large amount of redundant data, such as nearly identical pixel position information, read timestamps, etc., if the reading circuit 220 or 320 continues to use the event-stream-based reading mode, the amount of read data will surge. Therefore, in this situation, the control circuit 230 or 330 sends a mode switching signal to the reading circuit 220 or 320 to switch from the first reading mode to the second reading mode. After the switch, Figure 6-d As shown, the readout circuit 220 or 320 operates in a frame-scan-based mode, reading the data signal in a readout mode with a lower representation cost per pixel, thus alleviating the pressure of storing and transmitting the data signal.

[0507] In some possible embodiments, the vision sensor 200 or 300 may further include a parsing circuit that can be configured to parse the data signal output by the readout circuit 220 or 320. In some possible embodiments, the parsing circuit may parse the data signal using a parsing mode adapted to the current data readout mode of the readout circuit 220 or 320. This will be described in detail below.

[0508] It should be understood that other existing or future data reading modes, data parsing modes, etc., are also applicable to possible embodiments of this application, and all values ​​in the embodiments of this application are illustrative rather than restrictive. For example, possible embodiments of this application may switch between more than two data reading modes.

[0509] According to possible embodiments of this application, a visual sensor chip is provided that can adaptively switch between multiple readout modes based on historical statistics of light intensity change events generated in the pixel array circuit. In this way, the visual sensor chip can always achieve good readout and resolution performance in both dynamic and static scenes, avoiding the generation of redundant data and alleviating the pressure on image processing, transmission, and storage.

[0510] Figure 7 A flowchart illustrating a method for operating a vision sensor chip according to a possible embodiment of this application is shown. In some possible embodiments, the method can be... Figure 4-a The visual sensor 200 shown or Figure 4-b The visual sensor 300 shown below and the following Figure 9 This can be implemented in the electronic device shown, or it can be implemented using any suitable device, including various devices currently known or to be developed in the future. For ease of discussion, the following will be combined with... Figure 4-a The method is described using the visual sensor 200 shown.

[0511] See Figure 7 The present application provides a method for operating a vision sensor chip, which may include the following steps:

[0512] 501. Generate multiple data signals corresponding to multiple pixels in the pixel array circuit.

[0513] The pixel array circuit 210 generates multiple data signals corresponding to multiple pixels in the pixel array circuit 210 by measuring the amount of light intensity change. In the context of this document, the data signals may indicate, but are not limited to, light intensity polarity, absolute light intensity value, light intensity change value, etc.

[0514] 502. Read at least one of a plurality of data signals from the pixel array circuit in a first read mode.

[0515] The readout circuit 220 reads at least one of a plurality of data signals from the pixel array circuit 210 in a first readout mode, and these data signals occupy certain storage and transmission resources within the vision sensor 200 after being read. Depending on the specific readout mode, the vision sensor chip 200 may read the data signals in different ways. In some possible embodiments, for example, in an event-stream-based readout mode, the readout circuit 220 determines the position information (x, y) of pixels related to a light intensity change event by scanning the pixel array circuit 210. Based on this position information, the readout circuit 220 can read out the first data signal from the plurality of data signals. In this embodiment, the readout circuit 220 obtains characteristic information of light intensity, the position information (x, y) of the pixel that generated the light intensity change event, the timestamp t of the readout data signal, etc., by reading the data signal.

[0516] In some other possible embodiments, the first readout mode may be a frame-scan-based readout mode. In this mode, the vision sensor 200 scans the pixel array circuit 210 at a frame frequency associated with the frame-scan-based readout mode to read all data signals generated by the pixel array circuit 210. In this embodiment, the readout circuit 220 obtains characteristic information of light intensity by reading the data signals.

[0517] 503. Provide at least one data signal to the control circuit.

[0518] The reading circuit 220 provides at least one read data signal to the control circuit 230 for data statistics and analysis. In some embodiments, the control circuit 230 can determine statistical data related to at least one light intensity change event based on the at least one data signal. The control circuit 230 can use a switching strategy module to analyze the statistical data. If it is determined that the statistical data meets predetermined switching conditions, the control circuit 230 sends a mode switching signal to the reading circuit 220.

[0519] In some embodiments where the first read mode is a frame-scan-based read mode and the second read mode is an event-stream-based read mode, the control circuit 230 may determine the number of light intensity change events based on the number of first data signals among a plurality of data signals. Furthermore, the control circuit 230 compares the number of light intensity change events with a first threshold number. If statistical data indicates that the number of light intensity change events is less than or equal to the first threshold number, the control circuit 230 determines that the statistical data of the light intensity change events meets a predetermined switching condition and sends a mode switching signal. In this embodiment, the control circuit 230 may determine or adjust the first threshold number based on the number of pixels in the pixel array circuit, the frame rate and pixel data volume associated with the frame-scan-based read mode, and the event data volume associated with the event-stream-based read mode.

[0520] In some embodiments where the first read mode is an event-stream-based read mode and the second read mode is a frame-scan-based read mode, the control circuit 230 may determine statistical data related to light intensity change events based on the first data signal received from the read circuit 220. Furthermore, the control circuit 230 compares the number of light intensity change events with a second threshold number. If the number of light intensity change events is greater than or equal to the second threshold number, the control circuit 230 determines that the statistical data of the light intensity change events meets a predetermined switching condition and sends a mode switching signal. In this embodiment, the control circuit 230 may determine or adjust the second threshold number based on the number of pixels in the pixel array circuit, the frame rate and pixel data volume associated with the frame-scan-based read mode, and the event data volume associated with the event-stream-based read mode.

[0521] 504. Based on the mode switching signal, switch the first reading mode to the second reading mode.

[0522] Based on a mode switching signal received from the control circuit 220, the readout circuit 220 switches from a first readout mode to a second readout mode. Then, the readout circuit 220 reads at least one data signal generated by the pixel array circuit 210 in the second readout mode. The control circuit 230 can then continue to perform historical statistics on light intensity change events generated by the pixel array circuit 210, and when the switching conditions are met, send a mode switching signal to switch the readout circuit 220 from the second readout mode back to the first readout mode.

[0523] According to the method provided in a possible embodiment of this application, the control circuit continuously performs historical statistics and real-time analysis on light intensity change events generated in the pixel array circuit throughout the entire reading and parsing process. Once the switching conditions are met, a mode switching signal is sent to switch the reading circuit from the current reading mode to a more suitable alternative switching mode. This adaptive switching process is repeated continuously until all data signals have been read.

[0524] Figure 8 A block diagram of a control circuit according to a possible embodiment of this application is shown. The control circuit can be used to implement... Figure 4-a Control circuit 230 in Figure 5 The control circuit 330, etc., can also be implemented using other suitable devices. It should be understood that the control circuit is for illustrative purposes only and does not imply any limitation on the scope of this application. Embodiments of this application can also be embodied in different control circuits. Furthermore, it should be understood that the control circuit may also include other elements, modules, or entities not shown for clarity, but this does not mean that embodiments of this application do not possess these elements or entities.

[0525] like Figure 8 As shown, the control circuit includes at least one processor 602, at least one memory 604 coupled to the processor 602, and a communication mechanism 612 coupled to the processor 602. The memory 604 is used to store at least a computer program and data signals acquired from a read circuit. A statistical model 606 and a strategy module 608 are pre-configured on the processor 602. The control circuit 630 can be communicatively coupled to, for example, a computer program or a data acquisition circuit. Figure 4-a The reading circuit 220 of the visual sensor 200 shown, or a reading circuit external to the visual sensor, is used to implement control functions. For ease of description, the following refers to... Figure 4-a The read circuit 220 in the present application is also applicable to the configuration of the peripheral read circuit.

[0526] and Figure 4-a Similar to the control circuit 230 shown, in some possible embodiments, the control circuit can be configured to control the readout circuit 220 to read multiple data signals generated by the pixel array circuit 210 in a specific data readout mode (e.g., a synchronous readout mode based on frame scanning, an asynchronous readout mode based on event streams, etc.). Additionally, the control circuit can be configured to acquire data signals from the readout circuit 220, which can indicate, but are not limited to, light intensity polarity, absolute light intensity value, changes in light intensity, etc. For example, light intensity polarity can represent a trend in light intensity change, such as increasing or decreasing, typically represented by +1 / -1. An absolute light intensity value can represent the light intensity value measured at the current moment. Depending on the structure, application, and type of sensor, information about light intensity or changes in light intensity can have different physical meanings.

[0527] The control circuit determines statistical data related to at least one light intensity change event based on data signals acquired from the readout circuit 220. In some embodiments, the control circuit may acquire data signals generated by the pixel array circuit 210 over a period of time from the readout circuit 220 and store these data signals in memory 604 for historical statistics and analysis. In the context of this application, the first readout mode and the second readout mode may be one of an asynchronous readout mode based on event streams and a synchronous readout mode based on frame scanning, respectively. However, it should be noted that all the features described herein with respect to adaptively switching readout modes are equally applicable to other types of sensors and data readout modes known now or to be developed in the future, as well as switching between more than two data readout modes.

[0528] In some possible embodiments, the control circuit may utilize one or more pre-configured statistical models 606 to perform historical statistics on light intensity change events generated by the pixel array circuit 210 provided by the readout circuit 220 over a period of time. The statistical model 606 can then transmit the statistical data to the strategy module 608 as output. As described above, the statistical data may indicate the number of light intensity change events or the total amount of data related to these events. It should be understood that any suitable statistical model or algorithm can be applied to possible embodiments of this application, and the scope of this application is not limited in this respect.

[0529] Since the statistical data is a historical record of light intensity change events generated by the visual sensor over a period of time, it can be used by the strategy module 608 to analyze and predict the rate of event occurrence in the next time period. The strategy module 608 can be pre-configured with one or more switching decisions. When multiple switching decisions exist, the control circuit can select one for analysis and decision-making as needed, for example, based on factors such as the type of visual sensor 200, the characteristics of the light intensity change event, the attributes of the external environment, and the motion state. In possible embodiments of this application, other suitable strategy modules and mode switching conditions or strategies can also be employed, and the scope of this application is not limited in this respect.

[0530] In some embodiments, if the strategy module 608 determines that the statistical data meets the mode switching conditions, it outputs an indication to the read circuit 220 regarding switching the read mode. In another embodiment, if the strategy module 608 determines that the statistical data does not meet the mode switching conditions, it does not output an indication to the read circuit 220 regarding switching the read mode. In some embodiments, the indication regarding switching the read mode may be in an explicit form as described in the above embodiments, for example, notifying the read circuit 220 to switch the read mode in the form of a switching signal or a flag bit.

[0531] Figure 9A block diagram of an electronic device according to a possible embodiment of this application is shown. Figure 9 As shown, the electronic device includes a vision sensor chip 901, communication interfaces 902 and 903, a control circuit 930, and a resolution circuit 904. It should be understood that the electronic device is for illustrative purposes and can be implemented using any suitable device, including various sensor devices currently known and those developed in the future. Embodiments of this application can also be embodied in different sensor systems. Furthermore, it should be understood that the electronic device may also include other elements, modules, or entities not shown for clarity, but this does not mean that embodiments of this application do not possess these elements, modules, or entities.

[0532] like Figure 9 As shown, the vision sensor includes a pixel array circuit 710 and a readout circuit 720, wherein readout components 720-1 and 720-2 of the readout circuit 720 are coupled to the control circuit 730 via communication interfaces 702 and 703, respectively. In embodiments of this application, readout components 720-1 and 720-2 can be implemented using separate devices or integrated into the same device. For example, Figure 4-a The read circuit 220 shown is an integrated example implementation. For ease of description, read components 720-1 and 720-2 can be configured to implement data reading functions in a frame scan-based read mode and an event stream-based read mode, respectively.

[0533] The pixel array circuit 710 can utilize Figure 4-a The pixel array circuit 210 or Figure 5 The pixel array circuit 310 in this application can be used, but it can also be implemented using any other suitable device. This application is not limited in this respect. The features of the pixel array circuit 710 will not be described in detail here.

[0534] The readout circuit 720 can read the data signal generated by the pixel array circuit 710 in a specific readout mode. For example, in the example where readout component 720-1 is turned on and readout component 720-2 is turned off, the readout circuit 720 initially reads the data signal using a frame scan-based readout mode. In the example where readout component 720-2 is turned on and readout component 720-1 is turned off, the readout circuit 720 initially reads the data signal using an event stream-based readout mode. The readout circuit 720 can utilize... Figure 4-a The reading circuit 220 or Figure 5 The reading circuit 320 in the middle can be used to implement this, or it can be implemented using any other suitable device. The features of the reading circuit 720 will not be described in detail here.

[0535] In embodiments of this application, the control circuit 730 can instruct the read circuit 720 to switch from a first read mode to a second read mode via an indication signal or a flag bit. In this case, the read circuit 720 can receive an instruction from the control circuit 730 regarding the switching of read modes, for example, turning on read component 720-1 and turning off read component 720-2, or turning on read component 720-2 and turning off read component 720-1.

[0536] As described above, the electronic device may also include a parsing circuit 704. The parsing circuit 704 can be configured to parse the data signal read by the reading circuit 720. In possible embodiments of this application, the parsing circuit may employ a parsing mode adapted to the current data reading mode of the reading circuit 720. As an example, if the reading circuit 720 initially reads the data signal in an event-stream-based reading mode, the parsing circuit accordingly bases its parsing on a first data quantity B associated with that reading mode. ev ·N ev To parse the data. When the read circuit 720 switches from an event-stream-based read mode to a frame-scan-based read mode based on the instruction of the control circuit 730, the parsing circuit begins to parse the data according to the second data volume, i.e., the size of one frame of data, M·B. p To analyze data signals, and vice versa.

[0537] In some embodiments, the parsing circuit 704 can switch the parsing mode of the parsing circuit without explicit switching signals or flag bits. For example, the parsing circuit 704 can employ the same or corresponding statistical model and switching strategy as the control circuit 730 to perform the same statistical analysis on the data signal provided by the reading circuit 720 and make consistent switching predictions as the control circuit 730. As an example, if the reading circuit 720 initially reads the data signal in an event-stream-based reading mode, the parsing circuit initially bases its parsing on a first data quantity B associated with that reading mode. ev ·N ev To analyze the data. For example, the first b values ​​analyzed by the analytical circuit. x The first bit indicates the x-coordinate of the pixel, followed by the second bit. y The first bit indicates the pixel's coordinate y, followed by the second bit. t Each bit indicates the read time, and the last bit is taken. f Each bit indicates characteristic information of light intensity. The parsing circuit acquires at least one data signal from the reading circuit 720 and determines statistical data related to at least one light intensity change event. If the parsing circuit 704 determines that the statistical data meets the switching condition, it switches to the parsing mode corresponding to the frame-scan-based reading mode, with a frame data size of M·B. p To analyze data signals.

[0538] As another example, if the read circuit 720 initially reads the data signal in a frame-based scan read mode, the parsing circuit 704 reads the data signal in a parsing mode corresponding to that read mode, according to each B... p The bits sequentially extract the value of each pixel position within the frame, with the value of pixels where no light intensity change event occurred being 0. The parsing circuit 704 can count the number of non-zero values ​​within a frame based on the data signal, that is, the number of light intensity change events within that frame.

[0539] In some possible embodiments, the parsing circuit 704 acquires at least one data signal from the reading circuit 720, and determines, based at least on the at least one data signal, which of the current parsing mode and alternative parsing modes corresponds to the reading mode of the reading circuit 720. Furthermore, in some embodiments, the parsing circuit 704 may switch from the current parsing mode to another parsing mode based on this determination.

[0540] In some possible embodiments, the parsing circuit 704 may determine whether to switch parsing modes based on historical statistics of light intensity change events. For example, the parsing circuit 704 may determine statistical data related to at least one light intensity change event based on at least one data signal received from the readout circuit 720. If the statistical data is determined to meet the switching conditions, the parsing circuit 704 switches from the current parsing mode to an alternative parsing mode. For ease of comparison, the statistical data may be used to measure the readout data rate of the first readout mode and the second readout mode of the readout circuit 720, respectively.

[0541] In some embodiments, the statistical data may include the total amount of data representing the number of events measured by the pixel array circuit 710 per unit time. If the parsing circuit 704 determines, based on at least one data signal, that the total amount of data representing light intensity change events read by the reading circuit 720 in the first reading mode is greater than or equal to the total amount of data representing light intensity change events in its second reading mode, it indicates that the reading circuit 720 has switched from the first reading mode to the second reading mode. In this case, the parsing circuit 704 should accordingly switch to the parsing mode corresponding to the current reading mode.

[0542] In some embodiments, a first read mode is a frame-scan-based read mode and a second read mode is an event-stream-based read mode. In this embodiment, the parsing circuit 704 initially parses the data signal acquired from the read circuit 720 using a frame-based parsing mode corresponding to the first read mode. The parsing circuit 704 can parse the data signal based on the number of pixels M, frame rate f, and pixel data volume B of the pixel array circuit 710. p The total amount of data M·B of light intensity change events read by the readout circuit 720 in the first readout mode is determined. p • f. The analytical circuit 704 can be based on the number N of light intensity change events.ev And the amount of event data B associated with the event stream-based reading pattern ev To determine the total amount of data B of light intensity change events read by the reading circuit 720 in the second reading mode. ev ·N ev In some embodiments, switching parameters can be used to adjust the relationship between the total data volume in the two reading modes. Furthermore, the parsing circuit 704 can determine the total data volume M·B of the light intensity change event read by the reading circuit 720 in the first reading mode based on, for example, formula (1) above. p Is f greater than or equal to the total data volume B of the light intensity change event in the second reading mode? ev ·N ev If so, the parsing circuit 704 determines that the reading circuit 720 has switched to the event stream-based reading mode, and accordingly switches from the frame-based parsing mode to the event stream-based parsing mode.

[0543] As an alternative embodiment of the above embodiments, the parsing circuit 704 can determine the total amount of light intensity change events read by the reading circuit 720 in the first reading mode according to formula (2) above. p Is f greater than or equal to the total amount of data B of light intensity change events read in the second reading mode? ev ·N ev Similarly, the total amount of data M·B of light intensity change events read by the reading circuit 720 in the first reading mode is determined. p • f is greater than or equal to the total data volume B of light intensity change events in the second reading mode ev ·N ev In the case of this, the parsing circuit 704 determines that the reading circuit 720 has switched to the event stream-based reading mode, and accordingly switches from the frame-based parsing mode to the event stream-based parsing mode.

[0544] In some embodiments, the first reading mode is an event-stream-based reading mode and the second reading mode is a frame-scan-based reading mode. In this embodiment, the parsing circuit 704 initially parses the data signal acquired from the reading circuit 720 using an event-stream-based parsing mode corresponding to the first reading mode. As mentioned above, the parsing circuit 704 can directly determine the number N of light intensity change events generated in the pixel array circuit 710 based on the number of first data signals provided by the reading circuit 720. ev The analytical circuit 704 can be based on the number of events N. ev And the amount of event data B associated with the event stream-based reading pattern ev Determine the total amount of event data B read by the reading circuit 720 in the first reading mode. ev ·Nev Similarly, the parsing circuit 704 can also be based on the number of pixels M, frame rate f, and pixel data volume B of the pixel array circuit. p The total amount of data M·B of light intensity change events read by the reading circuit 720 in the second reading mode is determined. p ·f. Then, the parsing circuit 704 can, for example, determine the total amount of data B of the light intensity change events read in the first reading mode according to formula (3) above. ev ·N ev Is it greater than or equal to the total data volume M·B of the light intensity change event in the second reading mode? p ·f. Similarly, when the parsing circuit 704 determines the total amount of data B of the light intensity change events read by the reading circuit 720 in the first reading mode... ev ·N ev The total data volume M·B of light intensity change events greater than or equal to that of the second reading mode p When f is reached, the parsing circuit 704 determines that the reading circuit 720 has switched to the frame-scan-based reading mode, and accordingly switches from the event-stream-based parsing mode to the frame-based parsing mode.

[0545] As an alternative embodiment of the above embodiments, the parsing circuit 704 can determine the total amount of data B of the light intensity change event read by the reading circuit 720 in the first reading mode according to the formula (4) above. ev ·N ev Is it greater than or equal to the total data volume M·B of light intensity change events read in the second reading mode? p ·f. Similarly, the total amount of data B of light intensity change events read by the reading circuit 720 in the first reading mode is determined. ev ·N ev The total data volume M·B of light intensity change events greater than or equal to that of the second reading mode p In the case of f, the parsing circuit 704 determines that the reading circuit 720 has switched to the frame scan-based reading mode, and accordingly switches from the event stream-based parsing mode to the frame scan-based parsing mode.

[0546] For the event reading time t in frame-scan-based reading mode, it is assumed that all events within the same frame have the same reading time t. When higher precision is required for event reading time, the reading time of each event can be further determined as follows. Taking the above embodiment as an example, in frame-scan-based reading mode, the frequency of the reading circuit 720 scanning the pixel array circuit is f Hz, then the time interval for reading data from two adjacent frames is S = 1 / f, and the start time of each frame is given as:

[0547] T k =T0+kS(9)

[0548] Where T0 is the start time of the first frame and k is the frame number, the time required for digital-to-analog conversion of one pixel out of M pixels can be determined by the following formula (10):

[0549]

[0550] The time of the light intensity change event at the i-th pixel in the k-th frame can be determined by the following formula (11):

[0551]

[0552] Where i is a positive integer. If the current read mode is synchronous read mode, then switch to asynchronous read mode, according to each event B. ev Bit-based data parsing. In the above embodiments, the switching of parsing modes can be achieved without explicit switching signals or flag bits. For other currently known or future-developed data reading modes, the parsing circuit can also use a similar method adapted to the data reading mode to parse the data, which will not be elaborated here.

[0553] Figure 10 A schematic diagram illustrating the variation of data volume over time in a single data read mode and an adaptive switching read mode according to a possible embodiment of this application is shown. Figure 10 The left half of the diagram illustrates the change in the amount of data read over time for traditional vision sensors or sensor systems using either synchronous or asynchronous read modes. In the case of a purely synchronous read mode, as shown by curve 1001, the amount of data read remains constant over time because each frame has a fixed amount of data; that is, the data read rate (the amount of data read per unit time) is stable. As mentioned earlier, when a large number of events are generated in the pixel array circuit, it is more reasonable to use a frame-based scan read mode to read the data signal, as the majority of the frame data represents valid data indicating the occurrence of events, with minimal redundancy. However, when fewer events are generated in the pixel array circuit, a frame contains a large amount of invalid data representing and generating events. In this case, representing and reading the light intensity information at the pixel using the frame data structure would create redundancy, wasting transmission bandwidth and storage resources.

[0554] In the case of simply using asynchronous reading mode, as shown by curve 1002, the amount of data read varies with the rate of event generation, and therefore the data reading rate is not fixed. When few events are generated in the pixel array circuit, only a small number of bits are needed to represent the pixel coordinate information (x, y), the timestamp t of the data signal being read, and the characteristic information f of the light intensity. The total amount of data to be read is small, and asynchronous reading mode is reasonable in this case. When a large number of events are generated in the pixel array circuit in a short period of time, a large number of bits need to be allocated to represent these events. However, these pixel coordinates are almost adjacent, and the data signal reading time is almost the same. In other words, there is a large amount of duplicate data in the read event data, so there is also a redundancy problem in asynchronous reading mode. In this case, the data reading rate may even exceed the data reading rate in synchronous reading mode, and it is unreasonable to still use asynchronous reading mode.

[0555] Figure 10 The right half of the diagram illustrates the variation of data volume over time in an adaptive data readout mode according to a possible embodiment of this application. The adaptive data readout mode can utilize... Figure 4-a The visual sensor 200 shown Figure 4-b The visual sensor 300 shown or Figure 9 The electronic devices shown can be used to achieve this, or traditional vision sensors or sensor systems can be used by utilizing... Figure 8 The control circuit shown implements an adaptive data reading mode. For ease of description, please refer to the following... Figure 4-a The visual sensor 200 shown is used to describe the characteristics of the adaptive data readout mode. As shown in curve 1003, the visual sensor 200 selects, for example, an asynchronous readout mode in the initialization state. This is because the number of bits B used to represent each event in this mode is... ev It is pre-arranged (e.g., B) ev =b x +b y +b t +b f As events are generated and read, the vision sensor 200 can calculate the data reading rate in the current mode. On the other hand, in synchronous reading mode, the number of bits B used to represent each pixel in each frame... pThis is also predetermined, so the data reading rate using synchronous reading mode during that period can be calculated. The vision sensor 200 can then determine whether the relationship between the data rates of the two reading modes satisfies a mode switching condition. For example, the vision sensor 200 can compare which of the two reading modes has a lower data reading rate based on a predefined threshold. Once it is determined that the mode switching condition is met, the vision sensor 200 switches to another reading mode, for example, from the initial asynchronous reading mode to the synchronous reading mode. These steps continue during the reading and parsing of the data signal until all data is output. As shown in curve 1003, the vision sensor 200 adaptively selects the optimal reading mode throughout the data reading process, with the two reading modes alternating, ensuring that the data reading rate of the vision sensor 200 never exceeds the data reading rate of the synchronous reading mode, thereby reducing the cost of data transmission, parsing, and storage for the vision sensor.

[0556] In addition, according to the adaptive data reading method proposed in the embodiments of this application, the visual sensor 200 can perform historical data statistics on events to predict the possible event generation rate in the next time period, thus enabling the selection of a reading mode that is more suitable for the application scenario and motion state.

[0557] Through the above scheme, the visual sensor can adaptively switch between multiple data reading modes, ensuring that the data reading rate always remains within a predetermined data reading rate threshold. This reduces the cost of data transmission, parsing, and storage for the visual sensor, significantly improving its performance. Furthermore, such a visual sensor can perform data statistics on events generated over a period of time to predict the possible event generation rate in the next time period, thus enabling it to select a reading mode more suitable for the current external environment, application scenario, and motion state.

[0558] The previous section introduced how pixel array circuits can be used to measure changes in light intensity and generate multiple data signals corresponding to multiple pixels. These data signals can indicate, but are not limited to, light intensity polarity, absolute light intensity value, and changes in light intensity. The following section provides a detailed explanation of the data signals output by the pixel array circuit.

[0559] Figure 11 A schematic diagram of a pixel circuit 900 provided in this application is shown. Each pixel array circuit in pixel array circuits 210, 310, and 710 may include one or more pixel arrays, and each pixel array includes multiple pixels. Each pixel can be considered as a pixel circuit, and each pixel circuit is used to generate a data signal corresponding to that pixel. See also... Figure 11This is a schematic diagram of a preferred pixel circuit provided in an embodiment of this application. In this application, a pixel circuit is sometimes simply referred to as a pixel. Figure 11 As shown, a preferred pixel circuit in this application includes a light intensity detection unit 901, a threshold comparison unit 902, a readout control unit 903, and a light intensity acquisition unit 904.

[0560] A light intensity detection unit 901 is used to convert the acquired light signal into a first electrical signal. The light intensity detection unit 901 can monitor the light intensity information illuminating the pixel circuit in real time and convert the acquired light signal into an electrical signal and output it in real time. In some possible embodiments, the light intensity detection unit 901 can convert the acquired light signal into a voltage signal. This application does not limit the specific structure of the light intensity detection unit; any structure capable of converting the light signal into an electrical signal can be adopted in the embodiments of this application. For example, the light intensity detection unit may include a photodiode and a transistor. The anode of the photodiode is grounded, the cathode of the photodiode is connected to the source of the transistor, and the drain and gate of the transistor are connected to the power supply.

[0561] The threshold comparison unit 902 is used to determine whether the first electrical signal is greater than a first target threshold or less than a second target threshold. When the first electrical signal is greater than the first target threshold or less than the second target threshold, the threshold comparison unit 902 outputs a first data signal, which indicates that the pixel has a light intensity change event. The threshold comparison unit 902 is used to compare whether the difference between the current light intensity and the light intensity at the time of the last event exceeds a predetermined threshold, which can be understood with reference to Formula 1-1. The first target threshold can be understood as the sum of the first predetermined threshold and the second electrical signal, and the second target threshold can be understood as the sum of the second predetermined threshold and the second electrical signal. The second electrical signal is the electrical signal output by the light intensity detection unit 901 when the last event occurred. The threshold comparison unit in this application embodiment can be implemented in hardware or software, and this application embodiment does not limit this. The type of the first data signal output by the threshold comparison unit 902 may be different: in some possible embodiments, the first data signal includes polarity information, such as +1 or -1, to indicate light intensity enhancement or light intensity reduction. In some possible embodiments, the first data signal can be an activation signal, used to instruct the readout control unit 903 to control the light intensity acquisition unit 904 to acquire the first electrical signal and buffer the first electrical signal. When the first data signal is an activation signal, the first data signal can also be polarity information, in which case when the readout control unit 903 acquires the first data signal, it controls the light intensity acquisition unit 904 to acquire the first electrical signal.

[0562] The readout control unit 903 is also used to instruct the readout circuit to read the first electrical signal stored in the light intensity acquisition unit 904. Alternatively, it may instruct the readout circuit to read the first data signal output by the threshold comparison unit 902, which is polarity information.

[0563] The readout circuit 905 can be configured to scan pixels in the pixel array circuit in a predetermined order to read the data signals generated by the corresponding pixels. In some possible embodiments, the readout circuit 905 can be understood with reference to readout circuits 220, 320, and 720, i.e., the readout circuit 905 is configured to read the data signals output by the pixel circuit in more than one signal readout mode. For example, the readout circuit 905 can read in one of a first readout mode and a second readout mode, the first readout mode and the second readout mode corresponding to one of a frame-scan-based readout mode and an event-stream-based readout mode, respectively. In some possible embodiments, the readout circuit 905 can also read the data signals output by the pixel circuit in only one signal readout mode, such as being configured to read the data signals output by the pixel circuit only in a frame-scan-based readout mode, or being configured to read the data signals output by the pixel circuit only in an event-stream-based readout mode. Figure 11 In corresponding embodiments, the data signal read by the reading circuit 905 is represented in different ways. In some possible embodiments, the data signal read by the reading circuit is represented by polarity information. For example, the reading circuit can read the polarity information output by the threshold comparison unit. In some possible embodiments, the data signal read by the reading circuit can be represented by light intensity information. For example, the reading circuit can read the electrical signal buffered by the light intensity acquisition unit.

[0564] See Figure 11-a Taking the reading of data signals output by pixel circuits based on event stream as an example, this paper explains how to represent events using light intensity information and how to represent events using polarity information. Figure 11-a As shown in the upper half, the black dots represent pixels that generate light intensity change events. Figure 11-a The system comprises eight events in total. The first five events are represented using light intensity information, and the last three events are represented using polarity information. For example... Figure 11-aAs shown in the lower half, both events represented by light intensity information and events represented by polarity information need to include coordinate information (x, y) and time information t. The difference lies in that, for events represented by light intensity information, the characteristic information m of light intensity is light intensity information a, while for events represented by polarity information, the characteristic information m of light intensity is polarity information p. The difference between light intensity information and polarity information has already been introduced above and will not be repeated here. It is only emphasized that the data volume of events represented by polarity information is smaller than that of events represented by light intensity information.

[0565] How to determine what kind of information the data signal read by the reading circuit is represented by needs to be determined according to the instructions issued by the control circuit, which will be explained in detail below.

[0566] In some implementations, the readout circuit 905 can be configured to provide at least one read data signal to the control circuit 906. For example, the readout circuit 905 can provide the control circuit 906 with the total amount of data signals read over a period of time, for the control circuit 906 to perform historical data statistics and analysis. In one implementation, the readout circuit 906 can calculate the number of events N generated per second by the pixel array circuit based on the number of light intensity change events generated by each pixel circuit 900 in the pixel array circuit over a period of time. ev Among them, N ev It can be obtained through either the frame-scan-based reading mode or the event-stream-based reading mode.

[0567] Control circuit 906 is coupled to read circuit 905 and configured to control read circuit 906 to read data signals generated by pixel circuit 900 in a specific event representation mode. In some possible embodiments, control circuit 906 may acquire at least one data signal from read circuit 905 and, based at least on the at least one data signal, determine which of the current event representation mode and alternative event representation modes is more suitable for the current application scenario and motion state. Furthermore, in some embodiments, control circuit 906 may, based on this determination, instruct read circuit 905 to switch from the current event representation mode to another event representation mode.

[0568] In some possible embodiments, the control circuit 906 may send an instruction to the reading circuit 905 regarding the conversion of the event representation method based on historical statistics of light intensity change events. For example, the control circuit 906 may determine statistical data related to at least one light intensity change event based on at least one data signal received from the reading circuit 905. If the statistical data is determined to meet predetermined conversion conditions, the control circuit 906 sends an instruction signal to the reading circuit 905 to cause the reading circuit 905 to convert the format of the read events.

[0569] In some possible embodiments, assuming the readout circuit 905 is configured to read the data signal output by the pixel circuit only in an event-stream-based readout mode, the data provided by the readout circuit 905 to the control circuit 906 is the total number of events (light intensity change events) measured by the pixel array circuit per unit time. Assuming the current control circuit 906 controls the readout circuit 905 to read the data output by the threshold comparison unit 902, i.e., events are represented by polarity information, the readout circuit 905 can then read the data based on the number N of light intensity change events. ev The bit width H of the data format determines the total data volume N of the light intensity change event. ev ×H. Where the bit width H of the data format is b. x +b y +b t +b p b p One bit is used to represent the polarity information of the light intensity indicated by the data signal, usually 1 to 2 bits. Since the polarity information of light intensity is usually represented by 1 to 2 bits, the total amount of data of the event represented by the polarity information is always less than the bandwidth. In order to ensure that the event data with higher precision can be transmitted as much as possible without exceeding the bandwidth limit, if the total amount of data of the event represented by the light intensity information is also less than or equal to the bandwidth, then it is converted to represent the event by the light intensity information. In some embodiments, the conversion parameter can be used to adjust the relationship between the amount of data and the bandwidth K under an event representation method, as shown in the following formula (12), where the total amount of data N of the event represented by the light intensity information is... ev ×H is less than or equal to the bandwidth.

[0570] N ev ×H≤α×K(12)

[0571] Where α is the conversion parameter used for adjustment. From the above formula (12), it can be further concluded that if the total amount of event data represented by light intensity information is less than or equal to the bandwidth, the control circuit 906 can determine that the statistical data of the light intensity change event meets the predetermined switching conditions. Some possible application scenarios include when the pixel acquisition circuit generates fewer events within a certain period of time, or when the rate at which the pixel acquisition circuit generates events is slow within a certain period of time. In these cases, events can be represented by light intensity information. Since events represented by light intensity information can carry more information, it is beneficial for the subsequent processing and analysis of events, such as improving the quality of image reconstruction.

[0572] In some implementations, assuming that the current control circuit 906 controls the reading circuit 905 to read the electrical signal buffered by the light intensity acquisition unit 904, that is, to represent events through light intensity information, the reading circuit 905 can read the number N of light intensity change events. ev The bit width H of the data format determines the total data volume N of the light intensity change event.ev ×H. Wherein, when using an event-stream based reading mode, the bit width H of the data format is b. x +b y +b t +b a b a Each bit is used to represent the light intensity information indicated by the data signal, and is usually multiple bits, such as 8 bits to 12 bits. In some embodiments, a conversion parameter can be used to adjust the relationship between the amount of data and the bandwidth K under an event representation method, as shown in the following formula (13), where the total amount of data N of the event represented by the light intensity information is... ev If ×H is greater than 1, the reading circuit 220 should read the data output by the threshold comparison unit 902, that is, convert it into an event represented by polarity information:

[0573] N ev ×H>β×K(13)

[0574] Where β is the conversion parameter used for adjustment. From the above formula (13), it can be further concluded that if the total amount of data N of the light intensity change event... ev If ×H is greater than the threshold data volume β×K, it indicates that the total data volume representing light intensity change events using light intensity information is greater than or equal to the bandwidth. The control circuit 905 can then determine that the statistical data of the light intensity change events meets the predetermined conversion conditions. Some possible application scenarios include situations where the pixel acquisition circuit generates a large number of events over a period of time, or when the rate of event generation by the pixel acquisition circuit is relatively fast over a period of time. In these cases, if light intensity information is continued to represent events, event loss may occur. Therefore, polarity information can be used to represent events to alleviate the pressure on data transmission and reduce data loss.

[0575] In some implementations, the data provided by the readout circuit 905 to the control circuit 906 is the number N of events measured by the pixel array circuit per unit time. ev In some possible embodiments, assuming that the current control circuit 906 controls the reading circuit 905 to read the data output by the threshold comparison unit 902, that is, to represent events through polarity information, the control circuit can determine the number N of light intensity change events. ev and The relationship between N determines whether the predetermined transformation conditions are met. ev Less than or equal to The reading circuit 220 should read the electrical signal buffered in the light intensity acquisition unit 904, that is, convert it into an event represented by light intensity information, and convert the current event represented by polarity information into an event represented by light intensity information. For example, in the aforementioned embodiment, based on formula (12), the following formula (14) can be further obtained:

[0576]

[0577] In some implementations, assuming that the current control circuit 906 controls the reading circuit 905 to read the electrical signal buffered by the light intensity acquisition unit 904, that is, to represent events through light intensity information, the control circuit 906 can determine the number N of light intensity change events based on this information. ev and The relationship between N determines whether the predetermined transformation conditions are met. ev Greater than The reading circuit 220 should read the signal output from the threshold comparison unit 902, that is, convert it into an event represented by polarity information, and convert the current event represented by light intensity information into an event represented by polarity information. For example, in the aforementioned embodiment, based on formula (12), the following formula (15) can be further obtained:

[0578]

[0579] In some possible embodiments, it is assumed that the readout circuit 905 is configured to read the data signal output by the pixel circuit only in a frame-scan-based readout mode. The data provided by the readout circuit 905 to the control circuit 906 is the total number of events (light intensity change events) measured by the pixel array circuit per unit time. When using the frame-scan-based readout mode, the bit width of the data format is H = B. p B p To represent events using polarity information, in a frame-scan-based readout mode, the amount of pixel data (e.g., number of bits) allocated per pixel is used. p Typically, it is 1 to 2 bits, and when the event is represented by light intensity information, it is typically 8 to 12 bits. The reading circuit 905 can determine the total data amount M×H of the light intensity change event, where M represents the total number of pixels. Assume that the current control circuit 906 controls the reading circuit 905 to read the data output by the threshold comparison unit 902, that is, the event is represented by polarity information. The total data amount of the event represented by polarity information is always less than the bandwidth. In order to ensure that the event data with higher precision can be transmitted as much as possible without exceeding the bandwidth limit, if the total data amount of the event represented by light intensity information is also less than or equal to the bandwidth, then it is converted to represent the event by light intensity information. In some embodiments, the conversion parameter can be used to adjust the relationship between the data amount and the bandwidth K under an event representation method, as shown in the following formula (16), the total data amount N of the event represented by light intensity information ev ×H is less than or equal to the bandwidth.

[0580] M×H≤α×K(16)

[0581] In some implementations, assuming that the current control circuit 906 controls the reading circuit 905 to read the electrical signal buffered by the light intensity acquisition unit 904, that is, to represent the event through light intensity information, the reading circuit 905 can determine the total data amount M×H of the light intensity change event. In some implementations, the conversion parameter can be used to adjust the relationship between the data amount and the bandwidth K under an event representation method, as shown in the following formula (17). If the total data amount M×H of the event represented by the light intensity information is greater than the bandwidth, the reading circuit 220 should read the data output by the threshold comparison unit 902, that is, convert it to represent the event through polarity information:

[0582] M×H>α×K(17)

[0583] In some possible embodiments, it is assumed that the readout circuit 905 is configured to read in one of a first readout mode and a second readout mode, the first readout mode and the second readout mode corresponding to one of a frame-scan-based readout mode and an event-stream-based readout mode, respectively. For example, assuming the readout circuit 905 is currently reading the data signal output by the pixel circuit in the event-stream-based readout mode, and the control circuit 906 controls the readout circuit 905 to read the data output by the threshold comparison unit 902, i.e., a combination mode in which events are represented by polarity information, the following explanation illustrates how the control circuit determines whether the switching conditions are met:

[0584] In the initial state, any reading mode can be selected, such as a frame-scan-based reading mode or an event-stream-based reading mode. Furthermore, in the initial state, any event representation method can be selected. For example, control circuit 906 controls reading circuit 905 to read the electrical signal buffered by light intensity acquisition unit 904, i.e., representing the event through light intensity information; or control circuit 906 controls reading circuit 905 to read data output by threshold comparison unit 902, i.e., representing the event through polarity information. Assuming reading circuit 905 is currently reading the data signal output by the pixel circuit in event-stream-based reading mode, control circuit 906 controls reading circuit 905 to read the data output by threshold comparison unit 902, i.e., representing the event through polarity information. The data provided by reading circuit 905 to control circuit 906 can be the first total data amount of the number of events (light intensity change events) measured by the pixel array circuit per unit time. Since the total number of pixels M is known, the pixel data amount B allocated to each pixel in frame-scan-based reading mode... p It is known that the bit width H of the data format when representing an event using light intensity information is known. Based on the aforementioned known M and B... pH can acquire a second total data quantity representing the number of events measured by the pixel array circuit per unit time in a combined event mode, where the data signal is read from the pixel circuit output based on an event stream reading mode and the number of events measured by the pixel array circuit per unit time is represented by light intensity information; it can acquire a third total data quantity representing the number of events measured by the pixel array circuit per unit time in a combined event mode, where the data signal is read from the pixel circuit output based on a frame scan reading model and the number of events measured by the pixel array circuit per unit time is represented by polarity information; and it can acquire a fourth total data quantity representing the number of events measured by the pixel array circuit per unit time in a combined event mode, where the data signal is read from the pixel circuit output based on a frame scan reading model and the number of events measured by the pixel array circuit per unit time is represented by light intensity information. Specifically, this is based on M and B. p The methods for calculating the second, third, and fourth data quantities have been described above and will not be repeated here. The switching condition is determined by using the first total data quantity provided by the aforementioned reading circuit 905, the calculated second, third, and fourth total data quantities, and their relationship with the bandwidth K. If the current combination mode cannot guarantee the transmission of higher-precision event data within the bandwidth limit, then the switching condition is met, and the system switches to a combination mode that can guarantee the transmission of higher-precision event data within the bandwidth limit.

[0585] To better understand the above process, let's illustrate it with a specific example:

[0586] Assuming a bandwidth limit of K and a bandwidth adjustment factor of α, in event-stream-based reading mode, when events are represented using polarity information, the bit width H of the data format is H = b. x +b y +b t +b p When representing events using light intensity information, the bit width of the data format is H = b. x +b y +b t +b a Usually 1≤b p a For example, b p Typically 1 to 2 bits, b a Typically, it is 8 to 12 bits.

[0587] In frame-scan-based readout mode, events do not need to represent coordinates and time; events are determined based on the state of each pixel. Assuming the data bit width allocated to each pixel is b in polarity mode... sp In light intensity mode, it is b sa The total number of pixels is M. Assuming a bandwidth limit K = 1000bps, b x =5bit b y =4bit,b​t =10 bits, b p =1 bit, b a =8bit,b sp =1 bit, b sa =8 bits, total number of pixels M=100, bandwidth adjustment factor α=0.9. Assume 10 events are generated in the 1st second, 15 events in the 2nd second, and 30 events in the 3rd second.

[0588] Assuming that in the initial state, the default reading mode is based on event streams, and events are represented using polarity mode.

[0589] The following will refer to the event stream-based reading mode and the event represented by polarity information as asynchronous polarity mode, the event stream-based reading mode and the event represented by light intensity information as asynchronous light intensity mode, the frame scan-based reading mode and the event represented by polarity information as synchronous polarity mode, and the frame scan-based reading mode and the event represented by light intensity information as synchronous light intensity mode.

[0590] Second 1: 10 events are generated.

[0591] Asynchronous polarity mode: N ev =10,H=b x +b y +b t +b p =5 + 4 + 10 + 1 = 20 bits, estimated data size is N ev H = 200 bits, N ev H < α·K, thus satisfying the bandwidth limit.

[0592] Asynchronous light intensity mode: In this mode, H = b x +b y +b t +b a =5+4+10+8=27 bits, therefore the estimated data size in light intensity mode is N. ev H = 270 bits, N ev If H < α·K, the bandwidth limit is still satisfied.

[0593] Synchronous polarity mode: M=100, H=b sp =1 bit, at this time the estimated data volume is M·H = 100 bits, M·H < α·K, which still meets the bandwidth limit.

[0594] Synchronous light intensity mode: M=100, H=b sa =8 bits, at this time the estimated data volume is M·H = 800 bits, M·H < α·K, which still meets the bandwidth limit.

[0595] In summary, the asynchronous light intensity mode was selected in the first second, transmitting the light intensity information of all 10 events with a relatively small data volume (270 bits) without exceeding the bandwidth limit. If the control circuit 906 determines that the current combination mode cannot guarantee the transmission of higher-precision event data while not exceeding the bandwidth limit, then it determines that the switching conditions are met and controls the switch from the asynchronous polarity mode to the asynchronous light intensity mode. For example, an indication signal is sent to instruct the reading circuit 905 to switch from the current event representation mode to another event representation mode.

[0596] Second 2: 15 events are generated.

[0597] Asynchronous polarity mode: Estimated data size is N ev H = 15 × 20 = 300 bits, which meets the bandwidth limit.

[0598] Asynchronous light intensity mode: Estimated data volume is N ev H = 15 × 27 = 405 bits, which meets the bandwidth limit.

[0599] Synchronous polarity mode: The estimated data volume is M·H = 100 × 1 = 100 bits, which meets the bandwidth limit.

[0600] Synchronous optical intensity mode: The estimated data volume is M·H = 100 × 8 = 800 bits, which meets the bandwidth limit.

[0601] In summary, at the 2nd second, if the control circuit 906 determines that the current combination mode can guarantee the transmission of higher precision event data within the bandwidth limit, then it determines that the switching conditions are not met and the asynchronous optical intensity mode is still selected.

[0602] 3rd second: 30 events are generated.

[0603] Asynchronous polarity mode: Estimated data size is N ev H = 30 × 20 = 600 bits, which meets the bandwidth limit.

[0604] Asynchronous light intensity mode: Estimated data volume is N ev H = 30 × 27 = 810 bits, which meets the bandwidth limit.

[0605] Synchronous polarity mode: The estimated data volume is M·H = 100 × 1 = 100 bits, which meets the bandwidth limit.

[0606] Synchronous optical intensity mode: The estimated data volume is M·H = 100 × 8 = 800 bits, which meets the bandwidth limit.

[0607] At the 3rd second, in synchronous optical intensity mode, the optical intensity information of all 30 events can be transmitted with 800 bits of data. If the current combined mode (asynchronous optical intensity mode) cannot guarantee the transmission of higher-precision event data within bandwidth limitations at the 3rd second, and the switching conditions are met, then the control switches from asynchronous optical intensity mode to synchronous optical intensity mode. For example, an indication signal is sent to instruct the reading circuit 905 to switch from the current event reading mode to another event reading mode.

[0608] It should be understood that the formulas, conversion conditions and related calculation methods given above are merely an example implementation of the embodiments of this application. Conversion conditions, conversion strategies and calculation methods for other suitable event representation methods can also be adopted, and the scope of this application is not limited in this respect.

[0609] In some embodiments, the reading circuit 905 includes a data format control unit 9051, used to control the reading circuit to read the signal output from the threshold comparison unit 902, or to read the electrical signal buffered in the light intensity acquisition unit 904. The data format control unit 9051 will be described below with reference to two preferred embodiments, as exemplarily.

[0610] See Figure 12-a This is a schematic diagram of a data format control unit in the reading circuit of this application embodiment. The data format control unit may include AND gates 951 and 954, OR gate 953, and NOT gate 952. The input of AND gate 951 is used to receive the conversion signal sent by the control circuit 906 and the polarity information output by the threshold comparison unit 902. The input of AND gate 954 is used to receive the conversion signal sent by the control circuit 906 after passing through NOT gate 952, and the electrical signal (light intensity information) output by the light intensity acquisition unit 904. The outputs of AND gates 951 and 954 are connected to the input of OR gate 953, and the output of OR gate 953 is coupled to the control circuit 906. In one possible implementation, the conversion signal can be 0 or 1, then the data format control unit 9051 can control the reading of the polarity information output by the threshold comparison unit 902, or control the reading of the light intensity information output by the light intensity acquisition unit 904. For example, if the conversion signal is 0, the data format control unit 9051 can control the polarity information in the output threshold comparison unit 902; if the conversion signal is 1, the data format control unit 9051 can control the light intensity information in the output light intensity acquisition unit 904. In one possible implementation, the data format control unit 9051 can be connected to the control unit 906 via a format signal line, and receive the conversion signal sent by the control unit 906 via the format signal line.

[0611] It should be noted that, Figure 12-aThe data format control unit shown is only one possible structure; other logical structures capable of line switching can also be used in the embodiments of this application. For example... Figure 12-b As shown, the reading circuit 905 may include reading components 955 and 956, wherein reading components 955 and 956 may be implemented using separate devices or integrated into the same device. Reading component 955 can be used to read the data output by the threshold comparison 902, and reading component 956 can be used to read the electrical signal buffered by the light intensity acquisition unit.

[0612] The reading circuit 905 can read the data signal generated by the pixel array circuit in a specific event representation manner. For example, in an example where the control circuit can control the reading component 955 to be turned on and the reading component 956 to be turned off, the reading circuit 905 reads the event represented by polarity information by reading the data output by the interval reading threshold comparison unit 902. In an example where the reading component 956 is turned on and the reading component 955 is turned off, the reading circuit 905 reads the event represented by light intensity information by reading the electrical signal buffered in the light intensity acquisition unit 904.

[0613] Then, it should be noted that in some possible implementations, the reading circuit may also include other circuit structures, such as an analog-to-digital converter unit for converting analog signals into digital signals. For example, it may also include a statistics unit for counting the number N of events measured by the pixel array circuit per unit time. ev For example, it may also include a calculation unit for calculating the total amount of data related to the number of events (light intensity change events) measured by the pixel array circuit per unit time. Furthermore, it should be noted that the connections in this application can represent direct connections or couplings. For example, OR gate 953 and control circuit 906 are connected. In one possible implementation, OR gate 953 and control circuit 906 may be coupled, with OR gate 953 connected to the input of the statistics unit and control circuit 906 connected to the output of the statistics unit.

[0614] According to the method provided in a possible embodiment of this application, the control circuit 906 continuously performs historical statistics and real-time analysis on the light intensity change events generated in the pixel array circuit throughout the entire reading and parsing process. Once the conversion condition is met, a conversion signal is sent to convert the information in the reading threshold comparison unit 902 to the information in the reading light intensity acquisition unit 904, or to convert the information in the reading light intensity acquisition unit 904 to the information in the reading threshold comparison unit 902. This adaptive conversion process is repeated continuously until the reading of all data signals is completed.

[0615] Figure 13A block diagram of a control circuit according to a possible embodiment of this application is shown. The control circuit can be used to implement... Figure 11 , Figure 12-a Control circuits such as 906, etc. Figure 13 As shown, the control circuit includes at least one processor 1101, at least one memory 1102 coupled to the processor 1101, and a communication mechanism 1103 coupled to the processor 1101. The memory 1102 is used to store at least a computer program and data signals acquired from a read circuit. A statistical model 111 and a strategy module 112 are pre-configured on the processor 1101. The control circuit can be communicatively coupled to a computer via the communication mechanism 1103, such as... Figure 11 , Figure 12-a The reading circuit 905 of the vision sensor or the reading circuit outside the vision sensor is used to implement control functions.

[0616] In some possible embodiments, the control circuit can be configured to control the readout circuit 905 to read multiple data signals generated by the pixel array circuit in a specific event representation manner. Additionally, the control circuit can be configured to acquire data signals from the readout circuit 905, and when the control circuit controls the readout circuit 905 to read an event represented by light intensity information, the data signal can indicate an absolute light intensity value, which can represent the light intensity value measured at the current moment. When the control circuit controls the readout circuit 905 to read an event represented by polarity information, the data signal can indicate light intensity polarity, etc. For example, light intensity polarity can indicate a trend of light intensity change, such as increase or decrease, typically represented by +1 / -1.

[0617] The control circuit determines statistical data related to at least one light intensity change event based on data signals acquired from the readout circuit. For example, the statistical data mentioned above could be the total number of events (light intensity change events) measured by the pixel array circuit per unit time, or it could be the number of events N measured by the pixel array circuit per unit time. ev In some embodiments, the control circuit can acquire data signals generated by the pixel array circuit over a period of time from the reading circuit 905 and store these data signals in the memory 1102 for historical statistics and analysis.

[0618] In some possible embodiments, the control circuit may utilize one or more pre-configured statistical models 111 to perform historical statistics on light intensity change events generated by the pixel array circuit provided by the readout circuit 906 over a period of time. The statistical model 111 can then transmit the statistical data to the strategy module 112. As descr...

Claims

1. An image processing method, characterized in that, include: Acquire motion information, which includes information about the motion trajectory of the target object when it moves within the detection range of the motion sensor; At least one event image is generated based on the motion information, and the at least one event image includes an image representing the motion trajectory of the target object when it moves within the detection range; Obtain the target task, and determine the iteration duration based on the target task; The at least one frame of the event image is iteratively updated to obtain the updated at least one frame of the event image, and the duration of the iterative update of the at least one frame of the event image does not exceed the iteration duration. Any one of the iterative updates in the iterative update of the at least one frame of event image includes: Retrieve the value of the preset optimization model from the previous iteration update; The parameters of the inertial measurement unit (IMU) sensor are updated according to the value of the optimization model. The parameters of the IMU sensor are used to collect data. The data collected by the IMU sensor is used to calculate motion parameters, which represent the parameters of the relative motion between the motion sensor and the target object. The target event image in the at least one frame of event image is iteratively updated according to the motion parameters to obtain the updated target event image.

2. The method according to claim 1, characterized in that, The iterative update of the target event image in the at least one frame of event images based on the motion parameters includes: The motion trajectory of the target object in the target event image is compensated according to the motion parameters to obtain the target event image updated in the current iteration.

3. The method according to any one of claims 1-2, characterized in that, The motion parameters include one or more of the following: depth, optical flow information, acceleration of the motion sensor or angular velocity of the motion sensor, wherein the depth represents the distance between the motion sensor and the target object, and the optical flow information represents the relative motion speed between the motion sensor and the target object.

4. The method according to any one of claims 1-2, characterized in that, During any of the said iterative update processes, the method further includes: If the result of the current iteration meets a preset condition, the iteration is terminated. The preset condition includes at least one of the following: the number of iterations of the at least one event image reaches a preset number, or the value change of the optimization model during the update of the at least one event image is less than a preset value.

5. An image processing method, characterized in that, include: At least one event image is generated based on motion information, wherein the motion information includes information on the motion trajectory of the target object when it moves within the detection range of the motion sensor, and the at least one event image includes an image representing the motion trajectory of the target object when it moves within the detection range; Acquire motion parameters, which represent the parameters of the relative motion between the motion sensor and the target object; The values ​​of the preset optimization model are initialized based on the motion parameters to obtain the values ​​of the optimization model; The at least one frame of the event image is updated based on the value of the optimization model to obtain the updated at least one frame of the event image; After initializing the values ​​of the preset optimization model based on the motion parameters, the method further includes: The parameters of the inertial measurement unit (IMU) sensor are updated based on the values ​​of the optimized model. These parameters are used by the IMU sensor to acquire data.

6. The method according to claim 5, characterized in that, The motion parameters include one or more of the following: depth, optical flow information, acceleration of the motion sensor or angular velocity of the motion sensor, wherein the depth represents the distance between the motion sensor and the target object, and the optical flow information represents the relative motion speed between the motion sensor and the target object.

7. The method according to claim 5 or 6, characterized in that, The acquisition of motion parameters includes: Acquire data collected by the inertial measurement unit (IMU) sensors; The motion parameters are calculated based on the data collected by the IMU sensor.

8. An image processing apparatus, characterized in that, include: The acquisition module is used to acquire motion information, which includes information about the motion trajectory of the target object when it moves within the detection range of the motion sensor; The processing module is configured to generate at least one frame of event image based on the motion information, wherein the at least one frame of event image includes an image representing the motion trajectory of the target object when it moves within the detection range; The acquisition module is also used to acquire the target task and acquire the iteration duration based on the target task; The processing module is further configured to iteratively update the at least one frame of event image to obtain an updated at least one frame of event image, and the duration of iteratively updating the at least one frame of event image does not exceed the iteration duration. The processing module is specifically used for: Retrieve the value of the preset optimization model from the previous iteration update; The parameters of the inertial measurement unit (IMU) sensor are updated according to the value of the optimization model. The parameters of the IMU sensor are used to collect data. The data collected by the IMU sensor is used to calculate motion parameters, which represent the parameters of the relative motion between the motion sensor and the target object. The target event image in the at least one frame of event image is iteratively updated according to the motion parameters to obtain the updated target event image.

9. The apparatus according to claim 8, characterized in that, The processing module is specifically used for: The motion trajectory of the target object in the target event image is compensated according to the motion parameters to obtain the target event image updated in the current iteration.

10. The apparatus according to any one of claims 8-9, characterized in that, The motion parameters include one or more of the following: depth, optical flow information, acceleration of the motion sensor or angular velocity of the motion sensor, wherein the depth represents the distance between the motion sensor and the target object, and the optical flow information represents the relative motion speed between the motion sensor and the target object.

11. The apparatus according to any one of claims 8-9, characterized in that, The processing module is further configured to terminate the iteration if the result of the current iteration meets a preset condition during any iteration update process. The preset condition includes at least one of the following: the number of iteration updates of the at least one frame of event image reaches a preset number or the value change of the optimization model during the update of the at least one frame of event image is less than a preset value.

12. An image processing apparatus, characterized in that, include: The processing module is configured to generate at least one frame of event image based on motion information, wherein the motion information includes information on the motion trajectory of the target object when it moves within the detection range of the motion sensor, and the at least one frame of event image includes an image representing the motion trajectory of the target object when it moves within the detection range; An acquisition module is used to acquire motion parameters, which represent the parameters of the relative motion between the motion sensor and the target object; The processing module is also used to initialize the value of the preset optimization model according to the motion parameters to obtain the value of the optimization model; The processing module is further configured to update the at least one frame of the event image according to the value of the optimization model, so as to obtain the updated at least one frame of the event image; The processing module is further configured to, after initializing the value of the preset optimization model according to the motion parameters, update the parameters of the inertial measurement unit (IMU) sensor according to the value of the optimization model, wherein the parameters of the IMU sensor are used for the IMU sensor to acquire data.

13. The apparatus according to claim 12, characterized in that, The motion parameters include one or more of the following: depth, optical flow information, acceleration of the motion sensor or angular velocity of the motion sensor, wherein the depth represents the distance between the motion sensor and the target object, and the optical flow information represents the relative motion speed between the motion sensor and the target object.

14. The apparatus according to claim 12 or 13, characterized in that, The acquisition module is specifically used for: Acquire data collected by the inertial measurement unit (IMU) sensors; The motion parameters are calculated based on the data collected by the IMU sensor.

15. An image processing apparatus comprising a processor and a memory, the processor being coupled to the memory, characterized in that, The memory is used to store programs; The processor is configured to execute a program in the memory, causing the image processing apparatus to perform the method as described in any one of claims 1-4 or 5-7.

16. A computer-readable storage medium comprising a program, characterized in that, When it is run on a computer, it causes the computer to perform the method as described in any one of claims 1-4 or 5-7.

17. A computer program product containing instructions, characterized in that, When it is run on a computer, it causes the computer to perform the method as described in any one of claims 1-4 or 5-7.