An image processing method and apparatus

By selecting adaptive sensors and adaptive event representation methods in electronic devices, the efficiency and accuracy issues of removing moving objects during shooting are solved, and efficient and accurate image processing effects are achieved.

CN116113975BActive Publication Date: 2025-10-10HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080104269.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-31
Publication Date
2025-10-10
Estimated Expiration
2040-12-31

AI Technical Summary

Technical Problem

It is difficult to remove moving objects efficiently and accurately during the shooting process with existing technologies, especially high-speed moving objects, resulting in poor photography effects.

Method used

By selecting RGB sensors and motion sensors in electronic devices, different sensors are activated according to scene information, and the adaptive event representation and decoding circuit of the visual sensor chip are combined to dynamically adjust the data transmission and decoding methods, reduce power consumption and improve image processing efficiency.

Benefits of technology

It achieves efficient and accurate removal of moving objects in different scenarios, reduces the power consumption of electronic devices, and improves the accuracy and efficiency of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116113975B_ABST
    Figure CN116113975B_ABST
Patent Text Reader

Abstract

An image processing method and device are used for efficiently and accurately removing a moving object from a captured image. The method includes: acquiring an event stream and a first RGB image by a camera with a motion sensor (such as a DVS) and an RGB sensor (8201), constructing a mask according to the event stream (8202), and obtaining a second RGB image according to the event stream, the first RGB image and the mask, the second RGB image being a RGB image without a target object (8203). The method can remove a moving object based on only one RGB image and an event stream, thereby obtaining a RGB image without a moving object. Compared with the prior art which needs to remove a moving object by using multiple RGB images and event streams, the method only needs a user to capture one RGB image, and the user experience is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to an image processing method and device. Background Art

[0002] In daily photography, moving objects (referred to as moving foreground) often appear unexpectedly within the frame, affecting the quality of the photo. Currently, there are several methods available on the market for removing moving objects. For example, Lumia phones can remove moving objects in certain scenarios by capturing dynamic photos for a certain duration (e.g., 2 seconds) and stitching them together. However, this method requires high-quality capture timing and a stable capture time (e.g., the 2 seconds mentioned above). Furthermore, the removal effect is poor, and it is unable to identify and remove high-speed moving objects. Therefore, how to efficiently and accurately remove moving objects from captured images has become an urgent problem to be solved. Summary of the Invention

[0003] The embodiments of the present application provide an image processing method and apparatus for efficiently and accurately removing moving objects from captured images.

[0004] In a first aspect, the present application provides a switching method, which is applied to an electronic device, wherein the electronic device includes an RGB sensor and a motion sensor, the RGB (red green bule) sensor being used to capture images within a shooting range, and the motion sensor being used to capture information generated when an object moves relative to the motion sensor within a detection range of the motion sensor. The method includes: selecting at least one of the RGB sensor and the motion sensor based on scene information, and collecting data through the selected sensor, wherein the scene information includes at least one of status information of the electronic device, a type of application requesting image capture in the electronic device, or environmental information.

[0005] Therefore, in the embodiments of the present application, different sensors in the electronic device can be selectively activated according to different scenarios, adapting to more scenarios and having strong generalization capabilities. In addition, the corresponding sensors can be activated according to the actual scenario without activating all sensors, thereby reducing the power consumption of the electronic device.

[0006] In one possible implementation, the status information includes the remaining power and remaining storage capacity of the electronic device; the environmental information includes the change value of the light intensity within the shooting range of the color RGB sensor and the motion sensor or information about the moving objects within the shooting range.

[0007] Therefore, in the embodiments of the present application, the sensors to be activated can be selected according to the status of the electronic device or environmental information, which is adaptable to more scenarios and has strong generalization capabilities.

[0008] In addition, in the following different implementations, the activated sensors may be different. When a certain sensor mentioned above collects data, the sensor has been activated, which will not be described in detail below.

[0009] In a second aspect, the present application provides a visual sensor chip, which may include: a pixel array circuit configured to generate at least one data signal corresponding to a pixel in the pixel array circuit by measuring a change in light intensity, wherein the at least one data signal indicates a light intensity change event, where the light intensity change event indicates that the change in light intensity measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A readout circuit, coupled to the pixel array circuit, is configured to read the at least one data signal from the pixel array circuit using a first event representation. The readout circuit is also configured to provide the at least one data signal to a control circuit. The readout circuit is also configured to, upon receiving a conversion signal generated from the control circuit based on the at least one data signal, switch to reading the at least one data signal from the pixel array circuit using a second event representation. As can be seen from the first aspect, the visual sensor can adaptively switch between the two event representations, ensuring that the readout data rate always remains below a predetermined readout data rate threshold, thereby reducing the cost of data transmission, parsing, and storage in the visual sensor and significantly improving sensor performance. Furthermore, such a visual sensor can generate statistics on events generated within a period of time to predict the likely event rate within the next period of time, thereby selecting a reading mode that is more suitable for the current external environment, application scenario, and motion state.

[0010] In one possible embodiment, the first event representation method is to represent the event through polarity information. The pixel array circuit may include multiple pixels, and each pixel may include a threshold comparison unit. The threshold comparison unit is used to output polarity information when the light intensity change exceeds a predetermined threshold. The polarity information is used to indicate whether the light intensity change is increased or decreased. The reading circuit is specifically used to read the polarity information output by the threshold comparison unit. In this embodiment, the first event representation method is to represent the event through polarity information. The polarity information is usually represented by 1-2 bits and carries less information. This avoids the problem of a sudden increase in events faced by the visual sensor when there is a large amount of data and a large area of ​​object movement or light intensity fluctuations (such as entering and exiting a tunnel, turning on and off lights in a room, etc.). When the preset maximum bandwidth of the visual sensor (hereinafter referred to as bandwidth) is certain, it avoids the situation where event data cannot be read out, resulting in event loss.

[0011] In one possible embodiment, the first event is represented by light intensity information. The pixel array may include multiple pixels, each of which may include a threshold comparison unit, a readout control unit, and a light intensity acquisition unit. The light intensity detection unit is configured to output an electrical signal corresponding to the light signal incident thereon, the electrical signal being used to indicate light intensity. The threshold comparison unit is configured to output a first signal when it determines, based on the electrical signal, that the light intensity change exceeds a predetermined threshold. The readout control unit is configured to, in response to receiving the first signal, instruct the light intensity acquisition unit to acquire and cache the electrical signal corresponding to the moment the first signal was received. The reading circuit is specifically configured to read the electrical signal cached by the light intensity acquisition unit. In this embodiment, the first event is represented by light intensity information. When the amount of data transmitted does not exceed the bandwidth limit, light intensity information is used to represent the event. Typically, light intensity information is represented by multiple bits, such as 8-12 bits. Compared to polarity information, light intensity information can carry more information, which is beneficial for event processing and analysis, such as improving the quality of image reconstruction.

[0012] In one possible embodiment, the control circuit is further used to: determine statistical data based on at least one data signal received from the reading circuit. If it is determined that the statistical data meets the predetermined conversion condition, a conversion signal is sent to the reading circuit, and the predetermined conversion condition is determined based on the preset bandwidth of the visual sensor chip. In this embodiment, a method of converting two event representation methods is provided, and the conversion condition is obtained according to the amount of data to be transmitted. For example, when the amount of data transmitted is large, it is switched to represent the event through polarity information to ensure that the entire amount of data can be transmitted, avoiding the situation where the event data cannot be read out, resulting in the loss of the event. When the amount of data transmitted is small, it is switched to represent the event through light intensity information, so that the transmitted event can carry more information, which is beneficial to the processing and analysis of the event, such as improving the quality of image reconstruction.

[0013] In one possible embodiment, when the first event representation method is to represent an event through light intensity information and the second event representation method is to represent an event through polarity information, the predetermined conversion condition is that the total amount of data read from the pixel array circuit through the first event representation method is greater than the preset bandwidth, or the predetermined conversion condition is that the number of at least one data signal is greater than the ratio of the preset bandwidth to the first bit, where the first bit is a preset bit in the data format of the data signal. In this embodiment, a specific condition for switching from representing an event through light intensity information to representing an event through polarity information is provided. When the amount of data to be transmitted is greater than the preset bandwidth, the event is switched to being represented through polarity information to ensure that the entire amount of data can be transmitted, avoiding the situation where the event data cannot be read, resulting in event loss.

[0014] In a possible implementation, the first event representation is to represent events by polarity information, the second event representation is to represent events by light intensity information, the predetermined conversion condition is that if at least one data signal is read from the pixel array circuit by the second event representation, the total amount of data read is not greater than a preset bandwidth, or the predetermined conversion condition is that the number of the at least one data signal is not greater than a ratio of the preset bandwidth and a first bit, the first bit being a preset bit of a data format of the data signal. In this implementation, a specific condition for switching from representing events by polarity information to representing events by light intensity information is given. When the amount of data transmitted is not greater than the preset bandwidth, the switching to representing events by light intensity information is performed, so that the events transmitted can carry more information, which is beneficial to processing and analysis of the events, for example, the quality of image reconstruction can be improved.

[0015] In a third aspect, the present application provides a decoding circuit, which can include: a reading circuit configured to read a data signal from a visual sensor chip. The decoding circuit is configured to decode the data signal according to a first decoding mode. The decoding circuit is further configured to decode the data signal according to a second decoding mode when a conversion signal is received from a control circuit. The decoding circuit provided in the third aspect corresponds to the visual sensor chip provided in the second aspect, and is configured to decode the data signal output by the visual sensor chip provided in the second aspect. The decoding circuit provided in the third aspect can switch between different decoding modes for different event representations.

[0016] In a possible implementation, the control circuit is further configured to: determine statistical data based on the data signal read from the reading circuit. If it is determined that the statistical data satisfies a predetermined conversion condition, the conversion signal is sent to the encoding circuit, the predetermined conversion condition being determined based on a preset bandwidth of the visual sensor chip.

[0017] In a possible implementation, the first decoding mode is to decode the data signal according to a first bit corresponding to the first event representation, the first event representation is to represent events by light intensity information, the second decoding mode is to decode the data signal according to a second bit corresponding to the second event representation, the second event representation is to represent events by polarity information, the polarity information is used to indicate whether the light intensity change is an increase or a decrease, the conversion condition is that the total amount of data decoded according to the first decoding mode is greater than a preset bandwidth, or the predetermined conversion condition is that the number of the data signal is greater than a ratio of the preset bandwidth and a first bit, the first bit being a preset bit of a data format of the data signal.

[0018] In a possible implementation, the first decoding manner is decoding the data signal according to first bits corresponding to a first event representation manner, the first event representation manner is representing an event by polarity information, the polarity information is used to indicate that the light intensity change amount is increased or decreased, the second decoding manner is decoding the data signal by second bits corresponding to a second event representation manner, the second event representation manner is representing an event by light intensity information, and the conversion condition is that the total data amount is not greater than a preset bandwidth according to the second decoding manner, or the number of data signals is greater than a preset ratio of the preset bandwidth and the first bits, the first bits are preset bits of a data format of the data signal.

[0019] In a fourth aspect, the present application provides a method for operating a visual sensor chip, which can include: generating at least one data signal corresponding to a pixel in a pixel array circuit of the visual sensor chip by measuring a light intensity change amount by the pixel array circuit, the at least one data signal indicating a light intensity change event, the light intensity change event representing that the light intensity change amount measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. Reading the at least one data signal from the pixel array circuit in a first event representation manner by a reading circuit of the visual sensor chip. Providing the at least one data signal to a control circuit of the visual sensor chip by the reading circuit. When a conversion signal generated based on the at least one data signal is received from the control circuit by the reading circuit, converting to read the at least one data signal from the pixel array circuit in a second event representation manner.

[0020] In a possible implementation, the first event representation manner is representing an event by polarity information, and the pixel array circuit can include a plurality of pixels, each pixel can include a threshold comparison unit. Reading the at least one data signal from the pixel array circuit in the first event representation manner by the reading circuit of the visual sensor chip can include: outputting polarity information by the threshold comparison unit when the light intensity change amount exceeds the predetermined threshold, the polarity information being used to indicate that the light intensity change amount is increased or decreased. Reading the polarity information output by the threshold comparison unit by the reading circuit.

[0021] In one possible embodiment, the first event representation method is to represent the event through light intensity information. The pixel array may include multiple pixels, each pixel may include a threshold comparison unit, a readout control unit and a light intensity acquisition unit. The reading circuit of the visual sensor chip reads at least one data signal from the pixel array circuit in the first event representation method, which may include: outputting an electrical signal corresponding to the light signal irradiated thereon by the light intensity acquisition unit, and the electrical signal is used to indicate the light intensity. When the electrical signal determines that the light intensity change amount exceeds a predetermined threshold, the threshold comparison unit outputs a first signal. In response to receiving the first signal, the readout control unit instructs the light intensity acquisition unit to acquire and cache the electrical signal corresponding to the moment the first signal is received. The electrical signal cached by the light intensity acquisition unit is read by the reading circuit.

[0022] In one possible embodiment, the method may further include determining statistical data based on at least one data signal received from the readout circuit, and sending a conversion signal to the readout circuit if the statistical data satisfies a predetermined conversion condition, where the predetermined conversion condition is determined based on a preset bandwidth of the vision sensor chip.

[0023] In one possible embodiment, when the first event representation method is to represent the event through light intensity information and the second event representation method is to represent the event through polarity information, the predetermined conversion condition is that the total amount of data read from the pixel array circuit through the first event representation method is greater than the preset bandwidth, or the predetermined conversion condition is that the number of at least one data signal is greater than the ratio of the preset bandwidth to the first bit, and the first bit is a preset bit of the data format of the data signal.

[0024] In one possible embodiment, when the first event representation method is to represent the event through polarity information and the second event representation method is to represent the event through light intensity information, the predetermined conversion condition is that if at least one data signal is read from the pixel array circuit through the second event representation method, the total amount of data read is not greater than the preset bandwidth, or the predetermined conversion condition is that the number of at least one data signal is not greater than the ratio of the preset bandwidth to the first bit, and the first bit is a preset bit of the data format of the data signal.

[0025] In a fifth aspect, the present application provides a decoding method, comprising: reading a data signal from a visual sensor chip through a reading circuit; decoding the data signal according to a first decoding method through a decoding circuit; and decoding the data signal according to a second decoding method through a decoding circuit when a conversion signal is received from a control circuit.

[0026] In a possible implementation, the method further includes: determining statistical data based on a data signal read from a reading circuit; and sending a conversion signal to an encoding circuit if the statistical data is determined to meet a predetermined conversion condition, wherein the predetermined conversion condition is determined based on a preset bandwidth of the visual sensor chip.

[0027] In one possible embodiment, the first decoding method is to decode the data signal according to the first bit corresponding to the first event representation method, and the first event representation method is to represent the event through light intensity information. The second decoding method is to decode the data signal according to the second bit corresponding to the second event representation method, and the second event representation method is to represent the event through polarity information. The polarity information is used to indicate whether the change in light intensity is enhanced or weakened. The conversion condition is that the total amount of data decoded according to the first decoding method is greater than the preset bandwidth, or the predetermined conversion condition is that the number of data signals is greater than the ratio of the preset bandwidth to the first bit, and the first bit is the preset bit of the data format of the data signal.

[0028] In one possible embodiment, the first decoding method is to decode the data signal according to the first bit corresponding to the first event representation method, the first event representation method is to represent the event through polarity information, and the polarity information is used to indicate whether the change in light intensity is enhanced or weakened. The second decoding method is to decode the data signal through the second bit corresponding to the second event representation method, the second event representation method is to represent the event through light intensity information. The conversion condition is that if the data signal is decoded according to the second decoding method, the total data amount is not greater than the preset bandwidth, or the predetermined conversion condition is that the number of data signals is greater than the ratio of the preset bandwidth to the first bit, and the first bit is the preset bit of the data format of the data signal.

[0029] In a sixth aspect, the present application provides a visual sensor chip, which may include: a pixel array circuit configured to generate at least one data signal corresponding to a pixel in the pixel array circuit by measuring a change in light intensity, wherein the at least one data signal indicates a light intensity change event, wherein the light intensity change event indicates that the change in light intensity measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A first encoding unit configured to encode the at least one data signal based on a first bit to obtain first encoded data. The first encoding unit is further configured to, upon receiving a first control signal from a control circuit, encode the at least one data signal based on a second bit indicated by the first control signal, wherein the first control signal is determined by the control circuit based on the first encoded data. As can be seen from the solution provided in the sixth aspect, by dynamically adjusting the bit width representing light intensity characteristic information, when the event generation rate is low and has not yet reached the bandwidth limit, the events are quantized according to the maximum bit width and encoded. When the event generation rate is high, the bit width representing the light intensity characteristic information is gradually reduced to meet the bandwidth limit. Thereafter, if the event generation rate decreases again, the bit width representing the light intensity characteristic information can be incrementally increased without exceeding the bandwidth limit. Vision sensors can adaptively switch between multiple event representations to better achieve the goal of transmitting all events with greater representation accuracy.

[0030] In a possible implementation, the first control signal is determined by the control circuit according to the first coded data and a preset bandwidth of the vision sensor chip.

[0031] In one possible implementation, when the amount of the first encoded data is not less than the bandwidth, the second bit indicated by the control signal is smaller than the first bit, so that the total amount of data of the at least one data signal encoded by the second bit does not exceed the bandwidth. When the event generation rate is high, the bit width representing the light intensity characteristic information is gradually reduced to meet the bandwidth limitation.

[0032] In one possible implementation, when the amount of the first encoded data is less than the bandwidth, the second bit indicated by the control signal is larger than the first bit, and the total amount of data in the at least one data signal encoded by the second bit does not exceed the bandwidth. If the rate of events decreases, the bit width representing the light intensity characteristic information can be increased incrementally without exceeding the bandwidth limit, thereby better achieving the goal of transmitting all events with greater representation accuracy.

[0033] In one possible embodiment, a pixel array may include N regions, wherein at least two of the N regions have different maximum bit values, where the maximum bit value represents a preset maximum bit value for encoding at least one data signal generated by a region. A first encoding unit is specifically configured to encode the at least one data signal generated by the first region based on a first bit to obtain first encoded data, wherein the first bit value is not greater than the maximum bit value of the first region, and the first region is any one of the N regions. The first encoding unit is specifically configured to, upon receiving a first control signal from a control circuit, encode the at least one data signal generated by the first region based on a second bit indicated by the first control signal, wherein the first control signal is determined by the control circuit based on the first encoded data. In this embodiment, the pixel array may also be divided into regions, and different weights may be used to set the maximum bit widths of different regions to accommodate different regions of interest in a scene. For example, a larger weight may be assigned to a region that may contain a target object, thereby increasing the accuracy of the representation of events outputted in the region containing the target object, while a smaller weight may be assigned to a background region, thereby decreasing the accuracy of the representation of events outputted in the background region.

[0034] In one possible implementation, the control circuit is further configured to: upon determining that the total data volume of the at least one data signal encoded by the third bit exceeds the bandwidth and the total data volume of the at least one data signal encoded by the second bit does not exceed the bandwidth, send a first control signal to the first encoding unit, wherein the difference between the third bit and the second bit is one bit. In this implementation, all events can be transmitted with greater representation accuracy without exceeding the bandwidth limit.

[0035] In a seventh aspect, the present application provides a decoding device, which may include: a reading circuit for reading a data signal from a visual sensor chip. A decoding circuit for decoding the data signal according to a first bit. The decoding circuit is also used to decode the data signal according to a second bit indicated by the first control signal when receiving a first control signal from the control circuit. The decoding circuit provided in the seventh aspect corresponds to the visual sensor chip provided in the sixth aspect, and is used to decode the data signal output by the visual sensor chip provided in the sixth aspect. The decoding circuit provided in the seventh aspect can dynamically adjust the decoding method according to the coding bits used by the visual sensor.

[0036] In a possible implementation, the first control signal is determined by the control circuit according to the first coded data and a preset bandwidth of the vision sensor chip.

[0037] In a possible implementation manner, when the total data amount of the data signal decoded according to the first bit is not less than the bandwidth, the second bit is smaller than the first bit.

[0038] In a possible implementation, when the total data amount of the data signal decoded according to the first bit is less than the bandwidth, the second bit is greater than the first bit, and the total data amount of the data signal decoded according to the second bit is not greater than the bandwidth.

[0039] In one possible implementation, a reading circuit is specifically configured to read a data signal corresponding to a first region from a vision sensor chip, where the first region is any one of N regions that may be included in a pixel array of the vision sensor, and at least two of the N regions have different maximum bits, where the maximum bit represents a preset maximum bit for encoding at least one data signal generated by a region. A decoding circuit is specifically configured to decode the data signal corresponding to the first region based on the first bit.

[0040] In one possible embodiment, the control circuit is further used to: when it is determined that the total data amount of the data signal decoded by the third bit is greater than the bandwidth and the total data amount of the data signal decoded by the second bit is not greater than the bandwidth, send a first control signal to the first encoding unit, and the difference between the third bit and the second bit is 1 bit unit.

[0041] In an eighth aspect, the present application provides a method for operating a visual sensor chip, which may include: measuring a change in light intensity through a pixel array circuit of the visual sensor chip to generate at least one data signal corresponding to a pixel in the pixel array circuit, wherein the at least one data signal indicates a light intensity change event, and the light intensity change event indicates that the change in light intensity measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A first encoding unit of the visual sensor chip encodes the at least one data signal according to a first bit to obtain first encoded data. When the first encoding unit receives a first control signal from a control circuit of the visual sensor chip, the at least one data signal is encoded according to a second bit indicated by the first control signal, and the first control signal is determined by the control circuit based on the first encoded data.

[0042] In a possible implementation, the first control signal is determined by the control circuit according to the first coded data and a preset bandwidth of the vision sensor chip.

[0043] In a possible implementation, when the data volume of the first encoded data is not less than the bandwidth, the second bit indicated by the control signal is smaller than the first bit, so that the total data volume of at least one data signal encoded by the second bit is not greater than the bandwidth.

[0044] In a possible implementation, when the data volume of the first encoded data is smaller than the bandwidth, the second bit indicated by the control signal is larger than the first bit, and the total data volume of the at least one data signal encoded by the second bit is not larger than the bandwidth.

[0045] In one possible embodiment, a pixel array may include N regions, wherein at least two of the N regions have different maximum bits, the maximum bit representing a preset maximum bit for encoding at least one data signal generated by a region. Encoding the at least one data signal based on a first bit by a first encoding unit of a visual sensor chip may include: encoding the at least one data signal generated by a first region based on the first bit by the first encoding unit to obtain first encoded data, wherein the first bit is not greater than the maximum bit of the first region, and the first region is any region of the N regions. Encoding the at least one data signal based on a second bit indicated by the first control signal upon receiving a first control signal from a control circuit of the visual sensor chip by the first encoding unit may include: encoding the at least one data signal generated by the first region based on the second bit indicated by the first control signal upon receiving the first control signal from the control circuit by the first encoding unit, wherein the first control signal is determined by the control circuit based on the first encoded data.

[0046] In a possible implementation, the method may further include: when determining that the total data amount of at least one data signal encoded by the third bit is greater than the bandwidth and the total data amount of at least one data signal encoded by the second bit is not greater than the bandwidth, sending a first control signal to the first encoding unit through the control circuit, with the difference between the third bit and the second bit being 1 bit unit.

[0047] In a ninth aspect, the present application provides a decoding method, which may include: reading a data signal from a visual sensor chip via a reading circuit; decoding the data signal based on a first bit via a decoding circuit; and upon receiving a first control signal from a control circuit via the decoding circuit, decoding the data signal based on a second bit indicated by the first control signal.

[0048] In a possible implementation, the first control signal is determined by the control circuit according to the first coded data and a preset bandwidth of the vision sensor chip.

[0049] In a possible implementation manner, when the total data amount of the data signal decoded according to the first bit is not less than the bandwidth, the second bit is smaller than the first bit.

[0050] In a possible implementation, when the total data amount of the data signal decoded according to the first bit is less than the bandwidth, the second bit is greater than the first bit, and the total data amount of the data signal decoded according to the second bit is not greater than the bandwidth.

[0051] In one possible implementation, reading a data signal from a visual sensor chip using a readout circuit may include: reading, using the readout circuit, a data signal corresponding to a first region from the visual sensor chip, where the first region is any one of N regions that may be included in a pixel array of the visual sensor, at least two of the N regions having different maximum bits, the maximum bit representing a preset maximum bit for encoding at least one data signal generated by a region. Decoding, using a decoding circuit, the data signal based on the first bit may include: decoding, using the decoding circuit, the data signal corresponding to the first region based on the first bit.

[0052] In a possible embodiment, the method may further include: when it is determined that the total data amount of the data signal decoded by the third bit is greater than the bandwidth, and the total data amount of the data signal decoded by the second bit is not greater than the bandwidth, sending a first control signal to the first encoding unit, and the difference between the third bit and the second bit is 1 bit unit.

[0053] In a tenth aspect, the present application provides a visual sensor chip, which may include: a pixel array circuit configured to generate multiple data signals corresponding to multiple pixels in the pixel array circuit by measuring light intensity changes, the multiple data signals indicating at least one light intensity change event, wherein the at least one light intensity change event indicates that the light intensity change measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A third encoding unit configured to encode a first differential value based on a first preset bit, wherein the first differential value is the difference between the light intensity change and the predetermined threshold. Reducing the precision of event representation, i.e., reducing the bit width used to represent an event, reduces the amount of information that the event can carry, which is detrimental to event processing and analysis in some scenarios. Therefore, reducing the precision of event representation may not be suitable for all scenarios. In other words, in some scenarios, a high bit width is required to represent an event. However, while events represented by a high bit width can carry more data, the data volume is also larger. Given a preset maximum bandwidth for the visual sensor, it is possible that event data cannot be read, resulting in data loss. The solution provided in the tenth aspect adopts a method of encoding differential values, which reduces the cost of data transmission, analysis and storage of the visual sensor while transmitting events with the highest accuracy, significantly improving the performance of the sensor.

[0054] In one possible embodiment, the pixel array circuit may include multiple pixels, each of which may include a threshold comparison unit. The threshold comparison unit is configured to output polarity information when the light intensity change exceeds a predetermined threshold. The polarity information is used to indicate whether the light intensity change is increasing or decreasing. The third encoding unit is further configured to encode the polarity information based on a second preset bit. In this embodiment, the polarity information can also be encoded to indicate whether the light intensity is increasing or decreasing, which helps to obtain the current light intensity information based on the light intensity signal and polarity information obtained by the previous decoding.

[0055] In one possible embodiment, each pixel may also include a light intensity detection unit, a readout control unit and a light intensity collection unit. The light intensity detection unit is used to output an electrical signal corresponding to the light signal irradiated thereon, and the electrical signal is used to indicate the light intensity. The threshold comparison unit is specifically used to output polarity information when it determines that the light intensity change exceeds a predetermined threshold based on the electrical signal. The readout control unit is used to instruct the light intensity collection unit to collect and cache the electrical signal corresponding to the moment of receiving the polarity information in response to receiving the polarity signal. The third encoding unit is also used to encode the first electrical signal according to a third preset bit. The first electrical signal is the electrical signal collected by the light intensity collection unit at the first moment of receiving the corresponding polarity information. The third preset bit is the maximum bit preset by the visual sensor for representing the characteristic information of the light intensity. After the initial state is fully encoded, subsequent events only need to encode the polarity information and the difference between the light intensity change and the predetermined threshold, which can effectively reduce the amount of encoded data. Among them, full encoding refers to encoding an event using the maximum bit width predefined by the visual sensor. In addition, the light intensity information of the previous event, as well as the decoded polarity information and differential value, can be used to losslessly reconstruct the light intensity information at the current moment.

[0056] In one possible implementation, the third encoding unit is further configured to encode the electrical signal collected by the light intensity collection unit according to a third preset bit at every preset time interval. Full encoding is performed every preset time interval to reduce decoding dependency and prevent bit errors.

[0057] In a possible implementation manner, the third encoding unit is specifically configured to: when the first differential value is smaller than a predetermined threshold, encode the first differential value according to a first preset bit.

[0058] In a possible implementation, the third encoding unit is further used to: when the first differential value is not less than a predetermined threshold, encode the first residual differential value and the predetermined threshold according to a first preset bit, where the first residual differential value is the difference between the differential value and the predetermined threshold.

[0059] In one possible embodiment, the third encoding unit is specifically configured to: when the first residual differential value is not less than a predetermined threshold, encode a second residual differential value according to a first preset bit, where the second residual differential value is the difference between the first residual differential value and the predetermined threshold. The predetermined threshold is encoded for a first time according to the first preset bit. The predetermined threshold is encoded for a second time according to the first preset bit. Because visual sensors may have a certain delay, it may occur that the light intensity change is greater than the predetermined threshold twice or more times before an event is generated. This may result in the differential value being greater than or equal to the predetermined threshold, and the light intensity change being at least twice the predetermined threshold. For example, if the first residual differential value is not less than the predetermined threshold, the second residual differential value is encoded. If the second residual differential value is still not less than the predetermined threshold, a third residual differential value may be encoded, where the third differential value is the difference between the second residual differential value and the predetermined threshold. The predetermined threshold is then encoded for a third time, and the above process is repeated until the residual differential value is less than the predetermined threshold.

[0060] In the eleventh aspect, the present application provides a decoding device, which may include: an acquisition circuit for reading a data signal from a visual sensor chip. A decoding circuit for decoding the data signal according to the first bit to obtain a differential value, wherein the differential value is less than a predetermined threshold value, and the differential value is the difference between the light intensity change measured by the visual sensor and the predetermined threshold value. When the light intensity change exceeds the predetermined threshold value, the visual sensor generates at least one light intensity change event. The decoding circuit provided in the eleventh aspect corresponds to the visual sensor chip provided in the tenth aspect, and is used to decode the data signal output by the visual sensor chip provided in the tenth aspect. The decoding circuit provided in the eleventh aspect can adopt a corresponding differential decoding method for the differential encoding method adopted by the visual sensor.

[0061] In a possible implementation, the decoding circuit is further configured to: decode the data signal according to the second bit to obtain polarity information, where the polarity information is used to indicate whether the light intensity change is increased or decreased.

[0062] In one possible embodiment, the decoding circuit is also used to: decode the data signal received at the first moment according to the third bit to obtain an electrical signal corresponding to the light signal output by the visual sensor and irradiated thereon, wherein the third bit is the maximum bit preset by the visual sensor for representing characteristic information of light intensity.

[0063] In a possible implementation manner, the decoding circuit is further configured to: decode the data signal received at the first moment according to the third bit at every preset time duration.

[0064] In a possible implementation manner, the decoding circuit is specifically configured to: decode the data signal according to the first bit to obtain a differential value and at least one predetermined threshold.

[0065] In a twelfth aspect, the present application provides a method for operating a visual sensor chip, which may include: measuring light intensity changes via a pixel array circuit of the visual sensor chip to generate multiple data signals corresponding to multiple pixels in the pixel array circuit, the multiple data signals indicating at least one light intensity change event, the at least one light intensity change event indicating that the light intensity change measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. Encoding a first differential value using a third encoding unit of the visual sensor chip according to a first preset bit, the first differential value being the difference between the light intensity change and the predetermined threshold.

[0066] In one possible implementation, the pixel array circuit may include multiple pixels, each pixel may include a threshold comparison unit, and the method may further include: when a light intensity change exceeds a predetermined threshold, outputting polarity information via the threshold comparison unit, the polarity information indicating whether the light intensity change is an increase or decrease; and encoding the polarity information via a third encoding unit according to a second predetermined bit.

[0067] In a possible embodiment, each pixel may also include a light intensity detection unit, a readout control unit and a light intensity collection unit, and the method may further include: outputting an electrical signal corresponding to the light signal irradiated thereon through the light intensity detection unit, the electrical signal being used to indicate the light intensity. Outputting polarity information through the threshold comparison unit may include: outputting polarity information through the threshold comparison unit when determining, based on the electrical signal, that the light intensity conversion amount exceeds a predetermined threshold. The method may further include: in response to receiving the polarity signal, instructing the light intensity collection unit through the readout control unit to collect and cache the electrical signal corresponding to the moment of receiving the polarity information. The first electrical signal is encoded according to a third preset bit, the first electrical signal being the electrical signal collected by the light intensity collection unit at the first moment of receiving the corresponding polarity information, and the third preset bit being the maximum bit preset by the visual sensor for representing the characteristic information of light intensity.

[0068] In a possible implementation, the method may further include: encoding the electrical signal collected by the light intensity collection unit according to a third preset bit at every preset time length.

[0069] In a possible implementation, encoding the first differential value according to a first preset bit by a third encoding unit of the visual sensor chip may include: encoding the first differential value according to the first preset bit when the first differential value is less than a predetermined threshold.

[0070] In a possible implementation, encoding the first differential value according to a first preset bit by a third encoding unit of the visual sensor chip may also include: when the first differential value is not less than a predetermined threshold, encoding the first residual differential value and the predetermined threshold according to the first preset bit, where the first residual differential value is the difference between the differential value and the predetermined threshold.

[0071] In one possible implementation, when the first difference value is not less than a predetermined threshold, encoding the first residual difference value and the predetermined threshold according to a first preset bit may include: when the first residual difference value is not less than the predetermined threshold, encoding a second residual difference value according to the first preset bit, where the second residual difference value is the difference between the first residual difference value and the predetermined threshold; encoding the predetermined threshold for a first time according to the first preset bit; and encoding the predetermined threshold for a second time according to the first preset bit, where the first residual difference value may include the second residual difference value and two predetermined thresholds.

[0072] In a thirteenth aspect, the present application provides a decoding method, which may include: reading a data signal from a visual sensor chip via an acquisition circuit. The decoding circuit decodes the data signal based on a first bit to obtain a differential value, wherein the differential value is less than a predetermined threshold, and the differential value is the difference between a light intensity change measured by the visual sensor and the predetermined threshold. When the light intensity change exceeds the predetermined threshold, the visual sensor generates at least one light intensity change event.

[0073] In a possible implementation, the method may further include: decoding the data signal according to the second bit to obtain polarity information, where the polarity information is used to indicate whether the light intensity change is increased or decreased.

[0074] In a possible embodiment, it may also include: decoding the data signal received at the first moment according to the third bit to obtain an electrical signal corresponding to the light signal output by the visual sensor and irradiated thereon, the third bit being the maximum bit preset by the visual sensor for representing characteristic information of light intensity.

[0075] In a possible implementation manner, the method may further include: decoding the data signal received at the first moment according to the third bit at every preset time duration.

[0076] In a possible implementation, decoding the data signal according to the first bit by a decoding circuit to obtain a differential value may include: decoding the data signal according to the first bit to obtain the differential value and at least one predetermined threshold.

[0077] In a fourteenth aspect, the present application provides an image processing method, comprising: obtaining motion information, the motion information comprising information of a motion trajectory of a target object when the target object moves within a detection range of a motion sensor; generating at least one frame of event image according to the motion information, the at least one frame of event image being an image representing the motion trajectory of the target object when the target object moves within the detection range; obtaining a target task, and obtaining an iteration duration according to the target task; iteratively updating the at least one frame of event image to obtain updated at least one frame of event image, and the duration of iteratively updating the at least one frame of event image does not exceed the iteration duration.

[0078] Therefore, in the embodiments of the present application, the object in motion can be monitored by the motion sensor, and the information of the motion trajectory of the object when the object moves within the detection range can be collected by the motion sensor. After obtaining the target task, the iteration duration can be determined according to the target task, and the event image matching the target task can be obtained by iteratively updating the event image within the iteration duration.

[0079] In a possible implementation, any one of the iteratively updating the at least one frame of event image comprises: obtaining a motion parameter, the motion parameter representing a parameter of relative motion between the motion sensor and the target object; iteratively updating a target event image in the at least one frame of event image according to the motion parameter to obtain an updated target event image.

[0080] Therefore, in the embodiments of the present application, when iteratively updating the event image, the event image can be updated based on the parameter of relative motion between the object and the motion sensor, so that the event image is compensated to obtain a clearer event image.

[0081] In a possible implementation, the obtaining the motion parameter comprises: obtaining a value of a preset optimization model in a previous iteration update process; and calculating the motion parameter according to the value of the optimization model.

[0082] Therefore, in the embodiments of the present application, the event image can be updated based on the value of the optimization model, and the motion parameter can be calculated according to the optimization model to obtain a better motion parameter, and then the event image is updated using the motion parameter to obtain a clearer event image.

[0083] In a possible implementation, the iteratively updating the target event image according to the motion parameter comprises: compensating the motion trajectory of the target object in the target event image according to the motion parameter to obtain a target event image obtained in a current iteration update.

[0084] Therefore, in the embodiment of the present application, the motion parameters can be used to compensate for the motion trajectory of the target object in the event image, so that the motion trajectory of the target object in the event image is clearer, and thus the event image is clearer.

[0085] In a possible embodiment, the motion parameters include one or more of the following: depth, optical flow information, acceleration of the motion sensor, or angular velocity of the motion sensor, the depth represents the distance between the motion sensor and the target object, and the optical flow information represents information on the relative motion speed between the motion sensor and the target object.

[0086] Therefore, in the embodiment of the present application, motion compensation can be performed on the target object in the event image using a variety of motion parameters to improve the clarity of the event image.

[0087] In one possible embodiment, during any iterative update process, the method further includes: terminating the iteration if the result of the current iteration meets a preset condition, and the termination condition includes at least one of the following: the number of iterative updates of the at least one frame of event image reaches a preset number or the value change of the optimization model during the update of the at least one frame of event image is less than a preset value.

[0088] Therefore, in the embodiment of the present application, in addition to setting the iteration duration, convergence conditions related to the number of iterations or the value of the optimization model can also be set, so as to obtain an event image that meets the convergence conditions under the constraint of the iteration duration.

[0089] In the fifteenth aspect, the present application provides an image processing method, including: generating at least one frame of event image based on motion information, the motion information including information of a motion trajectory of a target object when it moves within a detection range of a motion sensor, and the at least one frame of event image is an image representing the motion trajectory of the target object when it moves within the detection range; obtaining motion parameters, the motion parameters representing parameters of relative motion between the motion sensor and the target object; initializing the value of a preset optimization model according to the motion parameters to obtain the value of the optimization model; updating the at least one frame of event image according to the value of the optimization model to obtain the updated at least one frame of event image.

[0090] In an embodiment of the present application, the parameters of the relative motion between the motion sensor and the target object can be used to initialize the optimization model, thereby reducing the number of initial iterations of the event image, accelerating the convergence speed of the iteration of the event image, and obtaining a clearer event image with fewer iterations.

[0091] In a possible embodiment, the motion parameters include one or more of the following: depth, optical flow information, acceleration of the motion sensor, or angular velocity of the motion sensor, the depth represents the distance between the motion sensor and the target object, and the optical flow information represents information on the relative motion speed between the motion sensor and the target object.

[0092] In one possible implementation, obtaining the motion parameters includes: obtaining data collected by an inertial measurement unit (IMU) sensor; and calculating the motion parameters based on the data collected by the IMU sensor. Therefore, in the implementation of the present application, the motion parameters can be calculated using the IMU, thereby obtaining more accurate motion parameters.

[0093] In a possible implementation, after initializing the value of the preset optimization model according to the motion parameters, the method further includes: updating the parameters of the IMU sensor according to the value of the optimization model, and the parameters of the IMU sensor are used for the IMU sensor to collect data.

[0094] Therefore, in the embodiment of the present application, the parameters of the IMU can also be updated according to the value of the optimization model to achieve correction of the IMU and make the data collected by the IMU more accurate.

[0095] In a sixteenth aspect, the present application provides an image processing device having the function of implementing the method of the fourteenth aspect or any possible implementation of the fourteenth aspect, or the image processing device having the function of implementing the method of the fifteenth aspect or any possible implementation of the fifteenth aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0096] In the seventeenth aspect, the present application provides a method for image processing, comprising: obtaining motion information, the motion information including information on a motion trajectory of a target object when it moves within a detection range of a motion sensor; generating an event image based on the motion information, the event image being an image representing the motion trajectory of the target object when it moves within the detection range; obtaining a first reconstructed image based on at least one event included in the event image, wherein a first pixel point and at least one second pixel point have different color types, the first pixel point is a pixel point corresponding to any one of the at least one events in the first reconstructed image, and the at least one second pixel point is included in a plurality of pixels adjacent to the first pixel point in the first reconstructed image.

[0097] Therefore, in the embodiment of the present application, when there is relative motion between the photographed object and the motion sensor, image reconstruction can be performed based on the data collected by the motion sensor to obtain a reconstructed image. Even if the RGB sensor does not capture clearly, a clear image can be obtained.

[0098] In a possible embodiment, determining the color type corresponding to each pixel in the event image based on at least one event included in the event image to obtain a first reconstructed image includes: scanning each pixel in the event image along a first direction, determining the color type corresponding to each pixel in the event image, and obtaining a first reconstructed image, wherein if the first pixel is scanned to have an event, the color type of the first pixel is determined to be the first color type, and if the second pixel arranged before the first pixel along the first direction does not have an event, the color type corresponding to the second pixel is the second color type, the first color type and the second color type are different color types, and the pixel with an event indicates the pixel in the event image corresponding to the position where a change is detected by the motion sensor.

[0099] In the embodiment of the present application, the event image can be scanned and image reconstruction can be performed based on the events of each pixel in the event image, thereby obtaining a clearer event image. Therefore, in the embodiment of the present application, the information collected by the motion sensor can be used to reconstruct the image, and the reconstructed image can be obtained efficiently and quickly, thereby improving the efficiency of subsequent image recognition, image classification, etc. on the reconstructed image. Even if a clear RGB image cannot be captured in some scenes such as shooting moving objects or there is shooting jitter, the image can be reconstructed by using the information collected by the motion sensor, and a clearer image can be reconstructed quickly and accurately to facilitate subsequent recognition or classification tasks.

[0100] In one possible implementation, the first direction is a preset direction, or the first direction is determined based on data collected by an IMU, or the first direction is determined based on an image captured by a color RGB camera. Therefore, in the implementation of the present application, the direction of the scanned event image can be determined in a variety of ways to accommodate a wider range of scenarios.

[0101] In one possible implementation, if a plurality of consecutive third pixels arranged after the first pixel in the first direction do not have an event, the color type corresponding to the plurality of third pixels is the first color type. Therefore, in the implementation of the present application, when there are multiple consecutive pixels without an event, the color types corresponding to the consecutive pixels are the same, thus avoiding the situation where edges are unclear due to the movement of the same object in the actual scene.

[0102] In a possible embodiment, if the fourth pixel point arranged behind the first pixel point and adjacent to the first pixel point in the first direction has an event, and the fifth pixel point arranged behind the fourth pixel point and adjacent to the fourth pixel point in the first direction does not have an event, then the color types corresponding to the fourth pixel point and the fifth pixel point are both the first color type.

[0103] Therefore, when there are at least two consecutive pixels in the event image that both have events, the reconstructed color type may not be changed when the second event is scanned, thereby avoiding unclear edges of the reconstructed image due to the target object's edges being too wide.

[0104] In a possible embodiment, after scanning each pixel in the event image in a first direction, determining the color type corresponding to each pixel in the event image, and obtaining a first reconstructed image, the method further includes: scanning the event image in a second direction, determining the color type corresponding to each pixel in the event image, and obtaining a second reconstructed image, wherein the second direction is different from the first direction; and fusing the first reconstructed image and the second reconstructed image to obtain an updated first reconstructed image.

[0105] In the embodiment of the present application, the event image may be scanned in different directions to obtain multiple reconstructed images from multiple directions, and then the multiple reconstructed images may be fused to obtain a more accurate reconstructed image.

[0106] In a possible implementation, the method further includes: if the first reconstructed image does not meet preset requirements, updating motion information, updating the event image according to the updated motion information, and obtaining an updated first reconstructed image according to the updated event image.

[0107] In the embodiment of the present application, the event image may be updated in combination with the information collected by the motion sensor, so that the updated event image is clearer.

[0108] In a possible implementation, before the color type corresponding to each pixel point in the event image is determined according to at least one event included in the event image, and a first reconstructed image is obtained, the method further includes: compensating the event image according to a motion parameter when the target object and the motion sensor perform relative motion, to obtain a compensated event image, the motion parameter including one or more of the following: depth, optical flow information, acceleration of the motion sensor performing motion, or angular velocity of the motion sensor performing motion, the depth representing a distance between the motion sensor and the target object, and the optical flow information representing information of a motion speed of relative motion between the motion sensor and the target object.

[0109] Therefore, in the embodiments of the present application, the event image can also be compensated according to the motion parameter, so that the event image is clearer, and the reconstructed image obtained through reconstruction is also clearer.

[0110] In a possible implementation, the color type of a pixel point in the reconstructed image is determined according to the color captured by the color RGB camera. In the embodiments of the present application, the color in the actual scene can be determined according to the RGB camera, so that the color of the reconstructed image matches the color in the actual scene, and the user experience is improved.

[0111] In a possible implementation, the method further includes: obtaining an RGB image according to data captured by the RGB camera; and fusing the RGB image and the first reconstructed image to obtain an updated first reconstructed image. Therefore, in the embodiments of the present application, the RGB image and the reconstructed image can be fused, so that the reconstructed image obtained finally is clearer.

[0112] In an eighteenth aspect, the present application further provides an image processing apparatus having a function of implementing the method of the eighteenth aspect or any one of the possible implementation manners of the eighteenth aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0113] In the nineteenth aspect, the present application provides an image processing method, including: obtaining a first event image (event image) and multiple captured first images, the first event image including information of objects moving within a preset range within a shooting time period of the multiple first images, the exposure times corresponding to the multiple first images are different, and the preset range is the shooting range of the camera; calculating a first jitter degree corresponding to each of the multiple first images according to the first event image, the first jitter degree being used to indicate the degree of camera jitter when shooting the multiple first images; determining a fusion weight of each of the multiple first images according to the first jitter degree corresponding to each first image, wherein the first jitter degree corresponding to the multiple first images and the fusion weight are negatively correlated; fusing the multiple first images according to the fusion weight of each first image to obtain a target image.

[0114] Therefore, in the embodiment of the present application, the event image can be used to quantify the degree of jitter when shooting the RGB image, and the fusion weight of each RGB image can be determined according to the degree of jitter of each RGB image. Generally, the fusion weight corresponding to the RGB image with a low degree of jitter is higher, so that the information included in the target image obtained in the end is more inclined to the clearer RGB image, thereby obtaining a clearer target image. Generally, the RGB image with a higher degree of jitter has a smaller corresponding weight value, and the RGB image with a lower degree of jitter has a larger corresponding weight value, so that the information included in the target image obtained in the end is more inclined to the information included in the clearer RGB image, making the target image obtained in the end clearer and improving the user experience. And if the target image is used for subsequent image recognition or feature extraction, the recognition result or extracted feature is also more accurate.

[0115] In a possible embodiment, before determining the fusion weight of each first image in the multiple first images based on the first jitter degree, the method further includes: if the first jitter degree is not higher than a first preset value and higher than a second preset value, de-jittering each first image to obtain each de-jittered first image.

[0116] Therefore, in the embodiment of the present application, the jitter conditions can be distinguished based on dynamic data, and direct fusion can be performed when there is no jitter. When the jitter is not strong, the RGB image can be adaptively de-jittered. When the jitter is strong, an additional RGB image can be taken. Scenarios with a variety of jitter levels are used, and the generalization ability is strong.

[0117] In a possible embodiment, determining the fusion weight of each first image in the multiple first images based on the first jitter degree includes: if the first jitter degree is higher than a first preset value, re-shooting to obtain a second image, and the second jitter degree of the second image is not higher than the first preset value; calculating the fusion weight of each first image based on the first jitter degree of each first image, and calculating the fusion weight of the second image based on the second jitter degree; fusing the multiple first images based on the fusion weight of each first image to obtain a target image includes: fusing the multiple first images and the second image based on the fusion weight of each first image and the fusion weight of the second image to obtain the target image.

[0118] Generally, RGB images with higher levels of jitter have smaller corresponding weights, while RGB images with lower levels of jitter have larger corresponding weights. This results in the information included in the final target image being more inclined towards that included in the clearer RGB image, making the final target image clearer and improving the user experience. Furthermore, if this target image is used for subsequent image recognition or feature extraction, the recognition results or extracted features will also be more accurate. For RGB images with higher levels of jitter, additional RBG images can be taken to obtain a clearer RGB image with lower levels of jitter. This allows for subsequent image fusion to use the clearer image, further enhancing the final target image.

[0119] In a possible embodiment, before reshooting the second image, the method further includes: acquiring a second event image, where the second event image is acquired before acquiring the first event image; and calculating exposure parameters based on information included in the second event image, where the exposure parameters are used to shoot the second image.

[0120] Therefore, in the embodiment of the present application, the exposure strategy is adaptively adjusted using the information collected by the dynamic perception camera (i.e., the motion sensor), that is, the high dynamic range perception characteristics of the texture within the shooting range are utilized by the dynamic perception information to adaptively supplement the shooting of images with appropriate exposure time, thereby improving the camera's ability to capture texture information in strong light areas or dark light areas.

[0121] In a possible embodiment, the reshooting to obtain the second image also includes: dividing the first event image into multiple areas, and dividing the third image into multiple areas, the third image is the first image with the smallest exposure value among the multiple first images, and the multiple areas included in the first event image correspond to the positions of the multiple areas included in the third image, and the exposure value includes at least one of exposure duration, exposure amount or exposure level; calculating whether each area in the first event image includes first texture information, and whether each area in the third image includes second texture information; if the first area in the first event image includes the first texture information, and the area corresponding to the first area in the third image does not include the second texture information, then shooting according to the exposure parameters to obtain the second image, and the first area is any area in the first dynamic area.

[0122] Therefore, in the embodiments of the present application, if a region in the first dynamic region contains texture information, and the same region in the RGB image with the minimum exposure value does not contain texture information, this indicates that the region in the RGB image is highly blurred, and a retake of the RGB image is necessary. If no region in the first event image contains texture information, then retake of the RGB image is not necessary.

[0123] In the twentieth aspect, the present application provides an image processing method, comprising: first, detecting motion information of a target object, which motion information may include information on a motion trajectory of the target object when it moves within a preset range, and the preset range is the shooting range of a camera; then, determining focus information based on the motion information, the focus information including parameters for focusing on the target object within the preset range; subsequently, focusing on the target object within the preset range according to the focus information, and capturing an image of the preset range.

[0124] Therefore, in the embodiments of the present application, the motion trajectory of the target object within the camera's shooting range can be detected, and then the focus information is determined and the focus is completed based on the target object's motion trajectory, thereby capturing a clearer image. Even if the target object is in motion, the target object can be accurately focused and clear images of the motion state can be captured, thereby improving the user experience.

[0125] In a possible implementation, the determining the focus information according to the motion information can include: predicting a motion trajectory of the target object in a preset time period according to the motion information, i.e., information of the motion trajectory of the target object when the target object moves in a preset range, to obtain a predicted region, the predicted region being a region in which the target object is located in the preset time period and which is predicted; determining a focus region according to the predicted region, the focus region including at least one focus point for focusing on the target object, and the focus information including position information of the at least one focus point.

[0126] Therefore, in the embodiments of the present application, the future motion trajectory of the target object can be predicted, and the focus region can be determined according to the predicted region, so that the focusing on the target object can be accurately completed. Even if the target object is in high-speed motion, the embodiments of the present application can focus on the target object in advance through prediction, so that the target object is in the focus region, and a clearer high-speed motion target object can be photographed.

[0127] In a possible implementation, the determining the focus region according to the predicted region can include: if the predicted region meets a preset condition, determining the predicted region as the focus region; and if the predicted region does not meet the preset condition, re-predicting the motion trajectory of the target object in the preset time period according to the motion information to obtain a new predicted region, and determining the focus region according to the new predicted region. The preset condition can be that the predicted region includes a complete target object, or that the area of the predicted region is greater than a preset value.

[0128] Therefore, in the embodiments of the present application, the focus region is determined according to the predicted region only when the predicted region meets the preset condition, and the camera is triggered to photograph, and when the predicted region does not meet the preset condition, the camera is not triggered to photograph, so that the target object in the photographed image can be complete, or meaningless photographing can be avoided. Moreover, when no photographing is performed, the camera can be in an unstarted state, and the camera is triggered to photograph only when the predicted region meets the preset condition, so that the power consumption of the camera can be reduced.

[0129] In a possible implementation, the motion information further includes at least one of a motion direction and a motion speed of the target object; and the predicting the motion trajectory of the target object in the preset time period according to the motion information to obtain the predicted region can include: predicting the motion trajectory of the target object in the preset time period according to the motion trajectory of the target object when the target object moves in the preset range, and the motion direction and / or the motion speed.

[0130] Therefore, in the embodiments of the present application, the motion trajectory of the target object within a preset time period in the future can be predicted based on the motion trajectory of the target object within a preset range, as well as the motion direction and / or motion speed, so that the area where the target object is located within the preset time period in the future can be accurately predicted, and the target object can be focused more accurately, and a clearer image can be captured.

[0131] In a possible embodiment, the above-mentioned prediction of the motion trajectory of the target object within a preset time period based on the motion trajectory, motion direction and / or motion speed of the target object when it moves within a preset range to obtain a predicted area may include: fitting a variation function of the center point of the area where the target object is located over time based on the motion trajectory, motion direction and / or motion speed of the target object when it moves within a preset range; then calculating the predicted center point based on the variation function, the predicted center point being the predicted center point of the area where the target object is located within the preset time period; and obtaining the predicted area based on the predicted center point.

[0132] Therefore, in the implementation mode of the present application, the change function of the center point of the area where the target object is located over time can be fitted according to the motion trajectory of the target object when it moves, and then the center point of the area where the target object is located at a certain moment in the future can be predicted based on the change function. The predicted area is determined based on the center point, and the target object can be focused more accurately, so that a clearer image can be captured.

[0133] In one possible embodiment, the image of the predicted range can be captured by an RGB camera, and the above-mentioned focusing on the target object in a preset range based on the focus information may include: focusing at least one point among the multiple focus points of the RGB camera that has the smallest norm distance to the center point of the focus area as the focus point.

[0134] Therefore, in the embodiment of the present application, at least one point having the shortest norm distance to the center point of the focus area may be selected as the focus point, and focusing may be performed to complete focusing on the target object.

[0135] In one possible embodiment, the motion information includes the current area where the target object is located. The above-mentioned determination of focus information based on the motion information may include: determining the current area where the target object is located as the focus area, the focus area includes at least one focus point for focusing on the target object, and the focus information includes position information of at least one focus point.

[0136] Therefore, in an embodiment of the present application, the information on the motion trajectory of the target object within a preset range may include the area where the target object is currently located and the area where the target object has historically located. The area where the target object is currently located can be used as the focus area to complete the focus on the target object, and a clearer image can be captured.

[0137] In a possible implementation, before capturing an image within a preset range, the method may further include: acquiring exposure parameters; and capturing an image within the preset range may include: capturing the image within the preset range according to the exposure parameters.

[0138] Therefore, in the embodiment of the present application, the exposure parameters can also be adjusted, so that the shooting is completed through the exposure parameters to obtain a clear image.

[0139] In a possible implementation, the above-mentioned acquisition of exposure parameters may include: determining the exposure parameters based on motion information, wherein the exposure parameters include exposure duration, the motion information includes the motion speed of the target object, and the exposure duration is negatively correlated with the motion speed of the target object.

[0140] Therefore, in the embodiments of the present application, the exposure duration can be determined by the target object's speed, so that the exposure duration matches the target object's speed. For example, the faster the speed, the shorter the exposure duration, and the slower the speed, the longer the exposure duration. This can avoid overexposure or underexposure, thereby allowing for clearer images to be captured later and improving the user experience.

[0141] In a possible implementation, the above-mentioned acquisition of exposure parameters may include: determining the exposure parameters according to light intensity, wherein the exposure parameters include exposure duration, and the light intensity within a preset range is negatively correlated with the exposure duration.

[0142] Therefore, in the embodiment of the present application, the exposure time can be determined based on the detected light intensity. When the light intensity is greater, the exposure time is shorter, and when the light intensity is smaller, the exposure time is longer, thereby ensuring an appropriate amount of exposure and capturing clearer images.

[0143] In a possible implementation, after capturing images within a preset range, the method may further include: fusing the images within the preset range based on information about the monitored target object and the corresponding motion of the image, to obtain a target image within the preset range.

[0144] Therefore, in an embodiment of the present application, while capturing an image, the movement of the target object within a preset range can also be monitored to obtain information corresponding to the movement of the target object in the image, such as the outline of the target object, the position of the target object within the preset range, and other information. The captured image can be enhanced using this information to obtain a clearer target image.

[0145] In a possible implementation, the detecting of the motion information of the target object within the preset range may include: monitoring the motion of the target object within the preset range by a dynamic vision sensor (DVS) to obtain the motion information.

[0146] Therefore, in the embodiment of the present application, the DVS can be used to monitor the moving objects within the shooting range of the camera, thereby obtaining accurate motion information. Even if the target object is in a high-speed moving state, the DVS can capture the motion information of the target object in a timely manner.

[0147] In a twenty-first aspect, the present application further provides an image processing device, which has the function of implementing the method of the nineteenth aspect or any possible implementation of the nineteenth aspect, or the image processing device has the function of implementing the method of the twentieth aspect or any possible implementation of the twentieth aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0148] In aspect 22, an embodiment of the present application provides a graphical user interface (GUI), characterized in that the graphical user interface is stored in an electronic device, the electronic device includes a display screen, a memory, and one or more processors, the one or more processors are used to execute one or more computer programs stored in the memory, the graphical user interface includes: responding to a trigger operation to shoot a target object, and shooting an image of a preset range according to focus information, and displaying the image of the preset range, the preset range is the camera shooting range, the focus information includes parameters for focusing on the target object within the preset range, the focus information is determined based on the motion information of the target object, and the motion information includes information on the motion trajectory of the target object when it moves within the preset range.

[0149] The beneficial effects produced by the 22nd aspect and any possible implementation method of the 22nd aspect can be referred to the description of the 20th aspect and any possible implementation method of the 20th aspect.

[0150] In a possible implementation, the graphical user interface can further include: predicting a motion trajectory of the target object in a preset time period in response to the motion information, obtaining a prediction region, the prediction region being a region in which the target object is located in the preset time period predicted, and determining the focus region according to the prediction region, displaying the focus region in the display screen, the focus region including at least one focus point for focusing on the target object, and the focus information including position information of the at least one focus point.

[0151] In a possible implementation, the graphical user interface can specifically include: if the prediction region meets a preset condition, displaying the focus region in the display screen in response to determining the focus region according to the prediction region; and if the prediction region does not meet the preset condition, displaying the focus region in the display screen in response to predicting a new prediction region in the preset time period in response to the motion information, and determining the focus region according to the new prediction region.

[0152] In a possible implementation, the motion information further includes at least one of a motion direction and a motion speed of the target object; and the graphical user interface can specifically include: predicting the motion trajectory of the target object in the preset time period in response to the motion trajectory of the target object when moving in a preset range, and the motion direction and / or the motion speed, and displaying the prediction region in the display screen.

[0153] In a possible implementation, the graphical user interface can specifically include: fitting a change function of a center point of a region in which the target object is located over time in response to the motion trajectory of the target object when moving in a preset range, and the motion direction and / or the motion speed, calculating a prediction center point from the change function, the prediction center point being a center point of the region in which the target object is located predicted, and obtaining the prediction region from the prediction center point, and displaying the prediction region in the display screen.

[0154] In a possible implementation, an image of the prediction region is captured by an RGB camera, and the graphical user interface can specifically include: displaying, in the display screen, an image captured based on at least one point with the smallest norm distance from the center point of the focus region in a plurality of focus points of the RGB camera as a focus point.

[0155] In one possible embodiment, the motion information includes the current area where the target object is located, and the graphical user interface may specifically include: in response to taking the current area where the target object is located as the focus area, the focus area includes at least one focus point for focusing on the target object, and the focus information includes position information of at least one focus point, and the focus area is displayed on the display screen.

[0156] In a possible embodiment, the graphical user interface may also include: in response to the information of the monitored movement of the target object and the image, fusing the images within the preset range to obtain the target image within the preset range, and displaying the target image on the display screen.

[0157] In a possible implementation, the motion information is obtained by monitoring the motion of the target object within the preset range through a dynamic vision sensor DVS.

[0158] In a possible embodiment, the graphical user interface may specifically include: in response to obtaining exposure parameters before taking the image of the preset range, displaying the exposure parameters on a display screen; in response to taking the image of the preset range according to the exposure parameters, displaying the image of the preset range taken according to the exposure parameters on a display screen.

[0159] In a possible implementation, an exposure parameter is determined according to the motion information, and the exposure parameter includes an exposure duration, and the exposure duration is negatively correlated with the motion speed of the target object.

[0160] In one possible embodiment, the exposure parameter is determined based on the light intensity, which can be the light intensity detected by a camera or the light intensity detected by a motion sensor. The exposure parameter includes the exposure duration, and the light intensity within the preset range is negatively correlated with the exposure duration.

[0161] In a twenty-third aspect, the present application provides an image processing method, which comprises: first, acquiring an event stream and a first RGB image by a camera with a motion sensor (e.g., a DVS) and an RGB sensor, wherein the acquired event stream comprises at least one event image, each event image in the at least one event image is generated from the motion trail information of a target object (i.e., a moving object) moving within the monitoring range of the motion sensor, and the first RGB image is the superposition of the scene at each time captured by the camera within the exposure time. After the event stream and the first RGB image are acquired, a mask can be constructed according to the event stream, the mask is used to determine the motion region of each event image in the event stream, i.e., to determine the position of the moving object in the RGB image. After the event stream, the first RGB image and the mask are obtained according to the above steps, a second RGB image can be obtained according to the event stream, the first RGB image and the mask, the second RGB image being the RGB image with the target object removed.

[0162] In the above-mentioned embodiments of the present application, the moving object can be removed based on only one RGB image and event stream, so as to obtain an RGB image without moving object. Compared with the prior art which requires multiple RGB images and event streams to remove the moving object, only one RGB image needs to be captured by the user, and the user experience is better.

[0163] In a possible implementation, before the mask is constructed according to the event stream, the method can further comprise: when the motion sensor detects a motion mutation in the monitoring range at a first time, triggering the camera to capture a third RGB image; and the second RGB image is obtained according to the event stream, the first RGB image and the mask, which can comprise: the second RGB image is obtained according to the event stream, the first RGB image, the third RGB image and the mask.

[0164] In the above-mentioned embodiments of the present application, whether the motion data collected by the motion sensor has a motion mutation can be determined, and when the motion mutation exists, the third RGB image is triggered to be captured by the camera, and then the event stream and the first RGB image are obtained in the above-mentioned similar manner, the mask is constructed according to the event stream, and finally the second RGB image without the motion foreground is obtained according to the event stream, the first RGB image, the third RGB image and the mask. The third RGB image obtained is automatically captured by the camera triggered under the condition of the motion mutation, and has high sensitivity, so that a frame of image can be obtained at the beginning of the user's perception of the change of the motion object, and the motion object can be removed more effectively based on the third RGB image and the first RGB image.

[0165] In a possible implementation, the motion sensor monitoring the motion mutation in the monitoring range at the first time includes: the overlapping part between the generated area of the first event stream collected by the motion sensor at the first time and the generated area of the second event stream collected by the motion sensor at the second time in the monitoring range is less than a preset value.

[0166] In the above-mentioned embodiments of the present application, the determination condition of the motion mutation is specifically described, and is feasible.

[0167] In a possible implementation, the way of constructing the mask according to the event stream can be: first, the monitoring range of the motion sensor can be divided into a plurality of preset neighborhoods (set as neighborhood k), and then in each neighborhood k, when the number of event images of the event stream in the preset time length Δt exceeds the threshold value P, the corresponding neighborhood is determined as a motion region, which can be marked as 0, and if the number of event images of the event stream in the preset time length Δt does not exceed the threshold value P, the corresponding neighborhood is determined as a background region, which can be marked as 1.

[0168] In the above-mentioned embodiments of the present application, a method of constructing the mask is specifically described, which is simple and easy to operate.

[0169] In a twenty-fourth aspect, the present application further provides an image processing device having a function of implementing the method of the above-mentioned twenty-second aspect or any one of the possible implementation manners of the twenty-second aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-mentioned functions.

[0170] In the twenty-fifth aspect, the present application provides a pose estimation method, which is applied to a simultaneous localization and mapping (SLAM) scenario. The method includes: the terminal acquires a first event image and a first RGB image, the first event image is time-series aligned with the first target image, and the first target image includes an RGB image or a depth image. The first event image is an image representing the motion trajectory of the target object when it moves within the detection range of the motion sensor. The terminal determines the integration time of the first event image. If the integration time is less than the first threshold, the terminal determines not to perform pose estimation through the first target image. The terminal performs pose estimation based on the first event image.

[0171] In this solution, when the terminal determines that the current scene is difficult for the RGB camera to collect effective environmental information based on the integration time of the event image being less than the threshold, the terminal determines not to perform pose estimation through the RGB image of poor quality to improve the accuracy of pose estimation.

[0172] Optionally, in one possible implementation, the method further includes: determining the acquisition time of the first event image and the acquisition time of the first target image; and determining that the first event image is time-series aligned with the first target image based on a time difference between the acquisition time of the first target image and the acquisition time of the first event image being less than a second threshold. The second threshold can be determined based on the accuracy of SLAM and the frequency of RGB image acquisition by the RGB camera. For example, the second threshold can be 5 milliseconds or 10 milliseconds.

[0173] Optionally, in a possible implementation, acquiring the first event image includes: acquiring N consecutive DVS events; integrating the N consecutive DVS events into a first event image; the method further includes: determining an acquisition time of the first event image based on the acquisition time of the N consecutive DVS events.

[0174] Optionally, in one possible implementation, determining the integration time of the first event image includes: determining N consecutive DVS events for integration into the first event image; and determining the integration time of the first event image based on the acquisition times of the first and last DVS events among the N consecutive DVS events. Because the first event image is obtained by integrating the N consecutive DVS events, the terminal can determine the acquisition time of the first event image based on the acquisition times corresponding to the N consecutive DVS events, that is, determining the acquisition time of the first event image as the time period from the acquisition of the first to the acquisition of the last DVS event among the N consecutive DVS events.

[0175] Optionally, in one possible implementation, the method further includes: acquiring a second event image, where the second event image is an image representing a motion trajectory of the target object when the target object moves within the detection range of the motion sensor. The time period during which the motion sensor detects and acquires the first event image is different from the time period during which the motion sensor detects and acquires the second event image. If there is no RGB image that is time-series aligned with the second event image, it is determined that the second event image does not have an RGB image for jointly performing pose estimation; and pose estimation is performed based on the second event image.

[0176] Optionally, in a possible implementation, before determining the posture based on the second event image, the method also includes: if it is determined that the second event image has time-aligned inertial measurement unit IMU data, then determining the posture based on the second event image and the IMU data corresponding to the second event image; if it is determined that the second event image does not have time-aligned inertial measurement unit IMU data, then determining the posture only based on the second event image.

[0177] Optionally, in a possible implementation, the method further includes: acquiring a second target image, the second target image including an RGB image or a depth image; if there is no event image time-series aligned with the second target image, determining that the second target image does not have an event image for jointly performing pose estimation; and determining the pose based on the second target image.

[0178] Optionally, in one possible implementation, the method further includes: performing loop detection based on the first event image and a dictionary, where the dictionary is a dictionary constructed based on the event image. That is, before performing loop detection, the terminal may pre-construct a dictionary based on the event image so that loop detection can be performed based on the dictionary during the loop detection process.

[0179] Optionally, in a possible implementation, the method further includes: acquiring a plurality of event images, wherein the plurality of event images are event images for training, and the plurality of event images may be event images taken by the terminal in different scenarios. Acquiring visual features of the plurality of event images, which may include, for example, features such as texture, pattern or grayscale statistics of the image. Clustering the visual features using a clustering algorithm to obtain clustered visual features, wherein the clustered visual features have corresponding descriptors. By clustering the visual features, similar visual features can be grouped into one category to facilitate subsequent visual feature matching. Finally, constructing the dictionary based on the clustered visual features.

[0180] Optionally, in a possible implementation, performing loop detection based on the first event image and a dictionary includes: determining a descriptor of the first event image; determining visual features corresponding to the descriptor of the first event image in the dictionary; determining a bag-of-words vector corresponding to the first event image based on the visual features; determining the similarity between the bag-of-words vector corresponding to the first event image and bag-of-words vectors of other event images to determine the event image matched by the first event image.

[0181] In aspect 26, the present application provides a key frame selection method, including: obtaining an event image; determining first information of the event image, the first information including events and / or features in the event image; if it is determined based on the first information that the event image satisfies at least a first condition, then determining that the event image is a key frame, the first condition is related to the number of events and / or the number of features.

[0182] In this solution, by determining the number of events, event distribution, number of features and / or feature distribution in the event image, it is determined whether the current event image is a key frame, which can achieve rapid selection of key frames. The algorithm is small and can meet the needs of rapid key frame selection in scenarios such as video analysis, video encoding and decoding, or security monitoring.

[0183] Optionally, in a possible implementation, the first condition includes: the number of events in the event image is greater than a first threshold, the number of event valid areas in the event image is greater than a second threshold, the number of features in the event image is greater than a third threshold, and the feature valid area in the event image is greater than a fourth threshold. One or more of these.

[0184] Optionally, in a possible implementation, the method further includes: acquiring a depth image time-aligned with the event image; if it is determined based on the first information that the event image satisfies at least a first condition, determining that the event image and the depth image are key frames.

[0185] Optionally, in a possible implementation, the method further includes: obtaining an RGB image that is time-aligned with the event image; obtaining the number of features and / or the feature valid area of ​​the RGB image; if it is determined based on the first information that the event image satisfies at least the first condition, and the number of features of the RGB image is greater than a fifth threshold and / or the number of feature valid areas of the RGB image is greater than a sixth threshold, then determining that the event image and the RGB image are key frames.

[0186] Optionally, in a possible implementation, if it is determined based on the first information that the event image satisfies at least a first condition, then the event image is determined to be a key frame, including: if it is determined based on the first information that the event image satisfies at least the first condition, then second information of the event image is determined, the second information including motion features and / or posture features in the event image; if it is determined based on the second information that the event image satisfies at least a second condition, then the event image is determined to be a key frame, and the second condition is related to the motion change amount and / or posture change amount.

[0187] Optionally, in a possible implementation, the method further includes: determining the clarity and / or brightness consistency index of the event image; if it is determined based on the second information that the event image at least meets the second condition, and the clarity of the event image is greater than the clarity threshold and / or the brightness consistency index of the event image is greater than a preset index threshold, then determining that the event image is a key frame.

[0188] Optionally, in a possible implementation, determining the brightness consistency index of the event image includes: if the pixels in the event image represent the polarity of light intensity changes, calculating the absolute value of the difference between the number of events in the event image and the number of events in the adjacent key frame, and dividing the absolute value by the number of pixels of the event image to obtain the brightness consistency index of the event image; if the pixels in the event image represent light intensity, taking the difference between the event image and the adjacent key frame pixel by pixel, and calculating the absolute value of the difference, summing the absolute values ​​corresponding to each group of pixels, and dividing the summed result by the number of pixels to obtain the brightness consistency index of the event image.

[0189] Optionally, in a possible implementation, the method further includes: obtaining an RGB image that is time-aligned with the event image; determining the clarity and / or brightness consistency index of the RGB image; if it is determined based on the second information that the event image satisfies at least the second condition, and the clarity of the RGB image is greater than a clarity threshold and / or the brightness consistency index of the RGB image is greater than a preset index threshold, then determining that the event image and the RGB image are key frames.

[0190] Optionally, in a possible implementation, the second condition includes: the distance between the event image and the previous key frame exceeds a preset distance value, the rotation angle between the event image and the previous key frame exceeds a preset angle value, and the distance between the event image and the previous key frame exceeds a preset distance value and the rotation angle between the event image and the previous key frame exceeds a preset angle value, or one or more of the following.

[0191] In aspect 27, the present application provides a pose estimation method, comprising: obtaining a first event image and a target image corresponding to the first event image, the first event image and the image capturing the same environmental information, the target image comprising a depth image or an RGB image; determining a first motion area in the first event image; determining a corresponding second motion area in the image based on the first motion area; and performing pose estimation based on the second motion area in the image.

[0192] In this solution, the dynamic area in the scene is captured by the event image, and the posture is determined based on the dynamic area, so that the posture information can be determined precisely.

[0193] Optionally, in a possible implementation, determining the first motion area in the first event image includes: if the dynamic vision sensor DVS that captures the first event image is stationary, obtaining pixel points in the first event image that have event responses; and determining the first motion area based on the pixel points that have event responses.

[0194] Optionally, in a possible implementation, determining the first motion area based on the pixel points that have event responses includes: determining a contour formed by pixel points that have event responses in the first event image; if the area enclosed by the contour is greater than a first threshold, determining that the area enclosed by the contour is the first motion area.

[0195] Optionally, in a possible implementation, determining the first motion area in the first event image includes: if the DVS that captures the first event image is moving, obtaining a second event image, where the second event image is an event image of the previous frame of the first event image; calculating the displacement size and displacement direction of the pixel in the first event image relative to the second event image; if the displacement direction of the pixel in the first event image is different from the displacement direction of the surrounding pixels, or the difference between the displacement size of the pixel in the first event image and the displacement size of the surrounding pixels is greater than a second threshold, determining that the pixel belongs to the first motion area.

[0196] Optionally, in a possible implementation, the method further includes: determining a corresponding static area in the image according to the first motion area; and determining a posture according to the static area in the image.

[0197] In aspect 28, the present application further provides a data processing device, which has the function of implementing the method of aspect 25 or any possible implementation of aspect 25, or the data processing device has the function of implementing the method of aspect 26 or any possible implementation of aspect 26, or the data processing device has the function of implementing the method of aspect 27 or any possible implementation of aspect 27. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0198] In a twenty-ninth aspect, an embodiment of the present application provides a device comprising: a processor and a memory, wherein the processor and the memory are interconnected via a circuit, and the processor invokes program code in the memory to execute the processing-related functions of the method described in any one of aspects 1 to 27 above. Optionally, the device may be a chip.

[0199] In aspect 30, the present application provides an electronic device comprising: a display module, a processing module and a storage module.

[0200] The display module is used to display a graphical user interface of an application stored in the storage module. The graphical user interface may be any of the aforementioned graphical user interfaces.

[0201] In aspect 31, an embodiment of the present application provides a device, which may also be referred to as a digital processing chip or chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to perform processing-related functions in any optional implementation of aspect 1 to aspect 27 above.

[0202] In aspect 32, an embodiment of the present application provides a computer-readable storage medium comprising instructions, which, when executed on a computer, enables the computer to execute a method in any optional implementation of aspect 1 to aspect 27 above.

[0203] In aspect 33, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute a method in any optional implementation of aspect 1 to aspect 27 above. BRIEF DESCRIPTION OF THE DRAWINGS

[0204] Figure 1A A schematic diagram of the system architecture provided for this application;

[0205] Figure 1B A schematic diagram of the structure of an electronic device provided in this application;

[0206] Figure 2 Another system architecture diagram provided for this application;

[0207] Figure 3-a A diagram showing the relationship between the amount of data read and time in an asynchronous reading mode based on event streams.

[0208] Figure 3-b A schematic diagram showing the relationship between the amount of data read and time in a synchronous reading mode based on frame scanning;

[0209] Figure 4-a A block diagram of a visual sensor provided in this application;

[0210] Figure 4-b A block diagram of another visual sensor provided for this application;

[0211] Figure 5 Schematic diagram of the principles of a synchronous reading mode based on frame scanning and an asynchronous reading mode based on event stream according to an embodiment of the present application;

[0212] Figure 6-a is a schematic diagram of a visual sensor according to an embodiment of the present application operating in a frame scanning-based reading mode;

[0213] Figure 6-b is a schematic diagram of a visual sensor operating in an event stream-based reading mode according to an embodiment of the present application;

[0214] Figure 6-c is a schematic diagram of a visual sensor operating in an event stream-based reading mode according to an embodiment of the present application;

[0215] Figure 6-d is a schematic diagram of a visual sensor according to an embodiment of the present application operating in a frame scanning-based reading mode;

[0216] Figure 7 is a flowchart of a method for operating a visual sensor chip according to a possible embodiment of the present application;

[0217] Figure 8 A block diagram of a control circuit provided in this application;

[0218] Figure 9 A block diagram of an electronic device provided in this application;

[0219] Figure 10 Schematic diagram of data volume changes over time in a single data reading mode and in an adaptive switching reading mode according to a possible embodiment of the present application;

[0220] Figure 11 A schematic diagram of a pixel circuit provided in this application;

[0221] Figure 11-a Schematic diagram of representing an event by light intensity information and representing an event by polarity information;

[0222] Figure 12-a A structural diagram of a data format control unit in a reading circuit in this application;

[0223] Figure 12-b This is another structural diagram of the data format control unit in the reading circuit in this application;

[0224] Figure 13 A block diagram of another control circuit provided by this application;

[0225] Figure 14 A block diagram of another control circuit provided by this application;

[0226] Figure 15 A block diagram of another control circuit provided by this application;

[0227] Figure 16 A block diagram of another control circuit provided by this application;

[0228] Figure 17 A block diagram of another control circuit provided by this application;

[0229] Figure 18 Schematic diagram showing the difference between a single event representation and an adaptive conversion event representation provided by the present application;

[0230] Figure 19 A block diagram of another electronic device provided for this application;

[0231] Figure 20 is a flowchart of a method for operating a visual sensor chip according to a possible embodiment of the present application;

[0232] Figure 21 A schematic diagram of another pixel circuit provided in this application;

[0233] Figure 22 A schematic diagram of a coding method provided in this application;

[0234] Figure 23 A block diagram of another visual sensor provided for this application;

[0235] Figure 24 A schematic diagram of dividing a pixel array into regions;

[0236] Figure 25 A block diagram of another control circuit provided by this application;

[0237] Figure 26 A block diagram of another electronic device provided for this application;

[0238] Figure 27 A schematic diagram of a binary data stream;

[0239] Figure 28 is a flowchart of a method for operating a visual sensor chip according to a possible embodiment of the present application;

[0240] Figure 29-a A block diagram of another visual sensor provided for this application;

[0241] Figure 29-b A block diagram of another visual sensor provided for this application;

[0242] Figure 29-c A block diagram of another visual sensor provided for this application;

[0243] Figure 30 A schematic diagram of another pixel circuit provided in this application;

[0244] Figure 31 A schematic block diagram of a third encoding unit provided in this application;

[0245] Figure 32 A flowchart of another encoding method provided for this application;

[0246] Figure 33 A block diagram of another electronic device provided for this application;

[0247] Figure 34 is a flowchart of a method for operating a visual sensor chip according to a possible embodiment of the present application;

[0248] Figure 35 An event diagram provided for this application;

[0249] Figure 36 A schematic diagram of events at a certain moment provided for this application;

[0250] Figure 37 A schematic diagram of a motion area provided for this application;

[0251] Figure 38 A flowchart of an image processing method provided in this application;

[0252] Figure 39 A flowchart of another image processing method provided in this application;

[0253] Figure 40 A flowchart of another image processing method provided in this application;

[0254] Figure 41 A schematic diagram of an event image provided for this application;

[0255] Figure 42 A flowchart of another image processing method provided in this application;

[0256] Figure 43 A flowchart of another image processing method provided in this application;

[0257] Figure 44 A flowchart of another image processing method provided in this application;

[0258] Figure 45 A flowchart of an image processing method provided in this application;

[0259] Figure 46A Another event image diagram provided for this application;

[0260] Figure 46B Another event image diagram provided for this application;

[0261] Figure 47A Another event image diagram provided for this application;

[0262] Figure 47B Another event image diagram provided for this application;

[0263] Figure 48 A flowchart of another image processing method provided by this application;

[0264] Figure 49 A flowchart of another image processing method provided by this application;

[0265] Figure 50 Another event image diagram provided for this application;

[0266] Figure 51 A schematic diagram of a reconstructed image provided in this application;

[0267] Figure 52 A flowchart of an image processing method provided in this application;

[0268] Figure 53 A schematic diagram of a motion trajectory fitting method provided in this application;

[0269] Figure 54 A schematic diagram of a method for determining focus provided in this application;

[0270] Figure 55 A schematic diagram of a method for determining a prediction center provided in this application;

[0271] Figure 56 A flowchart of another image processing method provided by the present application;

[0272] Figure 57 A schematic diagram of a shooting range provided by the present application;

[0273] Figure 58 A schematic diagram of a prediction area provided by the present application;

[0274] Figure 59 A schematic diagram of a focusing area provided by the present application;

[0275] Figure 60 A flowchart of another image processing method provided by the present application;

[0276] Figure 61 A schematic diagram of an image enhancement mode provided by the present application;

[0277] Figure 62 A flowchart of another image processing method provided by the present application;

[0278] Figure 63 A flowchart of another image processing method provided by the present application;

[0279] Figure 64 A schematic diagram of a scene applied by the present application;

[0280] Figure 65 A schematic diagram of another scene applied by the present application;

[0281] Figure 66 A display schematic diagram of a GUI provided by the present application;

[0282] Figure 67 A display schematic diagram of another GUI provided by the present application;

[0283] Figure 68 A display schematic diagram of another GUI provided by the present application;

[0284] Figure 69A A display schematic diagram of another GUI provided by the present application;

[0285] Figure 69B A display schematic diagram of another GUI provided by the present application;

[0286] Figure 69C A display schematic diagram of another GUI provided by the present application;

[0287] Figure 70 A display schematic diagram of another GUI provided by the present application;

[0288] Figure 71 A schematic diagram of another GUI display provided by this application;

[0289] Figure 72A A schematic diagram of another GUI display provided by this application;

[0290] Figure 72B A schematic diagram of another GUI display provided by this application;

[0291] Figure 73 A flowchart of another image processing method provided in this application;

[0292] Figure 74 This is a schematic diagram of an RGB image with low jitter provided by this application;

[0293] Figure 75 This is a schematic diagram of an RGB image with a high degree of dithering provided by this application;

[0294] Figure 76 A schematic diagram of an RGB image in a high-light-ratio scene provided by this application;

[0295] Figure 77 Another event image diagram provided for this application;

[0296] Figure 78 A schematic diagram of an RGB image provided in this application;

[0297] Figure 79 Another RGB image schematic provided for this application;

[0298] Figure 80 Another GUI diagram provided for this application;

[0299] Figure 81 A schematic diagram of the relationship between the photosensitive unit and the pixel value provided by this application;

[0300] Figure 82 A flowchart of the image processing method provided in this application;

[0301] Figure 83 A schematic diagram of the event flow provided for this application;

[0302] Figure 84 A schematic diagram of a blurred image obtained by superimposing exposures of multiple shooting scenes provided in this application;

[0303] Figure 85 A schematic diagram of the mask provided for this application;

[0304] Figure 86A schematic diagram of constructing a mask provided by the present application;

[0305] Figure 87 An effect diagram of removing moving objects from image I to obtain image I' provided by the present application;

[0306] Figure 88 A flowchart of removing moving objects from image I to obtain image I' provided by the present application;

[0307] Figure 89 A schematic diagram of a small motion of a moving object during photographing provided by the present application;

[0308] Figure 90 A schematic diagram of triggering the camera to capture a third RGB image provided by the present application;

[0309] Figure 91 A schematic diagram of an image B captured by the camera based on motion abruptness provided by the present application k And a schematic diagram of an image I actively captured by the user within a certain exposure time;

[0310] Figure 92 A flowchart of obtaining a second RGB image without moving objects based on a first RGB image and an event stream E provided by the present application;

[0311] Figure 93 A flowchart of obtaining a second RGB image without moving objects based on a first RGB image, a third RGB image and an event stream E provided by the present application;

[0312] Figure 94A Another GUI schematic diagram provided by the present application;

[0313] Figure 94B Another GUI schematic diagram provided by the present application;

[0314] Figure 95 A comparison schematic diagram of scenes captured by a traditional camera and a DVS provided by the present application

[0315] Figure 96 A comparison schematic diagram of scenes captured by a traditional camera and a DVS provided by the present application;

[0316] Figure 97 An outdoor navigation schematic diagram applied with a DVS provided by the present application;

[0317] Figure 98a A station navigation schematic diagram applied with a DVS provided by the present application;

[0318] Figure 98bA schematic diagram of a scenic spot navigation using DVS provided in this application;

[0319] Figure 99 A schematic diagram of a shopping mall navigation using DVS provided in this application;

[0320] Figure 100 A schematic diagram of a SLAM process provided by this application;

[0321] Figure 101 A flowchart of a pose estimation method 10100 provided in this application;

[0322] Figure 102 A schematic diagram of integrating DVS events into event images provided by this application;

[0323] Figure 103 A flowchart of a key frame selection method 10300 provided in this application;

[0324] Figure 104 A schematic diagram of the area division of an event image provided in this application;

[0325] Figure 105 A flowchart of a key frame selection method 10500 provided in this application;

[0326] Figure 106 A flowchart of a pose estimation method 1060 provided in this application;

[0327] Figure 107 A schematic diagram of a process for performing pose estimation based on a static area of ​​an image provided by this application;

[0328] Figure 108a A schematic diagram of a process for performing pose estimation based on a motion region of an image provided by this application;

[0329] Figure 108b A schematic diagram of a process for performing pose estimation based on the entire area of ​​an image provided by this application;

[0330] Figure 109 A schematic diagram of the structure of AR / VR glasses provided in this application;

[0331] Figure 110 A schematic diagram of a gaze perception structure provided in this application;

[0332] Figure 111 A schematic diagram of a network architecture provided for this application;

[0333] Figure 112 A schematic structural diagram of an image processing device provided in this application;

[0334] Figure 113 A schematic structural diagram of another image processing device provided by this application;

[0335] Figure 114 A schematic structural diagram of another image processing device provided by this application;

[0336] Figure 115 A schematic structural diagram of another image processing device provided by this application;

[0337] Figure 116 A schematic structural diagram of another image processing device provided by this application;

[0338] Figure 117 A schematic structural diagram of another image processing device provided by this application;

[0339] Figure 118 A schematic structural diagram of another image processing device provided by this application;

[0340] Figure 119 Another structural diagram of the data processing device provided by this application;

[0341] Figure 120 Another structural diagram of the data processing device provided by this application;

[0342] Figure 121 This is another structural diagram of the electronic device provided in this application. DETAILED DESCRIPTION

[0343] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.

[0344] The following is a detailed introduction to the electronic equipment, system architecture, method flow, etc. provided in this application from different perspectives.

[0345] 1. Electronic equipment

[0346] The method provided in this application can be applied to various electronic devices, or the method provided in this application can be executed by an electronic device. The electronic device can be applied to shooting scenarios, such as taking pictures, security, automatic driving, drone shooting, etc.

[0347] The electronic devices in this application may include, but are not limited to: smart mobile phones, televisions, tablet computers, wristbands, head-mounted display devices (HMDs), augmented reality (AR) devices, mixed reality (MR) devices, cellular phones, smart phones, personal digital assistants (PDAs), vehicle-mounted electronic devices, laptop computers, personal computers (PCs), monitoring equipment, robots, vehicle-mounted terminals, autonomous vehicles, etc. Of course, in the following embodiments, there is no limitation on the specific form of the electronic devices.

[0348] For example, the architecture of the electronic device application provided by this application is as follows: Figure 1A shown.

[0349] Among them, electronic equipment such as Figure 1A The devices mentioned above, such as cars, mobile phones, AR / VR glasses, security monitoring equipment, cameras, or other smart home terminals, can be connected to the cloud platform through wired or wireless networks. The cloud platform is equipped with a server, which can include a centralized server or a distributed server. Electronic devices can communicate with the cloud platform server through wired or wireless networks to achieve data transmission. For example, after the electronic device collects the device data, it can save or back up the data on the cloud platform to prevent data loss.

[0350] Electronic devices can access the cloud platform wirelessly or by wired connection via an access point or base station. For example, the access point can be a base station, and the electronic device can be equipped with a SIM card, which can be used to authenticate with the operator's network and access the wireless network. Alternatively, the access point can be a router, and the electronic device can access the cloud platform through the router via a 2.4GHz or 5GHz wireless network.

[0351] Furthermore, electronic devices can process data independently or in collaboration with the cloud, with the specific method being adjusted based on the actual application scenario. For example, a DVS can be set up in an electronic device, and the DVS can work in conjunction with a camera or other sensor in the electronic device, or it can work independently, with a processor set up in the DVS or the electronic device processing the data collected by the DVS or other sensors. The processor can also work in collaboration with cloud devices to process the data collected by the DVS or other sensors.

[0352] The specific structure of the electronic device is exemplarily introduced below.

[0353] For example, see Figure 1B , below, taking a specific structure as an example, the structure of the electronic device provided by this application is exemplarily described.

[0354] It should be noted that the electronic equipment provided in this application may include Figure 1B More or fewer parts, Figure 1B The electronic device shown in the figure is merely an example. Those skilled in the art may add or reduce components in the electronic device as required, and this application does not limit this.

[0355] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, an image sensor 180N, etc., wherein the image sensor 180N may include an independent color sensor 1801N and an independent motion sensor 1802N, or may include a photosensitive unit of a color sensor (which may be called a color sensor pixel, Figure 1B ) and the photosensitive units of the motion sensor (which may be referred to as motion sensor pixels, Figure 1B not shown).

[0356] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0357] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0358] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0359] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0360] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0361] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 can contain multiple sets of I2C bus. The processor 110 can be coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces respectively. For example, the processor 110 can be coupled to the touch sensor 180K through an I2C interface, so that the processor 110 and the touch sensor 180K communicate through the I2C bus interface, and the touch function of the electronic device 100 is realized.

[0362] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can contain multiple sets of I2S bus. The processor 110 can be coupled to the audio module 170 through the I2S bus, and communication between the processor 110 and the audio module 170 is realized. In some embodiments, the audio module 170 can deliver audio signals to the wireless communication module 160 through the I2S interface, and the function of answering a phone through a Bluetooth headset is realized.

[0363] The PCM interface can also be used for audio communication, sampling, quantizing and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In some embodiments, the audio module 170 can also deliver audio signals to the wireless communication module 160 through the PCM interface, and the function of answering a phone through a Bluetooth headset is realized. Both the I2S interface and the PCM interface can be used for audio communication.

[0364] The UART interface is a universal serial data bus, which is used for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is usually used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface, and the Bluetooth function is realized. In some embodiments, the audio module 170 can deliver audio signals to the wireless communication module 160 through the UART interface, and the function of playing music through a Bluetooth headset is realized.

[0365] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display 194 and the camera 193. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to implement the camera function of the electronic device 100. The processor 110 and the display 194 communicate via the DSI interface to implement the display function of the electronic device 100.

[0366] The GPIO interface can be configured via software. The GPIO interface can be configured as either a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, display 194, wireless communication module 160, audio module 170, sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0367] The USB interface 130 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the electronic device 100, or to transfer data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as augmented reality devices.

[0368] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0369] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also provide power to the electronic device via the power management module 141.

[0370] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.

[0371] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0372] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0373] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0374] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.

[0375] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0376] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150 , and antenna 2 is coupled to wireless communication module 160 , so that electronic device 100 can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include, but is not limited to, fifth-generation mobile communication technology (5th-Generation, 5G) system, global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long-term evolution (LTE), Bluetooth, global navigation satellite system (GNSS), wireless fidelity (WiFi), nearfield communication (NFC), FM (also known as FM radio), Zigbee protocol, radio frequency identification technology (RFID) and / or infrared (IR) technology, etc. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite-based augmentation system (SBAS), etc.

[0377] In some embodiments, the electronic device 100 may also include a wired communication module ( Figure 1B ), or the mobile communication module 150 or the wireless communication module 160 here can be replaced with a wired communication module ( Figure 1B (not shown), the wired communication module enables the electronic device to communicate with other devices via a wired network. The wired network may include, but is not limited to, one or more of the following: optical transport network (OTN), synchronous digital hierarchy (SDH), passive optical network (PON), Ethernet, or flexible Ethernet (FlexE).

[0378] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0379] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.

[0380] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0381] ISP is used to process the data feedback from the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and conversion into a visible image. ISP can also optimize the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, ISP can be provided in the camera 193.

[0382] The camera 193 is used to capture still images or videos. Objects generate optical images through lenses and project them onto photosensitive elements. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard RGB camera (or RGB sensor) 0, YUV, or other format image signal. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.

[0383] The digital signal processor is used to process digital signals, in addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0384] The video codec is used to compress or decompress digital video. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0385] NPU is a neural-network (NN) computing processor that learns from the structure of biological neural networks, such as the transmission mode between human brain neurons, and can quickly process input information and continuously self-learn. Through NPU, the electronic device 100 can achieve intelligent cognition applications such as image recognition, face recognition, voice recognition, and text understanding.

[0386] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0387] The internal memory 121 can be used to store computer executable program codes, which include instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory provided in the processor.

[0388] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0389] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0390] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.

[0391] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.

[0392] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.

[0393] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0394] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be located on display screen 194. There are many types of pressure sensors 180A, such as resistive, inductive, and capacitive. A capacitive pressure sensor can include at least two parallel plates made of conductive material. When force acts on pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the intensity of the pressure based on this change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the touch intensity based on pressure sensor 180A. Electronic device 100 can also calculate the touch location based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch location but with different touch intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a first pressure threshold is applied to a short message application icon, a command to view short messages is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to a short message application icon, a command to create a new short message is executed.

[0395] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the electronic device 100 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device 100 through reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and somatosensory game scenes.

[0396] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude using the air pressure value measured by the air pressure sensor 180C to assist in positioning and navigation.

[0397] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip case. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover based on the magnetic sensor 180D. Based on the detected opening and closing status of the case or flip cover, features such as automatic unlocking of the flip cover can be configured.

[0398] Accelerometer 180E can detect the magnitude of acceleration of electronic device 100 in all directions (generally three axes). It can also detect the magnitude and direction of gravity when electronic device 100 is stationary. It can also be used to identify the electronic device's posture, enabling applications such as switching between landscape and portrait modes and pedometers.

[0399] The distance sensor 180F is used to measure distance. The electronic device 100 can measure distance using infrared or laser. In some embodiments, when shooting a scene, the electronic device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.

[0400] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The electronic device 100 emits infrared light outward through the light emitting diode. The electronic device 100 uses a photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 can use the proximity light sensor 180G to detect that the user is holding the electronic device 100 close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor 180G can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.

[0401] Ambient light sensor 180L is used to sense ambient light brightness. Electronic device 100 can adaptively adjust display screen 194 brightness according to sensed ambient light brightness. Ambient light sensor 180L can also be used to automatically adjust white balance when taking a picture. Ambient light sensor 180L can also cooperate with proximity light sensor 180G to detect whether electronic device 100 is in a pocket to prevent accidental touch.

[0402] Fingerprint sensor 180H is used to collect a fingerprint. Electronic device 100 can use collected fingerprint characteristics to implement fingerprint unlocking, access application lock, take a picture with a fingerprint, answer an incoming call with a fingerprint, and so on.

[0403] Temperature sensor 180J is used to detect temperature. In some embodiments, electronic device 100 uses temperature detected by temperature sensor 180J to implement a temperature processing strategy. For example, when temperature reported by temperature sensor 180J exceeds a threshold value, electronic device 100 reduces performance of a processor located near temperature sensor 180J to reduce power consumption and implement thermal protection. In another embodiment, when temperature is lower than another threshold value, electronic device 100 heats battery 142 to avoid abnormal shutdown of electronic device 100 caused by low temperature. In other embodiments, when temperature is lower than yet another threshold value, electronic device 100 boosts output voltage of battery 142 to avoid abnormal shutdown caused by low temperature.

[0404] Touch sensor 180K, also referred to as a "touch device". Touch sensor 180K can be disposed on display screen 194, and touch sensor 180K and display screen 194 together form a touch screen, also referred to as a "touch panel". Touch sensor 180K is used to detect a touch operation acting on or near it. Touch sensor 180K can transmit the detected touch operation to an application processor to determine a touch event type. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K can also be disposed on a surface of electronic device 100, which is different from the position of display screen 194.

[0405] Bone conduction sensor 180M can obtain a vibration signal. In some embodiments, bone conduction sensor 180M can obtain a vibration signal of a human body sound part vibration bone block. Bone conduction sensor 180M can also contact a human body pulse to receive a blood pressure pulsation signal. In some embodiments, bone conduction sensor 180M can also be disposed in a headset to form a bone conduction headset. Audio module 170 can analyze a voice signal based on the vibration signal of the sound part vibration bone block obtained by bone conduction sensor 180M to implement a voice function. Application processor can analyze heart rate information based on the blood pressure pulsation signal obtained by bone conduction sensor 180M to implement a heart rate detection function.

[0406] The image sensor 180N, also known as a photosensitive device or element, is a device that converts optical images into electronic signals. It is widely used in digital cameras and other electro-optical devices. Image sensors utilize the photoelectric conversion function of photoelectric devices to convert the light image on a photosensitive surface into an electrical signal proportional to the light image. Compared to point-light-emitting photosensitive elements such as photodiodes and phototransistors, image sensors divide the light image on their light-receiving surface into many small units (i.e., pixels) and convert them into usable electrical signals. Each small unit corresponds to a photosensitive unit within the image sensor, also known as a sensor pixel. Image sensors are categorized into photoconductive tubes and solid-state image sensors. Compared to photoconductive tubes, solid-state image sensors offer advantages such as small size, light weight, high integration density, high resolution, low power consumption, long life, and low price. According to the different components, they can be divided into two categories: charge coupled device (CCD) and complementary metal-oxide semiconductor (CMOS); according to the different types of optical images taken, they can be divided into two categories: color sensor 1801N and motion sensor 1802N.

[0407] Specifically, the color sensor 1801N includes a traditional RGB image sensor, which can be used to detect objects within the camera's range. Each photosensitive cell corresponds to a pixel in the image sensor. Since photosensitive cells can only sense light intensity and cannot capture color information, they must be covered with color filters. As for how to cover the color filters, different sensor manufacturers have different solutions. The most common method is to cover the RGB red, green, and blue color filters, with four pixels forming a color pixel in a 1:2:1 ratio (that is, the red and blue filters each cover one pixel, and the remaining two pixels are covered by the green filter). The reason for adopting this ratio is that the human eye is more sensitive to green. After receiving light, the photosensitive cell generates a corresponding current. The current size corresponds to the light intensity. Therefore, the electrical signal directly output by the photosensitive cell is analog. The output analog electrical signal is then converted into a digital signal. The resulting digital signal is output in the form of a digital image matrix and is processed by a dedicated DSP processing chip. This traditional color sensor outputs a full-frame image of the captured area in a frame format.

[0408] Specifically, the motion sensor 1802N can include a plurality of different types of vision sensors, such as a motion detection vision sensor (MDVS) and an event-based motion detection vision sensor. The motion sensor 1802N can be used to detect a moving object within a range of the camera, capture a motion profile or a motion trajectory of the moving object, and the like.

[0409] In one possible scenario, the motion sensor 1802N can include a motion detection (MD) vision sensor, which is a type of vision sensor that detects motion information resulting from relative motion between the camera and a target. The relative motion can be motion of the camera, motion of the target, or motion of both the camera and the target. The motion detection vision sensor includes frame-based motion detection and event-based motion detection. The frame-based motion detection vision sensor requires exposure integration and obtains motion information through frame difference. The event-based motion detection vision sensor does not require integration and obtains motion information through asynchronous event detection.

[0410] In one possible scenario, the motion sensor 1802N can include a motion detection vision sensor (MDVS), a dynamic vision sensor (DVS), an active pixel sensor (APS), an infrared sensor, a laser sensor, or an inertial measurement unit (IMU), and the like. The DVS can specifically include a DAVIS (Dynamic and Active-pixel Vision Sensor), an ATIS (Asynchronous Time-based Image Sensor), or a CeleX sensor, and the like. The DVS is inspired by the characteristics of biological vision, and each pixel simulates a neuron and independently responds to a relative change in light intensity (hereinafter referred to as “light intensity”). For example, when the motion sensor is a DVS, when the relative change in light intensity exceeds a threshold, the pixel outputs an event signal including the position of the pixel, a timestamp, and feature information of the light intensity. It should be understood that in the following embodiments of the present application, the motion information, dynamic data, or dynamic image, and the like mentioned can be acquired by the motion sensor.

[0411] For example, the motion sensor 1802N may include an inertial measurement unit (IMU), which is a device that measures the three-axis angular velocity and acceleration of an object. An IMU is usually composed of three single-axis accelerometers and three single-axis gyroscopes, which respectively measure the acceleration signal of the object and the angular velocity signal relative to the navigation coordinate system, and use this to calculate the object's posture. For example, the aforementioned IMU may specifically include the aforementioned gyroscope sensor 180B and acceleration sensor 180E. The advantage of the IMU is its high acquisition frequency. The data acquisition frequency of the IMU can generally reach above 100 Hz, and consumer-grade IMUs can capture data up to 1600 Hz. In a relatively short period of time, the IMU can provide high-precision measurement results.

[0412] For example, motion sensor 1802N may include an active pixel sensor (APS). For example, it captures RGB images at a high frequency greater than 100 Hz and performs subtraction between two adjacent image frames to obtain a change value. If this change value is greater than a threshold (e.g., greater than 0), it is set to 1; if it is less than the threshold (e.g., less than 0), it is set to 0. The resulting data is similar to that obtained by a DVS, completing the capture of images of moving objects.

[0413] The buttons 190 include a power button, a volume button, and the like. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.

[0414] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0415] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.

[0416] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to or disconnected from the electronic device 100 by inserting it into or removing it from the SIM card interface 195. The electronic device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0417] 2. System Architecture

[0418] When electronic devices capture, read or save images, changes between multiple components are involved. The following application provides a detailed introduction to data acquisition, data encoding and decoding, image enhancement, image reconstruction or application scenarios.

[0419] For example, take a scenario of image acquisition and processing as an example, such as Figure 2 As shown, the processing flow of the electronic device is exemplarily described.

[0420] Data Acquisition: Data can be acquired using a brain-inspired camera, an RGB camera, or a combination thereof. A brain-inspired camera can include a biomimetic visual sensor, which uses integrated circuits to simulate a biological retina. Each pixel simulates a biological neuron, expressing changes in light intensity as events. Over time, various types of biomimetic visual sensors have emerged, all of which share the common feature of a pixel array that independently and asynchronously monitors changes in light intensity and outputs these changes as event signals, such as the aforementioned motion sensors DVS or DAVIS. An RGB camera converts analog signals into digital signals and then stores them on a storage medium. Data can also be acquired using a combination of a brain-inspired camera and an RGB camera. For example, data collected by the brain-inspired camera and the RGB camera can be projected onto the same canvas. The value of each pixel can be determined based on the values ​​fed back by the brain-inspired camera and / or the RGB camera, or the value of each pixel can include the values ​​of the brain-inspired camera and the RGB camera as independent channels. Using a brain-inspired camera, an RGB camera, or a combination thereof, optical signals can be converted into electrical signals, resulting in a data stream in frames or an event stream in events. Hereinafter, images captured by an RGB camera will be referred to as RGB images, and images captured by a brain-inspired camera will be referred to as event images.

[0421] Data coding and decoding: including data coding and data decoding. Data coding can include coding the collected data after data collection, and saving the coded data to the storage medium. Data decoding can include reading data from the storage medium and decoding the data, decoding the data into data that can be used for subsequent identification, detection, etc. In addition, the way of data coding and decoding can also be used to adjust the way of data collection, so as to realize more efficient data collection and data coding and decoding. Data coding and decoding can be divided into many kinds, including coding and decoding based on brain-like camera, coding and decoding based on brain-like camera and RGB camera, or coding and decoding based on RGB camera, etc. Specifically, in the process of coding, the data collected by the brain-like camera, the RGB camera or their combination can be coded to be stored in the storage medium in a certain format, and in the process of decoding, the data stored in the storage medium can be decoded into data that can be used subsequently. For example, the user can collect video or image data through the brain-like camera, the RGB camera or their combination on the first day, and code the video or image data and store it in the storage medium. The data can be read from the storage medium on the second day and decoded to obtain playable video or image.

[0422] Image optimization: that is, after the aforementioned brain-like camera or RGB camera collects the image, the collected image is read, and then the collected image is optimized by enhancement or reconstruction, etc. so as to facilitate subsequent processing based on the optimized image. Exemplarily, image enhancement and reconstruction can include image reconstruction or motion compensation, etc. Motion compensation, for example, the motion parameters of the moving object collected by the DVS, compensates the moving object in the event image or the RGB image, so that the obtained event image or RGB image is clearer. Image reconstruction, for example, the image collected by the brain-like vision camera, reconstructs the RGB image, so that even in a moving scene, a clear RGB image can be obtained from the data collected by the DVS.

[0423] Application scenario: after the optimized RGB image or event image is obtained by image optimization, the optimized RGB image or event image can be used for further application, of course, the collected RGB image or event image can also be used for further application, which can be adjusted according to the actual application scenario.

[0424] Specifically, application scenarios may include: motion photography enhancement, DVS image and RGB image fusion, detection and recognition, simultaneous localization and mapping (SLAM), eye tracking, keyframe selection, or pose estimation. For example, motion photography enhancement involves enhancing the captured image in a scene with moving objects, thereby capturing clearer moving objects. DVS image and RGB image fusion involves enhancing the RGB image of a moving object captured by DVS, compensating for moving objects or objects affected by large light ratios in the RGB image, thereby obtaining a clearer RGB image. Detection and recognition involves target detection or recognition based on RGB images or event images. Eye tracking involves tracking the user's eye movements based on captured RGB images, event images, or optimized RGB images or event images, to determine information such as the user's gaze point and gaze direction. Keyframe selection involves selecting certain frames from the video data captured by the RGB camera as keyframes, combining information captured by the brain-inspired camera.

[0425] In addition, in the following embodiments of the present application, different sensors may need to be activated in different embodiments. For example, when collecting data, when optimizing event images through motion compensation, the motion sensor can be activated, and optionally the IMU or gyroscope can also be activated. In the embodiment of image reconstruction, the motion sensor can be activated to collect event images and then optimized in combination with the event images. Or, in the embodiment of motion photography enhancement, the motion sensor and RGB sensor can be activated, etc. Therefore, in different embodiments, corresponding sensors can be selected to be activated.

[0426] Specifically, the method provided in the present application can be applied to an electronic device, which may include an RGB sensor and a motion sensor, etc. The RGB sensor is used to capture images within a shooting range, and the motion sensor is used to capture information generated when an object moves relative to the motion sensor within the detection range of the motion sensor. The method includes: selecting at least one from the RGB sensor and the motion sensor based on scene information, and collecting data through the selected sensor. The scene information includes status information of the electronic device, the type of application requesting image capture in the electronic device, or at least one of environmental information.

[0427] In a possible implementation, the aforementioned status information includes information such as the remaining power, remaining storage (or available storage), or CPU load of the electronic device.

[0428] In one possible implementation, the aforementioned environmental information may include changes in light intensity within the capture range of the color RGB sensor and the motion sensor, or information about moving objects within the capture range. For example, the environmental information may include changes in light intensity within the capture range of the RGB sensor or the DVS sensor, or the motion of objects within the capture range, such as the object's speed and direction, or abnormal motion of objects within the capture range, such as sudden changes in speed or direction.

[0429] The type of application requesting image capture in the aforementioned electronic device can be understood as that the electronic device carries a system such as Android, Linux or HarmonyOS, and applications can be run in the system, and the programs running in the system can be divided into multiple types, such as photo-taking applications or target detection applications.

[0430] Typically, motion sensors are sensitive to changes in motion but not to static scenes. They respond to motion changes by emitting events. However, since static areas rarely emit events, their data only represents the light intensity information in the areas of motion, not the complete light intensity information for the entire scene. RGB color cameras excel at recording the complete color of natural scenes and reproducing the texture details within them.

[0431] Taking the aforementioned electronic device as a mobile phone as an example, the default configuration is to disable the DVS camera (i.e., DVS sensor). When using the camera, depending on the type of application currently being called, for example, if a photo app is calling the camera and the camera is in high-speed motion, both the DVS camera and the RGB camera (i.e., RGB sensor) must be enabled. If the app requesting the camera is for object detection or motion detection and does not require object photography or face recognition, the DVS camera can be enabled, but the RGB camera can be disabled.

[0432] Optionally, you can also select a camera startup mode based on the current device status. For example, when the current battery level is lower than a certain threshold, the user activates the power saving mode and cannot take photos normally. You can only turn on the DVS camera. Although the DVS camera's photos are not clear, its power consumption is low, and high-definition imaging is not required when detecting moving objects.

[0433] Optionally, the device can sense the surrounding environment to decide whether to switch camera modes. For example, in night scenes or when the device is in high-speed motion, the DVS camera can be turned on. If the scene is static, the DVS camera can be turned off.

[0434] Based on the above application types, environmental information, and device status, the camera function startup mode is determined. During operation, it can decide whether to trigger the camera mode switch, thereby starting different sensors in different scenarios, with strong adaptability.

[0435] It can be understood that there are three startup modes: only the RGB camera, only the DVS camera, and both the RGB and DVS cameras. Furthermore, the reference factors for detection application type and environmental detection can vary for different products.

[0436] For example, a camera in a security scenario has a motion detection function. The camera will only store the video when it detects a moving object, thereby reducing storage space and extending the hard disk storage time. Specifically, when DVS and RGB cameras are used in home or security cameras, only the DVS camera is turned on by default for motion detection and analysis. When the DVS camera detects abnormal motion or abnormal behavior (such as sudden movement of an object, sudden change in movement direction, etc.), such as the approach of a person, or a significant change in light intensity, the RGB camera is turned on to shoot and record the full scene texture image during this period as monitoring evidence. When the abnormal motion ends, it switches back to DVS operation and the RGB camera enters standby mode, significantly saving data volume and power consumption of the monitoring equipment.

[0437] The intermittent recording method described above leverages the low power consumption of DVS. Furthermore, DVS's event-based motion detection is faster, providing a quicker response and more accurate detection than image-based motion detection. This enables continuous detection around the clock, resulting in more accurate, low-power, and storage-saving methods.

[0438] For example, when DVS and RGB cameras are used in assisted / autonomous driving, the RGB camera may be unable to capture valid scene information when encountering oncoming vehicles with high beams, direct sunlight, or when entering or exiting a tunnel. In these situations, while the DVS cannot capture texture information, it can obtain general outline information within the scene, which is extremely valuable in assisting the driver's judgment. Furthermore, in foggy weather, the outline information captured by the DVS can also assist in judging road conditions. Therefore, the DVS and RGB cameras can be triggered to switch between master and slave modes in specific scenarios, such as when light intensity changes dramatically or in extreme weather conditions.

[0439] The above process also applies when the DVS camera is used in AR / VR glasses. When using DVS for SLAM or eye tracking, the camera startup mode can be determined based on the device status and surrounding environment.

[0440] In the following embodiments of the present application, when data collected by a certain sensor is used, the sensor is turned on, which will not be described in detail below.

[0441] The following combines the above different working modes and the aforementioned Figure 2 Different implementations provided in this application are described.

[0442] 3. Methodology

[0443] The above is an exemplary description of the electronic device and system architecture provided by this application. Figure 1A-Figure 2 , the method provided in this application is described in detail. Specifically, in combination with the above Figure 2 The architecture of the present invention is described separately for each module. It should be understood that the method steps mentioned below in this application can be implemented separately or combined in one device, and can be adjusted according to the actual application scenario.

[0444] 1. Data collection and encoding and decoding

[0445] The following is an illustrative description of the data collection and data encoding and decoding processes.

[0446] In traditional technology, visual sensors (i.e., the aforementioned motion sensors) generally use an asynchronous reading mode based on event streams (hereinafter referred to as "event stream-based reading mode" or "asynchronous reading mode") and a synchronous reading mode based on frame scanning (hereinafter referred to as "frame scanning-based reading mode" or "synchronous reading mode"). For manufactured visual sensors, only one of the two modes can be used. Depending on the specific application scenario and motion state, the amount of signal data required to be read per unit time in these two reading modes may differ significantly, and the cost required to output the read data is also different. Figure 3-a and Figure 3-b Schematic diagrams showing the relationship between the amount of read data and time in an asynchronous reading mode based on event stream and a synchronous reading mode based on frame scanning are shown respectively.

[0447] On the one hand, bionic visual sensors are sensitive to motion, and static areas in the environment usually do not generate light intensity change events (also referred to as "events" in this article). Almost all of these sensors use an asynchronous reading mode based on event streams. Event streams refer to events arranged in a certain order. The following uses DVS as an example to illustrate the asynchronous reading mode. According to the sampling principle of DVS, by comparing the current light intensity with the light intensity when the last event occurred, when the change reaches a predetermined emission threshold C (hereinafter referred to as the predetermined threshold), an event is generated and output. That is, usually when the difference between the current light intensity and the light intensity when the last event occurred exceeds the predetermined threshold C, DVS will generate an event, which can be described by formula 1-1:

[0448] |LL′|≥C (1-1)

[0449] Where L and L' represent the light intensity at the current moment and the light intensity when the last event occurred, respectively.

[0450] For asynchronous read mode, each event can be expressed as<x,y,t,m> , (x, y) represents the pixel location where the event occurred, t represents the time when the event occurred, and m represents the characteristic information of the light intensity. Specifically, pixels in the pixel array circuit of a visual sensor measure the change in light intensity in the environment. If the measured light intensity change exceeds a predetermined threshold, the pixel can output a data signal indicating the event. Therefore, in the asynchronous reading mode based on the event stream, the pixels of the visual sensor are further divided into pixels that have generated light intensity change events and pixels that have not generated light intensity change events. Light intensity change events can be characterized by the coordinate information (x, y) of the pixel where the event occurred, the characteristic information of the light intensity at that pixel, and the time t when the characteristic information of the light intensity was read. The coordinate information (x, y) can be used to uniquely identify the pixel in the pixel array circuit. For example, x represents the row index of the pixel in the pixel array circuit, and y represents the column index of the pixel in the pixel array circuit. By identifying the coordinates and timestamps associated with the pixel, the spatiotemporal location of the light intensity change event can be uniquely determined, and all events can then be organized into an event stream in the order of their occurrence.

[0451] In some DVS sensors (e.g., DAVIS sensor, ATIS sensor, etc.), m represents the trend of light intensity change, also referred to as polarity information, and is usually represented by 1-bit to 2-bit, and can take ON / OFF values, where ON represents light intensity enhancement, and OFF represents light intensity reduction, i.e., when the light intensity increases and exceeds a predetermined threshold, an ON pulse is generated; when the light intensity decreases and exceeds a predetermined threshold, an OFF pulse is generated (in this application, "+" represents light intensity enhancement, and "-" represents light intensity reduction). In some DVS sensors, such as CeleX sensor, m represents absolute light intensity information, also referred to as light intensity information, and is usually represented by multiple bits, such as 8-bit to 12-bit.

[0452] In the asynchronous reading mode, only the data signal at the pixel where the light intensity change event occurs is read. Thus, for the biomimetic vision sensor, the event data required to be read has the characteristics of sparsity and asynchronicity. As shown in curve 101, when the rate of light intensity change events occurring in the pixel array circuit changes, the amount of data required to be read by the vision sensor also changes over time. Figure 3-a

[0453] On the other hand, conventional vision sensors, such as mobile phone cameras, digital video cameras, etc., usually adopt a synchronous reading mode based on frame scanning. This reading mode does not distinguish whether a light intensity change event occurs at a pixel of the vision sensor or not. Regardless of whether a light intensity change event occurs at a pixel or not, the data signal generated by the pixel is read. In reading the data signal, the vision sensor scans the pixel array circuit in a predetermined order, synchronously reads the feature information m (the feature information m of light intensity has been introduced above and will not be repeated here) indicating light intensity at each pixel, and outputs in order as frame 1 data, frame 2 data, etc. Thus, as shown in curve 102, in the synchronous reading mode, the amount of data read by the vision sensor for each frame has the same size, and the data amount remains unchanged over time. For example, assuming that 8 bits are used to represent the light intensity value of a pixel, and the total number of pixels in the vision sensor is 66, then the data amount of a frame of data is 528 bits. Usually, frames are output at equal time intervals, for example, at a rate of 30 frames per second, 60 frames per second, 120 frames per second, etc. Figure 3-b

[0454] The applicant finds that the current vision sensor still has defects, at least including the following aspects:

[0455] Firstly, a single reading mode cannot adapt to all scenarios, which is not conducive to relieving the pressure of data transmission and storage.

[0456] As​​ Figure 3-a As shown in the curve 101, the visual sensor operates in an asynchronous reading mode based on an event stream. When the rate of light intensity change events occurring in the pixel array circuit changes, the amount of data that the visual sensor needs to read also changes with time. In static scenes, fewer light intensity change events are generated, so the total amount of data that the visual sensor needs to read is also lower. In dynamic scenes, such as during intense exercise, a large number of light intensity change events are generated, and the total amount of data that the visual sensor needs to read also increases. In some scenes, a large number of light intensity change events are generated, causing the total amount of data to exceed the bandwidth limit, and event loss or delayed readout may occur. Figure 3-b As shown in the curve 102, the visual sensor operates in a frame-based synchronous readout mode. Regardless of whether a pixel changes, the state or intensity value of the pixel needs to be represented in one frame. When only a small number of pixels change, this representation cost is high.

[0457] The output and storage costs of the two modes can differ significantly in different application scenarios and motion conditions. For example, when capturing a static scene, only a small number of pixels generate light intensity change events over a period of time. By way of example, in a single scan, light intensity change events occur at only three pixels in the pixel array circuit. In asynchronous read mode, simply reading the coordinate information (x, y), time information t, and light intensity change of these three pixels can represent the three light intensity change events. Assuming that in asynchronous read mode, 4, 2, 2, and 2 bits are allocated for the coordinates, read timestamp, and light intensity change of a pixel, respectively, the total amount of data required to be read in this read mode is 30 bits. In synchronous read mode, although only three pixels generate valid data signals indicating light intensity change events, the data signals output by all pixels in the entire array must be read to form a complete frame of data. Assuming that in synchronous read mode, 8 bits are allocated to each pixel, and the pixel array circuit has a total of 66 pixels, the total amount of data required to be read is 528 bits. This shows that even if there are a large number of pixels in the pixel array circuit that have not generated events, the synchronous read mode still requires a large number of bits. This is uneconomical from a representation cost perspective and increases the pressure on data transmission and storage. Therefore, in this case, the asynchronous read mode is more economical.

[0458] In another example, when a scene experiences intense motion or a dramatic change in ambient light intensity, such as when a large number of people move around or lights are suddenly turned on or off, a large number of pixels in the visual sensor measure the light intensity change within a short period of time and generate data signals indicating the light intensity change event. Because the amount of data representing a single event in asynchronous reading mode is greater than that in synchronous reading mode, adopting asynchronous reading mode in this situation can incur a significant representation cost. Specifically, each row of the pixel array circuit may have multiple consecutive pixels generating light intensity change events. For each event, coordinate information (x, y), time information t, and characteristic light intensity information m must be transmitted. The coordinate changes between these events often differ by only one unit, and the read times are essentially the same. In this case, the asynchronous reading mode has a high representation cost for coordinate and time information, resulting in a surge in data volume. In contrast, in synchronous reading mode, regardless of the number of light intensity change events occurring in the pixel array circuit at any given moment, each pixel outputs a data signal indicating only the amount of light intensity change, eliminating the need to allocate bits for each pixel's coordinate and time information. Therefore, in scenarios with dense events, adopting synchronous reading mode is more economical.

[0459] Secondly, a single event representation method cannot adapt to all scenarios. Using light intensity information to represent events is not conducive to alleviating the pressure on data transmission and storage, and using polarity information to represent events affects the processing and analysis of events.

[0460] The above describes the synchronous and asynchronous read modes. All readout events must be represented by characteristic information m of light intensity, which includes polarity information and intensity information. This article refers to events represented by polarity information as polarity-formatted events, and events represented by light intensity information as intensity-formatted events. For manufactured vision sensors, only one of the two event formats can be used: representing events using polarity information or using light intensity information. The following uses the asynchronous read mode as an example to illustrate the advantages and disadvantages of polarity-formatted and intensity-formatted events.

[0461] In the asynchronous reading mode, the polarity information p is usually represented by 1-bit-2-bit, carrying less information, and only indicating whether the light intensity trend is increasing or weakening. Therefore, the event represented by the polarity information affects the processing and analysis of the event. For example, the event represented by the polarity information is difficult to reconstruct the image, and the accuracy for object recognition is poor. When the light intensity information is used to represent the event, it is usually represented by multiple bits, such as 8-bit-12-bit. Compared with the polarity information, the light intensity information can carry more information, which is beneficial to the processing and analysis of the event, such as improving the quality of image reconstruction. However, due to the large amount of data, it takes longer to obtain the event represented by the light intensity information. According to the DVS sampling principle, when the light intensity of the pixel changes by more than a predetermined threshold, an event will be generated. When a large area of object movement or light intensity fluctuation occurs in the scene (for example, entering and exiting a tunnel, turning on and off the light in a room, etc.), the visual sensor will face the problem of a sudden increase in events. In the case of a preset maximum bandwidth (hereinafter referred to as bandwidth) of the visual sensor, there is a situation that the event data cannot be read out. At present, the random discarding method is usually used for processing. If random discarding is used, although the amount of data transmitted can be ensured not to exceed the bandwidth, data loss is caused. In some special application scenarios (such as automatic driving, etc.), the discarded data may have high importance. In other words, when a large number of events are triggered, the data amount exceeds the bandwidth, and the light intensity format data cannot be completely output to the outside of the DVS, resulting in the loss of some events. These lost events may not be conducive to the processing and analysis of the event, such as causing trailing and incomplete contour in the brightness reconstruction process.

[0462] To solve the above problems, the embodiment of the present application provides a visual sensor. Based on the statistical results of the light intensity change events generated by the pixel array circuit, the data amount in the two reading modes is compared, so that the reading mode suitable for the current application scenario and motion state can be switched. In addition, based on the statistical results of the light intensity change events generated by the pixel array circuit, the relationship between the data amount of the event represented by the light intensity information and the bandwidth is compared, so as to adjust the representation precision of the event. Under the premise of meeting the bandwidth limitation, all events are transmitted in a suitable representation manner, and all events are transmitted with greater representation precision as much as possible.

[0463] A visual sensor provided by an embodiment of the present application is introduced below.

[0464] Figure 4-a A block diagram of a visual sensor provided by the present application is shown. The visual sensor can be implemented as a visual sensor chip and can read a data signal indicating an event in at least one of a frame scanning-based reading mode and an event stream-based reading mode. As shown in the figure, the visual sensor includes a pixel array circuit, a data processing circuit, and a data output circuit. Figure 4-aAs shown, the visual sensor 200 includes a pixel array circuit 210 and a reading circuit 220. The visual sensor is coupled to a control circuit 230. It should be understood that Figure 4-a The visual sensor shown is for illustrative purposes only and does not imply any limitation on the scope of the present application. The embodiments of the present application may also be embodied in different sensor architectures. In addition, it should be understood that the visual sensor may also include other elements or entities for achieving purposes such as image acquisition, image processing, and image transmission. For ease of description, these are not shown, but this does not mean that the embodiments of the present application do not have these elements or entities.

[0465] The pixel array circuit 210 may include one or more pixel arrays, and each pixel array includes multiple pixels, each pixel having location information for unique identification, such as coordinates (x, y). The pixel array circuit 210 can be used to measure the change in light intensity and generate multiple data signals corresponding to the multiple pixels. In some possible embodiments, each pixel is configured to independently respond to changes in light intensity in the environment. In some possible embodiments, the pixel compares the measured light intensity change with a predetermined threshold. If the measured light intensity change exceeds the predetermined threshold, the pixel generates a first data signal indicating the light intensity change event. For example, the first data signal includes polarity information, such as +1 or -1, or the first data signal can also be absolute light intensity information. In this example, the first data signal can indicate the light intensity change trend or absolute light intensity value at the corresponding pixel. In some possible embodiments, if the measured light intensity change does not exceed the predetermined threshold, the pixel generates a second data signal different from the first data signal, such as 0. In embodiments of the present application, the data signal can indicate, but is not limited to, light intensity polarity, absolute light intensity value, light intensity change value, etc. Light intensity polarity can indicate the trend of light intensity change, for example, an increase or decrease, and is typically represented by +1 and -1. An absolute light intensity value can represent the light intensity value measured at the current moment. Light intensity or light intensity change can have different physical meanings depending on the structure, application, and type of sensor. The scope of this application is not limited in this respect.

[0466] The readout circuit 220 is coupled to the pixel array circuit 210 and the control circuit 230 and can communicate with both. The readout circuit 220 is configured to read the data signal output by the pixel array circuit 210. This can be understood as the readout circuit 220 reading the data signal output by the pixel array 210 and transferring it to the control circuit 230. The control circuit 230 is configured to control the mode in which the readout circuit 220 reads the data signal. The control circuit 230 can also be configured to control the representation of the output data signal. In other words, it can control the representation accuracy of the data signal. For example, the control circuit can control the visual sensor to output events represented by polarity information, events represented by light intensity information, or events represented by a fixed number of bits, etc. This will be explained below in conjunction with specific embodiments.

[0467] According to a possible embodiment of the present application, the control circuit 230 may be as follows: Figure 4-a The control circuit 230 is shown as an independent circuit or chip outside the vision sensor 200, connected to the vision sensor 200 via a bus interface. In other possible embodiments, the control circuit 230 can also be a circuit or chip inside the vision sensor, integrated with the pixel array circuit and the reading circuit therein. Figure 4-b FIG. 3 is a block diagram of another visual sensor 300 according to a possible embodiment of the present application. The visual sensor 300 can be implemented as an example of the visual sensor 200. Figure 4-b FIG. 3 is another block diagram of a visual sensor provided by the present application. The visual sensor includes a pixel array circuit 310, a read circuit 320, and a control circuit 330. The pixel array circuit 310, the read circuit 320, and the control circuit 330 are functionally similar to the pixel array circuit 310. Figure 4-a The pixel array circuit 210, the reading circuit 220, and the control circuit 230 shown are the same and are not described in detail here. It should be understood that the visual sensor is used for exemplary purposes only and does not imply any limitation on the scope of the present application. The embodiments of the present application may also be embodied in different visual sensors. In addition, it should be understood that the visual sensor may also include other elements, modules, or entities, which are not shown for the purpose of clarity, but this does not mean that the embodiments of the present application do not have these elements or entities.

[0468] Based on the aforementioned architecture of the visual sensor, the visual sensor provided in this application is described in detail below.

[0469] The read circuit 220 can be configured to scan the pixels in the pixel array circuit 210 in a predetermined order to read the data signals generated by the corresponding pixels. In an embodiment of the present application, the read circuit 220 is configured to be able to read the data signals output by the pixel array circuit 210 in more than one signal reading mode. For example, the read circuit 220 can read in one of a first reading mode and a second reading mode. In the context of this document, the first reading mode and the second reading mode correspond to one of a frame scanning-based reading mode and an event stream-based reading mode, respectively. Further, the first reading mode can refer to the current reading mode of the read circuit 220, and the second reading mode can refer to a switchable alternative reading mode.

[0470] refer to Figure 5 , which shows a schematic diagram of the principles of the synchronous reading mode based on frame scanning and the asynchronous reading mode based on event stream according to an embodiment of the present application. Figure 5 As shown in the upper half of , the black dots represent pixels that generate light intensity change events, and the white dots represent pixels that do not generate light intensity change events. The dotted box on the left represents a synchronous reading mode based on frame scanning, in which all pixels generate voltage signals based on the received light signals, and then output data signals after analog-to-digital conversion. In this mode, the reading circuit 220 constructs a frame of data by reading the data signals generated by all pixels. The dotted box on the right represents an asynchronous reading mode based on an event stream. In this mode, when the reading circuit 220 scans a pixel that generates a light intensity change event, the coordinate information (x, y) of the pixel can be obtained. Then, only the data signal generated by the pixel that generates the light intensity change event is read, and the reading time t is recorded. In the case where there are multiple pixels that generate light intensity change events in the pixel array circuit, the reading circuit 220 reads the data signals generated by the multiple pixels in sequence according to the scanning order, and constructs an event stream as output.

[0471] Figure 5 The lower part of describes the two reading modes from the perspective of cost (e.g., the amount of data required to be read). Figure 5 As shown, in the synchronous reading mode, the amount of data read by the reading circuit 220 each time is the same, for example, 1 frame of data. Figure 5 The first frame data 401-1 and the second frame data 401-2 are shown in FIG. p ) and the total number of pixels M in the pixel array circuit can determine the amount of data to be read for one frame as M·B pIn the asynchronous reading mode, the reading circuit 220 reads the data signal indicating the light intensity change event, and then forms an event stream 402 according to the order in which all events occur. In this case, the amount of data read by the reading circuit 220 each time is equal to the event data amount B used to represent a single event. ev (For example, the sum of the number of bits representing the coordinates (x, y) of the pixel where the event occurred, the reading timestamp t, and the characteristic information of the light intensity) and the number of light intensity change events N ev related.

[0472] In some embodiments, the read circuit 220 may be configured to provide at least one read data signal to the control circuit 230. For example, the read circuit 220 may provide the control circuit 230 with data signals read over a period of time for the control circuit 230 to perform historical data statistics and analysis.

[0473] In some possible embodiments, when the first reading mode currently adopted is a reading mode based on an event stream, the reading circuit 220 reads the data signals generated by the pixels in the pixel array circuit 210 that generate the light intensity change event. For the sake of convenience, these data signals are also referred to as first data signals below. Specifically, the reading circuit 220 determines the position information (x, y) of the pixels related to the light intensity change event by scanning the pixel array circuit 210. Based on the position information (x, y) of the pixel, the reading circuit 220 reads the first data signal generated by the pixel from a plurality of data signals to obtain the characteristic information of the light intensity indicated by the first data signal and the reading time information t. By way of example, in the reading mode based on an event stream, the amount of event data read per second by the reading circuit 220 can be expressed as B ev ·N ev bits, that is, the read data rate of the read circuit 220 is B ev ·N ev Bit per second (bps), where B ev is the amount of event data (e.g., number of bits) allocated to each light intensity change event in the event stream-based reading mode, where the first b x and b y bits are used to represent the pixel coordinates (x, y), and the following b t bits are used to indicate the timestamp t at which the data signal is read, and finally b f bits are used to represent the characteristic information of the light intensity indicated by the data signal, that is, B ev =b x +b y +b t +b f , N evM·B is the average number of events per second generated by the reading circuit 220 based on the historical statistics of the number of light intensity change events generated in the pixel array circuit 210 over a period of time. Due to the frame scanning reading mode, the amount of data per frame read by the reading circuit 220 can be expressed as M·B p bits, and the amount of data read per second is M·B p f bits, that is, the read data rate of the read circuit 220 is M·B p f bps, where the total number of pixels in the given vision sensor 200 is M, B p is the amount of pixel data (e.g., the number of bits) allocated to each pixel in the frame-scanning-based reading mode, and f is the predetermined frame rate of the reading circuit 220 in the frame-scanning-based reading mode, that is, the reading circuit 220 scans the pixel array circuit 210 at the predetermined frame rate f Hz in this mode to read the data signals generated by all pixels in the pixel array circuit 210. Therefore, M, B p and f are known quantities, and the read data rate of the read circuit 220 in the frame scanning-based read mode can be directly obtained.

[0474] In some possible embodiments, when the first reading mode currently adopted is a frame scanning-based reading mode, the reading circuit 220 calculates the average number of events N generated per second based on historical statistics of the number of light intensity change events generated in the pixel array circuit 210 over a period of time. ev , according to the frame scan reading mode to obtain N ev , it can be calculated that in the event stream-based reading mode, the amount of event data read per second by the reading circuit 220 is B ev ·N ev bits, that is, in the event stream-based read mode, the read data rate of the read circuit 220 is B ev ·N ev bps.

[0475] It can be seen from the above two embodiments that the reading data rate of the reading circuit 220 in the frame scanning based reading mode can be directly calculated according to the predefined parameters, and the reading data rate of the reading circuit 220 in the event stream based reading mode can be obtained according to the N obtained in either mode. ev Calculated.

[0476] The control circuit 230 is coupled to the read circuit 220 and is configured to control the read circuit 220 to read the data signals generated by the pixel array circuit 210 in a specific read mode. In some possible embodiments, the control circuit 230 may obtain at least one data signal from the read circuit 220 and, based at least on the at least one data signal, determine which of the current read mode and the alternative read mode is more suitable for the current application scenario and motion state. Furthermore, in some embodiments, the control circuit 230 may instruct the read circuit 220 to switch from the current data read mode to another data read mode based on this determination.

[0477] In some possible embodiments, the control circuit 230 may send an instruction to the read circuit 220 regarding switching the read mode based on historical statistics of light intensity change events. For example, the control circuit 230 may determine statistical data related to at least one light intensity change event based on at least one data signal received from the read circuit 220. If the statistical data is determined to meet a predetermined switching condition, the control circuit 230 sends a mode switching signal to the read circuit 220, causing the read circuit 220 to switch to the second read mode. To facilitate comparison, the statistical data can be used to measure the read data rates of the first read mode and the second read mode, respectively.

[0478] In some embodiments, the statistical data may include the total data volume of the number of events measured by the pixel array circuit 210 per unit time. If the total data volume of the light intensity change events read by the reading circuit 220 in the first reading mode is greater than or equal to the total data volume of the light intensity change events in the second reading mode, it indicates that the reading circuit 220 should switch from the first reading mode to the second reading mode. In some embodiments, the first reading mode is a frame scanning-based reading mode and the second reading mode is an event stream-based reading mode. The control circuit 230 may be based on the number of pixels M, the frame rate f, and the pixel data volume B of the pixel array circuit. p To determine the total data volume M·B of the light intensity change events read in the first reading mode p f. The control circuit 230 may be based on the number N of light intensity change events ev and the amount of event data associated with the event stream-based reading mode B ev , to determine the total data volume B of the light intensity change event ev ·N ev , that is, the total data volume B of the light intensity change events read in the second reading mode ev ·N ev In some embodiments, the switching parameter can be used to adjust the relationship between the total data volume in the two reading modes, as shown in the following formula (1): the total data volume of the light intensity change event read in the first reading mode is M·B pf is greater than or equal to the total data volume B of the light intensity change events in the second reading mode ev ·N ev , the read circuit 220 should switch to the second read mode:

[0479] η·M·B P f ≥ B ev ·N ev (1)

[0480] Where η is the switching parameter used for adjustment. From the above formula (1), it can be further concluded that the first threshold data volume d1 = M·B p ·f·η. That is, if the total data volume of light intensity change events B ev ·N ev If the data amount is less than or equal to the threshold data amount d1, it indicates that the total data amount of the light intensity change events read in the first reading mode is greater than or equal to the total data amount of the light intensity change events in the second reading mode, and the control circuit 230 can determine that the statistical data of the light intensity change events meets the predetermined switching condition. In this embodiment, the switching condition can be determined based on at least the number of pixels M of the pixel array circuit, the frame rate f associated with the frame scanning reading mode, and the pixel data amount B. p To determine the threshold data volume d1.

[0481] As an alternative implementation of the above embodiment, the total data volume of the light intensity change event read in the first reading mode is M·B p f is greater than or equal to the total data volume B of the light intensity change events in the second reading mode ev ·N ev It can be expressed as the following formula (2):

[0482] M.B P ·fB ev ·N ev ≥θ (2)

[0483] Where θ is the switching parameter used for adjustment. From the above formula (2), it can be further concluded that the second threshold data volume

[0484] d2=M·B p f-θ

[0485] That is, if the total data volume of light intensity change events B ev ·N evIf the data amount is less than or equal to the second threshold data amount d2, it indicates that the total data amount of the light intensity change events read in the first reading mode is greater than or equal to the total data amount of the light intensity change events in the second reading mode, and the control circuit 230 can determine that the statistical data of the light intensity change events meets the predetermined switching condition. In this embodiment, the switching condition can be determined based on at least the number of pixels M of the pixel array circuit, the frame rate f associated with the frame scanning reading mode, and the pixel data amount B. p To determine the threshold data amount d2.

[0486] In some embodiments, the first reading mode is an event stream-based reading mode and the second reading mode is a frame scan-based reading mode. In the event stream-based reading mode, the reading circuit 220 only reads the data signals generated by the pixels that generate events. Therefore, the control circuit 230 can directly determine the number N of light intensity change events generated in the pixel array circuit 210 based on the number of data signals provided by the reading circuit 220. ev The control circuit 230 may be based on the number of events N ev and the amount of event data associated with the event stream-based reading mode B ev , determine the total data volume of the light intensity change event, that is, the total data volume B of the event read in the first reading mode ev ·N ev Similarly, the control circuit 230 may also be based on the number of pixels M, the frame rate f and the amount of pixel data B of the pixel array circuit. p To determine the total data volume M·B of the light intensity change events read in the second reading mode p As shown in the following formula (3), the total data volume B of the light intensity change event read in the first reading mode is ev ·N ev Greater than or equal to the total data volume M·B of the light intensity change events of the second reading mode p f, the read circuit 220 should switch to the second read mode:

[0487] B ev ·N ev ≥η·M·B P ·f (3)

[0488] Where η is the switching parameter used for adjustment. From the above formula (3), it can be further concluded that the first threshold data volume d1 = η·M·B P f. If the total data volume of light intensity change events is B ev ·N ev If the data amount is greater than or equal to the threshold data amount d1, the control circuit 230 determines that the statistical data of the light intensity change event meets the predetermined switching condition. In this embodiment, the switching condition can be determined based on at least the number of pixels M, the frame rate f and the pixel data amount B of the pixel array circuit.p To determine the threshold data volume d1.

[0489] As an alternative implementation of the above embodiment, the total data volume B of the light intensity change event read in the first reading mode is ev ·N ev Greater than or equal to the total data volume M·B of the light intensity change events of the second reading mode p f can be expressed as the following formula (4):

[0490] M.B P ·fB ev ·N ev ≤θ (4)

[0491] Where θ is the switching parameter used for adjustment. From the above formula (4), it can be further concluded that the second threshold data volume d2 = M·B P f-θ, if the total data volume of light intensity change events is B ev ·N ev If the data amount is greater than or equal to the threshold data amount d2, the control circuit 230 determines that the statistical data of the light intensity change event meets the predetermined switching condition. In this embodiment, the switching condition can be determined based on at least the number of pixels M, the frame rate f and the pixel data amount B of the pixel array circuit. p To determine the threshold data amount d2.

[0492] In other embodiments, the statistical data may include the number of events N measured by the pixel array circuit 210 per unit time. ev If the first reading mode is a frame scanning based reading mode and the second reading mode is an event stream based reading mode, the control circuit 230 determines the number N of light intensity change events based on the number of first data signals in the plurality of data signals provided by the reading circuit 220. ev If the statistics indicate the number of light intensity change events N ev If the number n1 is less than the first threshold value, the control circuit 230 determines that the statistical data of the light intensity change event meets the predetermined switching condition, which can be based on at least the number of pixels M of the pixel array circuit, the frame rate f associated with the frame scanning-based reading mode, and the pixel data amount B. p , and the event data volume B associated with the event stream-based reading mode ev To determine the first threshold number n1. For example, in the above embodiment, based on formula (1), the following formula (5) can be further obtained:

[0493]

[0494] That is, the first threshold number n1 can be determined as

[0495] As an alternative implementation of the above embodiment, the following formula (6) can be further obtained based on formula (2):

[0496]

[0497] Accordingly, the second threshold number n2 can be determined as

[0498] In some other embodiments, if the first reading mode is an event stream-based reading mode and the second reading mode is a frame scan-based reading mode, the control circuit 230 may directly determine the number N of light intensity change events based on the number of at least one data signal provided by the reading circuit 220. ev If the statistics indicate the number of light intensity change events N ev If the number of light intensity change events is greater than or equal to the first threshold number n1, the control circuit 230 determines that the statistical data of the light intensity change event meets the predetermined switching condition. The switching condition can be determined based on at least the number of pixels M of the pixel array circuit 210, the frame rate f associated with the frame scanning-based reading mode, and the pixel data amount B. p , and the event data volume B associated with the event stream-based reading mode ev To determine the first threshold number n1=M·B p ·f / (η·B ev ). For example, in the aforementioned embodiment, based on formula (3), the following formula (7) can be further obtained:

[0499]

[0500] That is, the first threshold number n1 can be determined as

[0501] As an alternative implementation of the above embodiment, the following formula (8) can be further obtained based on formula (4):

[0502]

[0503] Accordingly, the second threshold number n2 can be determined as

[0504] It should be understood that the formulas, switching conditions and related calculation methods given above are merely an example implementation of the embodiments of the present application. Other suitable mode switching conditions, switching strategies and calculation methods may also be adopted, and the scope of the present application is not limited in this respect.

[0505] Figure 6-a A schematic diagram showing a vision sensor according to an embodiment of the present application operating in a frame scanning-based reading mode. Figure 6-bFIG. 1 shows a schematic diagram of a visual sensor according to an embodiment of the present application operating in an event stream-based reading mode. Figure 6-a As shown, the reading circuit 220 or 320 currently operates in the first reading mode, that is, the reading mode based on frame scanning. Since the control circuit 230 or 330 determines based on historical statistics that the number of events generated in the current pixel array circuit 210 or 310 is small, for example, there are only four valid data in a frame of data, it then predicts that the possible event generation rate in the next time period is low. If the reading circuit 220 or 320 continues to use the reading mode based on frame scanning for reading, it will be necessary to repeatedly allocate bits to the pixels that generate events, thereby generating a large amount of redundant data. In this case, the control circuit 230 or 330 sends a mode switching signal to the reading circuit 220 or 320 to switch the reading circuit 220 or 320 from the first reading mode to the second reading mode. After switching, Figure 6-b As shown, the read circuit 220 or 320 operates in the second read mode and only reads valid data signals, thereby avoiding the transmission bandwidth and storage resources occupied by a large number of invalid data signals.

[0506] Figure 6-c A schematic diagram showing a vision sensor according to an embodiment of the present application operating in an event stream-based reading mode. Figure 6-d FIG. 1 shows a schematic diagram of a visual sensor according to an embodiment of the present application operating in a frame scanning-based reading mode. Figure 6-c As shown, the reading circuit 220 or 320 currently operates in the first reading mode, that is, the reading mode based on the event stream. Since the control circuit 230 or 330 determines based on historical statistics that the number of events generated in the current pixel array circuit 210 or 310 is large, for example, in a short period of time, almost all pixels in the pixel array circuit 210 or 310 generate data signals indicating that the change in light intensity is higher than a predetermined threshold. Subsequently, the reading circuit 220 or 320 can predict that the possible event generation rate in the next time period is high. Since there is a large amount of redundant data in the read data signal, for example, almost the same pixel position information, reading timestamps, etc., if the reading circuit 220 or 320 continues to use the reading mode based on the event stream for reading, it will cause a surge in the amount of read data. Therefore, in this case, the control circuit 230 or 330 sends a mode switching signal to the reading circuit 220 or 320 to switch the reading circuit 220 or 320 from the first reading mode to the second reading mode. After switching, Figure 6-d As shown, the reading circuit 220 or 320 operates in a frame-based scanning mode to read data signals in a reading mode with a lower representation cost of a single pixel, thereby alleviating the pressure of storing and transmitting data signals.

[0507] In some possible embodiments, the visual sensor 200 or 300 may further include a parsing circuit, which may be configured to parse the data signal output by the reading circuit 220 or 320. In some possible embodiments, the parsing circuit may parse the data signal using a parsing mode that is compatible with the current data reading mode of the reading circuit 220 or 320. This will be described in detail below.

[0508] It should be understood that other existing or future-developed data reading modes, data reading modes, data parsing modes, etc. are also applicable to possible embodiments of the present application, and all numerical values ​​in the embodiments of the present application are illustrative rather than restrictive. For example, possible embodiments of the present application can switch between more than two data reading modes.

[0509] According to a possible embodiment of the present application, a visual sensor chip is provided that can adaptively switch between multiple reading modes based on historical statistics of light intensity change events generated in the pixel array circuit. This allows the visual sensor chip to consistently achieve excellent reading and analysis performance in both dynamic and static scenes, avoiding the generation of redundant data and alleviating the pressure on image processing, transmission, and storage.

[0510] Figure 7 FIG. 1 shows a flow chart of a method for operating a visual sensor chip according to a possible embodiment of the present application. In some possible embodiments, the method may be Figure 4-a The visual sensor 200 shown or Figure 4-b The visual sensor 300 shown and the following Figure 9 The electronic device shown in FIG. 1 is implemented, or any suitable device may be used to implement the electronic device, including various devices currently known or to be developed in the future. Figure 4-a The method is described using a vision sensor 200 as shown.

[0511] See Figure 7 , a method for operating a visual sensor chip provided by an embodiment of the present application may include the following steps:

[0512] 501. Generate a plurality of data signals corresponding to a plurality of pixels in a pixel array circuit.

[0513] The pixel array circuit 210 generates a plurality of data signals corresponding to a plurality of pixels in the pixel array circuit 210 by measuring the light intensity variation. In the context of this document, the data signal may indicate, but is not limited to, light intensity polarity, absolute light intensity value, light intensity variation value, etc.

[0514] 502. reading at least one data signal of the plurality of data signals from the pixel array circuit in a first read mode.

[0515] The read circuit 220 reads at least one data signal of the plurality of data signals from the pixel array circuit 210 in a first read mode, and these data signals occupy certain storage and transmission resources within the vision sensor 200 after being read. Depending on the specific read mode, the way the vision sensor chip 200 reads the data signals can be different. In some possible embodiments, for example, in an event stream based read mode, the read circuit 220 determines the location information (x, y) of the pixels related to the light intensity change events by scanning the pixel array circuit 210. Based on the location information, the read circuit 220 can read out a first data signal of the plurality of data signals. In this embodiment, the read circuit 220 obtains the feature information of the light intensity, the location information (x, y) of the pixels generating the light intensity change events, the time stamp t of reading the data signal, etc. by reading the data signal.

[0516] In other possible embodiments, the first read mode can be a frame scanning based read mode. In this mode, the vision sensor 200 scans the pixel array circuit 210 at a frame frequency associated with the frame scanning based read mode to read all the data signals generated by the pixel array circuit 210. In this embodiment, the read circuit 220 obtains the feature information of the light intensity by reading the data signal.

[0517] 503. providing the at least one data signal to a control circuit.

[0518] The read circuit 220 provides the read at least one data signal to the control circuit 230 for data statistics and analysis by the control circuit 230. In some embodiments, the control circuit 230 can determine statistical data related to at least one light intensity change event based on the at least one data signal. The control circuit 230 can analyze the statistical data by using a switching strategy module. If it is determined that the statistical data meets a predetermined switching condition, the control circuit 230 sends a mode switching signal to the read circuit 220.

[0519] In the case where the first reading mode is a frame scan-based reading mode and the second reading mode is an event stream-based reading mode, in some embodiments, the control circuit 230 may determine the number of light intensity change events based on the number of first data signals in the plurality of data signals. Furthermore, the control circuit 230 compares the number of light intensity change events with a first threshold number. If the statistical data indicates that the number of light intensity change events is less than or equal to the first threshold number, the control circuit 230 determines that the statistical data of the light intensity change events meets a predetermined switching condition and sends a mode switching signal. In this embodiment, the control circuit 230 may determine or adjust the first threshold number based on the number of pixels of the pixel array circuit, the frame rate and pixel data volume associated with the frame scan-based reading mode, and the event data volume associated with the event stream-based reading mode.

[0520] In some embodiments, when the first reading mode is an event stream-based reading mode and the second reading mode is a frame scan-based reading mode, the control circuit 230 may determine statistical data related to light intensity change events based on the first data signal received from the reading circuit 220. Furthermore, the control circuit 230 compares the number of light intensity change events with a second threshold number. If the number of light intensity change events is greater than or equal to the second threshold number, the control circuit 230 determines that the statistical data of the light intensity change events meets a predetermined switching condition and sends a mode switching signal. In this embodiment, the control circuit 230 may determine or adjust the second threshold number based on the number of pixels of the pixel array circuit, the frame rate and pixel data volume associated with the frame scan-based reading mode, and the event data volume associated with the event stream-based reading mode.

[0521] 504. Switch the first reading mode to the second reading mode based on the mode switching signal.

[0522] The readout circuit 220 switches from the first readout mode to the second readout mode based on a mode switching signal received from the control circuit 220. Furthermore, the readout circuit 220 reads at least one data signal generated by the pixel array circuit 210 in the second readout mode. The control circuit 230 can then continue to collect historical statistics on light intensity change events generated by the pixel array circuit 210 and, when a switching condition is met, send a mode switching signal to switch the readout circuit 220 from the second readout mode to the first readout mode.

[0523] According to the method provided in a possible embodiment of the present application, the control circuit continuously performs historical statistics and real-time analysis of light intensity change events generated in the pixel array circuit throughout the entire reading and parsing process. Once a switching condition is met, the control circuit sends a mode switching signal to switch the reading circuit from the current reading mode to a more appropriate alternative switching mode. This adaptive switching process is repeated until all data signals are read.

[0524] Figure 8 The block diagram of the control circuit of a possible embodiment of the present application is shown. The control circuit can be used to implement Figure 4-a The control circuit 230, Figure 5 The control circuit 330 in FIG. 3 and the like may also be implemented using other suitable devices. It should be understood that the control circuit is for exemplary purposes only and does not imply any limitation on the scope of the present application. The embodiments of the present application may also be embodied in different control circuits. In addition, it should be understood that the control circuit may also include other elements, modules, or entities, which are not shown for the purpose of clarity, but do not mean that the embodiments of the present application do not have these elements or entities.

[0525] like Figure 8 As shown, the control circuit includes at least one processor 602, at least one memory 604 coupled to the processor 602, and a communication mechanism 612 coupled to the processor 602. The memory 604 is used to store at least a computer program and a data signal obtained from the reading circuit. The statistical model 606 and the strategy module 608 are pre-configured on the processor 602. The control circuit 630 can be communicatively coupled to the control circuit 630 through the communication mechanism 612. Figure 4-a The reading circuit 220 of the visual sensor 200 shown or the reading circuit outside the visual sensor is used to realize the control function. Figure 4-a However, the embodiments of the present application are also applicable to the configuration of a peripheral reading circuit.

[0526] and Figure 4-a Similar to the control circuit 230 shown, in some possible embodiments, the control circuit can be configured to control the reading circuit 220 to read the multiple data signals generated by the pixel array circuit 210 in a specific data reading mode (for example, a synchronous reading mode based on frame scanning, an asynchronous reading mode based on event stream, etc.). In addition, the control circuit can be configured to obtain a data signal from the reading circuit 220, and the data signal can indicate, but is not limited to, light intensity polarity, absolute light intensity value, light intensity change value, etc. For example, light intensity polarity can indicate the trend of light intensity change, such as increase or decrease, usually expressed as +1 / -1. The absolute light intensity value can indicate the light intensity value measured at the current moment. Depending on the structure, purpose and type of the sensor, information about light intensity or light intensity change can have different physical meanings.

[0527] The control circuit determines statistical data related to at least one light intensity change event based on the data signal obtained from the reading circuit 220. In some embodiments, the control circuit can obtain the data signal generated by the pixel array circuit 210 over a period of time from the reading circuit 220 and store the data signals in the memory 604 for historical statistics and analysis. In the context of the present application, the first reading mode and the second reading mode can be one of an asynchronous reading mode based on an event stream and a synchronous reading mode based on a frame scan, respectively. However, it should be noted that all features described herein regarding adaptively switching the reading mode are also applicable to other types of sensors and data reading modes currently known or to be developed in the future, as well as switching between more than two data reading modes.

[0528] In some possible embodiments, the control circuitry can utilize one or more preconfigured statistical models 606 to generate historical statistics of light intensity change events generated by the pixel array circuitry 210 provided by the readout circuitry 220 over a period of time. The statistical model 606 can then transmit the statistical data to the strategy module 608 as an output. As previously described, the statistical data can indicate the number of light intensity change events or the total amount of light intensity change events. It should be understood that any suitable statistical model or statistical algorithm can be applied to the possible embodiments of the present application, and the scope of the present application is not limited in this respect.

[0529] Since the statistical data is a historical statistical result of light intensity change events generated by the visual sensor over a period of time, it can be used by the strategy module 608 to analyze and predict the rate of event occurrence in the next time period. The strategy module 608 can be pre-configured with one or more switching decisions. When multiple switching decisions exist, the control circuit can select one from the multiple switching decisions for analysis and decision-making as needed, for example, based on factors such as the type of visual sensor 200, the characteristics of the light intensity change event, the properties of the external environment, the motion state, etc. In possible embodiments of the present application, other suitable strategy modules and mode switching conditions or strategies may also be adopted, and the scope of the present application is not limited in this respect.

[0530] In some embodiments, if the policy module 608 determines that the statistical data meets the mode switching condition, an instruction to switch the read mode is output to the read circuit 220. In another embodiment, if the policy module 608 determines that the statistical data does not meet the mode switching condition, no instruction to switch the read mode is output to the read circuit 220. In some embodiments, the instruction to switch the read mode can be in an explicit form as described in the above embodiments, for example, in the form of a switching signal or a flag bit to notify the read circuit 220 to switch the read mode.

[0531] Figure 9A block diagram of an electronic device according to possible embodiments of the present application is shown. As Figure 9 shown, the electronic device includes a visual sensor chip 901, communication interfaces 902 and 903, a control circuit 930, and a parsing circuit 904. It should be understood that the electronic device is for exemplary purposes, which can be implemented with any suitable device, including various sensor devices currently known and developed in the future. Embodiments of the present application can also be embodied in different sensor systems. In addition, it should also be understood that the electronic device can also include other elements, modules or entities, which are not shown for the purpose of clarity, but do not mean that embodiments of the present application do not have these elements, modules or entities.

[0532] As Figure 9 shown, the visual sensor includes a pixel array circuit 710 and a read circuit 720, wherein read components 720-1 and 720-2 of the read circuit 720 are coupled to the control circuit 730 via communication interfaces 702 and 703, respectively. In embodiments of the present application, the read components 720-1 and 720-2 can be implemented with independent devices, respectively, or can be integrated in the same device. For example, Figure 4-a The read circuit 220 shown is an integrated example implementation. For ease of description, the read components 720-1 and 720-2 can be configured to implement data reading functions in a frame scan-based reading mode and an event stream-based reading mode, respectively.

[0533] The pixel array circuit 710 can be implemented with the pixel array circuit 210 in Figure 4-a or the pixel array circuit 310 in Figure 5 , or any suitable other device, which is not limited in this regard by the present application. Features regarding the pixel array circuit 710 are not described herein again.

[0534] The read circuit 720 can read data signals generated by the pixel array circuit 710 in a specific reading mode. For example, in an example in which the read component 720-1 is turned on and the read component 720-2 is turned off, the read circuit 720 initially reads data signals in a frame scan-based reading mode. In an example in which the read component 720-2 is turned on and the read component 720-1 is turned off, the read circuit 720 initially reads data signals in an event stream-based reading mode. The read circuit 720 is implemented with the read circuit 220 in Figure 4-a or the read circuit 320 in Figure 5 , or any suitable other device, which is not limited in this regard by the present application. Features regarding the read circuit 720 are not described herein again.

[0535] In an embodiment of the present application, the control circuit 730 may instruct the read circuit 720 to switch from the first read mode to the second read mode by means of an indication signal or a flag bit. In this case, the read circuit 720 may receive an instruction from the control circuit 730 regarding switching the read mode, for example, turning on the read component 720-1 and turning off the read component 720-2, or turning on the read component 720-2 and turning off the read component 720-1.

[0536] As mentioned above, the electronic device may further include a parsing circuit 704. The parsing circuit 704 may be configured to parse the data signal read by the reading circuit 720. In a possible embodiment of the present application, the parsing circuit may adopt a parsing mode adapted to the current data reading mode of the reading circuit 720. For example, if the reading circuit 720 initially reads the data signal in a reading mode based on an event stream, the parsing circuit may accordingly determine the first data amount B associated with the reading mode. ev ·N ev When the reading circuit 720 switches from the event stream-based reading mode to the frame scan-based reading mode based on the instruction of the control circuit 730, the parsing circuit starts to read the data in accordance with the second data size, that is, the size of one frame of data M·B. p to analyze the data signal and vice versa.

[0537] In some embodiments, the parsing circuit 704 can implement the switching of the parsing mode of the parsing circuit without the need for an explicit switching signal or flag bit. For example, the parsing circuit 704 can use the same or corresponding statistical model and switching strategy as the control circuit 730 to perform the same statistical analysis on the data signal provided by the reading circuit 720 as the control circuit 730 and make consistent switching predictions. For example, if the reading circuit 720 initially reads the data signal in the event stream-based reading mode, the parsing circuit initially selects the first data amount B associated with the reading mode. ev ·N ev To parse the data. For example, the first b parsed by the parsing circuit x bits indicate the pixel's coordinate x, followed by b y bits indicate the y coordinate of the pixel, followed by b t bits indicate the reading time, and finally take b f The parsing circuit 704 obtains at least one data signal from the reading circuit 720 and determines statistical data related to at least one light intensity change event. If the parsing circuit 704 determines that the statistical data meets the switching condition, it switches to the parsing mode corresponding to the frame scanning-based reading mode, with the frame data size M·B. p To analyze the data signal.

[0538] As another example, if the read circuit 720 initially reads the data signal in a frame-scan based read mode, the parsing circuit 704 parses the data signal in a frame-based parsing mode corresponding to the read mode, taking out the value of each pixel position in the frame in sequence, where the value of a pixel position in which no light intensity change event occurs is 0. The parsing circuit 704 can count the number of non-0 values in a frame based on the data signal, i.e., the number of light intensity change events in the frame. p As another example, if the read circuit 720 initially reads the data signal in a frame-scan based read mode, the parsing circuit 704 parses the data signal in a frame-based parsing mode corresponding to the read mode, taking out the value of each pixel position in the frame in sequence, where the value of a pixel position in which no light intensity change event occurs is 0. The parsing circuit 704 can count the number of non-0 values in a frame based on the data signal, i.e., the number of light intensity change events in the frame.

[0539] In some possible embodiments, the parsing circuit 704 obtains at least one data signal from the read circuit 720, and determines which one of the current parsing mode and the alternative parsing mode corresponds to the read mode of the read circuit 720 based on at least the at least one data signal. In turn, in some embodiments, the parsing circuit 704 can switch from the current parsing mode to the other parsing mode based on the determination.

[0540] In some possible embodiments, the parsing circuit 704 can determine whether to switch the parsing mode based on historical statistics of the light intensity change events. For example, the parsing circuit 704 can determine statistics data related to at least one light intensity change event based on at least one data signal received from the read circuit 720. If the statistics data is determined to satisfy a switching condition, the parsing circuit 704 switches from the current parsing mode to the alternative parsing mode. For ease of comparison, the statistics data can be used to measure the read data rate of the first read mode and the second read mode of the read circuit 720, respectively.

[0541] In some embodiments, the statistics data can include the total amount of data of the number of events measured by the pixel array circuit 710 per unit time. If the parsing circuit 704 determines based on the at least one data signal that the total amount of data of the light intensity change events read by the read circuit 720 in the first read mode has been greater than or equal to that of the light intensity change events in the second read mode, it indicates that the read circuit 720 has switched from the first read mode to the second read mode. In this case, the parsing circuit 704 should accordingly switch to the parsing mode corresponding to the current read mode.

[0542] In some embodiments, the first read mode is a frame-scan based read mode and the second read mode is an event-stream based read mode. In this embodiment, the parsing circuit 704 initially parses the data signal obtained from the read circuit 720 in a frame-based parsing mode corresponding to the first read mode. The parsing circuit 704 can determine the total amount of data of the light intensity change events read by the read circuit 720 in the first read mode based on the number of pixels M of the pixel array circuit 710, the frame rate f, and the pixel data amount B p p ·f. The parsing circuit 704 can determine the total amount of data of the light intensity change events read by the read circuit 720 in the first read mode based on the number of pixels M of the pixel array circuit 710, the frame rate f, and the pixel data amount B​ev and the amount of event data associated with the event stream-based reading mode B ev , to determine the total data amount B of the light intensity change event read by the reading circuit 720 in the second reading mode ev ·N ev In some embodiments, the switching parameter can be used to adjust the relationship between the total data amounts in the two reading modes. Further, the analysis circuit 704 can determine the total data amount M·B of the light intensity change event read by the reading circuit 720 in the first reading mode according to, for example, the formula (1) above. p Is f greater than or equal to the total data volume B of the light intensity change events in the second reading mode? ev ·N ev If so, the parsing circuit 704 determines that the reading circuit 720 has switched to the event stream-based reading mode, and accordingly switches from the frame-based parsing mode to the event stream-based parsing mode.

[0543] As an alternative implementation of the above embodiment, the analysis circuit 704 can determine the total data volume M·B of the light intensity change event read by the reading circuit 720 in the first reading mode according to the above formula (2): p Is f greater than or equal to the total data volume B of the light intensity change events read in the second reading mode? ev ·N ev Similarly, when determining the total data amount M·B of the light intensity change events read by the reading circuit 720 in the first reading mode, p f is greater than or equal to the total data volume B of the light intensity change events in the second reading mode ev ·N ev In the case of , the parsing circuit 704 determines that the reading circuit 720 has switched to the event stream-based reading mode, and accordingly switches from the frame-based parsing mode to the event stream-based parsing mode.

[0544] In some embodiments, the first reading mode is an event-stream-based reading mode and the second reading mode is a frame-scan-based reading mode. In this embodiment, the parsing circuit 704 initially parses the data signal obtained from the reading circuit 720 in the event-stream-based parsing mode corresponding to the first reading mode. As previously described, the parsing circuit 704 can directly determine the number N of light intensity change events generated in the pixel array circuit 710 based on the number of first data signals provided by the reading circuit 720. ev The parsing circuit 704 can be based on the number of events N ev and the amount of event data associated with the event stream-based reading mode B ev , determine the total data amount B of the event read by the read circuit 720 in the first read mode ev ·Nev Similarly, the analysis circuit 704 can also be based on the number of pixels M, the frame rate f and the amount of pixel data B of the pixel array circuit. p To determine the total data amount M·B of the light intensity change events read by the reading circuit 720 in the second reading mode p Then, the analysis circuit 704 can determine the total data volume B of the light intensity change events read in the first reading mode, for example, according to the above formula (3): ev ·N ev Is it greater than or equal to the total data volume M·B of the light intensity change events in the second reading mode? p Similarly, when the parsing circuit 704 determines the total data volume B of the light intensity change events read by the reading circuit 720 in the first reading mode, ev ·N ev Greater than or equal to the total data volume M·B of the light intensity change events of the second reading mode p When f, the parsing circuit 704 determines that the reading circuit 720 has switched to the frame-scanning-based reading mode, and accordingly switches from the event-stream-based parsing mode to the frame-based parsing mode.

[0545] As an alternative implementation of the above embodiment, the analysis circuit 704 can determine the total data volume B of the light intensity change event read by the reading circuit 720 in the first reading mode according to the formula (4) above. ev ·N ev Is it greater than or equal to the total data volume M·B of the light intensity change events read in the second reading mode? p Similarly, when determining the total data amount B of the light intensity change events read by the reading circuit 720 in the first reading mode, ev ·N ev Greater than or equal to the total data volume M·B of the light intensity change events of the second reading mode p In the case of f, the parsing circuit 704 determines that the reading circuit 720 has switched to the frame-scanning-based reading mode, and accordingly switches from the event-stream-based parsing mode to the frame-scanning-based parsing mode.

[0546] For the read time t of an event in the frame-scanning read mode, it is assumed that all events in the same frame have the same read time t. When the accuracy of the event read time is required to be high, the read time of each event can be further determined in the following manner. Taking the above embodiment as an example, in the frame-scanning read mode, the frequency of the read circuit 720 scanning the pixel array circuit is f Hz, then the time interval between reading two adjacent frames of data is S = 1 / f, and the start time of each frame is given as:

[0547] T k =T0+kS(9)

[0548] Where T0 is the start time of the first frame, k is the frame number, and the time required for digital-to-analog conversion of one pixel in M ​​pixels can be determined by the following formula (10):

[0549]

[0550] The time when the light intensity change event occurs at the i-th pixel in the k-th frame can be determined by the following formula (11):

[0551]

[0552] Where i is a positive integer. If the current reading mode is synchronous reading mode, it will switch to asynchronous reading mode. ev Bit parsing data. In the above embodiment, the switching of the parsing mode can be achieved without an explicit switching signal or flag bit. For other currently known or future developed data reading modes, the parsing circuit can also parse the data in a similar manner that is compatible with the data reading mode, which will not be described in detail here.

[0553] Figure 10 A schematic diagram showing the change in data volume over time in a single data reading mode and an adaptive switching reading mode according to a possible embodiment of the present application is shown. Figure 10 The left half of FIG depicts a schematic diagram of the change in the amount of read data over time for a conventional visual sensor or sensor system that uses only a synchronous read mode or an asynchronous read mode. In the case of using only a synchronous read mode, as shown by curve 1001, since each frame has a fixed amount of data, the amount of data read remains unchanged over time, that is, the read data rate (the amount of data read per unit time) is stable. As mentioned above, when a large number of events are generated in the pixel array circuit, it is more reasonable to use a frame scanning-based read mode to read the data signal. The vast majority of the frame data is valid data representing the generated events, and there is less redundancy. However, when fewer events are generated in the pixel array circuit, a large amount of invalid data representing and generating events exists in a frame. At this time, still using the frame data structure to represent and read the light intensity information at the pixel will generate redundancy, wasting transmission bandwidth and storage resources.

[0554] In the case of using only the asynchronous reading mode, as shown in curve 1002, the amount of data read varies with the rate at which events are generated, and thus the reading data rate is not fixed. When there are fewer events generated in the pixel array circuit, only bits representing the pixel coordinate information (x, y), the timestamp t at which the data signal is read, and the characteristic information f of the light intensity need to be allocated for a small number of events. The total amount of data required to be read is small, and in this case, it is more reasonable to use the asynchronous reading mode. When a large number of events are generated in the pixel array circuit in a short period of time, a large number of bits need to be allocated to represent these events. However, these pixel coordinates are almost adjacent, and the reading time of the data signal is almost the same. In other words, there is a large amount of duplicate data in the read event data, and thus the problem of redundancy also exists in the asynchronous reading mode. In this case, the reading data rate even exceeds the reading data rate in the synchronous reading mode, and it is unreasonable to still use the asynchronous reading mode.

[0555] Figure 10 The right half of FIG depicts a schematic diagram of the data amount changing over time in the adaptive data reading mode according to a possible embodiment of the present application. The adaptive data reading mode can utilize Figure 4-a The visual sensor 200 shown, Figure 4-b The visual sensor 300 shown or Figure 9 The electronic devices shown can be implemented, or a conventional visual sensor or sensor system can be implemented by utilizing Figure 8 The control circuit shown in FIG. 1 realizes the adaptive data reading mode. For the convenience of description, the following reference Figure 4-a The characteristics of the adaptive data reading mode are described in the visual sensor 200 shown in FIG. As shown in curve 1003, the visual sensor 200 selects, for example, the asynchronous reading mode in the initialization state. Since the number of bits B used to represent each event in this mode is ev is predetermined (e.g., B ev =b x +b y +b t +b f ), as events are generated and read, the visual sensor 200 can calculate the reading data rate in the current mode. On the other hand, the number of bits B used to represent each pixel of each frame in the synchronous reading mode is pis also predetermined, so the read data rate using the synchronous read mode in the time period can be calculated. The vision sensor 200 can then determine whether the relationship between the data rates in the two read modes satisfies a mode switching condition. For example, the vision sensor 200 can compare the read data rates in the two read modes based on a predefined threshold to determine which mode has a smaller read data rate. Once it is determined that the mode switching condition is satisfied, the vision sensor 200 switches to the other read mode, e.g., from the initial asynchronous read mode to the synchronous read mode. The above steps continue during the reading and parsing of the data signal until the output of all data is completed. As shown in the curve 1003, the vision sensor 200 adaptively selects the optimal read mode throughout the data reading process, and the two read modes alternate, so that the read data rate of the vision sensor 200 always remains below the read data rate of the synchronous read mode, thereby reducing the cost of data transmission, parsing and storage of the vision sensor.

[0556] In addition, according to the adaptive data reading mode proposed by the embodiments of the present application, the vision sensor 200 can statistically analyze the historical data of events to predict the possible event generation rate in the next time period, so as to select a read mode that is more suitable for the application scenario and the motion state.

[0557] Through the above scheme, the vision sensor can adaptively switch between multiple data reading modes, so that the read data rate always remains below the predetermined read data rate threshold, thereby reducing the cost of data transmission, parsing and storage of the vision sensor, and significantly improving the performance of the sensor. In addition, such a vision sensor can statistically analyze the events generated in a time period to predict the possible event generation rate in the next time period, so as to select a read mode that is more suitable for the current external environment, application scenario and motion state.

[0558] The above describes that the pixel array circuit can be used to measure the light intensity change amount, and generate a plurality of data signals corresponding to a plurality of pixels. The data signal can indicate, but is not limited to, the light intensity polarity, the absolute light intensity value, the light intensity change value, and the like. The data signal output by the pixel array circuit is described in detail below.

[0559] Figure 11 A schematic diagram of a pixel circuit 900 provided by the present application is shown. Each of the pixel array circuit 210, the pixel array circuit 310 and the pixel array circuit 710 can include one or more pixel arrays, and each pixel array includes a plurality of pixels, each of which can be regarded as a pixel circuit, and each pixel circuit is used to generate a data signal corresponding to the pixel. Referring to Figure 11, is a schematic diagram of a preferred pixel circuit provided in an embodiment of the present application. In this application, a pixel circuit is sometimes referred to as a pixel. Figure 11 As shown, a preferred pixel circuit in the present application includes a light intensity detection unit 901 , a threshold comparison unit 902 , a readout control unit 903 and a light intensity collection unit 904 .

[0560] The light intensity detection unit 901 is used to convert the acquired light signal into a first electrical signal. The light intensity detection unit 901 can monitor the light intensity information irradiated on the pixel circuit in real time, and convert the acquired light signal into an electrical signal and output it in real time. In some possible embodiments, the light intensity detection unit 901 can convert the acquired light signal into a voltage signal. The present application does not limit the specific structure of the light intensity detection unit. Any structure that can convert the light signal into an electrical signal can be adopted in the embodiments of the present application. For example, the light intensity detection unit may include a photodiode and a transistor. The anode of the photodiode is grounded, the cathode of the photodiode is connected to the source of the transistor, and the drain and gate of the transistor are connected to a power supply.

[0561] Threshold comparison unit 902 is used to determine whether the first electrical signal is greater than a first target threshold or less than a second target threshold. When the first electrical signal is greater than the first target threshold or less than the second target threshold, threshold comparison unit 902 outputs a first data signal, which indicates that a light intensity change event has occurred at the pixel. Threshold comparison unit 902 compares whether the difference between the current light intensity and the light intensity at the time of the previous event exceeds a predetermined threshold. This can be understood with reference to Formula 1-1. The first target threshold can be understood as the sum of the first predetermined threshold and the second electrical signal, and the second target threshold can be understood as the sum of the second predetermined threshold and the second electrical signal. The second electrical signal is the electrical signal output by light intensity detection unit 901 when the previous event occurred. The threshold comparison unit in the embodiments of the present application can be implemented in hardware or software, and this is not limited in the embodiments of the present application. The type of the first data signal output by threshold comparison unit 902 may vary: in some possible embodiments, the first data signal includes polarity information, such as +1 or -1, to indicate an increase or decrease in light intensity. In some possible embodiments, the first data signal may be an activation signal, used to instruct the readout control unit 903 to control the light intensity acquisition unit 904 to acquire the first electrical signal and buffer the first electrical signal. When the first data signal is an activation signal, the first data signal may also be polarity information. In this case, when the readout control unit 903 obtains the first data signal, it controls the light intensity acquisition unit 904 to acquire the first electrical signal.

[0562] The readout control unit 903 is further configured to notify the readout circuit to read the first electrical signal stored in the light intensity acquisition unit 904, or to notify the readout circuit to read the first data signal output by the threshold comparison unit 902, where the first data signal is polarity information.

[0563] The reading circuit 905 can be configured to scan the pixels in the pixel array circuit in a predetermined order to read the data signals generated by the corresponding pixels. In some possible embodiments, the reading circuit 905 can be understood with reference to the reading circuit 220, the reading circuit 320, and the reading circuit 720, that is, the reading circuit 905 is configured to be able to read the data signal output by the pixel circuit in more than one signal reading mode. For example, the reading circuit 905 can read in one of the first reading mode and the second reading mode, and the first reading mode and the second reading mode correspond to one of the frame scanning-based reading mode and the event stream-based reading mode, respectively. In some possible embodiments, the reading circuit 905 can also read the data signal output by the pixel circuit in only one signal reading mode, such as the reading circuit 905 is configured to read the data signal output by the pixel circuit in only the frame scanning-based reading mode, or the reading circuit 905 is configured to read the data signal output by the pixel circuit in only the event stream-based reading mode. Figure 11 In corresponding embodiments, the data signal read by the reading circuit 905 is represented in different ways, that is, in some possible embodiments, the data signal read by the reading circuit is represented by polarity information, for example, the reading circuit can read the polarity information output by the threshold comparison unit; in some possible embodiments, the data signal read by the reading circuit can be represented by light intensity information, for example, the reading circuit can read the electrical signal cached by the light intensity acquisition unit.

[0564] See Figure 11-a , taking the reading mode based on event stream to read the data signal output by the pixel circuit as an example, the events represented by light intensity information and the events represented by polarity information are explained. Figure 11-a As shown in the upper part of the figure, the black dots represent the pixels that generate light intensity change events. Figure 11-a There are 8 events in total, of which the first 5 events are represented by light intensity information, and the last 3 events are represented by polarity information. Figure 11-aAs shown in the lower half of FIG. 9, both the event represented by the light intensity information and the event represented by the polarity information need to include coordinate information (x, y), time information t, and the difference is that the characteristic information m of the light intensity of the event represented by the light intensity information is the light intensity information a, and the characteristic information m of the light intensity of the event represented by the polarity information is the polarity information p. The difference between the light intensity information and the polarity information has been described above and will not be repeated here, and only it is emphasized that the amount of data of the event represented by the polarity information is less than the amount of data represented by the light intensity information.

[0565] As to how to determine which information is used to represent the data signal read by the read circuit, it needs to be determined according to the indication issued by the control circuit, which will be described in detail below.

[0566] In some embodiments, the read circuit 905 can be configured to provide the control circuit 906 with at least one data signal read. For example, the read circuit 905 can provide the control circuit 906 with the total amount of data of the data signal read in a period of time for the control circuit 906 to perform historical data statistics and analysis. In one embodiment, the control circuit 906 can obtain the number of events N generated per second by the pixel array circuit based on the number of light intensity change events generated by each pixel circuit 900 in the pixel array circuit in a period of time. ev ev The N can be obtained by one of the frame scanning based read mode and the event stream based read mode.

[0567] The control circuit 906 is coupled to the read circuit 905 and is configured to control the read circuit 906 to read the data signal generated by the pixel circuit 900 in a specific event representation manner. In some possible embodiments, the control circuit 906 can obtain at least one data signal from the read circuit 905, and determine which of the current event representation manner and the alternative event representation manner is more suitable for the current application scenario and motion state based on at least the at least one data signal. Further, in some embodiments, the control circuit 906 can instruct the read circuit 905 to switch from the current event representation manner to another event representation manner based on the determination.

[0568] In some possible embodiments, the control circuit 906 can send an indication to the read circuit 905 about switching the event representation manner based on the historical statistics of the light intensity change events. For example, the control circuit 906 can determine statistical data related to at least one light intensity change event based on at least one data signal received from the read circuit 905. If the statistical data is determined to satisfy a predetermined switching condition, the control circuit 906 sends an indication signal to the read circuit 905 to make the read circuit 905 switch the read event format.

[0569] ​In some possible embodiments, assuming that the reading circuit 905 is configured to read the data signal output by the pixel circuit only in a reading mode based on an event stream, the data provided by the reading circuit 905 to the control circuit 906 is the total data amount of the number of events (light intensity change events) measured by the pixel array circuit per unit time. Assuming that the current control circuit 906 controls the reading circuit 905 to read the data output by the threshold comparison unit 902, that is, the event is represented by polarity information, then the reading circuit 905 can be based on the number N of light intensity change events. ev , the bit width H of the data format determines the total data volume N of the light intensity change event ev ×H. Among them, the bit width of the data format is H = b x +b y +b t +b p , b p Bits are used to represent the polarity information of the light intensity indicated by the data signal, usually 1 to 2 bits. Since the polarity information of the light intensity is usually represented by 1 to 2 bits, the total data volume of the event represented by the polarity information must be less than the bandwidth. In order to ensure that the event data with higher precision can be transmitted as much as possible without exceeding the bandwidth limit, if the total data volume of the event represented by the light intensity information is also less than or equal to the bandwidth, it is converted to represent the event by the light intensity information. In some embodiments, the conversion parameter can be used to adjust the relationship between the data volume and the bandwidth K under a certain event representation method, as shown in the following formula (12), the total data volume N of the event represented by the light intensity information is ev ×H is less than or equal to the bandwidth.

[0570] N ev ×H≤α×K(12)

[0571] Wherein α is the conversion parameter used for adjustment. It can be further concluded from the above formula (12) that if the total amount of event data represented by the light intensity information is less than or equal to the bandwidth, the control circuit 906 can determine that the statistical data of the light intensity change event meets the predetermined switching condition. Some possible application scenarios include when the pixel acquisition circuit generates fewer events within a period of time, or when the rate at which the pixel acquisition circuit generates events within a certain period of time is slow. In these cases, the events can be represented by light intensity information. Since the events represented by the light intensity information can carry more information, it is beneficial to the subsequent processing and analysis of the events, for example, it can improve the quality of image reconstruction.

[0572] In some embodiments, assuming that the current control circuit 906 controls the reading circuit 905 to read the electrical signal cached by the light intensity acquisition unit 904, that is, the event is represented by the light intensity information, the reading circuit 905 can be based on the number N of light intensity change events. ev , the bit width H of the data format determines the total data volume N of the light intensity change eventev ×H. When the event stream-based reading mode is used, the data format bit width H = b x +b y +b t +b a , b a Bits are used to represent the light intensity information indicated by the data signal, usually multiple bits, such as 8 bits to 12 bits. In some embodiments, the conversion parameter can be used to adjust the relationship between the data volume and bandwidth K under an event representation mode, as shown in the following formula (13): the total data volume N of the event represented by the light intensity information is ev ×H is greater than, the reading circuit 220 should read the data output by the threshold comparison unit 902, that is, convert it into an event represented by polarity information:

[0573] N ev ×H>β×K(13)

[0574] Where β is the conversion parameter used for adjustment. From the above formula (13), it can be further concluded that if the total data volume N of the light intensity change event ev If ×H is greater than the threshold data volume β × K, it indicates that the total data volume of the light intensity change event represented by the light intensity information is greater than or equal to the bandwidth, and the control circuit 905 can determine that the statistical data of the light intensity change event meets the predetermined conversion condition. Some possible application scenarios include when the pixel acquisition circuit generates a large number of events over a period of time, or when the pixel acquisition circuit generates events at a high rate over a period of time. In these cases, if light intensity information is continued to be used to represent events, events may be lost. Therefore, using polarity information to represent events can alleviate the pressure of data transmission and reduce data loss.

[0575] In some embodiments, the data provided by the reading circuit 905 to the control circuit 906 is the number of events N measured by the pixel array circuit per unit time. ev In some possible embodiments, assuming that the current control circuit 906 controls the reading circuit 905 to read the data output by the threshold comparison unit 902, that is, the event is represented by the polarity information, the control circuit can determine the number N of light intensity change events. ev and The relationship between N and N is used to determine whether the predetermined conversion condition is met. ev Less than or equal to The reading circuit 220 should read the electrical signal cached in the light intensity acquisition unit 904, that is, convert it into a passing light intensity information representation event, and convert the current passing polarity information representation event into a passing light intensity information representation event. For example, in the above embodiment, based on formula (12), the following formula (14) can be further obtained:

[0576]

[0577] In some embodiments, assuming that the current control circuit 906 controls the reading circuit 905 to read the electrical signal cached by the light intensity acquisition unit 904, that is, the event is represented by the light intensity information, the control circuit 906 can be based on the number N of light intensity change events ev and The relationship between N and N is used to determine whether the predetermined conversion condition is met. ev Greater than The reading circuit 220 should read the signal output from the threshold comparison unit 902, that is, convert it into a passing polarity information indicating an event, and convert the current passing light intensity information indicating an event into a passing polarity information indicating an event. For example, in the aforementioned embodiment, based on formula (12), the following formula (15) can be further obtained:

[0578]

[0579] In some possible embodiments, assuming that the reading circuit 905 is configured to read the data signal output by the pixel circuit only in a frame-scanning reading mode, the data provided by the reading circuit 905 to the control circuit 906 is the total data amount of the number of events (light intensity change events) measured by the pixel array circuit per unit time. When the frame-scanning reading mode is adopted, the bit width of the data format is H=B p , B p The amount of pixel data (eg, the number of bits) allocated to each pixel in the frame scan-based reading mode. When an event is represented by polarity information, B p It is usually 1 to 2 bits, and when the event is represented by light intensity information, it is usually 8 to 12 bits. The reading circuit 905 can determine the total data volume M×H of the light intensity change event, where M represents the total number of pixels. Assume that the current control circuit 906 controls the reading circuit 905 to read the data output by the threshold comparison unit 902, that is, the event is represented by polarity information. The total data volume of the event represented by the polarity information must be less than the bandwidth. In order to ensure that higher precision event data can be transmitted as much as possible without exceeding the bandwidth limit, if the total data volume of the event represented by the light intensity information is also less than or equal to the bandwidth, it is converted to represent the event by light intensity information. In some embodiments, the conversion parameter can be used to adjust the relationship between the data volume and the bandwidth K under a certain event representation method, as shown in the following formula (16), the total data volume N of the event represented by the light intensity information ev ×H is less than or equal to the bandwidth.

[0580] M×H≤α×K(16)

[0581] In some embodiments, assuming that the current control circuit 906 controls the read circuit 905 to read the electrical signal cached by the light intensity collection unit 904, i.e., the event is represented by the light intensity information, the read circuit 905 can determine the total data amount MxH of the light intensity change event. In some embodiments, a conversion parameter can be used to adjust the relationship between the data amount and the bandwidth K in one event representation mode, as shown in the following equation (17). The total data amount MxH of the event represented by the light intensity information is greater than the bandwidth, and the read circuit 220 should read the data output by the threshold comparison unit 902, i.e., convert it to an event represented by the polarity information:

[0582] MxH > a x K (17)

[0583] In some possible embodiments, assuming that the read circuit 905 is configured to read in one of a first read mode and a second read mode, the first read mode and the second read mode correspond to one of a frame scanning-based read mode and an event stream-based read mode, respectively. For example, the following describes how the control circuit determines whether the switching condition is met in the combination mode in which the read circuit 905 currently reads the data signal output by the pixel circuit in the event stream-based read mode and the control circuit 906 controls the read circuit 905 to read the data output by the threshold comparison unit 902, i.e., the event is represented by the polarity information:

[0584] In the initial state, one read mode can be arbitrarily selected, such as the frame scanning-based read mode or the event stream-based read mode. In addition, in the initial state, one event representation mode can be arbitrarily selected, such as the control circuit 906 controlling the read circuit 905 to read the electrical signal cached by the light intensity collection unit 904, i.e., the event is represented by the light intensity information, or the control circuit 906 controlling the read circuit 905 to read the data output by the threshold comparison unit 902, i.e., the event is represented by the polarity information. Assuming that the read circuit 905 currently reads the data signal output by the pixel circuit in the event stream-based read mode and the control circuit 906 controls the read circuit 905 to read the data output by the threshold comparison unit 902, i.e., the event is represented by the polarity information. The data provided by the read circuit 905 to the control circuit 906 can be a first total data amount of the number of events (light intensity change events) measured by the pixel array circuit per unit time. Since the total number of pixels M is known, the pixel data amount B allocated to each pixel in the frame scanning-based read mode is known, and the bit width H of the data format when the event is represented by the light intensity information is known. According to the known M, B p p ​, H can obtain a reading mode based on event flow to read the data signal output by the pixel circuit, and the second total data volume of the number of events measured by the pixel array circuit per unit time in this combination mode of representing the event by light intensity information; can obtain a reading model based on frame scanning to read the data signal output by the pixel circuit, and the third total data volume of the number of events measured by the pixel array circuit per unit time in this combination mode of representing the event by polarity information; can obtain a reading model based on frame scanning to read the data signal output by the pixel circuit, and the fourth total data volume of the number of events measured by the pixel array circuit per unit time in this combination mode of representing the event by light intensity information. Specifically according to M, B p , H calculate the second data amount, the third data amount and the fourth data amount have been introduced above and will not be repeated here. The first total data amount provided by the above-mentioned reading circuit 905, the second total data amount, the third total data amount and the fourth total amount obtained by calculation, and the relationship between them and the bandwidth K are used to determine whether the switching condition is met. If the current combination mode cannot guarantee that event data with higher accuracy is transmitted as much as possible without exceeding the bandwidth limit, it is determined that the switching condition is met and the combination mode is switched to a combination mode that can guarantee that event data with higher accuracy is transmitted as much as possible without exceeding the bandwidth limit.

[0585] To better understand the above process, let's take a specific example to illustrate:

[0586] Assume that the bandwidth limit is K and the bandwidth adjustment factor is α. In the event stream-based reading mode, when the event is represented by polarity information, the bit width of the data format is H = b x +b y +b t +b p When the event is represented by light intensity information, the bit width of the data format is H = b x +b y +b t +b a . Usually 1≤b p a , such as b p Usually 1 to 2 bits, b a Typically 8 to 12 bits.

[0587] In the frame scan-based reading mode, events do not need to represent coordinates and time. Events are determined based on the state of each pixel. Assume that the data bit width allocated to each pixel is b in polarity mode. sp , b in light intensity mode sa , the total number of pixels is M. Assume that the bandwidth is limited to K = 1000 bps, b x =5bit b y =4bit,b​t = 10 bit, b p = 1 bit, b a = 8 bit, b sp = 1 bit, b sa = 8 bit, M = 100, bandwidth adjustment factor a = 0.9. Assume that 10 events are generated in the first second, 15 events are generated in the second second, and 30 events are generated in the third second.

[0588] Assume that the initial state is to use the event stream-based reading mode by default, and the event is represented in the polarity mode.

[0589] The event stream-based reading mode and the event represented by the polarity information are referred to as the asynchronous polarity mode, the event stream-based reading mode and the event represented by the light intensity information are referred to as the asynchronous light intensity mode, the frame scanning-based reading mode and the event represented by the polarity information are referred to as the synchronous polarity mode, and the frame scanning-based reading mode and the event represented by the light intensity information are referred to as the synchronous light intensity mode.

[0590] First second: 10 events are generated

[0591] Asynchronous polarity mode: N ev = 10, H = b x + b y + b t + b p = 5 + 4 + 10 + 1 = 20 bits, estimated data amount N ev · H = 200 bits, N ev · H < a · K, bandwidth limit is met.

[0592] Asynchronous light intensity mode: H = b x + b y + b t + b a = 5 + 4 + 10 + 8 = 27 bits, then the estimated data amount N in the light intensity mode ev · H = 270 bits, N ev · H < a · K, bandwidth limit is still met.

[0593] Synchronous polarity mode: M = 100, H = b sp = 1 bit, the estimated data amount is M · H = 100 bits at this time, M · H < a · K, bandwidth limit is still met.

[0594] Synchronous light intensity mode: M = 100, H = b sa = 8 bits, the estimated data amount is M · H = 800 bits at this time, M · H < a · K, bandwidth limit is still met.

[0595] In summary, the asynchronous intensity mode was selected in the first second, transmitting the intensity information for all 10 events with a relatively small data size (270 bits) without exceeding the bandwidth limit. If the control circuit 906 determines that the current combination mode cannot guarantee the transmission of event data with the highest possible accuracy without exceeding the bandwidth limit, then it determines that the switching condition is met and controls the switching from the asynchronous polarity mode to the asynchronous intensity mode. For example, an indication signal is sent to instruct the read circuit 905 to switch from the current event representation mode to another event representation mode.

[0596] Second 2: 15 events are generated

[0597] Asynchronous polarity mode: estimated data volume is N ev H = 15 × 20 = 300 bits, which meets the bandwidth limitation.

[0598] Asynchronous light intensity mode: estimated data volume is N ev H = 15 × 27 = 405 bits, which meets the bandwidth limitation.

[0599] Synchronous polarity mode: The estimated data volume is M·H = 100 × 1 = 100 bits, which meets the bandwidth limitation.

[0600] Synchronous optical intensity mode: The estimated data volume is M·H = 100 × 8 = 800 bits, which meets the bandwidth limitation.

[0601] In summary, at the second second, the control circuit 906 determines that the current combination mode can ensure that event data with higher precision is transmitted as much as possible without exceeding the bandwidth limit, and determines that the switching condition is not met, and still selects the asynchronous light intensity mode.

[0602] Second 3: 30 events are generated

[0603] Asynchronous polarity mode: estimated data volume is N ev H = 30 × 20 = 600 bits, which meets the bandwidth limitation.

[0604] Asynchronous light intensity mode: estimated data volume is N ev H = 30 × 27 = 810 bits, which meets the bandwidth limit.

[0605] Synchronous polarity mode: The estimated data volume is M·H = 100 × 1 = 100 bits, which meets the bandwidth limitation.

[0606] Synchronous optical intensity mode: The estimated data volume is M·H = 100 × 8 = 800 bits, which meets the bandwidth limitation.

[0607] At the 3rd second, the synchronous light intensity mode can transmit the light intensity information of all 30 events at a data rate of 800 bits. If the current combined mode (asynchronous light intensity mode) cannot guarantee the transmission of higher-precision event data without exceeding the bandwidth limit at the 3rd second, then the switching condition is determined to be met, and the asynchronous light intensity mode is controlled to switch to the synchronous light intensity mode. For example, an indication signal is sent to instruct the read circuit 905 to switch from the current event read mode to another event read mode.

[0608] It should be understood that the formulas, conversion conditions and related calculation methods given above are merely an example implementation of the embodiments of the present application. Conversion conditions, conversion strategies and calculation methods of other suitable event representations may also be adopted, and the scope of the present application is not limited in this respect.

[0609] In some embodiments, the reading circuit 905 includes a data format control unit 9051, which is used to control the reading circuit to read the signal output by the threshold comparison unit 902, or to read the electrical signal buffered in the light intensity acquisition unit 904. Exemplarily, the data format control unit 9051 is described below in conjunction with two preferred embodiments.

[0610] See Figure 12-a , is a structural diagram of a data format control unit in a reading circuit in an embodiment of the present application. The data format control unit may include an AND gate 951 and an AND gate 954, an OR gate 953, and a NOT gate 952. The input end of the AND gate 951 is used to receive the conversion signal sent by the control circuit 906 and the polarity information output by the threshold comparison unit 902, and the input end of the AND gate 954 is used to receive the conversion signal sent by the control circuit 906 after passing through the NOT gate 952, and the electrical signal (light intensity information) output by the light intensity acquisition unit 904. The output ends of the AND gate 951 and the AND gate 954 are connected to the input end of the OR gate 953, and the output end of the OR gate 953 is coupled to the control circuit 906. In one possible embodiment, the conversion signal can be 0 or 1, then the data format control unit 9051 can control the polarity information output by the reading threshold comparison unit 902, or control the reading of the light intensity information output by the light intensity acquisition unit 904. For example, if the conversion signal is 0, the data format control unit 9051 can control the output of polarity information in the threshold comparison unit 902, and if the conversion signal is 1, the data format control unit 9051 can control the output of light intensity information in the light intensity acquisition unit 904. In one possible implementation, the data format control unit 9051 can be connected to the control unit 906 via a format signal line and receive the conversion signal sent by the control unit 906 via the format signal line.

[0611] It should be noted that Figure 12-aThe data format control unit shown is only one possible structure, and other logical structures that can achieve line switching can a...

Claims

1. A method for image processing, characterized in that: include: Acquire an event stream and a first RGB image, wherein the event stream includes at least one event image, each of which is generated by motion trajectory information of a target object as it moves within a monitoring range of a motion sensor, and the first RGB image is a superposition of scenes captured at each moment by the camera within an exposure time. When the motion sensor detects a sudden movement within the monitoring range at a first moment, the camera is triggered to capture a third RGB image; Constructing a mask according to the event stream, wherein the mask is used to determine a motion area of ​​each frame of the event image; A second RGB image is obtained according to the event stream, the first RGB image, the third RGB image, and the mask, where the second RGB image is the RGB image without the target object.

2. The method according to claim 1, characterized in that The motion sensor detecting a sudden motion change within the monitoring range at the first moment includes: Within the monitoring range, an overlapping portion between a generation area of ​​a first event stream collected by the motion sensor at a first moment and a generation area of ​​a second event stream collected by the motion sensor at a second moment is smaller than a preset value.

3. The method according to any one of claims 1 to 2, characterized in that The constructing of a mask according to the event stream comprises: Dividing the monitoring range of the motion sensor into a plurality of preset neighborhoods; Within the target preset neighborhood, when the number of event images in the event stream within a preset time length exceeds a threshold, the target preset area is determined to be a motion sub-area, and the target preset neighborhood is any one of the multiple preset neighborhoods, and each motion sub-area constitutes the mask.

4. An image processing device, characterized in that include: an acquisition module, configured to acquire an event stream and a first RGB image, wherein the event stream includes at least one event image, each of the at least one event image frame being generated by motion trajectory information of a target object as it moves within a monitoring range of a motion sensor, and the first RGB image being a superposition of scenes captured at each moment by the camera within an exposure time; The acquisition module is further configured to trigger the camera to capture a third RGB image when the motion sensor detects a sudden motion within the monitoring range at the first moment; A construction module, configured to construct a mask according to the event stream, wherein the mask is used to determine a motion region of each frame of the event image; A processing module is configured to obtain a second RGB image according to the event stream, the first RGB image, the third RGB image, and the mask, where the second RGB image is an RGB image with the target object removed.

5. The device according to claim 4, characterized in that The motion sensor detecting a sudden motion change within the monitoring range at the first moment includes: Within the monitoring range, an overlapping portion between a generation area of ​​a first event stream collected by the motion sensor at a first moment and a generation area of ​​a second event stream collected by the motion sensor at a second moment is smaller than a preset value.

6. The device according to any one of claims 4-5, characterized in that The building blocks are specifically used for: Dividing the monitoring range of the motion sensor into a plurality of preset neighborhoods; Within the target preset neighborhood, when the number of event images in the event stream within a preset time length exceeds a threshold, the target preset area is determined to be a motion sub-area, and the target preset neighborhood is any one of the multiple preset neighborhoods, and each motion sub-area constitutes the mask.

7. An image processing device comprising a processor and a memory, wherein the processor is coupled to the memory, wherein: The memory is used to store programs; The processor is configured to execute the program in the memory so that the image processing apparatus performs the method according to any one of claims 1 to 3.

8. A computer-readable storage medium comprising a program, characterized in that When the method is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 3.

9. A computer program product comprising instructions, characterized in that When the method is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 3.