A pose estimation method and related apparatus
By selecting sensors adapted to different scenarios and dynamically switching data acquisition methods in electronic devices, the accuracy and power consumption issues of pose estimation in dynamic scenarios are solved, achieving more efficient pose estimation and image reconstruction.
Patent Information
- Application Number
- CN202080103758.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-31
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2040-12-31
AI Technical Summary
Existing pose estimation algorithms struggle to achieve accurate pose estimation in dynamic scenarios and consume significant power.
By selecting sensors suitable for different scenarios in electronic devices, and combining RGB sensors and motion sensors, data acquisition and decoding methods can be dynamically switched to reduce power consumption and improve estimation accuracy.
This enables more efficient pose estimation in different scenarios, reduces the power consumption of electronic devices, and improves the accuracy of pose estimation and the quality of image reconstruction.
Smart Images

Figure CN115997234B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer, and particularly relates to a pose estimation method and related device. BACKGROUND
[0002] Simultaneous Localization and Mapping (SLAM) is a technology that a host carrying a specific sensor realizes establishment of a surrounding environment map during movement and performs self-positioning according to the established environment map. The SLAM technology has a wide application prospect in the fields of robots, automatic driving, and virtual and augmented reality.
[0003] In the SLAM technology, performing pose estimation is an important process. At present, algorithms for realizing pose estimation are all suitable for static scenes, and it is usually difficult to realize accurate pose estimation in a dynamic scene. SUMMARY
[0004] Embodiments of the present application provide an image processing method and device, which are used to obtain clearer images.
[0005] In a first aspect, the present application provides a switching method applied to an electronic device, the electronic device comprising an RGB sensor and a motion sensor, the RGB (red green blue) sensor being used to collect images in a shooting range, and the motion sensor being used to collect information generated when an object moves relative to the motion sensor in a detection range of the motion sensor, the method comprising: selecting at least one from the RGB sensor and the motion sensor based on scene information, and collecting data through the selected sensor, the scene information comprising at least one of state information of the electronic device, a type of an application program in the electronic device requesting to collect images, or environment information.
[0006] Therefore, in the embodiments of the present application, different sensors in the electronic device can be selected to be started according to different scenes, more scenes can be adapted, and the generalization ability is strong. Moreover, the corresponding sensor can be started according to the actual scene, without starting all the sensors, so that the power consumption of the electronic device is reduced.
[0007] In a possible implementation, the state information comprises a remaining power of the electronic device, a remaining storage amount; and the environment information comprises a change value of illumination intensity in a shooting range of the color RGB sensor and the motion sensor or information of a moving object in the shooting range.
[0008] Therefore, in the embodiments of the present application, the sensor to be started can be selected according to the state or the environment information of the electronic device, more scenes can be adapted, and the generalization ability is strong.
[0009] In addition, in the following different embodiments, the activated sensors can not be the same, and the data collected by a certain sensor mentioned below is not described in detail.
[0010] In a second aspect, the present application provides a visual sensor chip, which can include: a pixel array circuit, configured to generate at least one data signal corresponding to a pixel in the pixel array circuit by measuring a light intensity change amount, the at least one data signal indicating a light intensity change event, the light intensity change event representing that the light intensity change amount measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A reading circuit, coupled with the pixel array circuit, configured to read the at least one data signal from the pixel array circuit in a first event representation manner. The reading circuit is further configured to provide the at least one data signal to a control circuit. The reading circuit is further configured to switch to reading the at least one data signal from the pixel array circuit in a second event representation manner when receiving a conversion signal generated based on the at least one data signal from the control circuit. As known from the first aspect, the visual sensor can adaptively switch between the two event representation manners, so that the reading data rate always remains below a predetermined reading data rate threshold, thereby reducing the cost of data transmission, analysis and storage of the visual sensor, and significantly improving the performance of the sensor. In addition, such a visual sensor can perform data statistics on events generated in a period of time for predicting the possible event generation rate in the next period of time, so as to select a reading mode more suitable for the current external environment, application scenario and motion state.
[0011] In a possible implementation, the first event representation manner is to represent events by polarity information. The pixel array circuit can include a plurality of pixels, each pixel can include a threshold comparison unit, the threshold comparison unit is configured to output polarity information when the light intensity change amount exceeds the predetermined threshold, the polarity information is used to indicate whether the light intensity change amount is enhanced or weakened. The reading circuit is specifically configured to read the polarity information output by the threshold comparison unit. In this implementation, the first event representation manner is to represent events by polarity information, the polarity information is usually represented by 1-bit-2-bit, and the information carried is less. In the case that the data amount is large, the visual sensor will face the problem of sudden increase of events when a large area of object moves or the light intensity fluctuates (for example, entering or exiting a tunnel, turning on or off a light in a room, etc.). In the case that the preset maximum bandwidth (hereinafter referred to as bandwidth) of the visual sensor is certain, the situation that the event data cannot be read out is avoided, and the event loss is avoided.
[0012] In a possible implementation, the first event representation manner is to represent events by light intensity information. The pixel array can include a plurality of pixels. Each pixel can include a threshold comparison unit, a readout control unit, and a light intensity acquisition unit. The light intensity detection unit is configured to output an electrical signal corresponding to a light signal irradiated thereon, and the electrical signal is configured to indicate the light intensity. The threshold comparison unit is configured to output a first signal when the light intensity variation exceeds a predetermined threshold according to the electrical signal. The readout control unit is configured to instruct the light intensity acquisition unit to acquire and cache the electrical signal corresponding to the time when the first signal is received in response to receiving the first signal. The readout circuit is specifically configured to read the electrical signal cached by the light intensity acquisition unit. In this implementation, the first event representation manner is to represent events by light intensity information. When the amount of data to be transmitted does not exceed the bandwidth limit, the light intensity information is used to represent events. Generally, the light intensity information is represented by a plurality of bits, for example, 8 bits-12 bits. Compared with the polarity information, the light intensity information can carry more information, which is beneficial to the processing and analysis of events, for example, can improve the quality of image reconstruction.
[0013] In a possible implementation, the control circuit is further configured to: determine the statistical data based on the at least one data signal received from the readout circuit. If it is determined that the statistical data satisfies a predetermined conversion condition, the control circuit is configured to send a conversion signal to the readout circuit. The predetermined conversion condition is determined based on a preset bandwidth of the visual sensor chip. In this implementation, a manner of converting the two event representation manners is given. The conversion condition is obtained according to the amount of data to be transmitted. For example, when the amount of data to be transmitted is large, the event is represented by the polarity information, so that the amount of data can be completely transmitted, and the situation that the event data cannot be read out and the event is lost is avoided. When the amount of data to be transmitted is small, the event is represented by the light intensity information, so that the event to be transmitted can carry more information, which is beneficial to the processing and analysis of events, for example, can improve the quality of image reconstruction.
[0014] In a possible implementation, the first event representation manner is to represent events by light intensity information, and the second event representation manner is to represent events by polarity information. The predetermined conversion condition is that the total amount of data read from the pixel array circuit by the first event representation manner is greater than the preset bandwidth, or the predetermined conversion condition is that the number of the at least one data signal is greater than the ratio of the preset bandwidth and a first bit. The first bit is a preset bit of the data format of the data signal. In this implementation, a specific condition for switching from the event represented by the light intensity information to the event represented by the polarity information is given. When the amount of data to be transmitted is greater than the preset bandwidth, the event is represented by the polarity information, so that the amount of data can be completely transmitted, and the situation that the event data cannot be read out and the event is lost is avoided.
[0015] In a possible implementation, the first event representation is to represent events by polarity information, the second event representation is to represent events by light intensity information, the predetermined conversion condition is that if at least one data signal is read from the pixel array circuit by the second event representation, the total amount of data read is not greater than a preset bandwidth, or the predetermined conversion condition is that the number of the at least one data signal is not greater than a ratio of the preset bandwidth and a first bit, the first bit being a preset bit of a data format of the data signal. In this implementation, a specific condition for switching from representing events by polarity information to representing events by light intensity information is given. When the amount of data transmitted is not greater than the preset bandwidth, the switching to representing events by light intensity information is performed, so that the events transmitted can carry more information, which is beneficial to processing and analysis of the events, for example, the quality of image reconstruction can be improved.
[0016] In a third aspect, the present application provides a decoding circuit, which can include: a reading circuit configured to read a data signal from a visual sensor chip. The decoding circuit is configured to decode the data signal according to a first decoding mode. The decoding circuit is further configured to decode the data signal according to a second decoding mode when a conversion signal is received from a control circuit. The decoding circuit provided in the third aspect corresponds to the visual sensor chip provided in the second aspect, and is configured to decode the data signal output by the visual sensor chip provided in the second aspect. The decoding circuit provided in the third aspect can switch between different decoding modes for different event representations.
[0017] In a possible implementation, the control circuit is further configured to: determine statistical data based on the data signal read from the reading circuit. If it is determined that the statistical data satisfies a predetermined conversion condition, the conversion signal is sent to the encoding circuit, the predetermined conversion condition being determined based on a preset bandwidth of the visual sensor chip.
[0018] In a possible implementation, the first decoding mode is to decode the data signal according to a first bit corresponding to the first event representation, the first event representation is to represent events by light intensity information, the second decoding mode is to decode the data signal according to a second bit corresponding to the second event representation, the second event representation is to represent events by polarity information, the polarity information is used to indicate whether the light intensity change is an increase or a decrease, the conversion condition is that the total amount of data decoded according to the first decoding mode is greater than a preset bandwidth, or the predetermined conversion condition is that the number of the data signal is greater than a ratio of the preset bandwidth and a first bit, the first bit being a preset bit of a data format of the data signal.
[0019] In a possible implementation, the first decoding manner is decoding the data signal according to first bits corresponding to a first event representation manner, the first event representation manner is representing an event by polarity information, the polarity information is used to indicate that the light intensity change amount is increased or decreased, the second decoding manner is decoding the data signal by second bits corresponding to a second event representation manner, the second event representation manner is representing an event by light intensity information, and the conversion condition is that the total data amount is not greater than a preset bandwidth according to the second decoding manner, or the number of data signals is greater than a preset ratio of the preset bandwidth and the first bits, the first bits are preset bits of a data format of the data signal.
[0020] In a fourth aspect, the present application provides a method for operating a visual sensor chip, which can include: generating at least one data signal corresponding to a pixel in a pixel array circuit of the visual sensor chip by measuring a light intensity change amount by the pixel array circuit, the at least one data signal indicating a light intensity change event, the light intensity change event representing that the light intensity change amount measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. Reading the at least one data signal from the pixel array circuit in a first event representation manner by a reading circuit of the visual sensor chip. Providing the at least one data signal to a control circuit of the visual sensor chip by the reading circuit. When a conversion signal generated based on the at least one data signal is received from the control circuit by the reading circuit, converting to read the at least one data signal from the pixel array circuit in a second event representation manner.
[0021] In a possible implementation, the first event representation manner is representing an event by polarity information, and the pixel array circuit can include a plurality of pixels, each pixel can include a threshold comparison unit. Reading the at least one data signal from the pixel array circuit in the first event representation manner by the reading circuit of the visual sensor chip can include: outputting polarity information by the threshold comparison unit when the light intensity change amount exceeds the predetermined threshold, the polarity information being used to indicate that the light intensity change amount is increased or decreased. Reading the polarity information output by the threshold comparison unit by the reading circuit.
[0022] In a possible implementation, the first event representation is to represent events by light intensity information, the pixel array can include a plurality of pixels, each pixel can include a threshold comparison unit, a readout control unit and a light intensity acquisition unit, reading at least one data signal from the pixel array circuit in the first event representation by the read circuit of the visual sensor chip can include: outputting, by the light intensity acquisition unit, an electrical signal corresponding to a light signal irradiated thereon, the electrical signal being used to indicate the light intensity. When the light intensity variation exceeds a predetermined threshold, outputting, by the threshold comparison unit, a first signal. In response to receiving the first signal, instructing, by the readout control unit, the light intensity acquisition unit to acquire and cache the electrical signal corresponding to the time when the first signal is received. Reading, by the read circuit, the electrical signal cached by the light intensity acquisition unit.
[0023] In a possible implementation, the method can further include: determining statistical data based on the at least one data signal received from the read circuit. If it is determined that the statistical data satisfies a predetermined conversion condition, sending a conversion signal to the read circuit, the predetermined conversion condition being determined based on a preset bandwidth of the visual sensor chip.
[0024] In a possible implementation, the first event representation is to represent events by light intensity information, the second event representation is to represent events by polarity information, and the predetermined conversion condition is that a total amount of data read from the pixel array circuit in the first event representation is greater than the preset bandwidth, or the predetermined conversion condition is that a number of the at least one data signal is greater than a ratio of the preset bandwidth and a first bit, the first bit being a preset bit of a data format of the data signal.
[0025] In a possible implementation, the first event representation is to represent events by polarity information, and the second event representation is to represent events by light intensity information, and the predetermined conversion condition is that if at least one data signal is read from the pixel array circuit in the second event representation, a total amount of data read is not greater than the preset bandwidth, or the predetermined conversion condition is that a number of the at least one data signal is not greater than a ratio of the preset bandwidth and a first bit, the first bit being a preset bit of a data format of the data signal.
[0026] In a fifth aspect, the present application provides a decoding method, including: reading, by a read circuit, a data signal from a visual sensor chip; decoding, by a decoding circuit, the data signal according to a first decoding mode; when a conversion signal is received from a control circuit, decoding, by the decoding circuit, the data signal according to a second decoding mode.
[0027] In a possible implementation, the method can further include: determining statistical data based on the at least one data signal received from the read circuit. If it is determined that the statistical data satisfies a predetermined conversion condition, sending a conversion signal to the read circuit, the predetermined conversion condition being determined based on a preset bandwidth of the visual sensor chip.
[0028] In a possible implementation, the first decoding manner is decoding the data signal according to first bits corresponding to a first event representation manner, the first event representation manner is representing events by light intensity information, the second decoding manner is decoding the data signal according to second bits corresponding to a second event representation manner, the second event representation manner is representing events by polarity information, the polarity information is used to indicate whether the light intensity change is an increase or a decrease, the conversion condition is that the total data quantity decoded according to the first decoding manner is greater than a preset bandwidth, or the predetermined conversion condition is that the number of data signals is greater than the ratio of the preset bandwidth and the first bits, the first bits are preset bits of a data format of the data signal.
[0029] In a possible implementation, the first decoding manner is decoding the data signal according to first bits corresponding to a first event representation manner, the first event representation manner is representing events by polarity information, the polarity information is used to indicate whether the light intensity change is an increase or a decrease, the second decoding manner is decoding the data signal by second bits corresponding to a second event representation manner, the second event representation manner is representing events by light intensity information, the conversion condition is that if the data signal is decoded according to the second decoding manner, the total data quantity is not greater than a preset bandwidth, or the predetermined conversion condition is that the number of data signals is greater than the ratio of the preset bandwidth and the first bits, the first bits are preset bits of a data format of the data signal.
[0030] In a sixth aspect, the present application provides a visual sensor chip, which can include: a pixel array circuit, configured to generate at least one data signal corresponding to a pixel in the pixel array circuit by measuring a light intensity change, the at least one data signal indicating a light intensity change event, the light intensity change event representing that the light intensity change measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A first encoding unit, configured to encode the at least one data signal according to first bits to obtain first encoded data. The first encoding unit is further configured to, when a first control signal is received from a control circuit, encode the at least one data signal according to second bits indicated by the first control signal, the first control signal being determined by the control circuit according to the first encoded data. According to the scheme provided by the sixth aspect, by dynamically adjusting the bit width of the light intensity feature information, when the event generation rate is small, the bandwidth limit has not been reached, the event is quantized according to the maximum bit width, and when the event generation rate is large, the bit width of the light intensity feature information is gradually reduced to meet the bandwidth limit. Thereafter, if the event generation rate becomes small again, the bit width of the light intensity feature information can be increased without exceeding the bandwidth limit. The visual sensor can adaptively switch between multiple event representation manners to better achieve the purpose of transmitting all events with greater representation accuracy.
[0031] In a possible implementation, the first control signal is determined by the control circuit according to the first encoded data and a preset bandwidth of the visual sensor chip.
[0032] In a possible implementation, when the data amount of the first encoded data is not less than the bandwidth, the second bit indicated by the control signal is less than the first bit, so that the total data amount of the at least one data signal encoded by the second bit is not greater than the bandwidth. When the event generation rate is large, the bit width representing the light intensity feature information is gradually reduced to meet the bandwidth limit.
[0033] In a possible implementation, when the data amount of the first encoded data is less than the bandwidth, the second bit indicated by the control signal is greater than the first bit, and the total data amount of the at least one data signal encoded by the second bit is not greater than the bandwidth. When the event generation rate is small, the bit width representing the light intensity feature information is gradually increased without exceeding the bandwidth limit, so as to better achieve the purpose of transmitting all events with greater representation accuracy.
[0034] In a possible implementation, the pixel array can include N regions, and the maximum bit of at least two regions of the N regions is different, the maximum bit representing a preset maximum bit for encoding the at least one data signal generated by one region, the first encoding unit is specifically configured to encode the at least one data signal generated by the first region according to the first bit to obtain the first encoded data, the first bit being not greater than the maximum bit of the first region, the first region being any one of the N regions. The first encoding unit is specifically configured to encode the at least one data signal generated by the first region according to the second bit indicated by the first control signal when the first control signal is received from the control circuit, the first control signal being determined by the control circuit according to the first encoded data. In this implementation, the pixel array can be regionally divided, and different maximum bit widths of different regions can be set by using different weights to adapt to different regions of interest in a scene. For example, a larger weight is set in a region that can include a target object, so that the representation accuracy of events output by the region that includes the target object is higher, and a smaller weight is set in a background region, so that the representation accuracy of events output by the background region is lower.
[0035] In a possible implementation, the control circuit is further configured to: when the total data amount of the at least one data signal encoded by the third bit is greater than the bandwidth, and the total data amount of the at least one data signal encoded by the second bit is not greater than the bandwidth, send the first control signal to the first encoding unit, the third bit and the second bit being different by one bit unit. In this implementation, all events can be transmitted with greater representation accuracy without exceeding the bandwidth limit.
[0036] In a seventh aspect, the present application provides a decoding device, which can include: a reading circuit configured to read a data signal from a visual sensor chip; and a decoding circuit configured to decode the data signal according to a first bit, and configured to decode the data signal according to a second bit indicated by a first control signal received from a control circuit. The decoding circuit of the seventh aspect corresponds to the visual sensor chip of the sixth aspect, and is configured to decode the data signal output by the visual sensor chip of the sixth aspect. The decoding circuit of the seventh aspect is configured to dynamically adjust the decoding manner according to the encoding bits used by the visual sensor.
[0037] In a possible implementation, the first control signal is determined by the control circuit according to the first encoding data and a preset bandwidth of the visual sensor chip.
[0038] In a possible implementation, when a total data amount of the data signal decoded according to the first bit is not less than the bandwidth, the second bit is less than the first bit.
[0039] In a possible implementation, when a total data amount of the data signal decoded according to the first bit is less than the bandwidth, the second bit is greater than the first bit, and a total data amount of the data signal decoded according to the second bit is not greater than the bandwidth.
[0040] In a possible implementation, the reading circuit is specifically configured to read a data signal corresponding to a first region from the visual sensor chip, the first region being any one of N regions included in a pixel array of the visual sensor, at least two regions of the N regions having different maximum bits, the maximum bit representing a preset maximum bit used for encoding at least one data signal generated by a region. The decoding circuit is specifically configured to decode the data signal corresponding to the first region according to the first bit.
[0041] In a possible implementation, the control circuit is further configured to: determine that a total data amount of the data signal decoded according to a third bit is greater than the bandwidth, and a total data amount of the data signal decoded according to the second bit is not greater than the bandwidth, and send the first control signal to the first encoding unit, the third bit and the second bit being different by one bit unit.
[0042] In an eighth aspect, the present application provides a method for operating a visual sensor chip, which can include: generating at least one data signal corresponding to a pixel in a pixel array circuit of the visual sensor chip by measuring a light intensity variation amount through the pixel array circuit, the at least one data signal indicating a light intensity variation event, the light intensity variation event representing that the light intensity variation amount measured by the corresponding pixel in the pixel array circuit exceeds a predetermined threshold; and encoding the at least one data signal according to a first bit through a first encoding unit of the visual sensor chip to obtain first encoded data. When a first control signal is received from a control circuit of the visual sensor chip through the first encoding unit, the at least one data signal is encoded according to a second bit indicated by the first control signal, the first control signal being determined by the control circuit according to the first encoded data.
[0043] In a possible implementation, the first control signal is determined by the control circuit according to the first encoded data and a preset bandwidth of the visual sensor chip.
[0044] In a possible implementation, when a data amount of the first encoded data is not less than the bandwidth, the second bit indicated by the control signal is less than the first bit, so that a total data amount of the at least one data signal encoded by the second bit is not greater than the bandwidth.
[0045] In a possible implementation, when the data amount of the first encoded data is less than the bandwidth, the second bit indicated by the control signal is greater than the first bit, and a total data amount of the at least one data signal encoded by the second bit is not greater than the bandwidth.
[0046] In a possible implementation, the pixel array can include N regions, and a maximum bit of at least two regions of the N regions is different, the maximum bit representing a preset maximum bit for encoding the at least one data signal generated by one region. The encoding the at least one data signal according to the first bit through the first encoding unit of the visual sensor chip can include: encoding the at least one data signal generated by a first region according to the first bit through the first encoding unit to obtain the first encoded data, the first bit being not greater than a maximum bit of the first region, the first region being any one of the N regions. When the first control signal is received from the control circuit of the visual sensor chip through the first encoding unit, the at least one data signal is encoded according to the second bit indicated by the first control signal, which can include: when the first control signal is received from the control circuit through the first encoding unit, the at least one data signal generated by the first region is encoded according to the second bit indicated by the first control signal, the first control signal being determined by the control circuit according to the first encoded data.
[0047] In a possible implementation, the method further includes: determining that the total data amount of the at least one data signal encoded by the third bits is greater than the bandwidth, and the total data amount of the at least one data signal encoded by the second bits is not greater than the bandwidth, and sending, by the control circuit, the first control signal to the first encoding unit, the third bits and the second bits being different by one bit unit.
[0048] In a ninth aspect, the present application provides a decoding method, which can include: reading, by a reading circuit, a data signal from a visual sensor chip. Decoding, by a decoding circuit, the data signal according to first bits. Decoding, by the decoding circuit, the data signal according to second bits indicated by a first control signal received from a control circuit.
[0049] In a possible implementation, the first control signal is determined by the control circuit according to the first encoding data and a preset bandwidth of the visual sensor chip.
[0050] In a possible implementation, when the total data amount of the data signal decoded according to the first bits is not less than the bandwidth, the second bits are less than the first bits.
[0051] In a possible implementation, when the total data amount of the data signal decoded according to the first bits is less than the bandwidth, the second bits are greater than the first bits, and the total data amount of the data signal decoded by the second bits is not greater than the bandwidth.
[0052] In a possible implementation, the reading, by the reading circuit, of the data signal from the visual sensor chip can include: reading, by the reading circuit, a data signal corresponding to a first region from the visual sensor chip, the first region being any one of N regions that can be included in a pixel array of the visual sensor, at least two of the N regions having different maximum bits, the maximum bit representing a preset maximum bit for encoding at least one data signal generated for a region. The decoding, by the decoding circuit, of the data signal according to the first bits can include: decoding, by the decoding circuit, the data signal corresponding to the first region according to the first bits.
[0053] In a possible implementation, the method further includes: determining that the total data amount of the data signal decoded by the third bits is greater than the bandwidth, and the total data amount of the data signal decoded by the second bits is not greater than the bandwidth, and sending, by the control circuit, the first control signal to the first encoding unit, the third bits and the second bits being different by one bit unit.
[0054] In a tenth aspect, the present application provides a visual sensor chip, which can include: a pixel array circuit configured to generate a plurality of data signals corresponding to a plurality of pixels in the pixel array circuit by measuring a light intensity change amount, the plurality of data signals indicating at least one light intensity change event, the at least one light intensity change event representing that the light intensity change amount measured by a corresponding pixel in the pixel array circuit exceeds a predetermined threshold. A third encoding unit is configured to encode a first difference value according to a first preset bit, the first difference value being a difference between the light intensity change amount and the predetermined threshold. Reducing the accuracy of event representation, i.e., reducing the bit width of event representation, makes the event carry less information, which is not conducive to the processing and analysis of the event in some scenarios. Therefore, the way of reducing the accuracy of event representation may not be applicable to all scenarios, i.e., in some scenarios, a high-bit width is required to represent the event, but the event represented by the high-bit width can carry more data, and the amount of data is also large, and in the case that the maximum bandwidth of the visual sensor is fixed, the event data may not be read out, resulting in data loss. The tenth aspect provides a scheme of encoding the difference value, which reduces the cost of data transmission, analysis and storage of the visual sensor while transmitting the event with the highest accuracy, and significantly improves the performance of the sensor.
[0055] In a possible implementation, the pixel array circuit can include a plurality of pixels, each pixel can include a threshold comparison unit configured to output polarity information when the light intensity change amount exceeds the predetermined threshold, the polarity information being used to indicate whether the light intensity change amount is enhanced or weakened. The third encoding unit is further configured to encode the polarity information according to a second preset bit. In this implementation, the polarity information can also be encoded, and the polarity information is used to indicate whether the light intensity is enhanced or weakened, which is helpful to obtain the current light intensity information according to the light intensity signal and the polarity information obtained by the last decoding.
[0056] In a possible implementation, each pixel can include a light intensity detection unit, a readout control unit and a light intensity acquisition unit. The light intensity detection unit is configured to output an electrical signal corresponding to a light signal irradiated thereon, the electrical signal being used to indicate the light intensity. A threshold comparison unit is configured to output a polarity information when the light intensity variation exceeds a predetermined threshold according to the electrical signal. The readout control unit is configured to instruct the light intensity acquisition unit to acquire and cache the electrical signal at the time when the polarity information is received in response to receiving the polarity signal. A third encoding unit is configured to encode the first electrical signal according to a third preset bit, the first electrical signal being the electrical signal at the time when the polarity information is received for the first time, and the third preset bit being a maximum bit of the characteristic information of the light intensity preset by the vision sensor. After the initial state full-quantity encoding, only the polarity information and the difference between the light intensity variation and the predetermined threshold need to be encoded in subsequent events, so that the amount of encoded data can be effectively reduced. The full-quantity encoding refers to encoding an event by using the maximum bit width predefined by the vision sensor. In addition, the light intensity information at the current time can be losslessly reconstructed by using the light intensity information of the last event, the decoded polarity information and the difference value.
[0057] In a possible implementation, the third encoding unit is further configured to encode the electrical signal acquired by the light intensity acquisition unit according to the third preset bit every preset time length. The full-quantity encoding is performed every preset time length, so as to reduce the decoding dependency and prevent errors.
[0058] In a possible implementation, the third encoding unit is specifically configured to encode the first difference value according to the first preset bit when the first difference value is less than the predetermined threshold.
[0059] In a possible implementation, the third encoding unit is further configured to encode the first remaining difference value and the predetermined threshold according to the first preset bit when the first difference value is not less than the predetermined threshold, the first remaining difference value being a difference between the difference value and the predetermined threshold.
[0060] In a possible implementation, the third encoding unit is specifically configured to: when the first residual difference value is not less than the predetermined threshold value, encode the second residual difference value according to the first preset bit, the second residual difference value being a difference between the first residual difference value and the predetermined threshold value. The predetermined threshold value is encoded for the first time according to the first preset bit. The predetermined threshold value is encoded for the second time according to the first preset bit. Because the visual sensor may have a certain delay, the light intensity change may be greater than the predetermined threshold value twice or more than twice, and only one event is generated. This may cause a problem that the difference value is greater than or equal to the predetermined threshold value, and the light intensity change is at least twice the predetermined threshold value. For example, the first residual difference value may not be less than the predetermined threshold value, the second residual difference value is encoded, and if the second residual difference value is still not less than the predetermined threshold value, the third residual difference value is encoded, the third residual difference value being a difference between the second residual difference value and the predetermined threshold value, and the predetermined threshold value is encoded for the third time. The above process is repeated until the residual difference value is less than the predetermined threshold value.
[0061] In a possible implementation, the decoding device provided in the eleventh aspect can include: an acquisition circuit configured to read a data signal from a visual sensor chip; and a decoding circuit configured to decode the data signal according to a first bit to obtain a difference value, the difference value being less than a predetermined threshold value, the difference value being a difference between a light intensity change measured by the visual sensor and the predetermined threshold value, the light intensity change exceeding the predetermined threshold value, and the visual sensor generating at least one light intensity change event. The decoding circuit provided in the eleventh aspect corresponds to the visual sensor chip provided in the tenth aspect, and is configured to decode the data signal output by the visual sensor chip provided in the tenth aspect. The decoding circuit provided in the eleventh aspect can adopt a corresponding differential decoding manner for the differential encoding manner adopted by the visual sensor.
[0062] In a possible implementation, the decoding circuit is further configured to: decode the data signal according to a second bit to obtain polarity information, the polarity information being used to indicate whether the light intensity change is enhanced or weakened.
[0063] In a possible implementation, the decoding circuit is further configured to: decode the data signal received at the first time according to a third bit to obtain an electrical signal corresponding to an optical signal output by the visual sensor and irradiated on the visual sensor, the third bit being a maximum bit of the characteristic information of the light intensity preset by the visual sensor.
[0064] In a possible implementation, the decoding circuit is further configured to: decode the data signal received at the first time according to the third bit every preset time length.
[0065] In a possible implementation, the decoding circuit is specifically configured to: decode the data signal according to the first bit to obtain the difference value and the at least one predetermined threshold value.
[0066] In a twelfth aspect, the application provides a method for operating a visual sensor chip, which can include: generating a plurality of data signals corresponding to a plurality of pixels in a pixel array circuit of the visual sensor chip by measuring a light intensity variation amount through the pixel array circuit, the plurality of data signals indicating at least one light intensity variation event, the at least one light intensity variation event representing that the light intensity variation amount measured by a corresponding pixel in the pixel array circuit exceeds a predetermined threshold value. Encoding a first difference value according to a first preset bit through a third encoding unit of the visual sensor chip, the first difference value being a difference between the light intensity variation amount and the predetermined threshold value.
[0067] In a possible implementation, the pixel array circuit can include a plurality of pixels, each pixel can include a threshold comparison unit, and the method can further include: outputting polarity information through the threshold comparison unit when the light intensity variation amount exceeds the predetermined threshold value, the polarity information being used to indicate whether the light intensity variation amount is an increase or a decrease. Encoding the polarity information according to a second preset bit through the third encoding unit.
[0068] In a possible implementation, each pixel can further include a light intensity detection unit, a readout control unit and a light intensity acquisition unit, and the method can further include: outputting an electrical signal corresponding to a light signal irradiated thereon through the light intensity detection unit, the electrical signal being used to indicate the light intensity. Outputting the polarity information through the threshold comparison unit can include: outputting the polarity information through the threshold comparison unit according to the electrical signal when the light intensity variation amount exceeds the predetermined threshold value. The method can further include: in response to receiving the polarity signal, instructing the light intensity acquisition unit to acquire and cache an electrical signal corresponding to a time when the polarity information is received through the readout control unit. Encoding a first electrical signal according to a third preset bit, the first electrical signal being an electrical signal corresponding to a time when the polarity information is first received by the light intensity acquisition unit, and the third preset bit being a maximum bit of feature information of the visual sensor preset for indicating the light intensity.
[0069] In a possible implementation, the method can further include: encoding the electrical signal acquired by the light intensity acquisition unit according to the third preset bit every preset time length.
[0070] In a possible implementation, encoding the first difference value according to the first preset bit through the third encoding unit of the visual sensor chip can include: encoding the first difference value according to the first preset bit when the first difference value is less than the predetermined threshold value.
[0071] In a possible implementation, the third encoding unit of the visual sensor chip encodes the first difference value according to the first preset bit, and can further include: when the first difference value is not less than a predetermined threshold, encoding the first residual difference value and the predetermined threshold according to the first preset bit, the first residual difference value being a difference between the difference value and the predetermined threshold.
[0072] In a possible implementation, when the first difference value is not less than a predetermined threshold, encoding the first residual difference value and the predetermined threshold according to the first preset bit can include: when the first residual difference value is not less than the predetermined threshold, encoding the second residual difference value according to the first preset bit, the second residual difference value being a difference between the first residual difference value and the predetermined threshold. The predetermined threshold is encoded for the first time according to the first preset bit. The predetermined threshold is encoded for the second time according to the first preset bit, and the first residual difference value can include the second residual difference value and two predetermined thresholds.
[0073] In a thirteenth aspect, the present application provides a decoding method, which can include: reading a data signal from a visual sensor chip by an acquisition circuit. Decoding the data signal according to a first bit by a decoding circuit to obtain a difference value, the difference value being less than a predetermined threshold, the difference value being a difference between a light intensity change amount measured by the visual sensor and the predetermined threshold, the light intensity change amount exceeding the predetermined threshold, the visual sensor generating at least one light intensity change event.
[0074] In a possible implementation, the method can further include: decoding the data signal according to a second bit to obtain polarity information, the polarity information being used to indicate whether the light intensity change amount is enhanced or weakened.
[0075] In a possible implementation, the method can further include: decoding the data signal received at a first time according to a third bit to obtain an electrical signal corresponding to an optical signal output by the visual sensor, the third bit being a maximum bit of feature information of the visual sensor preset to represent light intensity.
[0076] In a possible implementation, the method can further include: decoding the data signal received at the first time according to the third bit every preset time length.
[0077] In a possible implementation, the decoding circuit decodes the data signal according to the first bit to obtain the difference value can include: decoding the data signal according to the first bit to obtain the difference value and at least one predetermined threshold.
[0078] In a fourteenth aspect, the present application provides an image processing method, comprising: obtaining motion information, the motion information comprising information of a motion trajectory of a target object when the target object moves within a detection range of a motion sensor; generating at least one frame of event image according to the motion information, the at least one frame of event image being an image representing the motion trajectory of the target object when the target object moves within the detection range; obtaining a target task, and obtaining an iteration duration according to the target task; iteratively updating the at least one frame of event image to obtain updated at least one frame of event image, and the iteration duration of iteratively updating the at least one frame of event image does not exceed the iteration duration.
[0079] Therefore, in the embodiments of the present application, the object in motion can be monitored by the motion sensor, and the information of the motion trajectory of the object when the object moves within the detection range can be collected by the motion sensor. After the target task is obtained, the iteration duration can be determined according to the target task, and the event image matching the target task can be obtained by iteratively updating the event image within the iteration duration.
[0080] In a possible implementation, any one of the iteratively updating the at least one frame of event image comprises: obtaining a motion parameter, the motion parameter representing a parameter of relative motion between the motion sensor and the target object; iteratively updating a target event image in the at least one frame of event image according to the motion parameter to obtain an updated target event image.
[0081] Therefore, in the embodiments of the present application, when the event image is iteratively updated, the event image can be updated based on the parameter of the relative motion between the object and the motion sensor, so that the event image is compensated to obtain a clearer event image.
[0082] In a possible implementation, the obtaining the motion parameter comprises: obtaining a value of a preset optimization model in a previous iteration update process; and calculating the motion parameter according to the value of the optimization model.
[0083] Therefore, in the embodiments of the present application, the event image can be updated based on the value of the optimization model, and the motion parameter can be calculated according to the optimization model to obtain a better motion parameter, and then the event image is updated using the motion parameter to obtain a clearer event image.
[0084] In a possible implementation, the iteratively updating the target event image according to the motion parameter comprises: compensating the motion trajectory of the target object in the target event image according to the motion parameter to obtain a target event image obtained in a current iteration update.
[0085] Therefore, in the embodiments of the present application, the motion trajectory of the target object in the event image can be compensated by using the motion parameters, so that the motion trajectory of the target object in the event image is clearer, and thus the event image is clearer.
[0086] In a possible implementation, the motion parameters include one or more of the following: a depth, optical flow information, acceleration of the motion sensor, or angular velocity of the motion sensor.
[0087] Therefore, in the embodiments of the present application, the target object in the event image can be compensated by using various motion parameters, so that the clarity of the event image is improved.
[0088] In a possible implementation, in the iterative updating process, the method further includes: if the result of the current iteration meets a preset condition, terminating the iteration, and the termination condition includes at least one of the following: the number of times of updating the at least one frame of event image reaches a preset number of times, or the change of the value of the optimization model in the updating process of the at least one frame of event image is less than a preset value.
[0089] Therefore, in the embodiments of the present application, in addition to setting the iteration duration, a convergence condition related to the number of iterations or the value of the optimization model can also be set, so that the event image meeting the convergence condition is obtained under the constraint of the iteration duration.
[0090] In a fifteenth aspect, the present application provides an image processing method, including: generating at least one frame of event image according to motion information, the motion information including information of a motion trajectory of a target object when the target object generates motion in a detection range of a motion sensor, the at least one frame of event image being an image representing the motion trajectory of the target object when the target object generates motion in the detection range; obtaining motion parameters, the motion parameters representing parameters of relative motion between the motion sensor and the target object; initializing a value of a preset optimization model according to the motion parameters, to obtain a value of the optimization model; and updating the at least one frame of event image according to the value of the optimization model, to obtain updated at least one frame of event image.
[0091] In the embodiments of the present application, the parameters of the relative motion between the motion sensor and the target object can be used to initialize the optimization model, so that the initial number of iterations of the event image is reduced, the convergence speed of the iteration of the event image is accelerated, and a clearer event image is obtained with fewer iterations.
[0092] In a possible implementation, the motion parameter comprises one or more of the following: a depth, optical flow information, acceleration of the motion sensor, or angular velocity of the motion sensor.
[0093] In a possible implementation, the acquiring the motion parameter comprises: acquiring data collected by an inertial measurement unit (IMU) sensor; and calculating the motion parameter according to the data collected by the IMU sensor. Thus, in the embodiments of the present application, the motion parameter can be calculated by the IMU, so that a more accurate motion parameter can be obtained.
[0094] In a possible implementation, after the value of the preset optimization model is initialized according to the motion parameter, the method further comprises: updating a parameter of the IMU sensor according to the value of the optimization model, the parameter of the IMU sensor being used for data collection of the IMU sensor.
[0095] Thus, in the embodiments of the present application, the parameter of the IMU can also be updated according to the value of the optimization model, so that the IMU can be corrected, and the data collected by the IMU can be more accurate.
[0096] In a sixteenth aspect, the present application provides an image processing apparatus, which has a function of implementing the method of the fourteenth aspect or any one of the possible implementation manners of the fourteenth aspect, or which has a function of implementing the method of the fifteenth aspect or any one of the possible implementation manners of the fifteenth aspect. The function can be implemented by hardware, or the function can be implemented by hardware executing corresponding software. The hardware or software comprises one or more modules corresponding to the above functions.
[0097] In a seventeenth aspect, the present application provides an image processing method, comprising: acquiring motion information, the motion information comprising information of a motion trajectory of a target object when the target object moves in a detection range of a motion sensor; generating an event image according to the motion information, the event image being an image representing the motion trajectory of the target object when the target object moves in the detection range; and obtaining a first reconstructed image according to at least one event included in the event image, wherein a color type of a first pixel point is different from color types of at least one second pixel point, the first pixel point being a pixel point corresponding to any one of the at least one event in the first reconstructed image, and the at least one second pixel point being included in a plurality of pixel points adjacent to the first pixel point in the first reconstructed image.
[0098] Therefore, in the embodiments of the present application, when there is relative motion between the shooting object and the motion sensor, the image reconstruction can be performed based on the data collected by the motion sensor to obtain the reconstructed image, and even when the RGB sensor cannot shoot a clear image, a clear image can be obtained.
[0099] In a possible implementation, the determining, according to at least one event included in the event image, a color type corresponding to each pixel point in the event image to obtain a first reconstructed image comprises: scanning each pixel point in the event image in a first direction, determining a color type corresponding to each pixel point in the event image to obtain a first reconstructed image, wherein if the first pixel point has an event, the color type of the first pixel point is determined to be a first color type, and if a second pixel point arranged before the first pixel point in the first direction does not have an event, the color type corresponding to the second pixel point is a second color type, the first color type and the second color type are different color types, and a pixel point having an event represents a pixel point corresponding to a position where a change is monitored by the motion sensor in the event image.
[0100] In the embodiments of the present application, the image reconstruction can be performed based on the event of each pixel point in the event image by scanning the event image, so that a clearer event image is obtained. Therefore, in the embodiments of the present application, the information collected by the motion sensor can be used for image reconstruction, and the reconstructed image can be obtained efficiently and quickly, thereby improving the efficiency of subsequent image recognition, image classification and the like of the reconstructed image. Even in some scenes where a moving object is shot or there is shooting jitter, a clear RGB image cannot be shot, and the information collected by the motion sensor can be used for image reconstruction, so that a clearer image can be quickly and accurately reconstructed for subsequent identification or classification tasks.
[0101] In a possible implementation, the first direction is a direction set in advance, or the first direction is determined according to the data collected by the IMU, or the first direction is determined according to the image shot by the color RGB camera. Therefore, in the embodiments of the present application, the direction of scanning the event image can be determined in multiple ways to adapt to more scenes.
[0102] In a possible implementation, if a plurality of third pixel points arranged after the first pixel point in the first direction do not have events, the color types corresponding to the plurality of third pixel points are the first color type. Therefore, in the embodiments of the present application, when there are a plurality of continuous pixel points without events, the color types corresponding to the continuous pixel points are the same, thereby avoiding the case that the edge is not clear due to the movement of the same object in the actual scene.
[0103] In a possible implementation, if a fourth pixel point arranged after the first pixel point and adjacent to the first pixel point in the first direction has an event, and a fifth pixel point arranged after the fourth pixel point and adjacent to the fourth pixel point in the first direction does not have an event, the color types corresponding to the fourth pixel point and the fifth pixel point are both the first color type.
[0104] Therefore, when there are at least two continuous pixel points having events in the event image, the color type can not be changed when the second event is scanned, so that the edge of the reconstructed image is not clear due to the edge of the target object being too wide.
[0105] In a possible implementation, after the scanning of each pixel point in the event image in the first direction, the determination of the color type corresponding to each pixel point in the event image, and the obtaining of the first reconstructed image, the method further includes: scanning the event image in a second direction, determining the color type corresponding to each pixel point in the event image, and obtaining a second reconstructed image, the second direction being different from the first direction; and fusing the first reconstructed image and the second reconstructed image to obtain an updated first reconstructed image.
[0106] In the embodiments of the present application, the event image can be scanned in different directions, so that multiple reconstructed images are obtained from multiple directions, and then the multiple reconstructed images are fused to obtain a more accurate reconstructed image.
[0107] In a possible implementation, the method further includes: if the first reconstructed image does not meet the preset requirement, updating the motion information, updating the event image according to the updated motion information, and obtaining an updated first reconstructed image according to the updated event image.
[0108] In the embodiments of the present application, the event image can be updated in combination with the information collected by the motion sensor, so that the updated event image is clearer.
[0109] In a possible implementation, before the color type corresponding to each pixel point in the event image is determined according to at least one event included in the event image, and a first reconstructed image is obtained, the method further includes: compensating the event image according to a motion parameter when the target object and the motion sensor perform relative motion, to obtain a compensated event image, the motion parameter including one or more of the following: depth, optical flow information, acceleration of the motion sensor performing motion, or angular velocity of the motion sensor performing motion, the depth representing a distance between the motion sensor and the target object, and the optical flow information representing information of a motion speed of relative motion between the motion sensor and the target object.
[0110] Therefore, in the embodiments of the present application, the event image can also be compensated according to the motion parameter, so that the event image is clearer, and the reconstructed image obtained through reconstruction is also clearer.
[0111] In a possible implementation, the color type of a pixel point in the reconstructed image is determined according to a color captured by a color RGB camera. In the embodiments of the present application, the color in the actual scene can be determined according to the RGB camera, so that the color of the reconstructed image matches the color in the actual scene, and user experience is improved.
[0112] In a possible implementation, the method further includes: obtaining an RGB image according to data captured by the RGB camera; and fusing the RGB image and the first reconstructed image to obtain an updated first reconstructed image. Therefore, in the embodiments of the present application, the RGB image and the reconstructed image can be fused, so that the reconstructed image obtained finally is clearer.
[0113] In an eighteenth aspect, the present application further provides an image processing apparatus having a function of implementing the method of the eighteenth aspect or any one of the possible implementation manners of the eighteenth aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0114] In a nineteenth aspect, the present application provides a method for image processing, comprising: obtaining a first event image (event image) and a plurality of first images captured, the first event image comprising information of an object moving in a preset range within a capturing time period of the plurality of first images, the plurality of first images corresponding to different exposure time lengths, and the preset range being a capturing range of a camera; calculating a first shake degree corresponding to each first image in the plurality of first images according to the first event image, the first shake degree being used to represent a degree of shake of the camera when capturing the plurality of first images; determining a fusion weight of each first image in the plurality of first images according to the first shake degree corresponding to the each first image, wherein the first shake degree and the fusion weight corresponding to the plurality of first images are in a negative correlation relationship; and fusing the plurality of first images according to the fusion weight of the each first image to obtain a target image.
[0115] Therefore, in the embodiments of the present application, the degree of shake when capturing the RGB images can be quantified through the event image, and the fusion weight of each RGB image can be determined according to the degree of shake of the RGB image. Generally, the fusion weight corresponding to the RGB image with a low degree of shake is higher, so that the information included in the final target image is more inclined to the information included in the clearer RGB image, thereby obtaining a clearer target image. Generally, the weight value corresponding to the RGB image with a high degree of shake is smaller, and the weight value corresponding to the RGB image with a low degree of shake is larger, so that the information included in the final target image is more inclined to the information included in the clearer RGB image, so that the final target image is clearer, and the user experience is improved. Moreover, if the target image is used for subsequent image recognition or feature extraction, the obtained recognition result or extracted feature is more accurate.
[0116] In a possible implementation, before the fusion weight of each first image in the plurality of first images is determined according to the first shake degree, the method further comprises: if the first shake degree is not higher than a first preset value and higher than a second preset value, performing de-shake processing on the each first image to obtain a de-shake first image.
[0117] Therefore, in the embodiments of the present application, the shake conditions can be distinguished based on dynamic data, the RGB images are directly fused when there is no shake, the RGB images are adaptively de-shaken when the shake is not strong, and the RGB images are supplemented when the shake is strong, so that various shake degrees are used, and the generalization ability is strong.
[0118] In a possible implementation, the determining the fusion weight of each of the first images according to the first shake degree comprises: if the first shake degree is higher than a first preset value, taking a second image again, the second shake degree of the second image is not higher than the first preset value; calculating the fusion weight of each of the first images according to the first shake degree of each of the first images, and calculating the fusion weight of the second image according to the second shake degree; and the fusing the first images to obtain the target image according to the fusion weight of each of the first images comprises: fusing the first images and the second image according to the fusion weight of each of the first images and the fusion weight of the second image to obtain the target image.
[0119] Generally, the higher the shake degree of the RGB image, the smaller the corresponding weight value, and the lower the shake degree of the RGB image, the larger the corresponding weight value, so that the information included in the final target image is more inclined to the information included in the clearer RGB image, the final target image is clearer, and the user experience is improved. If the target image is used for subsequent image recognition or feature extraction, the obtained recognition result or extracted feature is more accurate. For the RGB image with high shake degree, the RBG image can be taken again to obtain the RGB image with lower shake degree and higher clarity, so that more clear images can be used for image fusion in subsequent image fusion, and the final target image is also clearer.
[0120] In a possible implementation, before the second image is taken again, the method further comprises: acquiring a second event image, the second event image being obtained before the first event image is acquired; and calculating an exposure parameter according to the information included in the second event image, the exposure parameter being used for taking the second image.
[0121] Therefore, in the embodiments of the present application, the information collected by the dynamic perception camera (i.e., the motion sensor) is used to adaptively adjust the exposure strategy, that is, the high dynamic range perception characteristics of the texture in the shooting range are used to adaptively supplement the image with appropriate exposure time, and the ability of the camera to capture texture information in strong light or dark light areas is improved.
[0122] In a possible implementation, the rephotographing to obtain the second image further includes: dividing the first event image into a plurality of regions, and dividing a third image into a plurality of regions, the third image being a first image with the minimum exposure value in the plurality of first images, and the plurality of regions included in the first event image corresponding to positions of the plurality of regions included in the third image, the exposure value including at least one of an exposure time, an exposure amount, or an exposure index; determining whether each region in the first event image includes first texture information, and whether each region in the third image includes second texture information; if a first region in the first event image includes the first texture information, and a region corresponding to the first region in the third image does not include the second texture information, photographing to obtain the second image according to the exposure parameter, the first region being any region in the first dynamic region.
[0123] Therefore, in the embodiment of the present application, if a region in the first dynamic region includes texture information, and a region corresponding to the region in the RGB image with the minimum exposure value does not include texture information, it indicates that the region in the RGB image has a high blur degree, and the RGB image can be rephotographed. If no region in the first event image includes texture information, the RGB image does not need to be rephotographed.
[0124] In a twentieth aspect, the present application provides an image processing method, including: first, detecting motion information of a target object, the motion information including information of a motion trajectory of the target object when moving in a preset range, the preset range being a camera shooting range; then, determining focusing information according to the motion information, the focusing information including a parameter for focusing on the target object in the preset range; subsequently, focusing on the target object in the preset range according to the focusing information, and photographing an image of the preset range.
[0125] Therefore, in the embodiment of the present application, the motion trajectory of the target object in the shooting range of the camera can be detected, and then the focusing information is determined according to the motion trajectory of the target object and the focusing is completed, so that a clearer image can be photographed. Even if the target object is in motion, the target object can be accurately focused, and a clear image of the motion state can be photographed, improving the user experience.
[0126] In a possible implementation, the determining the focus information according to the motion information can include: predicting a motion trajectory of the target object in a preset time period according to the motion information, i.e., information of the motion trajectory of the target object when the target object moves in a preset range, to obtain a predicted region, the predicted region being a region in which the target object is located in the preset time period and which is predicted; determining a focus region according to the predicted region, the focus region including at least one focus point for focusing on the target object, and the focus information including position information of the at least one focus point.
[0127] Therefore, in the embodiments of the present application, the future motion trajectory of the target object can be predicted, and the focus region can be determined according to the predicted region, so that the focusing on the target object can be accurately completed. Even if the target object is in high-speed motion, the embodiments of the present application can focus on the target object in advance through prediction, so that the target object is in the focus region, and a clearer high-speed motion target object can be photographed.
[0128] In a possible implementation, the determining the focus region according to the predicted region can include: if the predicted region meets a preset condition, determining the predicted region as the focus region; and if the predicted region does not meet the preset condition, re-predicting the motion trajectory of the target object in the preset time period according to the motion information to obtain a new predicted region, and determining the focus region according to the new predicted region. The preset condition can be that the predicted region includes a complete target object, or that the area of the predicted region is greater than a preset value.
[0129] Therefore, in the embodiments of the present application, the focus region is determined according to the predicted region only when the predicted region meets the preset condition, and the camera is triggered to photograph, and when the predicted region does not meet the preset condition, the camera is not triggered to photograph, so that the target object in the photographed image can be complete, or meaningless photographing can be avoided. Moreover, when no photographing is performed, the camera can be in an unstarted state, and the camera is triggered to photograph only when the predicted region meets the preset condition, so that the power consumption of the camera can be reduced.
[0130] In a possible implementation, the motion information further includes at least one of a motion direction and a motion speed of the target object; and the predicting the motion trajectory of the target object in the preset time period according to the motion information to obtain the predicted region can include: predicting the motion trajectory of the target object in the preset time period according to the motion trajectory of the target object when the target object moves in the preset range, and the motion direction and / or the motion speed.
[0131] Therefore, in the embodiments of the present application, the motion trajectory of the target object in the future preset time period can be predicted according to the motion trajectory, the motion direction and / or the motion speed of the target object in the preset range, so that the region where the target object is located in the future preset time period can be accurately predicted, and the target object can be focused more accurately, and a clearer image can be captured.
[0132] In a possible implementation, the prediction of the motion trajectory of the target object in the preset time period according to the motion trajectory, the motion direction and / or the motion speed of the target object in the preset range can include: fitting a change function of a center point of the region where the target object is located with respect to time according to the motion trajectory, the motion direction and / or the motion speed of the target object in the preset range; then calculating a predicted center point according to the change function, the predicted center point being a center point of the region where the target object is located in the predicted time period; and obtaining the predicted region according to the predicted center point.
[0133] Therefore, in the embodiments of the present application, the change function of the center point of the region where the target object is located with respect to time can be fitted according to the motion trajectory of the target object, and then the center point of the region where the target object is located at a certain moment in the future can be predicted according to the change function, and the predicted region can be determined according to the center point, so that the target object can be focused more accurately, and a clearer image can be captured.
[0134] In a possible implementation, the image of the prediction range can be captured by an RGB camera, and the focusing of the target object in the preset range according to the focusing information can include: taking at least one point with the smallest norm distance from the center point of the focusing region as a focusing point for focusing.
[0135] Therefore, in the embodiments of the present application, at least one point with the smallest norm distance from the center point of the focusing region can be selected as a focusing point, and focusing can be performed, so that the focusing of the target object is completed.
[0136] In a possible implementation, the motion information includes the region where the target object is currently located, and the determination of the focusing information according to the motion information can include: determining the region where the target object is currently located as a focusing region, the focusing region including at least one focusing point for focusing on the target object, and the focusing information including position information of the at least one focusing point.
[0137] Therefore, in the embodiments of the present application, the information of the movement track of the target object in the preset range can include the region where the target object is currently located and the region where the target object has historically located, and the region where the target object is currently located can be taken as the focusing region, so that the focusing on the target object is completed, and a clearer image can be captured.
[0138] In a possible implementation, before capturing the image of the preset range, the method can further include: obtaining an exposure parameter; and the capturing of the image of the preset range can include: capturing the image of the preset range according to the exposure parameter.
[0139] Therefore, in the embodiments of the present application, the exposure parameter can be adjusted, so that the capturing is completed through the exposure parameter, and a clear image is obtained.
[0140] In a possible implementation, the obtaining of the exposure parameter can include: determining the exposure parameter according to the movement information, wherein the exposure parameter includes an exposure duration, the movement information includes a movement speed of the target object, and the exposure duration is negatively correlated with the movement speed of the target object.
[0141] Therefore, in the embodiments of the present application, the exposure duration can be determined according to the movement speed of the target object, so that the exposure duration is matched with the movement speed of the target object, for example, the faster the movement speed, the shorter the exposure duration, and the slower the movement speed, the longer the exposure duration. The overexposure or underexposure can be avoided, so that a clearer image can be captured subsequently, and the user experience is improved.
[0142] In a possible implementation, the obtaining of the exposure parameter can include: determining the exposure parameter according to the light intensity, wherein the exposure parameter includes an exposure duration, and the size of the light intensity in the preset range is negatively correlated with the exposure duration.
[0143] Therefore, in the embodiments of the present application, the exposure duration can be determined according to the detected light intensity, for example, the greater the light intensity, the shorter the exposure duration, and the smaller the light intensity, the longer the exposure duration, so that the appropriate exposure amount can be ensured, and a clearer image can be captured.
[0144] In a possible implementation, after the image of the preset range is captured, the method can further include: fusing the images in the preset range according to the monitored information of the movement of the target object corresponding to the images, to obtain a target image in the preset range.
[0145] Therefore, in the embodiments of the present application, the motion of the target object in the preset range can be monitored while the image is being captured, and the corresponding motion information of the target object in the image, such as the contour of the target object, the position of the target object in the preset range, and the like, can be obtained, and the captured image can be enhanced based on the information to obtain a clearer target image.
[0146] In a possible implementation, the above-mentioned detecting the motion information of the target object in the preset range can include: monitoring the motion of the target object in the preset range by a dynamic vision sensor (DVS) to obtain the motion information.
[0147] Therefore, in the embodiments of the present application, the motion of the target object in the preset range can be monitored while the image is being captured, and the corresponding motion information of the target object in the image, such as the contour of the target object, the position of the target object in the preset range, and the like, can be obtained, and the captured image can be enhanced based on the information to obtain a clearer target image.
[0148] In a possible implementation, the above-mentioned detecting the motion information of the target object in the preset range can include: monitoring the motion of the target object in the preset range by a dynamic vision sensor (DVS) to obtain the motion information.
[0149] In a possible implementation, the above-mentioned detecting the motion information of the target object in the preset range can include: monitoring the motion of the target object in the preset range by a dynamic vision sensor (DVS) to obtain the motion information.
[0150] The beneficial effects of the twenty-second aspect and any possible implementation of the twenty-second aspect can refer to the description of the twentieth aspect and any possible implementation of the twentieth aspect.
[0151] In a possible implementation, the graphical user interface can further include: predicting a motion trajectory of the target object in a preset time period in response to the motion information, obtaining a prediction region, the prediction region being a region in which the target object is located in the preset time period predicted, and determining the focus region according to the prediction region, displaying the focus region in the display screen, the focus region including at least one focus point for focusing on the target object, and the focus information including position information of the at least one focus point.
[0152] In a possible implementation, the graphical user interface can specifically include: if the prediction region meets a preset condition, displaying the focus region in the display screen in response to determining the focus region according to the prediction region; and if the prediction region does not meet the preset condition, displaying the focus region in the display screen in response to predicting a new prediction region in the preset time period in response to the motion information, and determining the focus region according to the new prediction region.
[0153] In a possible implementation, the motion information further includes at least one of a motion direction and a motion speed of the target object; and the graphical user interface can specifically include: predicting the motion trajectory of the target object in the preset time period in response to the motion trajectory of the target object when moving in a preset range, and the motion direction and / or the motion speed, and displaying the prediction region in the display screen.
[0154] In a possible implementation, the graphical user interface can specifically include: fitting a change function of a center point of a region in which the target object is located over time in response to the motion trajectory of the target object when moving in a preset range, and the motion direction and / or the motion speed, calculating a prediction center point from the change function, the prediction center point being a center point of the region in which the target object is located predicted, and obtaining the prediction region from the prediction center point, and displaying the prediction region in the display screen.
[0155] In a possible implementation, an image of the prediction region is captured by an RGB camera, and the graphical user interface can specifically include: displaying, in the display screen, an image captured based on at least one point with the smallest norm distance from the center point of the focus region in a plurality of focus points of the RGB camera as a focus point.
[0156] In a possible implementation, the motion information includes a region where the target object is currently located, and the graphical user interface specifically can include: in response to taking the region where the target object is currently located as the focus region, the focus region including at least one focus point for focusing on the target object, the focus information including position information of the at least one focus point, and displaying the focus region in the display screen.
[0157] In a possible implementation, the graphical user interface further can include: in response to fusing images in the preset range according to the monitored motion information of the target object corresponding to the images, obtaining a target image in the preset range, and displaying the target image in the display screen.
[0158] In a possible implementation, the motion information is obtained by monitoring a motion condition of the target object in the preset range by using a dynamic vision sensor (DVS).
[0159] In a possible implementation, the graphical user interface specifically can include: in response to obtaining an exposure parameter before capturing the image of the preset range, displaying the exposure parameter in the display screen; and in response to capturing the image of the preset range according to the exposure parameter, displaying the image of the preset range captured according to the exposure parameter in the display screen.
[0160] In a possible implementation, the exposure parameter is determined according to the motion information, and the exposure parameter includes an exposure duration, and the exposure duration is negatively correlated with a motion speed of the target object.
[0161] In a possible implementation, the exposure parameter is determined according to an illumination intensity, which can be an illumination intensity detected by a camera or an illumination intensity detected by a motion sensor, and the exposure parameter includes an exposure duration, and a size of the illumination intensity in the preset range is negatively correlated with the exposure duration.
[0162] In a twenty-third aspect, the present application provides an image processing method, which comprises: first, acquiring an event stream and a first RGB image by a camera with a motion sensor (e.g., a DVS) and an RGB sensor, wherein the acquired event stream comprises at least one event image, each event image in the at least one event image is generated from the motion trail information of a target object (i.e., a moving object) moving within the monitoring range of the motion sensor, and the first RGB image is the superposition of the scene at each time captured by the camera within the exposure time. After the event stream and the first RGB image are acquired, a mask can be constructed according to the event stream, the mask is used to determine the motion region of each event image in the event stream, i.e., to determine the position of the moving object in the RGB image. After the event stream, the first RGB image and the mask are obtained according to the above steps, a second RGB image can be obtained according to the event stream, the first RGB image and the mask, the second RGB image being the RGB image with the target object removed.
[0163] In the above-mentioned embodiments of the present application, the moving object can be removed based on only one RGB image and event stream, so as to obtain an RGB image without moving object. Compared with the prior art which requires multiple RGB images and event streams to remove the moving object, only one RGB image needs to be captured by the user, and the user experience is better.
[0164] In a possible implementation, before the mask is constructed according to the event stream, the method can further comprise: when the motion sensor detects a motion mutation in the monitoring range at a first time, triggering the camera to capture a third RGB image; and the second RGB image is obtained according to the event stream, the first RGB image and the mask, which can comprise: the second RGB image is obtained according to the event stream, the first RGB image, the third RGB image and the mask.
[0165] In the above-mentioned embodiments of the present application, whether the motion data collected by the motion sensor has a motion mutation can be determined, and when the motion mutation exists, the third RGB image is triggered to be captured by the camera, and then the event stream and the first RGB image are obtained in the above-mentioned similar manner, the mask is constructed according to the event stream, and finally the second RGB image without the motion foreground is obtained according to the event stream, the first RGB image, the third RGB image and the mask. The third RGB image obtained is automatically captured by the camera triggered under the condition of the motion mutation, and has high sensitivity, so that a frame of image can be obtained at the beginning of the user's perception of the change of the motion object, and the motion object can be removed more effectively based on the third RGB image and the first RGB image.
[0166] In a possible implementation, the motion sensor monitoring the motion mutation in the monitoring range at the first time includes: the overlapping part between the generated area of the first event stream collected by the motion sensor at the first time and the generated area of the second event stream collected by the motion sensor at the second time in the monitoring range is less than a preset value.
[0167] In the above-mentioned embodiments of the present application, the determination condition of the motion mutation is specifically described, and is feasible.
[0168] In a possible implementation, the mask can be constructed according to the event stream in the following manner: first, the monitoring range of the motion sensor can be divided into a plurality of preset neighborhoods (set as neighborhood k), and then in each neighborhood k, when the number of event images of the event stream in the preset time length Δt exceeds the threshold value P, the corresponding neighborhood is determined as a motion region, and the motion region can be marked as 0, and if the number of event images of the event stream in the preset time length Δt does not exceed the threshold value P, the corresponding neighborhood is determined as a background region, and the background region can be marked as 1.
[0169] In the above-mentioned embodiments of the present application, a method for constructing the mask is specifically described, which is simple and easy to operate.
[0170] In a twenty-fourth aspect, the present application further provides an image processing device having a function of implementing the method of the above-mentioned twenty-second aspect or any one of the possible implementation manners of the twenty-second aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-mentioned functions.
[0171] In a twenty-fifth aspect, the present application provides a pose estimation method applied to a simultaneous localization and mapping (SLAM) scenario. The method comprises: obtaining, by a terminal, a first event image and a first RGB image, the first event image being time-aligned with the first target image, the first target image comprising an RGB image or a depth image. The first event image is an image representing a motion trajectory of a target object when the target object generates motion within a detection range of a motion sensor. The terminal determines an integration time of the first event image. If the integration time is less than a first threshold, the terminal determines not to perform pose estimation by using the first target image. The terminal performs pose estimation according to the first event image.
[0172] In the present solution, when the terminal determines that the integration time of the event image is less than a threshold value based on the fact that the terminal is currently in a scenario in which an RGB camera is difficult to collect effective environmental information, the terminal determines not to perform pose estimation by using the RGB image with poor quality, so as to improve the accuracy of pose estimation.
[0173] Optionally, in a possible implementation, the method further comprises: determining an acquisition time of the first event image and an acquisition time of the first target image; and determining that the first event image is time-aligned with the first target image according to the fact that a time difference between the acquisition time of the first target image and the acquisition time of the first event image is less than a second threshold. The second threshold can be determined according to the accuracy of SLAM and the frequency of the RGB camera in collecting the RGB image, for example, the second threshold can be 5 milliseconds or 10 milliseconds.
[0174] Optionally, in a possible implementation, the obtaining of the first event image comprises: obtaining N continuous DVS events; and integrating the N continuous DVS events into the first event image. The method further comprises: determining the acquisition time of the first event image according to the acquisition time of the N continuous DVS events.
[0175] Optionally, in a possible implementation, the determining of the integration time of the first event image comprises: determining N continuous DVS events used for integration into the first event image; and determining the integration time of the first event image according to the acquisition time of a first DVS event and a last DVS event in the N continuous DVS events. Since the first event image is obtained by integrating the N continuous DVS events, the terminal can determine the acquisition time of the first event image according to the acquisition time of the N continuous DVS events, that is, determine the acquisition time of the first event image as a time period from the acquisition of the first DVS event to the acquisition of the last DVS event in the N continuous DVS events.
[0176] Optionally, in a possible implementation, the method further includes: obtaining a second event image, the second event image being an image representing a motion trajectory of the target object when the target object generates motion within a detection range of a motion sensor. Wherein, a time period during which the motion sensor detects to obtain the first event image is different from a time period during which the motion sensor detects to obtain the second event image. If there is no RGB image that is time-aligned with the second event image, it is determined that the second event image does not have an RGB image for jointly performing pose estimation; and pose estimation is performed according to the second event image.
[0177] Optionally, in a possible implementation, before the pose is determined according to the second event image, the method further includes: if it is determined that the second event image has time-aligned inertial measurement unit (IMU) data, the pose is determined according to the second event image and the IMU data corresponding to the second event image; and if it is determined that the second event image does not have time-aligned inertial measurement unit (IMU) data, the pose is determined only according to the second event image.
[0178] Optionally, in a possible implementation, the method further includes: obtaining a second target image, the second target image including an RGB image or a depth image; if there is no event image that is time-aligned with the second target image, it is determined that the second target image does not have an event image for jointly performing pose estimation; and pose estimation is performed according to the second target image.
[0179] Optionally, in a possible implementation, the method further includes: performing loop detection according to the first event image and a dictionary, the dictionary being a dictionary constructed based on event images. That is, before performing loop detection, the terminal can construct a dictionary based on event images in advance, so as to perform loop detection based on the dictionary in the process of performing loop detection.
[0180] Optionally, in a possible implementation, the method further includes: obtaining a plurality of event images, the plurality of event images being event images for training, and the plurality of event images can be event images captured by the terminal in different scenes. Visual features of the plurality of event images are obtained, and the visual features can include, for example, texture, pattern, or gray scale statistics of the images, and the like. The visual features are clustered by a clustering algorithm to obtain clustered visual features, and the clustered visual features have corresponding descriptors. By clustering the visual features, similar visual features can be classified into a category, so as to facilitate subsequent matching of the visual features. Finally, the dictionary is constructed according to the clustered visual features.
[0181] Optionally, in a possible implementation, the performing loop detection according to the first event image and the dictionary comprises: determining a descriptor of the first event image; determining a visual feature corresponding to the descriptor of the first event image in the dictionary; determining a bag-of-words vector corresponding to the first event image based on the visual feature; and determining a similarity between the bag-of-words vector corresponding to the first event image and bag-of-words vectors of other event images, to determine an event image matched by the first event image.
[0182] In a twenty-sixth aspect, the present application provides a key frame selection method, comprising: obtaining an event image; determining first information of the event image, the first information comprising events and / or features in the event image; and determining the event image as a key frame if it is determined that the event image at least satisfies a first condition based on the first information, the first condition being related to a number of events and / or a number of features.
[0183] In this solution, whether the current event image is a key frame is determined by determining information such as the number of events, the distribution of events, the number of features, and / or the distribution of features in the event image, which can realize fast selection of key frames, has a small amount of calculation, and can meet the requirement of fast selection of key frames in scenarios such as video analysis, video encoding and decoding, or security monitoring.
[0184] Optionally, in a possible implementation, the first condition comprises one or more of the number of events in the event image being greater than a first threshold, the number of event valid regions in the event image being greater than a second threshold, the number of features in the event image being greater than a third threshold, and the number of feature valid regions in the event image being greater than a fourth threshold.
[0185] Optionally, in a possible implementation, the method further comprises: obtaining a depth image that is time-sequentially aligned with the event image; and determining the event image and the depth image as key frames if it is determined that the event image at least satisfies the first condition based on the first information.
[0186] Optionally, in a possible implementation, the method further comprises: obtaining an RGB image that is time-sequentially aligned with the event image; obtaining a number of features and / or a feature valid region of the RGB image; and determining the event image and the RGB image as key frames if it is determined that the event image at least satisfies the first condition based on the first information, and the number of features of the RGB image is greater than a fifth threshold and / or the number of feature valid regions of the RGB image is greater than a sixth threshold.
[0187] Optionally, in a possible implementation, if it is determined based on the first information that the event image at least meets the first condition, the method further includes: determining second information of the event image, the second information including motion features and / or pose features in the event image; and if it is determined based on the second information that the event image at least meets a second condition, determining the event image as a key frame, the second condition being related to a motion change amount and / or a pose change amount.
[0188] Optionally, in a possible implementation, the method further includes: determining a definition and / or a brightness consistency index of the event image; and if it is determined based on the second information that the event image at least meets the second condition, and the definition of the event image is greater than a definition threshold and / or the brightness consistency index of the event image is greater than a preset index threshold, determining the event image as a key frame.
[0189] Optionally, in a possible implementation, the determining of the brightness consistency index of the event image includes: if a pixel in the event image represents a light intensity change polarity, calculating an absolute value of a difference between an event quantity of the event image and an event quantity of a neighboring key frame, and dividing the absolute value by a pixel quantity of the event image to obtain the brightness consistency index of the event image; or if the pixel in the event image represents a light intensity, performing pixel-by-pixel difference between the event image and the neighboring key frame, calculating an absolute value of the difference, performing summation operation on the absolute value corresponding to each group of pixels, and dividing a summation result obtained by the summation operation by the pixel quantity to obtain the brightness consistency index of the event image.
[0190] Optionally, in a possible implementation, the method further includes: obtaining an RGB image that is time-aligned with the event image; determining a definition and / or a brightness consistency index of the RGB image; and if it is determined based on the second information that the event image at least meets the second condition, and the definition of the RGB image is greater than a definition threshold and / or the brightness consistency index of the RGB image is greater than a preset index threshold, determining the event image and the RGB image as a key frame.
[0191] Optionally, in a possible implementation, the second condition includes one or more of the following: a distance between the event image and a previous key frame exceeds a preset distance value, a rotation angle between the event image and the previous key frame exceeds a preset angle value, and the distance between the event image and the previous key frame exceeds the preset distance value and the rotation angle between the event image and the previous key frame exceeds the preset angle value.
[0192] In a twenty-seventh aspect, the present application provides a pose estimation method, comprising: acquiring a first event image and a target image corresponding to the first event image, the first event image and the target image capturing the same environment information, the target image comprising a depth image or an RGB image; determining a first motion region in the first event image; determining a second motion region in the target image according to the first motion region; and performing pose estimation according to the second motion region in the target image.
[0193] In the present solution, a dynamic region in a scene is captured by an event image, and pose determination is performed based on the dynamic region, so that the pose information can be determined in preparation.
[0194] Optionally, in a possible implementation, the determining the first motion region in the first event image comprises: if a dynamic vision sensor (DVS) for acquiring the first event image is static, acquiring a pixel point with an event response in the first event image; and determining the first motion region according to the pixel point with the event response.
[0195] Optionally, in a possible implementation, the determining the first motion region according to the pixel point with the event response comprises: determining an outline formed by the pixel point with the event response in the first event image; and if an area surrounded by the outline is greater than a first threshold, determining a region surrounded by the outline as the first motion region.
[0196] Optionally, in a possible implementation, the determining the first motion region in the first event image comprises: if the DVS for acquiring the first event image is in motion, acquiring a second event image, the second event image being a previous frame event image of the first event image; calculating a displacement size and a displacement direction of a pixel in the first event image relative to the second event image; and if the displacement direction of the pixel in the first event image is different from that of surrounding pixels, or a difference between the displacement size of the pixel in the first event image and that of the surrounding pixels is greater than a second threshold, determining that the pixel belongs to the first motion region.
[0197] Optionally, in a possible implementation, the method further comprises: determining a static region in the target image according to the first motion region; and determining a pose according to the static region in the target image.
[0198] In a twenty-eighth aspect, the present application provides a data processing apparatus having a function of implementing the method of the twenty-fifth aspect or any possible implementation manner of the twenty-fifth aspect, or having a function of implementing the method of the twenty-sixth aspect or any possible implementation manner of the twenty-sixth aspect, or having a function of implementing the method of the twenty-seventh aspect or any possible implementation manner of the twenty-seventh aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0199] In a twenty-ninth aspect, the embodiments of the present application provide an apparatus, including a processor and a memory, wherein the processor and the memory are interconnected through a circuit, and the processor invokes program codes in the memory to execute functions related to processing in the method shown in any one of the first aspect to the twenty-seventh aspect. Optionally, the apparatus can be a chip.
[0200] In a thirtieth aspect, the present application provides an electronic device, including a display module, a processing module and a storage module.
[0201] The display module is configured to display a graphical user interface of an application stored in the storage module, and the graphical user interface can be the graphical user interface of any one of the preceding aspects.
[0202] In a thirty-first aspect, the embodiments of the present application provide an apparatus, which can also be referred to as a digital processing chip or a chip. The chip includes a processing unit and a communication interface, the processing unit acquires program instructions through the communication interface, the program instructions are executed by the processing unit, and the processing unit is configured to execute functions related to processing in any optional implementation manner of the first aspect to the twenty-seventh aspect.
[0203] In a thirty-second aspect, the embodiments of the present application provide a computer readable storage medium, including instructions, which, when running on a computer, cause the computer to execute the method in any optional implementation manner of the first aspect to the twenty-seventh aspect.
[0204] In a thirty-third aspect, the embodiments of the present application provide a computer program product including instructions, which, when running on a computer, cause the computer to execute the method in any optional implementation manner of the first aspect to the twenty-seventh aspect. BRIEF DESCRIPTION OF DRAWINGS
[0205] Figure 1A A system architecture schematic diagram is provided for the present application;
[0206] Figure 1B A structure schematic diagram of an electronic device is provided for the present application;
[0207] Figure 2 Another system architecture diagram provided for the present application;
[0208] Figure 3-a Diagram of data volume versus time for event stream based asynchronous read mode;
[0209] Figure 3-b Diagram of data volume versus time for frame scan based synchronous read mode;
[0210] Figure 4-a Block diagram of a vision sensor provided for the present application;
[0211] Figure 4-b Block diagram of another vision sensor provided for the present application;
[0212] Figure 5 Diagram of the principle of frame scan based synchronous read mode and event stream based asynchronous read mode according to embodiments of the present application;
[0213] Figure 6-a Diagram of a vision sensor operating in frame scan based read mode according to embodiments of the present application;
[0214] Figure 6-b Diagram of a vision sensor operating in event stream based read mode according to embodiments of the present application;
[0215] Figure 6-c Diagram of a vision sensor operating in event stream based read mode according to embodiments of the present application;
[0216] Figure 6-d Diagram of a vision sensor operating in frame scan based read mode according to embodiments of the present application;
[0217] Figure 7 Flow chart of a method for operating a vision sensor chip according to possible embodiments of the present application;
[0218] Figure 8 Block diagram of a control circuit provided for the present application;
[0219] Figure 9 Block diagram of an electronic device provided for the present application;
[0220] Figure 10 Diagram of data volume versus time for single data read mode and adaptive switched read mode according to possible embodiments of the present application;
[0221] Figure 11 Diagram of a pixel circuit provided for the present application;
[0222] Figure 12 Schematic diagram for representing events by light intensity information and events by polarity information;
[0223] Figure 12-a Schematic diagram for a structure of a data format control unit in a reading circuit in the present application;
[0224] Figure 12-b Schematic diagram for another structure of a data format control unit in a reading circuit in the present application;
[0225] Figure 13 Block diagram of another control circuit provided in the present application;
[0226] Figure 14 Block diagram of another control circuit provided in the present application;
[0227] Figure 15 Block diagram of another control circuit provided in the present application;
[0228] Figure 16 Block diagram of another control circuit provided in the present application;
[0229] Figure 17 Block diagram of another control circuit provided in the present application;
[0230] Figure 18 Schematic diagram for distinguishing between a single event representation and an adaptive converted event representation according to the present application;
[0231] Figure 19 Block diagram of another electronic device provided in the present application;
[0232] Figure 20 Flowchart of a method for operating a visual sensor chip according to a possible embodiment of the present application;
[0233] Figure 21 Schematic diagram of another pixel circuit provided in the present application;
[0234] Figure 22 Flowchart of an encoding method provided in the present application;
[0235] Figure 23 Block diagram of another visual sensor provided in the present application;
[0236] Figure 24 Schematic diagram for region division of a pixel array;
[0237] Figure 25 Block diagram of another control circuit provided in the present application;
[0238] Figure 26 Another block diagram of an electronic device provided for the present application;
[0239] Figure 27 A schematic diagram of a binary data stream provided for the present application;
[0240] Figure 28 A flowchart of a method for operating a vision sensor chip according to possible embodiments of the present application;
[0241] Figure 29-a Another block diagram of a vision sensor provided for the present application;
[0242] Figure 29-b Another block diagram of a vision sensor provided for the present application;
[0243] Figure 29-c Another block diagram of a vision sensor provided for the present application;
[0244] Figure 30 A schematic diagram of another pixel circuit provided for the present application;
[0245] Figure 31 A block diagram schematic of a third encoding unit provided for the present application;
[0246] Figure 32 A flowchart schematic of another encoding approach provided for the present application;
[0247] Figure 33 Another block diagram of an electronic device provided for the present application;
[0248] Figure 34 A flowchart of a method for operating a vision sensor chip according to possible embodiments of the present application;
[0249] Figure 35 An event schematic provided for the present application;
[0250] Figure 36 An event schematic at a certain time provided for the present application;
[0251] Figure 37 A motion area schematic provided for the present application;
[0252] Figure 38 A flowchart schematic of an image processing method provided for the present application;
[0253] Figure 39 A flowchart schematic of another image processing method provided for the present application;
[0254] Figure 40 A flowchart schematic of another image processing method provided for the present application;
[0255] Figure 41 An event image schematic diagram provided for the present application;
[0256] Figure 42 A flowchart schematic diagram of another image processing method provided for the present application;
[0257] Figure 43 A flowchart schematic diagram of another image processing method provided for the present application;
[0258] Figure 44 A flowchart schematic diagram of another image processing method provided for the present application;
[0259] Figure 45 A flowchart schematic diagram of an image processing method provided for the present application;
[0260] Figure 46A An event image schematic diagram provided for the present application;
[0261] Figure 46B An event image schematic diagram provided for the present application;
[0262] Figure 47A An event image schematic diagram provided for the present application;
[0263] Figure 47B An event image schematic diagram provided for the present application;
[0264] Figure 48 A flowchart schematic diagram of another image processing method provided for the present application;
[0265] Figure 49 A flowchart schematic diagram of another image processing method provided for the present application;
[0266] Figure 50 An event image schematic diagram provided for the present application;
[0267] Figure 51 A reconstructed image schematic diagram provided for the present application;
[0268] Figure 52 A flowchart schematic diagram of an image processing method provided for the present application;
[0269] Figure 53 A schematic diagram of a manner of fitting a motion trajectory provided for the present application;
[0270] Figure 54 A schematic diagram of a manner of determining a focus point provided for the present application;
[0271] Figure 55 A schematic diagram of a manner of determining a prediction center provided for the present application;
[0272] Figure 56 A flowchart of another image processing method provided by the present application;
[0273] Figure 57 A shooting range diagram provided by the present application;
[0274] Figure 58 A prediction area diagram provided by the present application;
[0275] Figure 59 A focusing area diagram provided by the present application;
[0276] Figure 60 A flowchart of another image processing method provided by the present application;
[0277] Figure 61 An image enhancement mode diagram provided by the present application;
[0278] Figure 62 A flowchart of another image processing method provided by the present application;
[0279] Figure 63 A flowchart of another image processing method provided by the present application;
[0280] Figure 64 A scene diagram applied by the present application;
[0281] Figure 65 Another scene diagram applied by the present application;
[0282] Figure 66 A GUI display diagram provided by the present application;
[0283] Figure 67 Another GUI display diagram provided by the present application;
[0284] Figure 68 Another GUI display diagram provided by the present application;
[0285] Figure 69A Another GUI display diagram provided by the present application;
[0286] Figure 69B Another GUI display diagram provided by the present application;
[0287] Figure 69C Another GUI display diagram provided by the present application;
[0288] Figure 70 Another GUI display diagram provided by the present application;
[0289] Figure 71 Another GUI provided by the present application;
[0290] Figure 72A Another GUI provided by the present application;
[0291] Figure 72B Another GUI provided by the present application;
[0292] Figure 73 Another image processing method provided by the present application;
[0293] Figure 74 An RGB image with low jitter provided by the present application;
[0294] Figure 75 An RGB image with high jitter provided by the present application;
[0295] Figure 76 An RGB image in a large light ratio scene provided by the present application;
[0296] Figure 77 Another event image provided by the present application;
[0297] Figure 78 An RGB image provided by the present application;
[0298] Figure 79 Another RGB image provided by the present application;
[0299] Figure 80 Another GUI provided by the present application;
[0300] Figure 81 A diagram showing the relationship between a photosensitive unit and a pixel value provided by the present application;
[0301] Figure 82 A flowchart of an image processing method provided by the present application;
[0302] Figure 83 An event stream provided by the present application;
[0303] Figure 84 A diagram showing a blurred image obtained by superimposing exposures of multiple shooting scenes provided by the present application;
[0304] Figure 85 A diagram showing a mask provided by the present application;
[0305] Figure 86A schematic diagram of constructing a mask provided by the present application;
[0306] Figure 87 An effect diagram of removing moving objects from image I to obtain image I' provided by the present application;
[0307] Figure 88 A flowchart of removing moving objects from image I to obtain image I' provided by the present application;
[0308] Figure 89 A schematic diagram of a small motion of a moving object during photographing provided by the present application;
[0309] Figure 90 A schematic diagram of triggering the camera to capture a third RGB image provided by the present application;
[0310] Figure 91 A schematic diagram of an image B captured by the camera triggered based on motion abruptness provided by the present application k And a schematic diagram of an image I actively captured by the user within a certain exposure time;
[0311] Figure 92 A flowchart of obtaining a second RGB image without moving objects based on a first RGB image and an event stream E provided by the present application;
[0312] Figure 93 A flowchart of obtaining a second RGB image without moving objects based on a first RGB image, a third RGB image and an event stream E provided by the present application;
[0313] Figure 94A Another GUI schematic diagram provided by the present application;
[0314] Figure 94B Another GUI schematic diagram provided by the present application;
[0315] Figure 95 A comparison schematic diagram of scenes captured by a traditional camera and a DVS provided by the present application
[0316] Figure 96 A comparison schematic diagram of scenes captured by a traditional camera and a DVS provided by the present application;
[0317] Figure 97 An outdoor navigation schematic diagram applied with a DVS provided by the present application;
[0318] Figure 98a A station navigation schematic diagram applied with a DVS provided by the present application;
[0319] Figure 98bA scenic spot navigation schematic diagram applying a DVS provided for the present application;
[0320] Figure 99 A shopping mall navigation schematic diagram applying a DVS provided for the present application;
[0321] Figure 100 A flowchart of a process of performing SLAM provided for the present application;
[0322] Figure 101 A flowchart of a pose estimation method 10100 provided for the present application;
[0323] Figure 102 A schematic diagram of integrating DVS events into an event image provided for the present application;
[0324] Figure 103 A flowchart of a key frame selection method 10300 provided for the present application;
[0325] Figure 104 A region division schematic diagram of an event image provided for the present application;
[0326] Figure 105 A flowchart of a key frame selection method 10500 provided for the present application;
[0327] Figure 106 A flowchart of a pose estimation method 1060 provided for the present application;
[0328] Figure 107 A flowchart of performing pose estimation based on a static region of an image provided for the present application;
[0329] Figure 108a A flowchart of performing pose estimation based on a motion region of an image provided for the present application;
[0330] Figure 108b A flowchart of performing pose estimation based on an overall region of an image provided for the present application;
[0331] Figure 109 An AR / VR glasses structure schematic diagram provided for the present application;
[0332] Figure 110 A gaze perception structure schematic diagram provided for the present application;
[0333] Figure 111 A network architecture schematic diagram provided for the present application;
[0334] Figure 112 A structure schematic diagram of an image processing device provided for the present application;
[0335] Figure 113 Another structural schematic diagram of an image processing apparatus provided in the present application is provided.
[0336] Figure 114 Another structural schematic diagram of an image processing apparatus provided in the present application is provided.
[0337] Figure 115 Another structural schematic diagram of an image processing apparatus provided in the present application is provided.
[0338] Figure 116 Another structural schematic diagram of an image processing apparatus provided in the present application is provided.
[0339] Figure 117 Another structural schematic diagram of an image processing apparatus provided in the present application is provided.
[0340] Figure 118 Another structural schematic diagram of an image processing apparatus provided in the present application is provided.
[0341] Figure 119 Another structural schematic diagram of a data processing apparatus provided in the present application is provided.
[0342] Figure 120 Another structural schematic diagram of a data processing apparatus provided in the present application is provided.
[0343] Figure 121 Another structural schematic diagram of an electronic device provided in the present application is provided. DETAILED DESCRIPTION
[0344] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0345] The electronic device, system architecture and method flow provided in the present application will be described in detail from different angles.
[0346] I. Electronic device
[0347] The method provided in the present application can be applied to various electronic devices, or said electronic devices execute the method provided in the present application. The electronic device can be applied to shooting scenes, such as photographing, security, automatic driving, unmanned aerial vehicle shooting, etc.
[0348] The electronic device in the present application can include but is not limited to: a smart mobile phone, a television, a tablet computer, a bracelet, a head-mounted display (HMD), an augmented reality (AR) device, a mixed reality (MR) device, a cellular phone, a smart phone, a personal digital assistant (PDA), a vehicle-mounted electronic device, a laptop computer, a personal computer (PC), a monitoring device, a robot, a vehicle-mounted terminal, an autonomous vehicle, etc. Of course, in the following embodiments, the specific form of the electronic device is not limited in any way.
[0349] Exemplarily, the architecture of the electronic device provided in the present application is as shown in Figure 1A
[0350] Among them, the electronic device, such as the car, mobile phone, AR / VR glasses, security monitoring device, camera or other smart home terminal and other devices as described in Figure 1A , can access the cloud platform through wired or wireless network, and the cloud platform is provided with a server, which can include a centralized server or a distributed server. The electronic device can communicate with the server of the cloud platform through wired or wireless network, so as to realize the transmission of data. For example, after the electronic device collects the device, it can be saved or backed up in the cloud platform to prevent data loss.
[0351] The electronic device can access the access point or base station to realize wireless or wired access to the cloud platform. For example, the access point can be a base station, and the electronic device is provided with a SIM card, which realizes the network authentication of the operator, so as to access the wireless network. Or, the access point can include a router, and the electronic device accesses the router through 2.4GHz or 5GHz wireless network, so as to access the cloud platform through the router.
[0352] In addition, the electronic device can process data alone, or process data through cooperation with the cloud, which can be adjusted according to the actual application scene. For example, the DVS can be arranged in the electronic device, and the DVS can work cooperatively with the camera or other sensors in the electronic device, or can work independently. The processor arranged in the DVS or the processor arranged in the electronic device processes the data collected by the DVS or other sensors, and the DVS or other sensors can also process the data collected by the DVS or other sensors cooperatively with the cloud device.
[0353] The following is an exemplary description of the specific structure of the electronic device.
[0354] For example, see Figure 1B The structure of the electronic device provided in this application will be illustrated below using a specific example.
[0355] It should be noted that the electronic device provided in this application may include, but is not limited to, those that are more advanced than those that are designed for use in electronic devices. Figure 1B More or fewer parts, Figure 1B The electronic device shown is merely an illustrative example. Those skilled in the art can add or remove components in the electronic device as needed, and this application does not limit this.
[0356] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a proximity sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, and an image sensor 180N. The image sensor 180N may include a separate color sensor 1801N and a separate motion sensor 1802N, or it may include a photosensitive unit (which can be called a color sensor pixel) of the color sensor. Figure 1B (not shown in the image) and the photosensitive unit of the motion sensor (which may be called a motion sensor pixel, Figure 1B (Not shown in the image).
[0357] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0358] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated in one or more processors.
[0359] The controller can generate operation control signals according to the instruction operation code and the timing signal, complete the control of fetching and executing instructions.
[0360] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that have just been used or are used repeatedly by the processor 110. If the processor 110 needs to use the instructions or data again, it can directly call from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system.
[0361] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0362] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 can contain multiple sets of I2C bus. The processor 110 can be coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces respectively. For example, the processor 110 can be coupled to the touch sensor 180K through an I2C interface, so that the processor 110 and the touch sensor 180K communicate through the I2C bus interface, and the touch function of the electronic device 100 is realized.
[0363] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can contain multiple sets of I2S bus. The processor 110 can be coupled to the audio module 170 through the I2S bus, and communication between the processor 110 and the audio module 170 is realized. In some embodiments, the audio module 170 can deliver audio signals to the wireless communication module 160 through the I2S interface, and the function of answering a phone through a Bluetooth headset is realized.
[0364] The PCM interface can also be used for audio communication, sampling, quantizing and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In some embodiments, the audio module 170 can also deliver audio signals to the wireless communication module 160 through the PCM interface, and the function of answering a phone through a Bluetooth headset is realized. Both the I2S interface and the PCM interface can be used for audio communication.
[0365] The UART interface is a universal serial data bus, which is used for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is usually used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface, and the Bluetooth function is realized. In some embodiments, the audio module 170 can deliver audio signals to the wireless communication module 160 through the UART interface, and the function of playing music through a Bluetooth headset is realized.
[0366] The MIPI interface can be used to connect the processor 110 and the display screen 194, the camera 193 and other peripheral devices. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), and the like. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface to implement the photographing function of the electronic device 100. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the electronic device 100.
[0367] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 and the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, and the like. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, and the like.
[0368] The USB interface 130 is an interface that conforms to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, or the like. The USB interface 130 can be used to connect a charger to charge the electronic device 100, or to transmit data between the electronic device 100 and a peripheral device. It can also be used to connect a headset to play audio through the headset. The interface can also be used to connect other electronic devices, such as AR devices and the like.
[0369] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation of the electronic device 100. In other embodiments of the present application, the electronic device 100 can also use different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0370] The charging management module 140 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input through the wireless charging coil of the electronic device 100. The charging management module 140 can charge the battery 142 while also providing power to the electronic device through the power management module 141.
[0371] The power management module 141 is configured to connect the battery 142 and the charging management module 140 to the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, the internal memory 121, the display 194, the camera 193, the wireless communication module 160, and the like. The power management module 141 can also be configured to monitor parameters such as battery capacity, battery cycle count, battery health status (leakage, impedance), and the like. In some embodiments, the power management module 141 can also be disposed in the processor 110. In some embodiments, the power management module 141 and the charging management module 140 can also be disposed in the same device.
[0372] The wireless communication function of the electronic device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, and the like.
[0373] The antenna 1 and the antenna 2 are configured to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be configured to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some embodiments, the antennas can be used in combination with a tuning switch.
[0374] The mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G and the like applied to the electronic device 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), and the like. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and perform filtering, amplification, and the like on the received electromagnetic waves, and transmit the processed electromagnetic waves to the modem processor for demodulation. The mobile communication module 150 can also amplify signals modulated by the modem processor, and radiate the amplified signals as electromagnetic waves through the antenna 1. In some embodiments, at least part of the function modules of the mobile communication module 150 can be disposed in the processor 110. In some embodiments, at least part of the function modules of the mobile communication module 150 and at least part of the modules of the processor 110 can be disposed in the same device.
[0375] The modem processor can include a modulator and a demodulator. The modulator is configured to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to a speaker 170A, a microphone 170B, etc.), or displays an image or a video through the display 194. In some embodiments, the modem processor can be a separate device. In other embodiments, the modem processor can be independent of the processor 110 and disposed in the same device as the mobile communication module 150 or other functional modules.
[0376] The wireless communication module 160 can provide a wireless communication solution including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, modulates and filters the electromagnetic wave signal, and transmits the processed signal to the processor 110. The wireless communication module 160 can also receive a signal to be transmitted from the processor 110, modulate it, amplify it, and radiate it as an electromagnetic wave via the antenna 2.
[0377] In some embodiments, the antenna 1 and the mobile communication module 150 of the electronic device 100 are coupled, and the antenna 2 and the wireless communication module 160 are coupled, so that the electronic device 100 can communicate with a network and other devices through wireless communication technology. The wireless communication technology can include, but is not limited to, a 5th-Generation (5G) system, a global system for mobile communications (GSM), a general packet radio service (GPRS), a code division multiple access (CDMA), a wideband code division multiple access (WCDMA), a time-division code division multiple access (TD-SCDMA), a long term evolution (LTE), a bluetooth, a global navigation satellite system (GNSS), a wireless fidelity (WiFi), a near field communication (NFC), an FM (also referred to as a frequency modulation broadcast), a Zigbee, a radio frequency identification (RFID), and / or an infrared (IR) technology, etc. The GNSS can include a global positioning system (GPS), a global navigation satellite system (GLONASS), a beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS), a satellite based augmentation systems (SBAS), etc.
[0378] In some embodiments, the electronic device 100 can also include a wired communication module (not shown in FIG. 1), or the mobile communication module 150 or the wireless communication module 160 can be replaced by the wired communication module (not shown in FIG. 1) in some embodiments. Figure 1B In some embodiments, the electronic device 100 can also include a wired communication module (not shown in FIG. 1), or the mobile communication module 150 or the wireless communication module 160 can be replaced by the wired communication module (not shown in FIG. 1) in some embodiments. Figure 1B The wired communication module can enable the electronic device to communicate with other devices through a wired network (not shown in FIG. 1). The wired network can include, but is not limited to, one or more of an optical transport network (OTN), a synchronous digital hierarchy (SDH), a passive optical network (PON), Ethernet, or flex Ethernet (FlexE), etc.
[0379] The electronic device 100 implements a display function through a GPU, a display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.
[0380] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 can include 1 or N display screens 194, N being a positive integer greater than 1.
[0381] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor, etc.
[0382] ISP is used to process the data feedback from the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and conversion into a visible image. ISP can also optimize the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, ISP can be provided in the camera 193.
[0383] The camera 193 is used to capture still images or videos. Objects generate optical images through lenses and project them onto photosensitive elements. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard RGB camera (or RGB sensor) 0, YUV, or other format image signal. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.
[0384] The digital signal processor is used to process digital signals, in addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0385] The video codec is used to compress or decompress digital video. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0386] The NPU is a neural-network (NN) computing processor that learns from the structure of biological neural networks, such as the transmission mode between human brain neurons, and can quickly process input information and continuously self-learn. Through the NPU, the electronic device 100 can achieve intelligent cognition applications such as image recognition, face recognition, voice recognition, and text understanding.
[0387] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to extend the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, files such as music and videos are stored in the external memory card.
[0388] The internal memory 121 can be used to store computer executable program codes including instructions. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required for a function (such as a sound play function, an image play function, etc.), and the like. The data storage area can store data created during use of the electronic device 100 (such as audio data, a phonebook, etc.), and the like. In addition, the internal memory 121 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one of a magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like. The processor 110 executes various function applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory disposed in the processor.
[0389] The electronic device 100 can implement an audio function through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, an application processor, and the like. For example, music play, recording, and the like.
[0390] The audio module 170 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode an audio signal. In some embodiments, the audio module 170 can be disposed in the processor 110, or part of the function modules of the audio module 170 can be disposed in the processor 110.
[0391] The speaker 170A, also referred to as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0392] The receiver 170B, also referred to as a "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the receiver 170B can be used to listen to the voice by being close to a human ear.
[0393] Microphone 170C, also called "microphone", "transducer", is used to convert sound signal into electric signal. When making a call or sending voice message, user can make sound by approaching microphone 170C with mouth, inputting sound signal into microphone 170C. Electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, electronic device 100 can be provided with two microphones 170C, which can realize noise reduction function in addition to sound signal collection. In some other embodiments, electronic device 100 can be provided with three, four or more microphones 170C, which can realize sound signal collection, noise reduction, sound source identification, directional recording and other functions.
[0394] Earphone interface 170D is used to connect wired earphone. Earphone interface 170D can be USB interface 130, or 3.5mm open mobile terminal platform (OMTP) standard interface, or cellular telecommunications industry association of the USA (CTIA) standard interface.
[0395] Pressure sensor 180A is used to sense pressure signal, and can convert pressure signal into electric signal. In some embodiments, pressure sensor 180A can be provided on display screen 194. There are many types of pressure sensor 180A, such as resistance type pressure sensor, inductance type pressure sensor, capacitance type pressure sensor, etc. Capacitance type pressure sensor can include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between electrodes changes. Electronic device 100 determines the intensity of pressure according to the change of capacitance. When touch operation is applied to display screen 194, electronic device 100 detects the intensity of touch operation according to pressure sensor 180A. Electronic device 100 can also calculate the position of touch according to the detection signal of pressure sensor 180A. In some embodiments, touch operation applied to the same touch position but with different touch operation intensity can correspond to different operation instructions. For example, when touch operation with touch operation intensity less than first pressure threshold is applied to short message application icon, the instruction of viewing short message is executed. When touch operation with touch operation intensity greater than or equal to first pressure threshold is applied to short message application icon, the instruction of creating new short message is executed.
[0396] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake photography. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of shaking of the electronic device 100, calculates the distance that the lens module needs to compensate according to the angle, and lets the lens offset the shaking of the electronic device 100 by reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and motion sensing game scenarios.
[0397] The barometer sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude, assists positioning and navigation by the air pressure value measured by the barometer sensor 180C.
[0398] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can detect the opening and closing of a flip cover with the magnetic sensor 180D. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover according to the magnetic sensor 180D. In turn, according to the detected opening and closing state of the cover or the opening and closing state of the flip cover, the electronic device 100 can set features such as automatic unlocking of the flip cover.
[0399] The acceleration sensor 180E can detect the magnitude of acceleration of the electronic device 100 in various directions (typically three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device, applied to landscape / portrait switching, pedometer, etc.
[0400] The distance sensor 180F is used to measure distance. The electronic device 100 can measure distance by infrared or laser. In some embodiments, in a shooting scenario, the electronic device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.
[0401] The proximity light sensor 180G can include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode can be an infrared light-emitting diode. The electronic device 100 emits infrared light outwardly through the light-emitting diode. The electronic device 100 detects infrared reflected light from nearby objects using the photodiode. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 can use the proximity light sensor 180G to detect that the user is holding the electronic device 100 close to the ear for a call, so as to automatically turn off the screen to achieve the purpose of power saving. The proximity light sensor 180G can also be used for automatic unlocking and locking of the cover mode and pocket mode.
[0402] Ambient light sensor 180L is used to sense ambient light brightness. Electronic device 100 can adaptively adjust display screen 194 brightness according to sensed ambient light brightness. Ambient light sensor 180L can also be used to automatically adjust white balance when taking a picture. Ambient light sensor 180L can also cooperate with proximity light sensor 180G to detect whether electronic device 100 is in a pocket to prevent accidental touch.
[0403] Fingerprint sensor 180H is used to collect a fingerprint. Electronic device 100 can use collected fingerprint characteristics to implement fingerprint unlocking, access application lock, take a picture with a fingerprint, answer an incoming call with a fingerprint, and so on.
[0404] Temperature sensor 180J is used to detect temperature. In some embodiments, electronic device 100 uses temperature detected by temperature sensor 180J to implement a temperature processing strategy. For example, when temperature reported by temperature sensor 180J exceeds a threshold value, electronic device 100 reduces performance of a processor located near temperature sensor 180J to reduce power consumption and implement thermal protection. In another embodiment, when temperature is lower than another threshold value, electronic device 100 heats battery 142 to avoid abnormal shutdown of electronic device 100 caused by low temperature. In other embodiments, when temperature is lower than yet another threshold value, electronic device 100 boosts output voltage of battery 142 to avoid abnormal shutdown caused by low temperature.
[0405] Touch sensor 180K, also referred to as a "touch device". Touch sensor 180K can be disposed on display screen 194, and touch sensor 180K and display screen 194 together form a touch screen, also referred to as a "touch panel". Touch sensor 180K is used to detect a touch operation acting on or near it. Touch sensor 180K can pass detected touch operation to an application processor to determine a touch event type. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K can also be disposed on a surface of electronic device 100, which is different from the location of display screen 194.
[0406] Bone conduction sensor 180M can obtain a vibration signal. In some embodiments, bone conduction sensor 180M can obtain a vibration signal of a human body sound part vibration bone block. Bone conduction sensor 180M can also contact a human body pulse to receive a blood pressure pulsation signal. In some embodiments, bone conduction sensor 180M can also be disposed in a headset to form a bone conduction headset. Audio module 170 can analyze a voice signal based on the vibration signal of the sound part vibration bone block obtained by bone conduction sensor 180M to implement a voice function. Application processor can analyze heart rate information based on the blood pressure pulsation signal obtained by bone conduction sensor 180M to implement a heart rate detection function.
[0407] The image sensor 180N, also known as a photosensitive device or a photosensitive element, is a device that converts an optical image into an electronic signal, which is widely used in digital cameras and other electronic optical devices. The image sensor uses the photoelectric conversion function of a photoelectric device to convert the optical image on the photosensitive surface into an electrical signal in a corresponding proportional relationship. Compared with point light-sensitive elements such as photosensitive diodes and photosensitive triodes, the image sensor divides the optical image on its light-receiving surface into many small units (i.e., pixels) and converts them into usable electrical signals. Each small unit corresponds to a photosensitive unit in the image sensor, which can also be referred to as a sensor pixel. Image sensors are divided into photoconductive image tubes and solid-state image sensors. Compared with photoconductive image tubes, solid-state image sensors have the characteristics of small size, light weight, high integration, high resolution, low power consumption, long service life, and low price. According to the different elements, they can be divided into two categories: charge-coupled devices (CCD) and complementary metal-oxide semiconductor elements (CMOS). According to the different types of optical images captured, they can be divided into two categories: color sensors 1801N and motion sensors 1802N.
[0408] Specifically, the color sensor 1801N, including the traditional RGB image sensor, can be used to detect objects within the range of the camera. Each photosensitive unit corresponds to a pixel in the image sensor. Since the photosensitive unit can only sense the intensity of light and cannot capture color information, a color filter must be overlaid on the photosensitive unit. As for how to overlay the color filter, different sensor manufacturers have different solutions. The most common method is to overlay RGB red, green, and blue color filters, with a composition of 1:2:1, to form a color pixel from four pixels (i.e., one pixel is covered by a red filter, one pixel is covered by a blue filter, and the remaining two pixels are covered by a green filter). The reason for this ratio is that the human eye is more sensitive to green. After receiving light, the photosensitive unit generates a corresponding current, and the current size corresponds to the light intensity. Therefore, the electrical signal output directly by the photosensitive unit is analog. The output analog electrical signal is then converted into a digital signal, and all the digital signals are finally output in the form of a digital image matrix to a dedicated DSP processing chip for processing. The traditional color sensor outputs the full-frame image of the shooting area in frame format.
[0409] Specifically, the motion sensor 1802N can include a plurality of different types of vision sensors, such as a motion detection vision sensor (MDVS) and an event-based motion detection vision sensor. The motion sensor 1802N can be used to detect a moving object within a range of the camera, capture a motion profile or a motion trajectory of the moving object, and the like.
[0410] In one possible scenario, the motion sensor 1802N can include a motion detection (MD) vision sensor, which is a type of vision sensor that detects motion information resulting from relative motion between the camera and a target. The relative motion can be motion of the camera, motion of the target, or motion of both the camera and the target. The motion detection vision sensor includes frame-based motion detection and event-based motion detection. The frame-based motion detection vision sensor requires exposure integration and obtains motion information through frame difference. The event-based motion detection vision sensor does not require integration and obtains motion information through asynchronous event detection.
[0411] In one possible scenario, the motion sensor 1802N can include a motion detection vision sensor (MDVS), a dynamic vision sensor (DVS), an active pixel sensor (APS), an infrared sensor, a laser sensor, or an inertial measurement unit (IMU), and the like. The DVS can specifically include a DAVIS (Dynamic and Active-pixel Vision Sensor), an ATIS (Asynchronous Time-based Image Sensor), or a CeleX sensor, and the like. The DVS is inspired by the characteristics of biological vision, and each pixel simulates a neuron and independently responds to a relative change in light intensity (hereinafter referred to as “light intensity”). For example, when the motion sensor is a DVS, when the relative change in light intensity exceeds a threshold, the pixel outputs an event signal including the position of the pixel, a timestamp, and characteristic information of the light intensity. It should be understood that in the following embodiments of the present application, the motion information, dynamic data, or dynamic image, and the like mentioned can be acquired by the motion sensor.
[0412] For example, the motion sensor 1802N can include an inertial measurement unit (IMU), which is a device that measures the angular rate and acceleration of an object in three axes. The IMU is usually composed of three single-axis accelerometers and three single-axis gyroscopes, which measure the acceleration signal of the object and the angular velocity signal relative to the navigation coordinate system, respectively, and calculate the attitude of the object based on the signals. For example, the aforementioned IMU can specifically include the aforementioned gyroscope sensor 180B and the acceleration sensor 180E. The advantage of the IMU is high acquisition frequency. The data acquisition frequency of the IMU can generally reach more than 100 HZ, and the consumer-level IMU can capture data up to 1600 HZ. In a short period of time, the IMU can give a high-precision measurement result.
[0413] For example, the motion sensor 1802N can include an active pixel sensor (APS). For example, RGB images are captured at a high frequency of >100 HZ, and the difference between two adjacent images is obtained. If the change value is greater than a threshold value, such as >0, it is set to 1, and if it is not greater than the threshold value, such as =0, it is set to 0. The final data is similar to the data obtained by the DVS, and the image capture of the moving object is completed.
[0414] The key 190 includes a power-on key, a volume key, and the like. The key 190 can be a mechanical key. It can also be a touch key. The electronic device 100 can receive a key input and generate a key signal input related to user settings and function control of the electronic device 100.
[0415] The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompts and also for touch vibration feedback. For example, touch operations on different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. The motor 191 can also correspond to different vibration feedback effects for touch operations on different regions of the display screen 194. Different application scenarios (such as time reminders, received messages, alarms, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0416] The indicator 192 can be an indicator light, which can be used to indicate the charging state, the power change, and also to indicate messages, missed calls, notifications, and the like.
[0417] The SIM card interface 195 is configured to connect a SIM card. The SIM card can be inserted into or pulled out of the SIM card interface 195 to realize contact and separation with the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support a Nano SIM card, a Micro SIM card, a SIM card, and the like. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external storage cards. The electronic device 100 interacts with a network through the SIM card to realize functions such as call and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0418] II. System architecture
[0419] In the process of implementing image shooting, reading or saving, the electronic device involves changes among multiple components. The present application will be described in detail from data acquisition, data coding and decoding, image enhancement, image reconstruction or application scenarios.
[0420] Exemplarily, the processing flow of the electronic device is exemplarily described in the scenario of image acquisition and processing, as shown in Figure 2
[0421] Data acquisition: data can be acquired through a brain-like camera, an RGB camera or a combination thereof. The brain-like camera can include a living visual sensor, which simulates a biological retina by using an integrated circuit, and each pixel simulates a biological neuron to express the change of light intensity in the form of an event. After development, various types of bionic visual sensors have appeared, and their common feature is that the pixel array independently and asynchronously monitors the change of light intensity and outputs the change as an event signal, such as the aforementioned motion sensor DVS or DAVIS. The RGB camera converts an analog signal into a digital signal, and then stores it in a storage medium. Data can also be acquired through a combination of a brain-like camera and an RGB camera, for example, the data acquired by the brain-like camera and the RGB camera can be projected into the same canvas, the value of each pixel point can be determined based on the values fed back by the brain-like camera and / or the RGB camera, or the value of each pixel point can include the values of the brain-like camera and the RGB as independent channels. Through the brain-like camera, the RGB camera or their combination, the optical signal can be converted into an electrical signal, so as to obtain a data stream in frames or an event stream in events. In the following, the image acquired by the RGB camera is referred to as an RGB image, and the image acquired by the brain-like camera is referred to as an event image.
[0422] Data coding and decoding: including data coding and data decoding. Data coding can include coding the collected data after data collection, and saving the coded data to the storage medium. Data decoding can include reading data from the storage medium and decoding the data, decoding the data into data that can be used for subsequent identification, detection, etc. In addition, the way of data coding and decoding can also be used to adjust the way of data collection, so as to realize more efficient data collection and data coding and decoding. Data coding and decoding can be divided into many kinds, including coding and decoding based on brain-like camera, coding and decoding based on brain-like camera and RGB camera, or coding and decoding based on RGB camera, etc. Specifically, in the process of coding, the data collected by the brain-like camera, the RGB camera or their combination can be coded to be stored in the storage medium in a certain format, and in the process of decoding, the data stored in the storage medium can be decoded into data that can be used subsequently. For example, the user can collect video or image data through the brain-like camera, the RGB camera or their combination on the first day, and code the video or image data and store it in the storage medium. The data can be read from the storage medium on the second day and decoded to obtain playable video or image.
[0423] Image optimization: that is, after the aforementioned brain-like camera or RGB camera collects the image, the collected image is read, and then the collected image is optimized by enhancement or reconstruction, etc. so as to facilitate subsequent processing based on the optimized image. Exemplarily, image enhancement and reconstruction can include image reconstruction or motion compensation, etc. Motion compensation, for example, the motion parameters of the moving object collected by the DVS, compensates the moving object in the event image or the RGB image, so that the obtained event image or RGB image is clearer. Image reconstruction, for example, the image collected by the brain-like vision camera, reconstructs the RGB image, so that even in a moving scene, a clear RGB image can be obtained from the data collected by the DVS.
[0424] Application scenario: after the optimized RGB image or event image is obtained through image optimization, the optimized RGB image or event image can be used for further application, of course, the collected RGB image or event image can also be used for further application, which can be adjusted according to the actual application scenario.
[0425] Specifically, the application scenarios can include motion photography enhancement, DVS image and RGB image fusion, detection and recognition, simultaneous localization and mapping (SLAM), eye tracking, key frame selection, or pose estimation, etc. For example, the motion photography enhancement is to enhance the captured image in the scene with moving objects, so as to capture clearer moving objects. The DVS image and RGB image fusion is to enhance the RGB image based on the moving objects captured by the DVS, and compensate the objects with motion or affected by large light ratio in the RGB image, so as to obtain a clearer RGB image. The detection and recognition is to perform target detection or target recognition based on the RGB image or event image. The eye tracking is to track the eye movement of the user according to the captured RGB image, event image, or optimized RGB image and event image, and determine the gaze point and gaze direction of the user. The key frame selection is to select some frames as key frames from the video data captured by the RGB camera in combination with the information captured by the brain-like camera.
[0426] In addition, in the following embodiments of the present application, different sensors can be activated in different embodiments. For example, when collecting data, the motion sensor can be activated, and optionally the IMU or gyroscope, etc. can also be activated when the event image is optimized by motion compensation. In the image reconstruction embodiment, the motion sensor can be activated to collect the event image, and then the event image is optimized in combination. Or, in the motion photography enhancement embodiment, the motion sensor and RGB sensor, etc. can be activated, so that in different embodiments, the corresponding sensors can be selected to be activated.
[0427] Specifically, the method provided in the present application can be applied to an electronic device, which can include an RGB sensor and a motion sensor, etc. The RGB sensor is used to collect images in the shooting range, and the motion sensor is used to collect information generated when the object moves relative to the motion sensor within the detection range of the motion sensor. The method comprises: selecting at least one from the RGB sensor and the motion sensor based on scene information, and collecting data through the selected sensor, wherein the scene information includes at least one of state information of the electronic device, type of an application program in the electronic device requesting to collect images, or environment information.
[0428] In a possible embodiment, the aforementioned state information includes information such as remaining power of the electronic device, remaining storage (or available storage) or CPU load, etc.
[0429] In a possible implementation, the aforementioned environmental information can include a change value of illumination intensity within a shooting range of a color RGB sensor and a motion sensor or information of a moving object within the shooting range. For example, the environmental information can include a change of illumination intensity within a shooting range of an RGB sensor or a DVS sensor or a motion of an object within the shooting range, such as a motion speed, a motion direction, or the like of the object, or an abnormal motion of the object within the shooting range, such as a sudden change of speed or direction of the object.
[0430] The type of the application program requesting to collect images in the aforementioned electronic device can be understood as that the electronic device carries a system such as Android, Linux, or Harmony, an application program can be run in the system, and the program running in the system can be divided into multiple types, such as a photographing type application program or an object detection application program.
[0431] Generally, a motion sensor is sensitive to motion changes and insensitive to static scenes, and responds to motion changes by issuing events, and since a static area almost does not issue events, data thereof only expresses light intensity information of a motion change area and is not complete full-scene light intensity information. An RGB color camera is good at complete color recording of a natural scene and recording texture details in the scene.
[0432] Taking a mobile phone as an example of the aforementioned electronic device, a default configuration is that a DVS camera (that is, a DVS sensor) is turned off. When the camera is used, according to a currently invoked application type, such as a photographing APP, if the camera is invoked, and if it is in a high-speed motion state, the DVS camera and the RGB camera (that is, an RGB sensor) need to be turned on at the same time; if the APP requesting to invoke the camera is an APP for object detection or motion detection, and object photographing and face recognition are not needed, the DVS camera can be turned on and the RGB camera can be turned off.
[0433] Optionally, a camera starting mode can also be selected according to a current device condition, for example, when a current power is lower than a certain threshold, a user starts a power saving mode and cannot normally take a photo, and only the DVS camera can be turned on, because although the DVS camera does not clearly image a photo, power consumption is low, and when a moving object is detected, high-definition imaging is not needed.
[0434] Optionally, the device can perceive a surrounding environment to determine whether to switch a camera mode, for example, when it is a night scene or the current device is in a high-speed motion, the DVS camera can be turned on. If it is a static scene, the DVS camera can not be turned on.
[0435] Based on the above application types, environmental information, and device status, the camera starting mode is determined, and during operation, it can be decided whether to trigger the camera mode switching, so that different sensors are started in different scenes, and the adaptability is strong.
[0436] It can be understood that the starting mode includes three types, only the RGB camera is started, only the DVS camera is started, and the RGB and DVS cameras are started at the same time. Moreover, for different products, the detection application type and the reference factor of environmental detection can be different.
[0437] For example, the camera in the security scene has a motion detection function, and the camera only stores the video when a moving object is detected, thereby reducing the storage space and prolonging the hard disk storage time. Specifically, when the DVS and RGB cameras are applied to a home or security camera, by default, only the DVS camera is started for motion detection and analysis. When the DVS camera detects abnormal motion or abnormal behavior (such as sudden movement of an object or sudden change of motion direction), for example, a person approaches or there is a significant change in light intensity, the RGB camera is started to take pictures, and the full-scene texture image of this period is recorded as a monitoring voucher. When the abnormal motion ends, the DVS is switched to work, and the RGB camera is in standby state, thereby significantly saving data volume and monitoring device power consumption.
[0438] The above intermittent camera method takes advantage of the low power consumption of DVS, and DVS is faster based on event motion detection, has a faster response than image-based motion detection, and detects more accurately. It realizes all-weather uninterrupted detection. It realizes a more accurate, low-power, and storage space-saving method.
[0439] For another example, when the DVS and RGB cameras are applied to vehicle auxiliary / automatic driving, during driving, when encountering oncoming vehicles with high beams, or sunset direct radiation, or entering and exiting tunnels, the RGB camera may not be able to take effective scene information. At this time, although the DVS cannot obtain texture information, it can obtain the general outline information of the scene, which has great auxiliary value for the judgment of the driver. In addition, in heavy fog weather, the outline information taken by the DVS can also assist in judging the road conditions. Therefore, the master-slave working state switching of the DVS and the RGB camera can be triggered in specific scenes, such as triggering when the light intensity changes dramatically or triggering in extreme weather.
[0440] When the DVS camera is applied to AR / VR glasses, the above process is also applicable. When the DVS is used for SLAM or eye tracking, the starting mode of the camera can be determined according to the device status and the surrounding environment.
[0441] In the following embodiments of the present application, when the data collected by a certain sensor is used, i.e. the sensor is turned on, the following will not be repeated.
[0442] The above different working modes and the foregoing Figure 2 The different embodiments provided by the present application are described.
[0443] III. Method flow
[0444] The foregoing electronic device and system architecture provided by the present application are exemplarily described, and the following Figure 1A-Figure 2 , the method provided by the present application is described in detail. Specifically, in combination with the foregoing Figure 2 architecture, the method corresponding to each module is described respectively. It should be understood that the method steps mentioned in the following embodiments of the present application can be implemented separately or combined in one device. The specific implementation can be adjusted according to the actual application scene.
[0445] 1. Data acquisition and codec
[0446] The process of data acquisition and data codec is exemplarily described below.
[0447] In the conventional technology, the visual sensor (i.e. the foregoing motion sensor) generally adopts an asynchronous reading mode based on event stream (hereinafter also referred to as "reading mode based on event stream" or "asynchronous reading mode") and a synchronous reading mode based on frame scanning (hereinafter also referred to as "reading mode based on frame scanning" or "synchronous reading mode"). For a visual sensor that has been manufactured, only one of the two modes can be adopted. According to the specific application scene and motion state, the amount of signal data required to be read in unit time by the above two reading modes can have a significant difference, and the cost required for outputting the read data is also different. Figure 3-a and Figure 3-b respectively show the schematic diagram of the relationship between the amount of read data and time in the asynchronous reading mode based on event stream and the synchronous reading mode based on frame scanning.
[0448] In one aspect, since the static region in the environment usually does not generate light intensity change events (also referred to as "events" herein) which are sensitive to motion, the biomimetic vision sensor almost all adopts an asynchronous reading mode based on an event stream, where the event stream refers to events arranged in a certain order. The asynchronous reading mode is exemplarily described below by taking a DVS as an example. According to the sampling principle of the DVS, when the change amount of the current light intensity and the light intensity at the last event occurrence reaches a predetermined threshold C (hereinafter referred to as a predetermined threshold), an event is generated and output. That is, the DVS will generate an event when the difference between the current light intensity and the light intensity at the last event occurrence exceeds the predetermined threshold C, which can be described by formula 1-1:
[0449] |L-L′|≥C(1-1)
[0450] wherein L and L' represent the light intensity at the current moment and the light intensity at the last event occurrence, respectively.
[0451] wherein for the asynchronous reading mode, each event can be represented as <x, y, t, m>, (x, y) represents the pixel position at which the event is generated, t represents the time at which the event is generated, and m represents the characteristic information of the light intensity. Specifically, the pixel in the pixel array circuit of the vision sensor measures the light intensity change amount in the environment. If the measured light intensity change amount exceeds the predetermined threshold, the pixel can output a data signal indicating the event. Therefore, in the asynchronous reading mode based on the event stream, the pixels of the vision sensor are further divided into pixels generating the light intensity change event and pixels not generating the light intensity change event. The light intensity change event can be characterized by the coordinate information (x, y) of the pixel generating the event, the characteristic information of the light intensity at the pixel, and the time t at which the characteristic information of the light intensity is read, etc. The coordinate information (x, y) can be used to uniquely identify the pixel in the pixel array circuit, for example, x represents the row index of the pixel in the pixel array circuit, and y represents the column index of the pixel in the pixel array circuit. By identifying the coordinates and the time stamp associated with the pixel, the spatiotemporal position at which the light intensity change event occurs can be uniquely determined, and then all the events can be arranged in the order of occurrence to form an event stream.
[0452] In some DVS sensors (e.g., DAVIS sensor, ATIS sensor, etc.), m represents the trend of light intensity change, also referred to as polarity information, and is usually represented by 1-bit to 2-bit, and can take ON / OFF values, where ON represents light intensity enhancement, and OFF represents light intensity reduction, i.e., when the light intensity increases and exceeds a predetermined threshold, an ON pulse is generated; when the light intensity decreases and exceeds a predetermined threshold, an OFF pulse is generated (in this application, "+" represents light intensity enhancement, and "-" represents light intensity reduction). In some DVS sensors, such as CeleX sensor, m represents absolute light intensity information, also referred to as light intensity information, and is usually represented by multiple bits, such as 8-bit to 12-bit.
[0453] In the asynchronous reading mode, only the data signal at the pixel where the light intensity change event occurs is read. Thus, for the biomimetic vision sensor, the event data required to be read has the characteristics of sparsity and asynchronicity. As shown in curve 101, when the rate of light intensity change events occurring in the pixel array circuit changes, the amount of data required to be read by the vision sensor also changes over time. Figure 3-a
[0454] On the other hand, conventional vision sensors, such as mobile phone cameras, digital video cameras, etc., usually adopt a synchronous reading mode based on frame scanning. This reading mode does not distinguish whether a light intensity change event occurs at a pixel of the vision sensor or not. Regardless of whether a light intensity change event occurs at a pixel or not, the data signal generated by the pixel is read. In reading the data signal, the vision sensor scans the pixel array circuit in a predetermined order, synchronously reads the feature information m (the feature information m of light intensity has been introduced above and will not be repeated here) indicating light intensity at each pixel, and outputs in order as frame 1 data, frame 2 data, etc. Thus, as shown in curve 102, in the synchronous reading mode, the amount of data read by the vision sensor for each frame has the same size, and the data amount remains unchanged over time. For example, assuming that 8 bits are used to represent the light intensity value of a pixel, and the total number of pixels in the vision sensor is 66, then the data amount of a frame of data is 528 bits. Usually, frames are output at equal time intervals, for example, at a rate of 30 frames per second, 60 frames per second, 120 frames per second, etc. Figure 3-b
[0455] The applicant finds that the current vision sensor still has defects, at least including the following aspects:
[0456] Firstly, a single reading mode cannot adapt to all scenarios, which is not conducive to relieving the pressure of data transmission and storage.
[0457] As Figure 3-a As shown in curve 101, when the visual sensor works in the event stream based asynchronous readout mode, the amount of data required to be read by the visual sensor varies over time as the rate of light intensity change events occurring in the pixel array circuit changes. In a static scene, there are fewer light intensity change events, and thus the total amount of data required to be read by the visual sensor is also lower. In a dynamic scene, such as a fast moving scene, there are a large number of light intensity change events, and thus the total amount of data required to be read by the visual sensor increases. In some cases, the large number of light intensity change events causes the total amount of data to exceed the bandwidth limit, and event loss or delayed readout can occur. Figure 3-b As shown in curve 102, when the visual sensor works in the frame based synchronous readout mode, the state or intensity value of each pixel needs to be represented in a frame regardless of whether the pixel has changed. This representation is costly when only a small number of pixels have changed.
[0458] The output and storage cost of the two modes can vary significantly in different application scenarios and motion states. For example, when a static scene is captured, only a small number of pixels produce light intensity change events in a period of time. By way of example, only three pixels in the pixel array circuit produce light intensity change events in a scan. In the asynchronous readout mode, only the coordinate information (x, y), time information t and light intensity change amount of the three pixels need to be read to represent the three light intensity change events. Assuming that in the asynchronous readout mode, 4, 2, 2 and 2 bits are allocated for the coordinate, read timestamp and light intensity change amount of a pixel, respectively, the total amount of data required to be read in the asynchronous readout mode is 30 bits. In the synchronous readout mode, although only three pixels produce valid data signals indicating light intensity change events, the data signals output by all pixels in the pixel array circuit need to be read to form a complete frame of data. Assuming that in the synchronous readout mode, 8 bits are allocated for each pixel, and the total number of pixels in the pixel array circuit is 66, the total amount of data required to be read is 528 bits. As can be seen, even though there are a large number of pixels in the pixel array circuit that do not produce events, a large number of bits are still allocated in the synchronous readout mode. This is not economical in terms of representation cost, and increases the pressure on data transmission and storage. Thus, in this case, the asynchronous readout mode is more economical.
[0459] In another example, when there is a drastic motion in the scene or a drastic change in the light intensity in the environment, such as a large number of people walking around, or a light being turned on or off suddenly, a large number of pixels in the vision sensor measure the change in light intensity in a short time and generate data signals indicating the light intensity change events. Since the amount of data characterizing a single event in the asynchronous reading mode is larger than the amount of data characterizing a single event in the synchronous reading mode, it can be very costly to use the asynchronous reading mode in this case. Specifically, there can be a number of consecutive pixels in each row of the pixel array circuit generating light intensity change events, and the coordinate information (x, y), the time information t, and the light intensity feature information m need to be transmitted for each event. The coordinate change between these events often has only a unit deviation, and the reading time is also basically the same. In this case, the asynchronous reading mode has a large cost for the representation of the coordinate and time information, which can cause a large increase in the amount of data. In the synchronous reading mode, no matter how many light intensity change events are generated in the pixel array circuit at any time, each pixel outputs a data signal indicating only the amount of change in light intensity, and there is no need to allocate bits for the coordinate information and the time information of each pixel. Thus, for the event dense case, it is more economical to use the synchronous reading mode.
[0460] Secondly, a single event representation method cannot adapt to all scenes, and using light intensity information to represent events is not conducive to relieving the pressure on data transmission and storage, and using polarity information to represent events affects the processing and analysis of events.
[0461] The above introduces the synchronous reading mode and the asynchronous reading mode, and the events read out need to be represented by the light intensity feature information m, which includes the polarity information and the light intensity information. In this paper, the event represented by the polarity information is referred to as the polarity format event, and the event represented by the light intensity information is referred to as the light intensity format event. For a vision sensor that has been manufactured, only one of the two event formats can be used, i.e., the event is represented by the polarity information or the event is represented by the light intensity information. The advantages and disadvantages of the polarity format event and the light intensity format event are described below taking the asynchronous reading mode as an example.
[0462] In the asynchronous reading mode, the polarity information p is usually represented by 1-bit-2-bit, carrying less information, and only indicating whether the light intensity trend is increasing or weakening. Therefore, the event represented by the polarity information affects the processing and analysis of the event. For example, the event represented by the polarity information is difficult to reconstruct the image, and the accuracy for object recognition is poor. When the light intensity information is used to represent the event, it is usually represented by multiple bits, such as 8-bit-12-bit. Compared with the polarity information, the light intensity information can carry more information, which is beneficial to the processing and analysis of the event, such as improving the quality of image reconstruction. However, due to the large amount of data, it takes longer to obtain the event represented by the light intensity information. According to the DVS sampling principle, when the light intensity of the pixel changes by more than a predetermined threshold, an event will be generated. When a large area of object movement or light intensity fluctuation occurs in the scene (for example, entering and exiting a tunnel, turning on and off the light in a room, etc.), the visual sensor will face the problem of a sudden increase in events. In the case of a preset maximum bandwidth (hereinafter referred to as bandwidth) of the visual sensor, there is a situation that the event data cannot be read out. At present, the random discarding method is usually used for processing. If random discarding is used, although the amount of data transmitted can be ensured not to exceed the bandwidth, data loss is caused. In some special application scenarios (such as automatic driving, etc.), the discarded data may have high importance. In other words, when a large number of events are triggered, the data amount exceeds the bandwidth, and the light intensity format data cannot be completely output to the outside of the DVS, resulting in the loss of some events. These lost events may not be conducive to the processing and analysis of the event, such as causing trailing and incomplete contour in the brightness reconstruction process.
[0463] To solve the above problems, the embodiment of the present application provides a visual sensor. Based on the statistical results of the light intensity change events generated by the pixel array circuit, the data amount in the two reading modes is compared, so that the reading mode suitable for the current application scenario and motion state can be switched. In addition, based on the statistical results of the light intensity change events generated by the pixel array circuit, the relationship between the data amount of the event represented by the light intensity information and the bandwidth is compared, so as to adjust the representation precision of the event. Under the premise of meeting the bandwidth limitation, all events are transmitted in a suitable representation manner, and all events are transmitted with greater representation precision as much as possible.
[0464] A visual sensor provided by an embodiment of the present application is introduced below.
[0465] Figure 4-a A block diagram of a visual sensor provided by the present application is shown. The visual sensor can be implemented as a visual sensor chip and can read a data signal indicating an event in at least one of a frame scanning-based reading mode and an event stream-based reading mode. As shown in the figure, the visual sensor includes a pixel array circuit, a data processing circuit, and a data output circuit. Figure 4-aAs shown, the vision sensor 200 includes a pixel array circuit 210 and a readout circuit 220. The vision sensor is coupled with a control circuit 230. It should be understood that Figure 4-a The vision sensor shown is for exemplary purposes only and does not imply any limitation on the scope of the present application. Embodiments of the present application can also be embodied in different sensor architectures. In addition, it should also be understood that the vision sensor can also include other elements or entities for purposes of image acquisition, image processing, image transmission, etc. which are not shown for ease of description but do not mean that embodiments of the present application do not possess these elements or entities.
[0466] The pixel array circuit 210 can include one or more pixel arrays, and each pixel array includes a plurality of pixels, each pixel having location information for unique identification, such as coordinates (x, y). The pixel array circuit 210 can be configured to measure a light intensity variation amount and generate a plurality of data signals corresponding to the plurality of pixels. In some possible embodiments, each pixel is configured to independently respond to a light intensity variation in the environment. In some possible embodiments, the pixel compares the measured light intensity variation amount with a predetermined threshold, and if the measured light intensity variation amount exceeds the predetermined threshold, the pixel generates a first data signal indicating a light intensity variation event, for example, the first data signal includes polarity information, such as +1 or -1, or the first data signal can also be an absolute light intensity information. In this example, the first data signal can indicate a light intensity variation trend or an absolute light intensity value at the corresponding pixel. In some possible embodiments, if the measured light intensity variation amount does not exceed the predetermined threshold, the pixel generates a second data signal different from the first data signal, for example, 0. In embodiments of the present application, the data signal can indicate, including but not limited to, light intensity polarity, absolute light intensity value, light intensity variation value, etc. The light intensity polarity can represent a trend of light intensity variation, for example, enhancement or weakening, usually represented by +1 and -1. The absolute light intensity value can represent a light intensity value measured at the current time. Depending on the structure, purpose and kind of the sensor, there can be different physical meanings about the light intensity or light intensity variation amount. The scope of the present application is not limited in this respect.
[0467] The readout circuit 220 is coupled to and can communicate with the pixel array circuit 210 and the control circuit 230. The readout circuit 220 is configured to read the data signal output by the pixel array circuit 210, which can be understood as the readout circuit 220 reading the data signal output by the pixel array 210 into the control circuit 230. The control circuit 230 is configured to control the mode of reading the data signal by the readout circuit 220, and the control circuit 230 can also be configured to control the representation of the output data signal, in other words, control the representation precision of the data signal, such as the control circuit can control the visual sensor to output the event represented by the polarity information, or the event represented by the light intensity information, or the event represented by a certain fixed number of bits, and the like, which will be described in detail below in conjunction with specific embodiments.
[0468] According to possible embodiments of the present application, the control circuit 230 can be connected to the visual sensor 200 as an independent circuit or chip outside the visual sensor 200 through a bus interface, as shown in Figure 4-a In other possible embodiments, the control circuit 230 can also be integrated with the pixel array circuit and the readout circuit as a circuit or chip inside the visual sensor. Figure 4-b A block diagram of another visual sensor 300 according to possible embodiments of the present application is shown. The visual sensor 300 can be implemented as an example of the visual sensor 200. As Figure 4-b shown, another block diagram of a visual sensor provided by the present application is shown. The visual sensor includes a pixel array circuit 310, a readout circuit 320, and a control circuit 330. The pixel array circuit 310, the readout circuit 320, and the control circuit 330 are the same in function as the pixel array circuit 210, the readout circuit 220, and the control circuit 230 shown in Figure 4-a Therefore, no further description is given here. It should be understood that the visual sensor is for exemplary purposes only, and does not imply any limitation on the scope of the present application. Embodiments of the present application can also be embodied in different visual sensors. In addition, it should also be understood that the visual sensor can also include other elements, modules or entities, which are not shown for the purpose of clarity, but do not mean that embodiments of the present application do not have these elements or entities.
[0469] Based on the foregoing architecture of the visual sensor, the visual sensor provided by the present application is described in detail below.
[0470] The readout circuit 220 can be configured to scan the pixels in the pixel array circuit 210 in a predetermined order to read the data signals generated by the corresponding pixels. In embodiments of the present application, the readout circuit 220 is configured to be capable of reading the data signals output by the pixel array circuit 210 in more than one signal readout mode. For example, the readout circuit 220 can read in one of a first readout mode and a second readout mode. In the context of the present document, the first readout mode and the second readout mode correspond to one of a frame-based readout mode and an event stream-based readout mode, respectively, and further, the first readout mode can refer to the current readout mode of the readout circuit 220, and the second readout mode can refer to an alternative readout mode that can be switched to.
[0471] Reference is made to Figure 5 which shows a schematic diagram illustrating the principle of a frame-based synchronous readout mode and an event stream-based asynchronous readout mode according to embodiments of the present application. As Figure 5 shown in the upper half of FIG. 4, the black dots represent the pixels that generate a light intensity change event, and the white dots represent the pixels that do not generate a light intensity change event. The left dashed box represents the frame-based synchronous readout mode, in which all the pixels generate a voltage signal based on the received light signal, which is then output as a data signal after analog-to-digital conversion. In this mode, the readout circuit 220 constitutes a frame of data by reading the data signals generated by all the pixels. The right dashed box represents the event stream-based asynchronous readout mode, in which the readout circuit 220 acquires the coordinate information (x, y) of the pixel that generates a light intensity change event when the pixel is scanned. Then, only the data signal generated by the pixel that generates a light intensity change event is read, and the readout time t is recorded. In the case where there are multiple pixels that generate a light intensity change event in the pixel array circuit, the readout circuit 220 reads the data signals generated by the multiple pixels in sequence according to the scanning order, and constitutes an event stream as output.
[0472] Figure 5 The lower half of FIG. 4 describes the two readout modes from the perspective of cost (e.g., the amount of data to be read). As Figure 5 shown in the lower half of FIG. 4, in the synchronous readout mode, the amount of data read by the readout circuit 220 is the same each time, for example, 1 frame of data. In Figure 5 , the amount of data to be read is shown as the 1st frame of data 401-1 and the 2nd frame of data 401-2. According to the amount of data (e.g., the number of bits B p ) representing a single pixel and the total number of pixels M in the pixel array circuit, the amount of data to be read for one frame is M·B pIn the asynchronous read mode, the read circuit 220 reads the data signals indicative of the light intensity change events, and then forms an event stream 402 of all the events in the order of their occurrence. In this case, the amount of data read by the read circuit 220 each time is the event data amount B ev (e.g., the sum of the bit number of the coordinate (x, y) of the pixel generating the event, the read time stamp t, and the characteristic information of the light intensity) for representing a single event and the number N ev of the light intensity change events
[0473] In some embodiments, the read circuit 220 can be configured to provide the read at least one data signal to the control circuit 230. For example, the read circuit 220 can provide the data signals read in a period of time to the control circuit 230 for historical data statistics and analysis by the control circuit 230.
[0474] In some possible embodiments, in the case that the currently adopted first read mode is the event stream based read mode, the read circuit 220 reads the data signals generated by the pixels in the pixel array circuit 210 generating the light intensity change events, which are also referred to as the first data signals for convenience of description below. Specifically, the read circuit 220 determines the position information (x, y) of the pixels related to the light intensity change events by scanning the pixel array circuit 210. Based on the position information (x, y) of the pixels, the read circuit 220 reads the first data signals generated by the pixels in the plurality of data signals to obtain the characteristic information of the light intensity and the read time information t indicated by the first data signals. By way of example, in the event stream based read mode, the event data amount read by the read circuit 220 each second can be represented as B ev ·N ev bits, i.e., the read data rate of the read circuit 220 is B ev ·N ev bits per second (bps), where B ev is the event data amount (e.g., the bit number) allocated for each light intensity change event in the event stream based read mode, where the first b x and b y bits are used to represent the pixel coordinate (x, y), the following b t bits are used to represent the time stamp t when the data signal is read, and the last b f bits are used to represent the characteristic information of the light intensity indicated by the data signal, i.e., B ev = b x + b y + b t + b f , N evThe average number of events generated per second can be derived by the reading circuit 220 based on the historical statistics of the number of light intensity change events generated in the pixel array circuit 210 over a period of time. Due to the frame scanning based reading mode, the amount of data read by the reading circuit 220 per frame can be represented as M B p bits, and the amount of data read per second is M B p bits, and the amount of data read per second is M B p bits, and the amount of data read per second is M B p bits, and the amount of data read per second is M B p bits, and the amount of data read per second is M B ev bits, and the amount of data read per second is M B ev bits, and the amount of data read per second is M B ev bits, and the amount of data read per second is M B ev bits, and the amount of data read per second is M B ev bits, and the amount of data read per second is M B ev bits, and the amount of data read per second is M B
[0475] In some possible embodiments, in the case that the first reading mode currently adopted is the frame scanning based reading mode, the average number of events generated per second N ev can be derived by the reading circuit 220 based on the historical statistics of the number of light intensity change events generated in the pixel array circuit 210 over a period of time. N ev obtained according to the frame scanning based reading mode, the amount of event data read by the reading circuit 220 per second in the event stream based reading mode can be calculated as B ev N ev bits, and the reading data rate of the reading circuit 220 in the event stream based reading mode is B ev N ev bps.
[0476] As can be seen from the above two embodiments, the reading data rate of the reading circuit 220 in the frame scanning based reading mode can be directly calculated according to the predefined parameters, and the reading data rate of the reading circuit 220 in the event stream based reading mode can be calculated according to N ev obtained according to either of the two modes.
[0477] Control circuit 230 is coupled to read circuit 220 and configured to control read circuit 220 to read data signals generated by pixel array circuit 210 in a specific read mode. In some possible embodiments, control circuit 230 may acquire at least one data signal from read circuit 220 and, based at least on the at least one data signal, determine which of the current read mode and alternative read modes is more suitable for the current application scenario and motion state. Furthermore, in some embodiments, control circuit 230 may, based on this determination, instruct read circuit 220 to switch from the current data read mode to another data read mode.
[0478] In some possible embodiments, control circuit 230 may send an instruction to readout circuit 220 regarding switching readout modes based on historical statistics of light intensity change events. For example, control circuit 230 may determine statistical data related to at least one light intensity change event based on at least one data signal received from readout circuit 220. If the statistical data is determined to meet predetermined switching conditions, control circuit 230 sends a mode switching signal to readout circuit 220 to switch readout circuit 220 to a second readout mode. For ease of comparison, the statistical data may be used to measure the readout data rate of the first readout mode and the second readout mode, respectively.
[0479] In some embodiments, the statistical data may include the total amount of data representing the number of events measured by the pixel array circuit 210 per unit time. If the total amount of data representing light intensity change events read by the readout circuit 220 in the first readout mode is greater than or equal to the total amount of data representing light intensity change events in the second readout mode, then the readout circuit 220 should switch from the first readout mode to the second readout mode. In some embodiments, the first readout mode is a frame-scan-based readout mode and the second readout mode is an event-stream-based readout mode. The control circuit 230 may be based on the number of pixels M, frame rate f, and pixel data amount B of the pixel array circuit. p To determine the total data volume M·B of light intensity change events read in the first reading mode. p • f. The control circuit 230 can be based on the number N of light intensity change events. ev And the amount of event data B associated with the event stream-based reading pattern ev To determine the total amount of data B for light intensity change events. ev ·N ev That is, the total amount of data B of light intensity change events read in the second reading mode. ev ·N ev In some embodiments, a switching parameter can be used to adjust the relationship between the total data volume in the two reading modes, as shown in the following formula (1), where the total data volume M·B of the light intensity change event read in the first reading mode is... p• f is greater than or equal to a total data amount B of the light intensity variation events in the second read mode ev • N ev , the read circuit 220 should switch to the second read mode:
[0480] η·M·B P • f > B ev • N ev (1)
[0481] where η is a switching parameter for adjustment. It can be further derived from the above formula (1) that a first threshold data amount d1 = M·B p • f·η. That is, if the total data amount B ev • N ev of the light intensity variation events in the first read mode is less than or equal to the threshold data amount d1, it indicates that the total data amount of the light intensity variation events read in the first read mode has been greater than or equal to the total data amount of the light intensity variation events in the second read mode, and the control circuit 230 can determine that the statistical data of the light intensity variation events meet the predetermined switching condition. In this embodiment, the threshold data amount d1 can be determined based on at least the number of pixels M of the pixel array circuit, the frame rate f associated with the read mode based on frame scanning, and the pixel data amount B p .
[0482] As an alternative implementation of the above embodiment, the total data amount M·B p • f of the light intensity variation events in the first read mode is greater than or equal to a total data amount B ev • N ev of the light intensity variation events in the second read mode, the read circuit 220 should switch to the second read mode:
[0483] M·B P • f - B ev • N ev ≥ θ (2)
[0484] where θ is a switching parameter for adjustment. It can be further derived from the above formula (2) that a second threshold data amount
[0485] d2 = M·B p • f - θ
[0486] That is, if the total data amount B ev • N evless than or equal to a second threshold data amount d2, it indicates that the total data amount of the light intensity change events read in the first read mode has been greater than or equal to the total data amount of the light intensity change events in the second read mode, and the control circuit 230 can determine that the statistical data of the light intensity change events satisfy a predetermined switching condition. In this embodiment, the threshold data amount d2 can be determined based on at least the number of pixels M of the pixel array circuit, the frame rate f associated with the read mode based on frame scanning, and the pixel data amount B p .
[0487] In some embodiments, the first read mode is an event stream based read mode and the second read mode is a frame scanning based read mode. Since in the event stream based read mode, the read circuit 220 only reads the data signals generated by the pixels that generate events. Thus, the control circuit 230 can directly determine the number N ev of the light intensity change events generated in the pixel array circuit 210 based on the number of data signals provided by the read circuit 220. ev The control circuit 230 can determine the total data amount of the light intensity change events, i.e., the total data amount B ev · N ev of the events read in the first read mode, based on the number N ev of the events and the event data amount B ev associated with the event stream based read mode. Similarly, the control circuit 230 can also determine the total data amount M·B p · f of the light intensity change events read in the second read mode based on the number of pixels M of the pixel array circuit, the frame rate f, and the pixel data amount B p . As shown in the following equation (3), the total data amount B ev · N ev of the light intensity change events read in the first read mode is greater than or equal to the total data amount M·B p · f of the light intensity change events in the second read mode, the read circuit 220 should switch to the second read mode:
[0488] B ev · N ev ≥ η· M· B P · f (3)
[0489] where η is a switching parameter for adjustment. It can be further derived from the above equation (3) that the first threshold data amount d1 = η· M· B P · f. If the total data amount B ev · N ev of the light intensity change events is greater than or equal to the threshold data amount d1, the control circuit 230 determines that the statistical data of the light intensity change events satisfy a predetermined switching condition. In this embodiment, the threshold data amount d1 can be determined based on at least the number of pixels M of the pixel array circuit, the frame rate f, and the pixel data amount B pto determine the threshold data amount d1.
[0490] As an alternative implementation of the above-described embodiment, the total data amount B of the light intensity change events read in the first read mode ev ·N ev is greater than or equal to the total data amount M-B of the light intensity change events in the second read mode p ·f can be determined as shown in the following equation (4):
[0491] M-B P ·f-B ev ·N ev ≤ θ (4)
[0492] where θ is a switching parameter for adjustment. It can be further derived from the above equation (4) that the second threshold data amount d2 = M-B P ·f-θ if the total data amount B of the light intensity change events ev ·N ev is greater than or equal to the threshold data amount d2, the control circuit 230 determines that the statistical data of the light intensity change events satisfy the predetermined switching condition. In this embodiment, the threshold data amount d2 can be determined based at least on the pixel number M of the pixel array circuit, the frame rate f and the pixel data amount B p .
[0493] In some other embodiments, the statistical data can include the number N of events measured by the pixel array circuit 210 per unit time ev . If the first read mode is a frame scan-based read mode and the second read mode is an event stream-based read mode, the control circuit 230 determines the number N of light intensity change events based on the number of the first data signals among the plurality of data signals provided by the read circuit 220 ev . If the statistical data indicates that the number N of light intensity change events ev is less than a first threshold number n1, the control circuit 230 determines that the statistical data of the light intensity change events satisfy the predetermined switching condition can be determined based at least on the pixel number M of the pixel array circuit, the frame rate f associated with the frame scan-based read mode and the pixel data amount B p , and the event data amount B ev associated with the event stream-based read mode. For example, in the aforementioned embodiment, based on equation (1) it can be further derived as the following equation (5):
[0494]
[0495] i.e., the first threshold number n1 can be determined as
[0496] As an alternative embodiment of the above embodiments, based on formula (2), the following formula (6) can be further obtained:
[0497]
[0498] Accordingly, the number of second thresholds n2 can be determined as
[0499] In some other embodiments, if the first reading mode is an event stream-based reading mode and the second reading mode is a frame scan-based reading mode, the control circuit 230 can directly determine the number N of light intensity change events based on the number of at least one data signal provided by the reading circuit 220. ev If the statistical data indicates the number N of light intensity change events... ev If the number of light intensity changes is greater than or equal to the first threshold number n1, then the control circuit 230 determines that the statistical data of the light intensity change event meets the predetermined switching conditions. This can be based at least on the number of pixels M of the pixel array circuit 210, the frame rate f associated with the frame scan-based readout mode, and the pixel data volume B. p And the amount of event data B associated with the event stream-based reading pattern. ev Determine the number of the first threshold n1 = M·B p ·f / (η·B ev For example, in the foregoing embodiments, based on formula (3), the following formula (7) can be further obtained:
[0500]
[0501] That is, the number of the first thresholds n1 can be determined as
[0502] As an alternative embodiment of the above embodiments, based on formula (4), the following formula (8) can be further obtained:
[0503]
[0504] Accordingly, the number of second thresholds n2 can be determined as
[0505] It should be understood that the formulas, switching conditions and related calculation methods given above are merely an example implementation of the embodiments of this application. Other suitable mode switching conditions, switching strategies and calculation methods may also be adopted, and the scope of this application is not limited in this respect.
[0506] Figure 6-a A schematic diagram of a vision sensor according to an embodiment of the present application operating in a frame-scan-based readout mode is shown. Figure 6-bA schematic diagram showing a visual sensor operating in a frame scan based read mode according to an embodiment of the application is shown. As shown, Figure 6-a The read circuit 220 or 320 is currently operating in a first read mode, i.e. a frame scan based read mode. Since the control circuit 230 or 330 determines that the number of events generated in the current pixel array circuit 210 or 310 is small, e.g. only four valid data in a frame of data, it then predicts that the rate of event generation in the next time period is likely to be low. If the read circuit 220 or 320 continues to read in the frame scan based read mode, it will need to repeatedly allocate bits for the pixels that generate events, resulting in a large amount of redundant data. In this case, the control circuit 230 or 330 sends a mode switching signal to the read circuit 220 or 320 to cause the read circuit 220 or 320 to switch from the first read mode to a second read mode. After the switch, Figure 6-b The read circuit 220 or 320 is currently operating in a first read mode, i.e. a frame scan based read mode. Since the control circuit 230 or 330 determines that the number of events generated in the current pixel array circuit 210 or 310 is small, e.g. only four valid data in a frame of data, it then predicts that the rate of event generation in the next time period is likely to be low. If the read circuit 220 or 320 continues to read in the frame scan based read mode, it will need to repeatedly allocate bits for the pixels that generate events, resulting in a large amount of redundant data. In this case, the control circuit 230 or 330 sends a mode switching signal to the read circuit 220 or 320 to cause the read circuit 220 or 320 to switch from the first read mode to a second read mode. After the switch,
[0507] Figure 6-c A schematic diagram showing a visual sensor operating in a frame scan based read mode according to an embodiment of the application is shown. As shown, Figure 6-d A schematic diagram showing a visual sensor operating in a frame scan based read mode according to an embodiment of the application is shown. As shown, Figure 6-c The read circuit 220 or 320 is currently operating in a first read mode, i.e. a frame scan based read mode. Since the control circuit 230 or 330 determines that the number of events generated in the current pixel array circuit 210 or 310 is small, e.g. only four valid data in a frame of data, it then predicts that the rate of event generation in the next time period is likely to be low. If the read circuit 220 or 320 continues to read in the frame scan based read mode, it will need to repeatedly allocate bits for the pixels that generate events, resulting in a large amount of redundant data. In this case, the control circuit 230 or 330 sends a mode switching signal to the read circuit 220 or 320 to cause the read circuit 220 or 320 to switch from the first read mode to a second read mode. After the switch, Figure 6-d The read circuit 220 or 320 is currently operating in a first read mode, i.e. a frame scan based read mode. Since the control circuit 230 or 330 determines that the number of events generated in the current pixel array circuit 210 or 310 is small, e.g. only four valid data in a frame of data, it then predicts that the rate of event generation in the next time period is likely to be low. If the read circuit 220 or 320 continues to read in the frame scan based read mode, it will need to repeatedly allocate bits for the pixels that generate events, resulting in a large amount of redundant data. In this case, the control circuit 230 or 330 sends a mode switching signal to the read circuit 220 or 320 to cause the read circuit 220 or 320 to switch from the first read mode to a second read mode. After the switch,
[0508] In some possible embodiments, the visual sensor 200 or 300 can further include a parsing circuit, which can be configured to parse the data signals output by the read circuit 220 or 320. In some possible embodiments, the parsing circuit can parse the data signals in a parsing mode that is adapted to the current data read mode of the read circuit 220 or 320. This will be described in detail below.
[0509] It should be understood that other existing or future to be developed data read modes, data read modes, data parsing modes, etc. are also applicable to possible embodiments of the present application, and all numerical values in the embodiments of the present application are illustrative rather than limiting, for example, the possible embodiments of the present application can switch between more than two data read modes.
[0510] According to possible embodiments of the present application, a visual sensor chip is provided, which can adaptively switch between multiple read modes according to the historical statistics of the light intensity change events generated in the pixel array circuit. In this way, whether in a dynamic scene or a static scene, the visual sensor chip can always achieve good reading and parsing performance, avoid the generation of redundant data, and alleviate the pressure of image processing, transmission and storage.
[0511] Figure 7 A flowchart of a method for operating a visual sensor chip according to possible embodiments of the present application is shown. In some possible embodiments, the method can be implemented in Figure 4-a the visual sensor 200 shown, Figure 4-b the visual sensor 300 shown and the Figure 9 electronic device shown below, or can also be implemented using any appropriate device, including various devices known at present or to be developed in the future. For the convenience of discussion, the method will be described below in conjunction with Figure 4-a the visual sensor 200 shown.
[0512] Referring to Figure 7 , a method for operating a visual sensor chip provided by the embodiments of the present application can include the following steps:
[0513] 501. Generate a plurality of data signals corresponding to a plurality of pixels in a pixel array circuit.
[0514] The pixel array circuit 210 generates a plurality of data signals corresponding to a plurality of pixels in the pixel array circuit 210 by measuring the light intensity change amount. In the context herein, the data signals can indicate, including but not limited to, light intensity polarity, absolute light intensity value, light intensity change value, etc.
[0515] 502. reading at least one data signal of the plurality of data signals from the pixel array circuit in a first read mode.
[0516] The read circuit 220 reads at least one data signal of the plurality of data signals from the pixel array circuit 210 in a first read mode, and these data signals occupy certain storage and transmission resources within the vision sensor 200 after being read. Depending on the specific read mode, the vision sensor 200 can read the data signals in different ways. In some possible embodiments, for example, in an event stream based read mode, the read circuit 220 determines the position information (x, y) of the pixels related to the light intensity change event by scanning the pixel array circuit 210. Based on the position information, the read circuit 220 can read out a first data signal of the plurality of data signals. In this embodiment, the read circuit 220 obtains the characteristic information of the light intensity, the position information (x, y) of the pixel generating the light intensity change event, the time stamp t of reading the data signal, etc. by reading the data signal.
[0517] In other possible embodiments, the first read mode can be a frame scanning based read mode. In this mode, the vision sensor 200 scans the pixel array circuit 210 at a frame frequency associated with the frame scanning based read mode to read all data signals generated by the pixel array circuit 210. In this embodiment, the read circuit 220 obtains the characteristic information of the light intensity by reading the data signal.
[0518] 503. providing the at least one data signal to a control circuit.
[0519] The read circuit 220 provides the read at least one data signal to the control circuit 230 for data statistics and analysis by the control circuit 230. In some embodiments, the control circuit 230 can determine statistical data related to at least one light intensity change event based on the at least one data signal. The control circuit 230 can analyze the statistical data by using a switching strategy module. If it is determined that the statistical data meets a predetermined switching condition, the control circuit 230 sends a mode switching signal to the read circuit 220.
[0520] In some embodiments, when the first read mode is the frame-based read mode and the second read mode is the event-stream-based read mode, the control circuit 230 can determine the number of light intensity change events based on the number of the first data signals in the plurality of data signals. Then, the control circuit 230 compares the number of light intensity change events with a first threshold number. If the number of light intensity change events is less than or equal to the first threshold number, the control circuit 230 determines that the statistical data of the light intensity change events satisfy the predetermined switching condition and sends the mode switching signal. In this embodiment, the control circuit 230 can determine or adjust the first threshold number based on the number of pixels of the pixel array circuit, the frame rate and the amount of pixel data associated with the frame-based read mode, and the amount of event data associated with the event-stream-based read mode.
[0521] In some embodiments, when the first read mode is the event-stream-based read mode and the second read mode is the frame-based read mode, the control circuit 230 can determine the statistical data related to the light intensity change events based on the first data signals received from the read circuit 220. Then, the control circuit 230 compares the number of light intensity change events with a second threshold number. If the number of light intensity change events is greater than or equal to the second threshold number, the control circuit 230 determines that the statistical data of the light intensity change events satisfy the predetermined switching condition and sends the mode switching signal. In this embodiment, the control circuit 230 can determine or adjust the second threshold number based on the number of pixels of the pixel array circuit, the frame rate and the amount of pixel data associated with the frame-based read mode, and the amount of event data associated with the event-stream-based read mode.
[0522] 504、based on the mode switching signal, switching the first read mode to the second read mode.
[0523] The read circuit 220 switches the first read mode to the second read mode based on the mode switching signal received from the control circuit 220. Then, the read circuit 220 reads at least one data signal generated by the pixel array circuit 210 in the second read mode. The control circuit 230 can then continue to statistically record the light intensity change events generated by the pixel array circuit 210 and analyze the light intensity change events in real time, and send the mode switching signal to switch the read circuit 220 from the second read mode to the first read mode when the switching condition is satisfied.
[0524] According to the method provided by the possible embodiments of the present application, the control circuit continuously statistically records and analyzes the light intensity change events generated by the pixel array circuit throughout the reading and resolving process, and sends the mode switching signal to switch the read circuit from the current read mode to a more suitable alternative switching mode once the switching condition is satisfied. This adaptive switching process is repeated until the reading of all data signals is completed.
[0525] Figure 8 A block diagram of a control circuit of possible embodiments of the present application is shown. The control circuit can be used to implement the control circuit 230 in Figure 4-a , the control circuit 330 in Figure 5 , etc., and can also be implemented with other suitable devices. It should be understood that the control circuit is for exemplary purposes only and does not imply any limitation on the scope of the present application. Embodiments of the present application can also be embodied in different control circuits. In addition, it should also be understood that the control circuit can also include other elements, modules or entities, which are not shown for the purpose of clarity, but do not mean that embodiments of the present application do not have these elements or entities.
[0526] As shown in Figure 8 , the control circuit includes at least one processor 602, at least one memory 604 coupled to the processor 602, and a communication mechanism 612 coupled to the processor 602. The memory 604 is at least used to store computer programs and data signals obtained from the reading circuit. The statistical model 606 and the policy module 608 are pre-configured on the processor 602. The control circuit 630 can be communicatively coupled to the reading circuit 220 of the visual sensor 200 as shown in Figure 4-a or the reading circuit outside the visual sensor to implement control functions thereon. For ease of description, the reading circuit 220 in Figure 4-a is referred to below, but embodiments of the present application are equally applicable to the configuration of the peripheral reading circuit.
[0527] Similar to the control circuit 230 shown in Figure 4-a , in some possible embodiments, the control circuit can be configured to control the reading circuit 220 to read a plurality of data signals generated by the pixel array circuit 210 in a specific data reading mode (e.g., a synchronous reading mode based on frame scanning, an asynchronous reading mode based on event stream, etc.). In addition, the control circuit can be configured to obtain data signals from the reading circuit 220, which can indicate, but are not limited to, light intensity polarity, absolute light intensity value, light intensity change value, etc. For example, the light intensity polarity can represent the trend of light intensity change, such as enhancement or weakening, usually represented by +1 / -1. The absolute light intensity value can represent the light intensity value measured at the current time. Depending on the structure, purpose and type of the sensor, the information about the light intensity or light intensity change can have different physical meanings.
[0528] The control circuit determines statistics related to the at least one light intensity change event based on the data signals obtained from the read circuit 220. In some embodiments, the control circuit can obtain the data signals generated by the pixel array circuit 210 over a period of time from the read circuit 220 and store these data signals in the memory 604 for historical statistics and analysis. In the context of the present application, the first read mode and the second read mode can be one of an event stream based asynchronous read mode and a frame scan based synchronous read mode, respectively. It should be noted, however, that all features described herein with respect to adaptively switching read modes are equally applicable to other types of sensors and data read modes that are presently known or developed in the future, as well as switching between more than two data read modes.
[0529] In some possible embodiments, the control circuit can utilize one or more preconfigured statistical models 606 to perform historical statistics on the light intensity change events generated by the pixel array circuit 210 over a period of time provided by the read circuit 220. The statistical models 606 can then transmit the statistics data to the policy module 608 as an output. As previously described, the statistics data can be indicative of the number of light intensity change events or the total amount of data of the light intensity change events. It should be understood that any suitable statistical models, statistical algorithms can be applied to the possible embodiments of the present application, and the scope of the present application is not limited in this respect.
[0530] Since the statistics data is a statistical result of the historical light intensity change events generated by the vision sensor over a period of time, the policy module 608 can analyze and predict the rate of events occurring in the next period of time. The policy module 608 can be preconfigured with one or more switching decisions. When there are multiple switching decisions, the control circuit can select one from the multiple switching decisions for analysis and decision making according to needs, for example, according to the type of the vision sensor 200, the characteristics of the light intensity change events, the properties of the external environment, the motion state, and other factors. In the possible embodiments of the present application, other suitable policy modules and mode switching conditions or strategies can also be employed, and the scope of the present application is not limited in this respect.
[0531] In some embodiments, if the policy module 608 determines that the statistics data satisfies the mode switching condition, an indication about switching the read mode is output to the read circuit 220. In another embodiment, if the policy module 608 determines that the statistics data does not satisfy the mode switching condition, no indication about switching the read mode is output to the read circuit 220. In some embodiments, the indication about switching the read mode can be in an explicit form as described in the above embodiments, for example, in the form of a switching signal or a flag bit to inform the read circuit 220 to switch the read mode.
[0532] Figure 9A block diagram of an electronic device according to possible embodiments of the present application is shown. It should be understood that the electronic device is for exemplary purposes, which can be implemented with any suitable device, including various sensor devices that are currently known and those to be developed in the future. Embodiments of the present application can also be embodied in different sensor systems. In addition, it should also be understood that the electronic device can also include other elements, modules or entities, which are not shown for the purpose of clarity, but do not mean that embodiments of the present application do not have these elements, modules or entities.
[0533] As shown in Figure 9 , the visual sensor includes a pixel array circuit 710 and a read circuit 720, in which read components 720-1 and 720-2 of the read circuit 720 are coupled to a control circuit 730 via communication interfaces 702 and 703, respectively. In embodiments of the present application, the read components 720-1 and 720-2 can be implemented with independent devices, respectively, or can be integrated in the same device. For example, Figure 4-a The read circuit 220 shown is an integrated example implementation. For ease of description, the read components 720-1 and 720-2 can be configured to implement data reading functions in a frame scan-based reading mode and an event stream-based reading mode, respectively.
[0534] The pixel array circuit 710 can be implemented with the pixel array circuit 210 in Figure 4-a or the pixel array circuit 310 in Figure 5 , or any suitable other device, which is not limited in this regard by the present application. Features regarding the pixel array circuit 710 are not repeated here.
[0535] The read circuit 720 can read data signals generated by the pixel array circuit 710 in a specific reading mode. For example, in an example in which the read component 720-1 is turned on and the read component 720-2 is turned off, the read circuit 720 initially reads data signals in a frame scan-based reading mode. In an example in which the read component 720-2 is turned on and the read component 720-1 is turned off, the read circuit 720 initially reads data signals in an event stream-based reading mode. The read circuit 720 is implemented with the read circuit 220 in Figure 4-a or the read circuit 320 in Figure 5 , or any suitable other device, which is not limited in this regard by the present application. Features regarding the read circuit 720 are not repeated here.
[0536] In embodiments of the present application, the control circuit 730 can instruct the read circuit 720 to switch from the first read mode to the second read mode by means of an instruction signal or a flag bit. In this case, the read circuit 720 can receive the instruction from the control circuit 730 about switching the read mode, for example, turning on the read component 720-1 and turning off the read component 720-2, or vice versa.
[0537] As mentioned previously, the electronic device can further include a parsing circuit 704. The parsing circuit 704 can be configured to parse the data signal read by the read circuit 720. In possible embodiments of the present application, the parsing circuit can adopt a parsing mode that is adapted to the current data read mode of the read circuit 720. As an example, if the read circuit 720 initially reads the data signal in the event stream based read mode, the parsing circuit accordingly parses the data based on the first data amount B ev ·N ev associated with the read mode. When the read circuit 720 switches from the event stream based read mode to the frame scanning based read mode based on the instruction of the control circuit 730, the parsing circuit starts to parse the data signal in the second data amount, i.e., a frame data size M·B p , and vice versa.
[0538] In some embodiments, the parsing circuit 704 can implement the switching of the parsing mode of the parsing circuit without the need of an explicit switching signal or flag bit. For example, the parsing circuit 704 can employ the same or corresponding statistical model and switching strategy as the control circuit 730 to make the same statistical analysis on the data signal provided by the read circuit 720 and make the same switching prediction as the control circuit 730. As an example, if the read circuit 720 initially reads the data signal in the event stream based read mode, the parsing circuit accordingly initially parses the data based on the first data amount B ev ·N ev associated with the read mode. For example, the first b x bits parsed by the parsing circuit indicate the coordinate x of the pixel, the next b y bits indicate the coordinate y of the pixel, the following b t bits indicate the reading time, and the last b f bits indicate the feature information of the light intensity. The parsing circuit obtains at least one data signal from the read circuit 720 and determines the statistical data related to at least one light intensity change event. If the parsing circuit 704 determines that the statistical data satisfies the switching condition, the parsing circuit switches to the parsing mode corresponding to the frame scanning based read mode to parse the data signal in the frame data size M·B p .
[0539] As another example, if the read circuit 720 initially reads the data signal in a frame-scan based read mode, the parsing circuit 704 parses the data signal in a frame-based parsing mode corresponding to the read mode, taking out the value of each pixel position in the frame in sequence, where the value of a pixel position in which no light intensity change event occurs is 0. The parsing circuit 704 can count the number of non-0 values in a frame based on the data signal, i.e., the number of light intensity change events in the frame. p As another example, if the read circuit 720 initially reads the data signal in a frame-scan based read mode, the parsing circuit 704 parses the data signal in a frame-based parsing mode corresponding to the read mode, taking out the value of each pixel position in the frame in sequence, where the value of a pixel position in which no light intensity change event occurs is 0. The parsing circuit 704 can count the number of non-0 values in a frame based on the data signal, i.e., the number of light intensity change events in the frame.
[0540] In some possible embodiments, the parsing circuit 704 obtains at least one data signal from the read circuit 720, and determines which one of the current parsing mode and the alternative parsing mode corresponds to the read mode of the read circuit 720 based on at least the at least one data signal. In turn, in some embodiments, the parsing circuit 704 can switch from the current parsing mode to the other parsing mode based on the determination.
[0541] In some possible embodiments, the parsing circuit 704 can determine whether to switch the parsing mode based on historical statistics of the light intensity change events. For example, the parsing circuit 704 can determine statistics data related to at least one light intensity change event based on at least one data signal received from the read circuit 720. If the statistics data is determined to satisfy a switching condition, the parsing circuit 704 switches from the current parsing mode to the alternative parsing mode. For ease of comparison, the statistics data can be used to measure the read data rate of the first read mode and the second read mode of the read circuit 720, respectively.
[0542] In some embodiments, the statistics data can include the total amount of data of the number of events measured by the pixel array circuit 710 per unit time. If the parsing circuit 704 determines based on the at least one data signal that the total amount of data of the light intensity change events read by the read circuit 720 in the first read mode has been greater than or equal to that of the light intensity change events in the second read mode, it indicates that the read circuit 720 has switched from the first read mode to the second read mode. In this case, the parsing circuit 704 should accordingly switch to the parsing mode corresponding to the current read mode.
[0543] In some embodiments, the first read mode is a frame-scan based read mode and the second read mode is an event-stream based read mode. In this embodiment, the parsing circuit 704 initially parses the data signal obtained from the read circuit 720 in a frame-based parsing mode corresponding to the first read mode. The parsing circuit 704 can determine the total amount of data of the light intensity change events read by the read circuit 720 in the first read mode based on the number of pixels M of the pixel array circuit 710, the frame rate f, and the pixel data amount B p p ·f. The parsing circuit 704 can determine the total amount of data of the light intensity change events read by the read circuit 720 in the first read mode based on the number of pixels M of the pixel array circuit 710, the frame rate f, and the pixel data amount Bev and an event data volume B associated with the event stream-based read mode ev to determine a total data volume B of light intensity change events read by the read circuit 720 in the second read mode ev • N ev In some embodiments, a switching parameter can be utilized to adjust the relationship between the total data volumes in the two read modes. In turn, the parsing circuit 704 can determine a total data volume M•B of light intensity change events read by the read circuit 720 in the first read mode according to, for example, equation (1) above p • f is greater than or equal to a total data volume B of light intensity change events in the second read mode ev • N ev If so, the parsing circuit 704 determines that the read circuit 720 has switched to the event stream-based read mode and accordingly switches from the frame-based parsing mode to the event stream-based parsing mode.
[0544] As an alternative implementation to the above-described embodiments, the parsing circuit 704 can determine a total data volume M•B of light intensity change events read by the read circuit 720 in the first read mode according to equation (2) above p • f is greater than or equal to a total data volume B of light intensity change events in the second read mode ev • N ev Similarly, in the case where the total data volume M•B of light intensity change events read by the read circuit 720 in the first read mode p • f is greater than or equal to a total data volume B of light intensity change events in the second read mode ev • N ev the parsing circuit 704 determines that the read circuit 720 has switched to the event stream-based read mode and accordingly switches from the frame-based parsing mode to the event stream-based parsing mode.
[0545] In some embodiments, the first read mode is an event stream-based read mode and the second read mode is a frame scan-based read mode. In this embodiment, the parsing circuit 704 initially parses the data signal obtained from the read circuit 720 in an event stream-based parsing mode corresponding to the first read mode. As previously described, the parsing circuit 704 can directly determine a number N ev of light intensity change events generated in the pixel array circuit 710 based on the number of first data signals provided by the read circuit 720. The parsing circuit 704 can determine a total data volume B of events read by the read circuit 720 in the first read mode based on the event number N ev and an event data volume B associated with the event stream-based read mode ev • N ev • Nev Similarly, the parsing circuit 704 can also be based on the number of pixels M, frame rate f, and pixel data volume B of the pixel array circuit. p The total amount of data M·B of light intensity change events read by the reading circuit 720 in the second reading mode is determined. p ·f. Then, the parsing circuit 704 can, for example, determine the total amount of data B of the light intensity change events read in the first reading mode according to formula (3) above. ev ·N ev Is it greater than or equal to the total data volume M·B of the light intensity change event in the second reading mode? p ·f. Similarly, when the parsing circuit 704 determines the total amount of data B of the light intensity change events read by the reading circuit 720 in the first reading mode... ev ·N ev The total data volume M·B of light intensity change events greater than or equal to that of the second reading mode p When f is reached, the parsing circuit 704 determines that the reading circuit 720 has switched to the frame-scan-based reading mode, and accordingly switches from the event-stream-based parsing mode to the frame-based parsing mode.
[0546] As an alternative embodiment of the above embodiments, the parsing circuit 704 can determine the total amount of data B of the light intensity change event read by the reading circuit 720 in the first reading mode according to the formula (4) above. ev ·N ev Is it greater than or equal to the total data volume M·B of light intensity change events read in the second reading mode? p ·f. Similarly, the total amount of data B of light intensity change events read by the reading circuit 720 in the first reading mode is determined. ev ·N ev The total data volume M·B of light intensity change events greater than or equal to that of the second reading mode p In the case of f, the parsing circuit 704 determines that the reading circuit 720 has switched to the frame scan-based reading mode, and accordingly switches from the event stream-based parsing mode to the frame scan-based parsing mode.
[0547] For the event reading time t in frame-scan-based reading mode, it is assumed that all events within the same frame have the same reading time t. When higher precision is required for event reading time, the reading time of each event can be further determined as follows. Taking the above embodiment as an example, in frame-scan-based reading mode, the frequency of the reading circuit 720 scanning the pixel array circuit is f Hz, then the time interval for reading data from two adjacent frames is S = 1 / f, and the start time of each frame is given as:
[0548] T k =T0+kS(9)
[0549] wherein T0 is the start time of the first frame, and k is the frame number, the time required for digital-to-analog conversion for one of the M pixels can be determined by equation (10) as follows:
[0550]
[0551] The time at which the light intensity change event occurs at the i-th pixel in the k-th frame can be determined by equation (11) as follows:
[0552]
[0553] wherein i is a positive integer. If the current read mode is the synchronous read mode, the read mode is switched to the asynchronous read mode, and the data is read according to each event B ev bit parsing data. In the above embodiments, the switching of the parsing mode can be implemented without explicit switching signals or flag bits. For other currently known or future developed data read modes, the parsing circuit can also parse data in a similar manner adapted to the data read mode, which will not be described here.
[0554] Figure 10 A schematic diagram of the data amount over time for a single data read mode and an adaptive switching read mode according to possible embodiments of the present application is shown. Figure 10 The left half of FIG. 10 depicts a schematic diagram of the read data amount over time for a conventional vision sensor or sensor system using only the synchronous read mode or the asynchronous read mode. In the case of using only the synchronous read mode, as shown by curve 1001, since each frame has a fixed data amount, the read data amount remains constant over time, i.e., the read data rate (the data amount read per unit time) is stable. As previously described, when a large number of events occur in the pixel array circuit, it is reasonable to read the data signal using a read mode based on frame scanning, and most of the frame data is valid data indicating the occurrence of events, and there is little redundancy. When the number of events occurring in the pixel array circuit is small, there are a large number of invalid data in a frame indicating the occurrence of events, and reading the light intensity information at the pixel in the frame data structure will generate redundancy and waste transmission bandwidth and storage resources.
[0555] In the case of simply using asynchronous reading mode, as shown by curve 1002, the amount of data read varies with the rate of event generation, and therefore the data reading rate is not fixed. When few events are generated in the pixel array circuit, only a small number of bits are needed to represent the pixel coordinate information (x, y), the timestamp t of the data signal being read, and the characteristic information f of the light intensity. The total amount of data to be read is small, and asynchronous reading mode is reasonable in this case. When a large number of events are generated in the pixel array circuit in a short period of time, a large number of bits need to be allocated to represent these events. However, these pixel coordinates are almost adjacent, and the data signal reading time is almost the same. In other words, there is a large amount of duplicate data in the read event data, so there is also a redundancy problem in asynchronous reading mode. In this case, the data reading rate may even exceed the data reading rate in synchronous reading mode, and it is unreasonable to still use asynchronous reading mode.
[0556] Figure 10 The right half of the diagram illustrates the variation of data volume over time in an adaptive data readout mode according to a possible embodiment of this application. The adaptive data readout mode can utilize... Figure 4-a The visual sensor 200 shown Figure 4-b The visual sensor 300 shown or Figure 9 The electronic devices shown can be used to achieve this, or traditional vision sensors or sensor systems can be used by utilizing... Figure 8 The control circuit shown implements an adaptive data reading mode. For ease of description, please refer to the following... Figure 4-a The visual sensor 200 shown is used to describe the characteristics of the adaptive data readout mode. As shown in curve 1003, the visual sensor 200 selects, for example, an asynchronous readout mode in the initialization state. This is because the number of bits B used to represent each event in this mode is... ev It is pre-arranged (e.g., B) ev =b x +b y +b t +b f As events are generated and read, the vision sensor 200 can calculate the data reading rate in the current mode. On the other hand, in synchronous reading mode, the number of bits B used to represent each pixel in each frame... pis also predetermined, so the read data rate using the synchronous read mode in the time period can be calculated. The vision sensor 200 can then determine whether the relationship between the data rates in the two read modes satisfies a mode switching condition. For example, the vision sensor 200 can compare the read data rates in the two read modes based on a predefined threshold to determine which mode has a smaller read data rate. Once it is determined that the mode switching condition is satisfied, the vision sensor 200 switches to the other read mode, e.g., from the initial asynchronous read mode to the synchronous read mode. The above steps continue during the reading and parsing of the data signal until the output of all data is completed. As shown in the curve 1003, the vision sensor 200 adaptively selects the optimal read mode throughout the data reading process, and the two read modes alternate, so that the read data rate of the vision sensor 200 always remains below the read data rate of the synchronous read mode, thereby reducing the cost of data transmission, parsing and storage of the vision sensor.
[0557] In addition, according to the adaptive data reading mode proposed by the embodiments of the present application, the vision sensor 200 can statistically analyze the historical data of events to predict the possible event generation rate in the next time period, so as to select a read mode that is more suitable for the application scenario and the motion state.
[0558] Through the above scheme, the vision sensor can adaptively switch between multiple data reading modes, so that the read data rate always remains below the predetermined read data rate threshold, thereby reducing the cost of data transmission, parsing and storage of the vision sensor, and significantly improving the performance of the sensor. In addition, such a vision sensor can statistically analyze the events generated in a time period to predict the possible event generation rate in the next time period, so as to select a read mode that is more suitable for the current external environment, application scenario and motion state.
[0559] The above describes that the pixel array circuit can be used to measure the light intensity change amount, and generate a plurality of data signals corresponding to a plurality of pixels. The data signal can indicate, but is not limited to, the light intensity polarity, the absolute light intensity value, the light intensity change value, and the like. The data signal output by the pixel array circuit is described in detail below.
[0560] Figure 11 A schematic diagram of a pixel circuit 900 provided by the present application is shown. Each of the pixel array circuit 210, the pixel array circuit 310 and the pixel array circuit 710 can include one or more pixel arrays, and each pixel array includes a plurality of pixels, each of which can be regarded as a pixel circuit, and each pixel circuit is used to generate a data signal corresponding to the pixel. Referring to Figure 11A schematic diagram of a preferred pixel circuit is provided in the embodiments of the present application. The pixel circuit is also referred to as a pixel in the embodiments of the present application. Figure 11 As shown in the figure, the preferred pixel circuit in the embodiments of the present application includes a light intensity detection unit 901, a threshold comparison unit 902, a readout control unit 903, and a light intensity collection unit 904.
[0561] The light intensity detection unit 901 is configured to convert the acquired light signal into a first electrical signal. The light intensity detection unit 901 can monitor the light intensity information irradiated on the pixel circuit in real time, and convert the acquired light signal into an electrical signal and output in real time. In some possible embodiments, the light intensity detection unit 901 can convert the acquired light signal into a voltage signal. The specific structure of the light intensity detection unit is not limited in the embodiments of the present application, and any structure that can convert the light signal into an electrical signal can be used in the embodiments of the present application, for example, the light intensity detection unit can include a photodiode and a transistor. The anode of the photodiode is grounded, the cathode of the photodiode is connected to the source of the transistor, and the drain and gate of the transistor are connected to a power supply.
[0562] The threshold comparison unit 902 is configured to determine whether the first electrical signal is greater than a first target threshold or whether the first electrical signal is less than a second target threshold. When the first electrical signal is greater than the first target threshold or the first electrical signal is less than the second target threshold, the threshold comparison unit 902 outputs a first data signal, which is used to indicate that the pixel has a light intensity change event. The threshold comparison unit 902 is configured to compare whether the difference between the current light intensity and the light intensity when the last event occurs exceeds a predetermined threshold, which can be understood with reference to formula 1-1. The first target threshold can be understood as the sum of the first predetermined threshold and the second electrical signal, and the second target threshold can be understood as the sum of the second predetermined threshold and the second electrical signal. The second electrical signal is the electrical signal output by the light intensity detection unit 901 when the last event occurs. The threshold comparison unit in the embodiments of the present application can be implemented in a hardware manner or in a software manner, which is not limited in the embodiments of the present application. The type of the first data signal output by the threshold comparison unit 902 can be different: in some possible embodiments, the first data signal includes polarity information, such as +1 or -1, which is used to indicate that the light intensity is enhanced or the light intensity is weakened. In some possible embodiments, the first data signal can be an activation signal, which is used to instruct the readout control unit 903 to control the light intensity collection unit 904 to collect the first electrical signal and buffer the first electrical signal. When the first data signal is the activation signal, the first data signal can also be the polarity information, and when the readout control unit 903 acquires the first data signal, the readout control unit 903 controls the light intensity collection unit 904 to collect the first electrical signal.
[0563] The readout control unit 903 is further configured to instruct the readout circuit to read the first electrical signal stored in the light intensity collection unit 904. Alternatively, the readout control unit 903 is configured to instruct the readout circuit to read the first data signal outputted by the threshold comparison unit 902, which is the polarity information.
[0564] The readout circuit 905 can be configured to scan the pixels in the pixel array circuit in a predetermined order to read the data signal generated by the corresponding pixels. In some possible embodiments, the readout circuit 905 can be understood with reference to the readout circuit 220, the readout circuit 320, and the readout circuit 720, i.e., the readout circuit 905 is configured to be able to read the data signal outputted by the pixel circuit in more than one signal reading mode. For example, the readout circuit 905 can be configured to read in one of a first reading mode and a second reading mode, which correspond to one of a frame scanning based reading mode and an event stream based reading mode, respectively. In some possible embodiments, the readout circuit 905 can also be configured to read the data signal outputted by the pixel circuit in only one signal reading mode, such as the readout circuit 905 being configured to read the data signal outputted by the pixel circuit in only the frame scanning based reading mode, or the readout circuit 905 being configured to read the data signal outputted by the pixel circuit in only the event stream based reading mode. In Figure 11 In corresponding embodiments, the readout circuit 905 reads the data signal in different manners, i.e., in some possible embodiments, the readout circuit reads the data signal by polarity information, such as the readout circuit reading the polarity information outputted by the threshold comparison unit; in some possible embodiments, the readout circuit reads the data signal by light intensity information, such as the readout circuit reading the electrical signal buffered by the light intensity collection unit.
[0565] Referring to Figure 12 For example, the data signal outputted by the pixel circuit is read in the event stream based reading mode, and the event represented by the light intensity information and the event represented by the polarity information are explained. As shown in the upper half of FIG. 8, the black dots represent the pixels generating the light intensity change events, Figure 12 As shown in the upper half of FIG. 8, the black dots represent the pixels generating the light intensity change events, Figure 12 In total, there are 8 events in FIG. 8, in which the first 5 events are represented by the light intensity information, and the last 3 events are represented by the polarity information. As shown in the lower half of FIG. 8, the black dots represent the pixels generating the light intensity change events, Figure 12As shown in the lower half of FIG. 9, both the event represented by the light intensity information and the event represented by the polarity information need to include coordinate information (x, y), time information t, and the difference is that the characteristic information m of the light intensity of the event represented by the light intensity information is the light intensity information a, and the characteristic information m of the light intensity of the event represented by the polarity information is the polarity information p. The difference between the light intensity information and the polarity information has been described above and will not be repeated here, and only it is emphasized that the amount of data of the event represented by the polarity information is less than the amount of data represented by the light intensity information.
[0566] As to how to determine which information is used to represent the data signal read by the read circuit, it needs to be determined according to the indication issued by the control circuit, which will be described in detail below.
[0567] In some embodiments, the read circuit 905 can be configured to provide the control circuit 906 with at least one data signal read. For example, the read circuit 905 can provide the control circuit 906 with the total amount of data of the data signal read in a period of time for the control circuit 906 to perform historical data statistics and analysis. In one embodiment, the control circuit 906 can obtain the number of events N generated per second by the pixel array circuit based on the number of light intensity change events generated by each pixel circuit 900 in the pixel array circuit in a period of time. ev ev The N can be obtained by one of the frame scanning based read mode and the event stream based read mode.
[0568] The control circuit 906 is coupled to the read circuit 905 and is configured to control the read circuit 906 to read the data signal generated by the pixel circuit 900 in a specific event representation manner. In some possible embodiments, the control circuit 906 can obtain at least one data signal from the read circuit 905, and determine which of the current event representation manner and the alternative event representation manner is more suitable for the current application scenario and motion state based on at least the at least one data signal. Further, in some embodiments, the control circuit 906 can instruct the read circuit 905 to switch from the current event representation manner to another event representation manner based on the determination.
[0569] In some possible embodiments, the control circuit 906 can send an indication to the read circuit 905 about switching the event representation manner based on the historical statistics of the light intensity change events. For example, the control circuit 906 can determine statistical data related to at least one light intensity change event based on at least one data signal received from the read circuit 905. If the statistical data is determined to satisfy a predetermined switching condition, the control circuit 906 sends an indication signal to the read circuit 905 to make the read circuit 905 switch the read event format.
[0570] In some possible embodiments, assuming the readout circuit 905 is configured to read the data signal output by the pixel circuit only in an event-stream-based readout mode, the data provided by the readout circuit 905 to the control circuit 906 is the total number of events (light intensity change events) measured by the pixel array circuit per unit time. Assuming the current control circuit 906 controls the readout circuit 905 to read the data output by the threshold comparison unit 902, i.e., events are represented by polarity information, the readout circuit 905 can then read the data based on the number N of light intensity change events. ev The bit width H of the data format determines the total data volume N of the light intensity change event. ev ×H. Where the bit width H of the data format is b. x +b y +b t +b p b p One bit is used to represent the polarity information of the light intensity indicated by the data signal, usually 1 to 2 bits. Since the polarity information of light intensity is usually represented by 1 to 2 bits, the total amount of data of the event represented by the polarity information is always less than the bandwidth. In order to ensure that the event data with higher precision can be transmitted as much as possible without exceeding the bandwidth limit, if the total amount of data of the event represented by the light intensity information is also less than or equal to the bandwidth, then it is converted to represent the event by the light intensity information. In some embodiments, the conversion parameter can be used to adjust the relationship between the amount of data and the bandwidth K under an event representation method, as shown in the following formula (12), where the total amount of data N of the event represented by the light intensity information is... ev ×H is less than or equal to the bandwidth.
[0571] N ev ×H≤α×K(12)
[0572] Where α is the conversion parameter used for adjustment. From the above formula (12), it can be further concluded that if the total amount of event data represented by light intensity information is less than or equal to the bandwidth, the control circuit 906 can determine that the statistical data of the light intensity change event meets the predetermined switching conditions. Some possible application scenarios include when the pixel acquisition circuit generates fewer events within a certain period of time, or when the rate at which the pixel acquisition circuit generates events is slow within a certain period of time. In these cases, events can be represented by light intensity information. Since events represented by light intensity information can carry more information, it is beneficial for the subsequent processing and analysis of events, such as improving the quality of image reconstruction.
[0573] In some implementations, assuming that the current control circuit 906 controls the reading circuit 905 to read the electrical signal buffered by the light intensity acquisition unit 904, that is, to represent events through light intensity information, the reading circuit 905 can read the number N of light intensity change events. ev The bit width H of the data format determines the total data volume N of the light intensity change event.ev x H, where, when the event stream-based reading mode is adopted, the bit width H = b of the data format x + b y + b t + b a , b a bits are used to represent the light intensity information indicated by the data signal, usually a plurality of bits, such as 8 bits-12 bits. In some embodiments, a conversion parameter can be used to adjust the relationship between the data amount and the bandwidth K in one event representation, as shown in the following formula (13), the total data amount N ev x H of the event represented by the light intensity information is greater than, the reading circuit 220 should read the data output by the threshold comparison unit 902, i.e., converted to represent the event by the polarity information:
[0574] N ev x H > β x K (13)
[0575] Where β is a conversion parameter for adjustment. From the above formula (13), it can be further concluded that if the total data amount N ev x H of the light intensity change event is greater than the threshold data amount β x K, it indicates that the total data amount of the light intensity change event represented by the light intensity information is greater than or equal to the bandwidth, and the reading circuit 905 can determine that the statistical data of the light intensity change event satisfies the predetermined conversion condition. Some possible application scenarios include a large number of events generated by the pixel acquisition circuit in a period of time, or a fast rate of events generated by the pixel acquisition circuit in a period of time, in these cases, if the event is represented by the light intensity information, the event loss may occur, so the event can be represented by the polarity information, to alleviate the pressure of data transmission and reduce the loss of data.
[0576] In some embodiments, the data provided by the reading circuit 905 to the control circuit 906 is the number of events N ev measured by the pixel array circuit in a unit of time. In some possible embodiments, assuming that the control circuit 906 currently controls the reading circuit 905 to read the data output by the threshold comparison unit 902, i.e., the event is represented by the polarity information, the control circuit can determine whether the number N ev of the light intensity change event satisfies the predetermined conversion condition by judging the relationship between N If N ev is less than or equal to , the reading circuit 220 should read the electrical signal buffered in the light intensity acquisition unit 904, i.e., convert the event to be represented by the light intensity information, and convert the current event represented by the polarity information to be represented by the light intensity information. For example, in the foregoing embodiment, based on formula (12), the following formula (14) can be further obtained:
[0577]
[0578] In some embodiments, assuming that the current control circuit 906 controls the read circuit 905 to read the electrical signal cached by the light intensity collection unit 904, i.e., to convert the event represented by the light intensity information, the control circuit 906 can determine whether the predetermined conversion condition is met according to the relationship between the number N of light intensity change events and the number M of events represented by the light intensity information within a unit time. ev With If N ev is greater than M, the read circuit 220 should read the signal output by the threshold comparison unit 902, i.e., to convert the event represented by the pass polarity information, to convert the event represented by the current light intensity information into the event represented by the pass polarity information. For example, in the foregoing embodiment, based on formula (12), the following formula (15) can be further obtained:
[0579]
[0580] In some possible embodiments, assuming that the read circuit 905 is configured to read the data signal output by the pixel circuit only in the read mode based on frame scanning, the data provided by the read circuit 905 to the control circuit 906 is the total data amount of the number of events (light intensity change events) measured by the pixel array circuit within a unit time. When the read mode based on frame scanning is adopted, the bit width H = B p of the data format is B p , which is the pixel data amount (e.g., the number of bits) allocated for each pixel in the read mode based on frame scanning. When the event is represented by the pass polarity information, B p is usually 1 bit to 2 bits, and when the event is represented by the light intensity information, it is usually 8 bits to 12 bits. The read circuit 905 can determine the total data amount M × H of the light intensity change events, where M represents the total number of pixels. Assuming that the current control circuit 906 controls the read circuit 905 to read the data output by the threshold comparison unit 902, i.e., to convert the event represented by the pass polarity information. The total data amount of the event represented by the pass polarity information is certainly less than the bandwidth. In order to ensure that the event data with higher precision can be transmitted as much as possible without exceeding the bandwidth limit, if the total data amount of the event represented by the light intensity information is also less than or equal to the bandwidth, the event represented by the light intensity information is converted. In some embodiments, a conversion parameter can be used to adjust the relationship between the data amount in one event representation and the bandwidth K, as shown in the following formula (16), the total data amount N ev × H of the event represented by the light intensity information is less than or equal to the bandwidth.
[0581] M × H ≤ α × K (16)
[0582] In some embodiments, assuming that the current control circuit 906 controls the read circuit 905 to read the electrical signal cached by the light intensity collection unit 904, i.e., the event is represented by the light intensity information, the read circuit 905 can determine the total data amount MxH of the light intensity change event. In some embodiments, a conversion parameter can be used to adjust the relationship between the data amount and the bandwidth K in one event representation mode, as shown in the following equation (17). The total data amount MxH of the event represented by the light intensity information is greater than the bandwidth, and the read circuit 220 should read the data output by the threshold comparison unit 902, i.e., converted to the event represented by the polarity information:
[0583] MxH > a x K (17)
[0584] In some possible embodiments, assuming that the read circuit 905 is configured to read in one of a first read mode and a second read mode, the first read mode and the second read mode correspond to one of a frame scanning-based read mode and an event stream-based read mode, respectively. For example, the following describes how the control circuit determines whether the switching condition is met in the combination mode in which the read circuit 905 currently reads the data signal output by the pixel circuit in the event stream-based read mode and the control circuit 906 controls the read circuit 905 to read the data output by the threshold comparison unit 902, i.e., the event is represented by the polarity information:
[0585] In the initial state, one read mode can be arbitrarily selected, such as the frame scanning-based read mode or the event stream-based read mode. In addition, in the initial state, one event representation mode can be arbitrarily selected, such as the control circuit 906 controlling the read circuit 905 to read the electrical signal cached by the light intensity collection unit 904, i.e., the event is represented by the light intensity information, or the control circuit 906 controlling the read circuit 905 to read the data output by the threshold comparison unit 902, i.e., the event is represented by the polarity information. Assuming that the read circuit 905 currently reads the data signal output by the pixel circuit in the event stream-based read mode and the control circuit 906 controls the read circuit 905 to read the data output by the threshold comparison unit 902, i.e., the event is represented by the polarity information. The data provided by the read circuit 905 to the control circuit 906 can be a first total data amount of the number of events (light intensity change events) measured by the pixel array circuit per unit time. Since the total number of pixels M is known, the pixel data amount B allocated to each pixel in the frame scanning-based read mode is known, and the bit width H of the data format when the event is represented by the light intensity information is known. According to the known M, B p p H can acquire a second total data quantity representing the number of events measured by the pixel array circuit per unit time in a combined event mode, where the data signal is read from the pixel circuit output based on an event stream reading mode and the number of events measured by the pixel array circuit per unit time is represented by light intensity information; it can acquire a third total data quantity representing the number of events measured by the pixel array circuit per unit time in a combined event mode, where the data signal is read from the pixel circuit output based on a frame scan reading model and the number of events measured by the pixel array circuit per unit time is represented by polarity information; and it can acquire a fourth total data quantity representing the number of events measured by the pixel array circuit per unit time in a combined event mode, where the data signal is read from the pixel circuit output based on a frame scan reading model and the number of events measured by the pixel array circuit per unit time is represented by light intensity information. Specifically, this is based on M and B. p The methods for calculating the second, third, and fourth data quantities have been described above and will not be repeated here. The switching condition is determined by using the first total data quantity provided by the aforementioned reading circuit 905, the calculated second, third, and fourth total data quantities, and their relationship with the bandwidth K. If the current combination mode cannot guarantee the transmission of higher-precision event data within the bandwidth limit, then the switching condition is met, and the system switches to a combination mode that can guarantee the transmission of higher-precision event data within the bandwidth limit.
[0586] To better understand the above process, let's illustrate it with a specific example:
[0587] Assuming a bandwidth limit of K and a bandwidth adjustment factor of α, in event-stream-based reading mode, when events are represented using polarity information, the bit width H of the data format is H = b. x +b y +b t +b p When representing events using light intensity information, the bit width of the data format is H = b. x +b y +b t +b a Usually 1≤b p a For example, b p Typically 1 to 2 bits, b a Typically, it is 8 to 12 bits.
[0588] In frame-scan-based readout mode, events do not need to represent coordinates and time; events are determined based on the state of each pixel. Assuming the data bit width allocated to each pixel is b in polarity mode... sp In light intensity mode, it is b sa The total number of pixels is M. Assuming a bandwidth limit K = 1000bps, b x =5 bits y =4bit,bt = 10 bit, b p = 1 bit, b a = 8 bit, b sp = 1 bit, b sa = 8 bit, M = 100, bandwidth adjustment factor a = 0.9. Assume that 10 events are generated in the first second, 15 events are generated in the second second, and 30 events are generated in the third second.
[0589] Assume that the initial state is to use the event stream-based reading mode by default, and the event is represented in the polarity mode.
[0590] The event stream-based reading mode and the event represented by the polarity information are referred to as the asynchronous polarity mode, the event stream-based reading mode and the event represented by the light intensity information are referred to as the asynchronous light intensity mode, the frame scanning-based reading mode and the event represented by the polarity information are referred to as the synchronous polarity mode, and the frame scanning-based reading mode and the event represented by the light intensity information are referred to as the synchronous light intensity mode.
[0591] First second: 10 events are generated
[0592] Asynchronous polarity mode: N ev = 10, H = b x + b y + b t + b p = 5 + 4 + 10 + 1 = 20 bits, estimated data amount N ev · H = 200 bits, N ev · H < a · K, bandwidth limit is met.
[0593] Asynchronous light intensity mode: H = b x + b y + b t + b a = 5 + 4 + 10 + 8 = 27 bits, then the estimated data amount N in the light intensity mode ev · H = 270 bits, N ev · H < a · K, bandwidth limit is still met.
[0594] Synchronous polarity mode: M = 100, H = b sp = 1 bit, the estimated data amount is M · H = 100 bits at this time, M · H < a · K, bandwidth limit is still met.
[0595] Synchronous light intensity mode: M = 100, H = b sa = 8 bits, the estimated data amount is M · H = 800 bits at this time, M · H < a · K, bandwidth limit is still met.
[0596] In summary, in the first second, the asynchronous intensity mode is selected, and all 10 events' intensity information is transmitted with a small amount of data (270 bits) without exceeding the bandwidth limit. The control circuit 906 determines that the current combination mode cannot guarantee that higher-precision event data is transmitted as much as possible without exceeding the bandwidth limit, and determines that the switching condition is met, and then controls the asynchronous polarity mode to be switched to the asynchronous intensity mode. For example, a signal indicating that the reading circuit 905 is switched from the current event representation mode to another event representation mode is sent.
[0597] Second 2: 15 events are generated
[0598] Asynchronous polarity mode: estimated data amount N ev H = 15 x 20 = 300 bits, which satisfies the bandwidth limit.
[0599] Asynchronous intensity mode: estimated data amount N ev H = 15 x 27 = 405 bits, which satisfies the bandwidth limit.
[0600] Synchronous polarity mode: estimated data amount M x H = 100 x 1 = 100 bits, which satisfies the bandwidth limit.
[0601] Synchronous intensity mode: estimated data amount M x H = 100 x 8 = 800 bits, which satisfies the bandwidth limit.
[0602] In summary, in the second second, the control circuit 906 determines that the current combination mode can guarantee that higher-precision event data is transmitted as much as possible without exceeding the bandwidth limit, and determines that the switching condition is not met, and still selects the asynchronous intensity mode.
[0603] Third second: 30 events are generated
[0604] Asynchronous polarity mode: estimated data amount N ev H = 30 x 20 = 600 bits, which satisfies the bandwidth limit.
[0605] Asynchronous intensity mode: estimated data amount N ev H = 30 x 27 = 810 bits, which satisfies the bandwidth limit.
[0606] Synchronous polarity mode: estimated data amount M x H = 100 x 1 = 100 bits, which satisfies the bandwidth limit.
[0607] Synchronous intensity mode: estimated data amount M x H = 100 x 8 = 800 bits, which satisfies the bandwidth limit.
[0608] At the 3rd second, the synchronous light intensity mode can transmit the light intensity information of all 30 events with a data amount of 800 bits. At the 3rd second, the current combined mode (asynchronous light intensity mode) cannot guarantee that the event data with higher precision is transmitted as much as possible without exceeding the bandwidth limit, and it is determined that the switching condition is met, so the asynchronous light intensity mode is switched to the synchronous light intensity mode. For example, a switching signal is sent to instruct the reading circuit 905 to switch from the current event reading mode to another event reading mode.
[0609] It should be understood that the above formula, conversion condition and related calculation method are only an example implementation of an embodiment of the present application, and other suitable conversion conditions, conversion strategies and calculation methods of event representation can also be used, and the scope of the present application is not limited in this regard.
[0610] In some embodiments, the reading circuit 905 includes a data format control unit 9051 for controlling the reading of the signals output by the threshold comparison unit 902 or the electrical signals buffered in the light intensity acquisition unit 904. For example, the data format control unit 9051 will be described in conjunction with two preferred embodiments.
[0611] Referring to Figure 12-a FIG. 6 shows a structure of the data format control unit in the reading circuit according to an embodiment of the present application. The data format control unit can include an AND gate 951 and an AND gate 954, or a gate 953, and a NOT gate 952. The input end of the AND gate 951 is used to receive the conversion signal sent by the control circuit 906 and the polarity information output by the threshold comparison unit 902, and the input end of the AND gate 954 is used to receive the conversion signal sent by the control circuit 906 after p...
Claims
1. A pose estimation method, characterized by, The method comprises: obtaining a first event image and a first target image, the first event image being time-aligned with the first target image, the first target image comprising an RGB image or a depth image, and the first event image comprising an image representing a motion trajectory of a target object when the target object generates motion within a detection range of a motion sensor; determining an integration time of the first event image; if the integration time is less than a first threshold, determining not to perform pose estimation by using the first target image; performing pose estimation according to the first event image.
2. The pose estimation method of claim 1, wherein, The method further comprises: determining a time of obtaining the first event image and a time of obtaining the first target image; determining that the first event image is time-aligned with the first target image according to a time difference between the time of obtaining the first target image and the time of obtaining the first event image being less than a second threshold.
3. The pose estimation method of claim 2, wherein, The method further comprises: obtaining N continuous DVS events detected by the motion sensor; integrating the N continuous DVS events into the first event image. The method further comprises: determining the time of obtaining the first event image according to a time of obtaining the N continuous DVS events.
4. The pose estimation method according to any one of claims 1-3, characterized in that, The method further comprises: determining N continuous DVS events used for integration into the first event image; determining the integration time of the first event image according to a time of obtaining a first DVS event and a last DVS event in the N continuous DVS events.
5. The pose estimation method of any one of claims 1-3, wherein, The method further comprises: obtaining a second event image, the second event image comprising an image representing a motion trajectory of the target object when the target object generates motion within the detection range of the motion sensor; if there is no target image time-aligned with the second event image, determining that the second event image does not have a target image for jointly performing pose estimation; performing pose estimation according to the second event image.
6. The pose estimation method of claim 5, wherein, Before determining the pose according to the second event image, the method further comprises: if it is determined that the second event image has time-aligned inertial measurement unit (IMU) data, determining the pose according to the second event image and the IMU data corresponding to the second event image; if it is determined that the second event image does not have time-aligned IMU data, determining the pose according to the second event image only.
7. The pose estimation method of any one of claims 1-3, wherein, The method further comprises: obtaining a second target image, the second target image comprising an RGB image or a depth image; if there is no event image time-aligned with the second target image, determining that the second target image does not have an event image for jointly performing pose estimation; determining the pose according to the second target image.
8. The pose estimation method of any one of claims 1-3, wherein, The method further comprises: performing loop closure detection according to the first event image and a dictionary, the dictionary being a dictionary constructed based on event images.
9. The pose estimation method of claim 8, wherein, The method further comprises: obtaining a plurality of event images, the plurality of event images being event images for training; obtaining visual features of the plurality of event images; The visual features are clustered by a clustering algorithm to obtain clustered visual features, the clustered visual features having corresponding descriptors; The dictionary is constructed according to the clustered visual features.
10. The pose estimation method of claim 9, wherein, The loop detection is performed according to the first event image and the dictionary, including: A descriptor of the first event image is determined; The descriptor of the first event image is determined in the dictionary to correspond to a visual feature; A bag-of-words vector corresponding to the first event image is determined based on the visual feature; Similarities between the bag-of-words vector corresponding to the first event image and bag-of-words vectors of other event images are determined to determine an event image matched by the first event image.
11. The pose estimation method of any one of claims 1-3, wherein, The method further includes: First information of the first event image is determined, the first information including events and / or features in the event image; If it is determined based on the first information that the first event image at least satisfies a first condition, the first event image is determined as a key frame, the first condition being related to an event quantity and / or a feature quantity.
12. The method of claim 11, wherein, The first condition includes one or more of the event quantity in the first event image being greater than a first threshold, the event effective area quantity in the first event image being greater than a second threshold, the feature quantity in the first event image being greater than a third threshold, and the feature effective area in the first event image being greater than a fourth threshold.
13. The method of claim 11, wherein, The method further includes: A depth image temporally aligned with the first event image is obtained; If it is determined based on the first information that the first event image at least satisfies a first condition, the first event image and the depth image are determined as key frames.
14. The method of claim 11, wherein, The method further includes: An RGB image temporally aligned with the first event image is obtained; A feature quantity and / or a feature effective area of the RGB image are obtained; If it is determined based on the first information that the first event image at least satisfies a first condition, and the feature quantity of the RGB image is greater than a fifth threshold and / or the feature effective area quantity of the RGB image is greater than a sixth threshold, the first event image and the RGB image are determined as key frames.
15. The method of claim 11, wherein, If it is determined based on the first information that the first event image at least satisfies a first condition, the first event image is determined as a key frame, including: If it is determined based on the first information that the first event image at least satisfies the first condition, second information of the first event image is determined, the second information including motion features and / or pose features in the first event image; If it is determined based on the second information that the first event image at least satisfies a second condition, the first event image is determined as a key frame, the second condition being related to a motion change quantity and / or a pose change quantity.
16. The method of claim 15, wherein, The method further includes: A definition and / or brightness consistency index of the first event image is determined; If it is determined based on the second information that the first event image at least satisfies the second condition, and the definition of the first event image is greater than a definition threshold and / or the brightness consistency index of the first event image is greater than a preset index threshold, the first event image is determined as a key frame.
17. The method of claim 16, wherein, The determining the brightness consistency index of the first event image comprises: if the pixel in the first event image represents a light intensity change polarity, calculating an absolute value of a difference between the number of events of the first event image and the number of events of the adjacent key frame, and dividing the absolute value by the number of pixels of the first event image to obtain the brightness consistency index of the first event image; if the pixel in the first event image represents light intensity, pixel-by-pixel difference is calculated between the first event image and the adjacent key frame, and the absolute value of the difference is calculated, and the sum of the absolute values corresponding to each group of pixels is calculated, and the sum is divided by the number of pixels to obtain the brightness consistency index of the first event image.
18. The method of claim 15, wherein, The method further comprises: obtaining an RGB image that is time-aligned with the first event image; determining the sharpness and / or brightness consistency index of the RGB image; if it is determined based on the second information that the first event image at least meets the second condition, and the sharpness of the RGB image is greater than a sharpness threshold and / or the brightness consistency index of the RGB image is greater than a preset index threshold, then the first event image and the RGB image are determined as a key frame.
19. The method of claim 15, wherein, The second condition comprises one or more of the distance between the first event image and the previous key frame exceeding a preset distance value, the rotation angle between the first event image and the previous key frame exceeding a preset angle value, and the distance between the first event image and the previous key frame exceeding a preset distance value and the rotation angle between the first event image and the previous key frame exceeding a preset angle value.
20. The method of any one of claims 1-3, wherein, The method further comprises: determining a first motion region in the first event image; determining a corresponding second motion region in the first target image according to the first motion region; performing pose estimation according to the second motion region in the first target image.
21. The method of claim 20, wherein, The determining the first motion region in the first event image comprises: if the DVS that collects the first event image is static, obtaining a pixel point with event response in the first event image; determining the first motion region according to the pixel point with event response.
22. The method of claim 21, wherein, The determining the first motion region according to the pixel point with event response comprises: determining an outline formed by the pixel point with event response in the first event image; if the area surrounded by the outline is greater than a first threshold, determining the region surrounded by the outline as the first motion region.
23. The method of claim 20, wherein, The determining the first motion region in the first event image comprises: if the DVS that collects the first event image is moving, obtaining a second event image, the second event image being a previous frame event image of the first event image; calculating the displacement size and displacement direction of the pixel in the first event image relative to the second event image; if the displacement direction of the pixel in the first event image is not the same as the displacement direction of the surrounding pixels, or the difference between the displacement size of the pixel in the first event image and the displacement size of the surrounding pixels is greater than a second threshold, then the pixel is determined to belong to the first motion region.
24. The method of claim 20, wherein, The method further comprises: determining a corresponding static region in the image according to the first motion region; determining a pose according to the static region in the first target image. 25.A data processing apparatus, comprising a processor and a memory, the processor being coupled with the memory, characterized in that, the memory is configured to store a program; the processor is configured to execute the program in the memory, so that the data processing apparatus performs the method according to any one of claims 1-24.
26. A computer readable storage medium comprising a program, characterized in that, when it is run on a computer, it makes the computer perform the method according to any one of claims 1-24.
27. A computer program product comprising instructions, wherein: when it is run on a computer, it makes the computer perform the method according to any one of claims 1-24. when it is run on a computer, it makes the computer perform the method according to any one of claims 1-24.
Citation Information
Patent Citations
Method of processing data for dynamic image sensor and dynamic image sensor
CN110536078A
SLAM method and system based on binocular event camera
CN111899276A