Electronic device, method for electronic device, computer readable storage medium and computer program product
Through the space-time segmentation and semantic recognition optimization encoding of event cameras, the problems of light changes and high computing power requirements in autonomous driving are solved, and efficient and accurate information transmission is achieved.
Patent Information
- Application Number
- CN202410160422.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-04
- Publication Date
- 2025-08-05
AI Technical Summary
In autonomous driving, existing vision algorithms have problems such as a huge impact on light and dark changes, large data volume and high computing power requirements, and traditional compilation and decoding fails to effectively utilize signal meaning, resulting in low information transmission efficiency.
The event camera is used to perform spatiotemporal segmentation and semantic recognition of pulsed video data, and optimize the codeword allocation process through semantic encoding, reducing redundant data and improving information transmission efficiency.
Through space-time segmentation and semantic recognition, the amount of data transmission is reduced, the accuracy of information transmission is improved, the computing power demand is reduced, the light changes are adapted to the environment and the safety of autonomous driving is improved.
Smart Images

Figure CN120431501A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wireless communication. Specifically, it relates to the technology of semantic communication for pulse video data, and more specifically, to an electronic device, a method for an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] An event camera can be simply understood as a sensor that only senses moving objects. At each pixel of the event camera, there is an independent optoelectronic sensing module. When the brightness change at this pixel exceeds a set threshold, event data (sometimes also called pulse data) will be generated and output. Additionally, since all pixels work independently, the data output of the event camera is asynchronous and sparse in space. This is also the biggest difference between the event camera and the standard camera, as well as the core innovation of the event camera. The advantage of this imaging paradigm is that it can greatly reduce redundant data, thereby improving the computational efficiency of post-processing algorithms.
[0003] In the development process of autonomous driving, the application of vision algorithms has become an indispensable part. However, current vision algorithms still have some limitations: on the one hand, cameras are easily affected by sudden changes in light brightness, backlighting, etc.; on the other hand, when the camera is running, the amount of data generated is very large, so the requirement for computing power is particularly high. Event cameras have advantages such as extremely fast response speed, reducing invalid information, reducing computing power and power consumption, and high dynamic range, which can help autonomous driving vehicles reduce the complexity of information processing, improve the driving safety of the vehicle, and can work normally in extremely bright or extremely dark environments.
[0004] On the other hand, the limitations of Shannon or classical information theory are specifically manifested in only considering signal transmission, rather than the meaning of the signal or the observer's understanding of the meaning of the signal. Based on traditional encoding and decoding, semantic encoding and decoding expands a new dimension, namely the semantic dimension. Semantic features (such as the communication scenario, purpose, context of the communication content, context, etc.) can change the probability distribution of symbols in transmission, and there may be a certain correlation between context symbols in transmission. Using semantic features as the background knowledge of the sender to optimize the codeword allocation process can achieve efficient information expression ability. When the same data is effectively encoded, the required expected codeword length is greatly reduced, thereby saving transmission and storage overheads and achieving concise expression. Summary of the Invention
[0005] A brief summary of the present disclosure is given below to provide a basic understanding of certain aspects of the present disclosure. It should be understood that this summary is not an exhaustive summary of the present disclosure. It is not intended to identify the key or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description discussed later.
[0006] According to one aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, cause the electronic device to perform: performing spatio-temporal segmentation on the pulsed video data obtained by an event camera to obtain pulsed video data corresponding to different categories of objects; respectively performing semantic recognition on different categories of objects based on the pulsed video data corresponding to different categories of objects; and encoding the results of the semantic recognition to obtain an encoded message to be provided to other electronic devices.
[0007] According to another aspect of the present disclosure, there is provided a method for an electronic device, comprising: performing spatio-temporal segmentation on the pulsed video data obtained by an event camera to obtain pulsed video data corresponding to different categories of objects; respectively performing semantic recognition on different categories of objects based on the pulsed video data corresponding to different categories of objects; and encoding the results of the semantic recognition to obtain an encoded message to be provided to other electronic devices.
[0008] According to one aspect of the present disclosure, there is provided an electronic device, comprising: at least one processor; and at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, cause the electronic device to perform: receiving an encoded message from other electronic devices, the encoded message being obtained by the other electronic devices by: performing spatio-temporal segmentation on the pulsed video data obtained by an event camera to obtain pulsed video data corresponding to different categories of objects; respectively performing semantic recognition on different categories of objects based on the pulsed video data corresponding to different categories of objects; and encoding the results of the semantic recognition; and performing semantic decoding and semantic fusion recombination on the encoded message to obtain the semantics of the original scene captured by the event camera.
[0009] According to another aspect of the present disclosure, a method for an electronic device is provided, including: receiving an encoded message from another electronic device, where the encoded message is obtained by the other electronic device through the following steps: performing spatio-temporal segmentation on the pulsed video data obtained by an event camera to obtain pulsed video data corresponding to different categories of objects; respectively performing semantic recognition on different categories of objects based on the pulsed video data corresponding to different categories of objects; and encoding the results of the semantic recognition; and performing semantic decoding and semantic fusion and recombination on the encoded message to obtain the semantics of the original scene captured by the event camera.
[0010] According to other aspects of the present disclosure, computer program code, a computer program product for implementing the above method for an electronic device, and a computer-readable storage medium having recorded thereon the computer program code for implementing the above method for an electronic device are also provided.
[0011] The electronic device and method according to the embodiments of the present application can greatly reduce the data transmission volume and improve the accuracy of information transmission by performing spatio-temporal segmentation on the pulsed video data obtained by an event camera to distinguish different categories of objects and performing semantic recognition and semantic encoding on different categories of objects.
[0012] These and other advantages of the present disclosure will become more apparent through the following detailed description of the preferred embodiments of the present disclosure in conjunction with the accompanying drawings. Description of the Drawings
[0013] To further elaborate the above and other advantages and features of the present disclosure, the specific embodiments of the present disclosure will be described in further detail below in conjunction with the accompanying drawings. The accompanying drawings are included in this specification and form a part of this specification together with the following detailed description. Elements having the same function and structure are denoted by the same reference numerals. It should be understood that these drawings only depict typical examples of the present disclosure and should not be regarded as limiting the scope of the present disclosure. In the drawings:
[0014] Figure 1 A functional module block diagram for an electronic device according to an embodiment of the present application is shown;
[0015] Figure 2 A schematic example of spatial segmentation is shown;
[0016] Figure 3 A schematic example of temporal segmentation is shown;
[0017] Figure 4 A schematic flowchart showing spatio-temporal segmentation of pulsed video data and the selection of a classification of a semantic recognition model is shown;
[0018] Figure 5Shows a functional module block diagram for an electronic device according to an embodiment of the present application;
[0019] Figure 6 Shows a schematic diagram of the codeword composition of each semantic feature item;
[0020] Figure 7 Shows a schematic example of a semantic coding tree;
[0021] Figure 8 Shows an example of extending the DayIII Ext Msg message body MsgFrameIII;
[0022] Figure 9 Shows a functional module block diagram for an electronic device according to another embodiment of the present application; r
[0023] Figure 10 Shows an exemplary operation process of a decoding and semantic recombination unit;
[0024] Figure 11 Shows a functional module block diagram for an electronic device according to another embodiment of the present application;
[0025] Figure 12 Shows a schematic architecture of scene semantic recombination;
[0026] Figure 13 Shows a schematic diagram of scene analysis prediction and warning;
[0027] Figure 14 Shows a schematic diagram of the operations of the transmitting end and the receiving end when applying the solution of the present application to the V2X scenario;
[0028] Figure 15 Shows a diagram of the pulse video data after spatial segmentation;
[0029] Figure 16 Shows an example of the time slices of each space obtained after time segmentation;
[0030] Figure 17 Shows a schematic diagram of semantic decoding;
[0031] Figure 18 Shows a schematic diagram of the overall architecture of scene semantic recombination and scene analysis prediction and warning;
[0032] Figure 19 Shows an example of semantic recombination;
[0033] Figure 20 Shows a flowchart of a method for an electronic device according to an embodiment of the present application;
[0034] Figure 21 Shows an example of the sub-steps of step S13;
[0035] Figure 22 Shows a flowchart of a method for an electronic device according to another embodiment of the present application;
[0036] Figure 23 Shows an example of a flowchart of step S22 including semantic fault tolerance processing;
[0037] Figure 24 Is a block diagram showing an example of a schematic configuration of a smart phone to which the technology of the present disclosure can be applied;
[0038] Figure 25 Is a block diagram showing an example of a schematic configuration of an in-vehicle navigation device to which the technology of the present disclosure can be applied; and
[0039] Figure 26 Is a block diagram of an exemplary structure of a general-purpose personal computer in which a method and / or apparatus and / or system according to an embodiment of the present disclosure can be implemented. Detailed implementation manners
[0040] Hereinafter, exemplary embodiments of the present disclosure will be described in conjunction with the accompanying drawings. For clarity and conciseness, not all features of the actual implementation manners are described in the specification. However, it should be understood that many implementation-specific decisions must be made during the development of any such actual implementation to achieve the specific goals of the developer, for example, to comply with those restrictions related to the system and the business, and these restrictions may vary with different implementation manners. In addition, it should also be understood that although the development work may be very complex and time-consuming, for those skilled in the art who benefit from the present disclosure, such development work is merely a routine task.
[0041] Here, it should also be noted that, in order to avoid obscuring the present disclosure due to unnecessary details, only the device structures and / or processing steps closely related to the solution according to the present disclosure are shown in the drawings, while other details less related to the present disclosure are omitted.
[0042] <First Embodiment>
[0043] As mentioned above, the characteristics of the event camera make it suitable for application in complex scenarios in the field of autonomous driving, and semantic communication can expand the coding dimension to transmit more information with less overhead. The amount of pulsed video data of the event camera varies greatly. At the same time, in the scenario of vehicle-to-everything (V2X) in the Internet of Vehicles, the communication environment is poor, and there may also be recognition errors at the perception end of the event camera. Therefore, how to improve the correct rate of information transmission is also an important issue.
[0044] In this embodiment, an electronic device 100 is provided, which is used to perform semantic recognition on pulsed video data and perform semantic encoding on the recognition result. In the following description, the autonomous driving scenario or the V2X application scenario will be mainly described as an exemplary application scenario, but this is not restrictive, only for the convenience and clarity of description. The embodiments of the present application can be applied to any occasion that can obtain data with similar characteristics and has similar requirements.
[0045] Figure 1 The functional module block diagram of the electronic device 100 according to this embodiment is shown, as Figure 1 shown, the electronic device 100 includes: a spatio-temporal segmentation unit 101, configured to perform spatio-temporal segmentation on the pulsed video data obtained by the event camera to obtain the pulsed video data corresponding to different categories of objects; a semantic recognition unit 102, configured to perform semantic recognition on different categories of objects respectively based on the pulsed video data corresponding to different categories of objects; and a semantic encoding unit 103, configured to encode the result of the semantic recognition to obtain an encoded message to be provided to other electronic devices.
[0046] Among them, the spatio-temporal segmentation unit 101, the semantic recognition unit 102 and the semantic encoding unit 103 can be implemented by one or more processing circuits and at least one memory. The processing circuit can be implemented as a chip, a processor, etc., and the at least one memory can be any form of storage device such as RAM, ROM, flash memory, etc. The at least one memory is used to store computer program codes and data required for the processing circuit to execute processing, etc. And it should be understood that Figure 1 each functional unit in the electronic device shown in is only a logical module divided according to the specific functions it implements, rather than for restricting the specific implementation manner. In addition, the functions of the spatio-temporal segmentation unit 101, the semantic recognition unit 102 and the semantic encoding unit 103 can also be implemented by a brain-like chip, which are not restrictive.
[0047] When the electronic device 100 is applied to the autonomous driving scenario, the electronic device 100 can be set in a roadside unit (RSU) or a vehicle capable of perception. The vehicle described here can be more generally various user devices located on the vehicle and capable of accessing various sensors. The electronic device 100 in this embodiment can work as a data provider and an encoding device.
[0048] It should also be noted that the electronic device 100 can be implemented at the chip level or at the device level. For example, the electronic device 100 can operate as an RSU, a vehicle, or a user device itself, and can also include external devices such as a memory, a transceiver (not shown in the figure), etc. The memory can be used to store programs and related data information required for an RSU, a vehicle, or a user device to implement various functions. The transceiver can include one or more communication interfaces to support communication with different devices (such as other RSUs, vehicles, or user devices, etc.), and the implementation form of the transceiver is not specifically limited here.
[0049] Among them, the event camera is compared with the traditional frame camera. The frame camera outputs pictures frame by frame at a fixed frame rate and finally forms a video stream; while the event camera only records the pixel points with brightness changes and calls the light intensity changes of these pixel points events. Due to these characteristics of the event camera, it is particularly suitable for many specific scenarios in the autonomous driving scenario.
[0050] These specific scenarios include, for example: scenarios with obvious sudden changes in light intensity, including sudden changes in light and darkness, backlighting, etc., such as being blinded by strong light from the oncoming vehicle during a meeting, and the high-exposure scenario faced after the vehicle comes out of the tunnel; scenarios with too bright or too dark light, for example, in a late-night environment, the frame camera cannot recognize the surrounding things due to the extremely dark light, while the event camera can still effectively recognize the surrounding things; lateral blind spot perception scenario: the traditional frame camera has an unsatisfactory effect in detecting lateral pedestrians / non-motor vehicles (the scenario of sudden appearance of pedestrians, especially electric vehicles), and it is also too late to respond to pedestrians suddenly appearing in the field of vision blind spot blocked by a vehicle or an obstacle in front, while the event camera can sense the danger signal faster; high-speed obstacle avoidance scenario, for example, when the vehicle is driving fast on the highway and encounters a tire on the front road surface, the event camera can quickly recognize the tire in front and take timely obstacle avoidance actions; low-power scenario (in-vehicle deployment scenario for electric vehicles), compared with the traditional method, the calculation method based on the event camera can greatly save the power consumption of the vehicle battery, thereby extending the vehicle's battery life.
[0051] The data of an event camera shows a certain sparsity in a two-dimensional structure. For example, if an object moves only at time t0 and then remains stationary, only one event will be shown at time t0, and no data will be generated thereafter. In other words, an event camera only generates asynchronous pulse signals for the changing parts, and the data volume may be only dozens of KB in 10 seconds. Since an event camera is different from a traditional frame camera and outputs an asynchronous event pulse video data stream, it is necessary to classify and process such a data stream in terms of time and space. Specifically, in the time dimension, due to the different timings of light intensity changes of objects at different levels, the time when pulse video event data is generated is different, so the pulse video event data stream captured by the event camera is distributed in time sequence; in the space dimension, the light intensity changes of static background objects, dynamic background objects, and moving objects in the target area in the same space are different, and the attentions are also different. Therefore, it is necessary to segment and classify and identify the pulse video data containing multi-level information of the event camera from the spatio-temporal dimension.
[0052] It can be understood that in the autonomous driving scenario, the captured images of the surrounding environment obtained by the RSU or the vehicle may include various different types of objects, such as stationary objects (buildings, street lights, etc.), low-speed moving objects (pedestrians, etc.), and high-speed moving objects (cars, etc.). Due to the characteristics of the event camera, the pulse video data corresponding to different types of objects has different characteristics. Therefore, it is necessary to treat them differently in the processing.
[0053] For example, the spatio-temporal segmentation unit 101 is configured to: segment the space corresponding to the pulse video data according to a predetermined separation principle, and the predetermined separation principle may include one or more of the following: global or region of interest, the movement status of the region of interest; and for each segmented space, perform event accumulation based on the pulse video data with the time period granularity corresponding to the space.
[0054] In other words, the spatio-temporal segmentation unit 101 can perform segmentation in the current space dimension first and then in the time dimension. For example, in the space dimension, perform spatial semantic segmentation according to different classifications such as static background objects, dynamic background objects, and moving objects in the target area, so as to divide the global static space, local high-definition static space, and each moving single-body space. The moving single-body space can be further divided into a high-speed moving single-body space and a low-speed moving single-body space according to the speed of movement. Therefore, the pulse video data of the event camera is segmented into the global static space and its affiliated objects, the local high-definition static space and its affiliated objects, and the moving single-body space and its affiliated objects, which respectively correspond to the entire photosensitive pixel matrix of the dynamic vision sensor of the event camera, the photosensitive pixel matrix of the static region of interest (ROI) part, and the photosensitive pixel matrix of the dynamic ROI part. For the sake of understanding, Figure 2 A schematic example of the space segmentation is shown.
[0055] In addition, the spatio-temporal segmentation unit 101 can determine the time period granularity corresponding to each segmented space based on the light intensity change in the space and the motion state of the object in the space. Among them, when the light intensity change in the space is greater and the object moves faster, the time period granularity corresponding to the space is determined to be smaller. Specifically, in the time dimension, the event camera pulse data corresponding to different spaces after spatial segmentation can be accumulated according to different time periods to form pixel matrix data, that is, event image frames.
[0056] Figure 3 A schematic example of time segmentation is shown. Among them, the time period granularity of the high-speed moving single-body space is the smallest, while the time period granularity of the global static space is the largest. For example, long-time period accumulation is suitable for the recognition of static objects with small light intensity changes and few pulse event data, short-time period accumulation is suitable for the recognition of low-speed moving objects with relatively large but slow light intensity changes and more pulse event data, and extremely short-time period accumulation is suitable for the recognition of high-speed moving objects with large and fast light intensity changes and a large amount of pulse event data. Correspondingly, the electronic device 100 will subsequently send encoded messages for different spaces at different time granularities, that is, send the encoded messages obtained by encoding the semantic recognition results of the objects in the space according to the time period granularity of each space.
[0057] After the spatio-temporal segmentation unit 101 obtains the pulse video data corresponding to different categories of objects, the semantic recognition unit 102 performs semantic recognition on different categories of objects based on the pulse video data corresponding to different categories of objects respectively.
[0058] For example, the semantic recognition unit 102 can use the semantic recognition model corresponding to the object of each category in different categories to perform semantic recognition. Among them, the semantic recognition model can be an artificial intelligence (AI) model, such as a spiking convolutional neural network (SCNN) model. The spiking neural network has the characteristics of event-driven, asynchronous operation, and extremely low power consumption, and the generation of spiking signals is very compatible with the timestamp-based event stream output mode of the event camera. Therefore, in the example of this application, the SCNN model will be used as an example of the semantic recognition model for description, but it is not limited to this.
[0059] From the perspective of the classification of semantic recognition models, for example, the semantic recognition model may include one or more of the following: static dark object recognition model, static luminous object recognition model, static oscillating object recognition model (static oscillating object recognition filtering model), low-speed moving object recognition model, and high-speed moving object recognition model. Corresponding SCNN models can be established for each of these recognition models respectively. The SCNN model structure includes a pulse coding layer, a pulse convolution layer, a pulse pooling layer, and a fully connected layer. The SCNN model fully integrates the advantages of the SNN (Spiking Neural Network) model and the CNN (Convolutional Neural Network) model, with fast training and recognition speeds, while also reducing computing power consumption and saving a large amount of computing costs.
[0060] Figure 4 Fig. shows a schematic flowchart for spatio-temporal segmentation of pulsed video data and classification selection of semantic recognition models. First, the event camera obtains the original pulsed video data (also referred to as pulsed event data). Then, the pulsed video data is spatially segmented. For example, according to separation principles such as "global / ROI", "whether the ROI moves", etc., it is segmented into a global static space and its affiliated objects, a local high-definition static space and its affiliated objects, and a moving single-body space and its affiliated objects. This is because the semantic recognition model can be different for static and moving objects.
[0061] Next, time segmentation is performed, that is, event accumulation is carried out according to different time period granularities. For example, for the recognition of dynamic objects in an extremely short period, a high-speed moving object recognition model is selected.
[0062] For low-speed moving objects such as pedestrians, a low-speed moving object recognition model is selected. In addition to recognizing the object category through the external contour, in some scenarios, the motion posture of the human body can be further recognized, such as standing still, walking slowly, running fast, etc. At this time, ordinary object category AI recognition models are not applicable and need to be further optimized through reinforcement training, etc. As shown in the figure, a low-speed moving object posture recognition model is further selected.
[0063] For the recognition of static objects in a long cycle, it is also possible to further distinguish whether the object itself emits light. Static objects that emit light by themselves will bring obvious changes in light intensity, and there is a large amount of accumulated pulse event data, so that a relatively clear event image frame can be formed, and an AI recognition model (static luminous object recognition model) trained conventionally can recognize this type of object. However, for static objects that do not emit light, when the event camera is stationary, there will basically be no obvious change in light intensity, so the formed event image frame is relatively blurred. The AI recognition model can be further optimized through enhanced training and other methods to recognize this type of object. As shown in the figure, a static dark object recognition model is used. For static objects that do not emit light, it is also necessary to continue to distinguish whether the object will be affected by the surrounding environment such as wind, rain, and snow, causing the object to vibrate and shake, thus bringing obvious changes in light intensity. The AI recognition model for such static oscillating objects can also be further optimized through enhanced training and other methods. As shown in the figure, a static oscillating object recognition and filtering model is further selected. For static objects that do not emit light and are not affected by the surrounding environment, finally, different AI recognition models need to be selected according to the motion / static state of the event camera itself: when the event camera itself is in a moving state (such as deployed on an autonomous vehicle), there will be obvious changes in light intensity on the surface of the static object, and an AI recognition model trained conventionally can recognize this type of object; when the event camera itself is in a stationary state (such as deployed on an RSU), there will basically be no obvious change in light intensity on the surface of the static object, and the AI recognition model can be further optimized through enhanced training and other methods to recognize this type of object. As shown in the figure, a static dark object fuzzy recognition model is further selected.
[0064] It should be understood that the classification and selection of the semantic recognition model referred to above Figure 4 are only exemplary, and can be appropriately increased, decreased, or adjusted in actual applications.
[0065] For example, the static dark object recognition model can be used to recognize static objects that do not emit light by themselves, such as surrounding street buildings, pedestrians waiting by the roadside, vehicles parked by the roadside, and road lane lines. In addition, due to the self-motion of the event camera, optical flow is generated in the surrounding static background. Therefore, according to whether the event camera itself is moving, recognition models with different clarity can be further adopted.
[0066] The static luminous object recognition model is used to recognize objects that emit light by themselves and produce changes in light intensity, such as traffic lights, flashing street lights, neon lights in street windows, etc.
[0067] The static oscillating object recognition model is used to recognize surrounding background objects with periodic oscillating motion, such as street trees blown by the wind, falling raindrops / snowflakes, fluttering roadside flags, etc.
[0068] The low-speed moving object recognition model recognizes low-speed moving objects, such as pedestrians crossing the road on a zebra crossing, vehicles moving at low speeds on a lane, non-motor vehicles on a non-motor vehicle lane, etc.; if it is necessary to recognize different motion postures of an object, it is also necessary to further define different posture recognition models.
[0069] The high-speed moving object recognition model is used to recognize high-speed moving objects, such as high-speed projectiles, falling rocks in mountainous areas, motor vehicles, non-motor vehicles, pedestrians or animals crossing at high speeds at a fork in the road, etc.
[0070] Table 1 below shows an example of a selection table of semantic recognition models in the case where an event camera takes pictures of a traffic scene. The left side shows the semantic models, where categories I and II represent traffic classifications, the recognition subjects are the various objects to be recognized, and the features represent the parameter items used to describe the corresponding objects. The right side shows the selection of recognition models, which are divided into three types: long period, short period, and extremely short period. The specific recognition models have been described above. The mark √ represents that the recognition subject corresponding to the row where it is located uses the recognition model corresponding to the column where it is located for recognition. It should be understood that Table 1 is only exemplary and not restrictive.
[0071] Table 1
[0072]
[0073]
[0074]
[0075]
[0076] The semantic recognition unit 102 can accurately recognize each object and the related features of the object (also referred to as semantic feature items in the following) by using the semantic recognition model of the corresponding category. The semantic encoding unit 103 uses the object for which semantic recognition has been performed as the semantic model recognition subject, and generates an encoding message for the semantic model recognition subject, and the encoding message includes the semantic feature items of the semantic model recognition subject. In other words, an encoding message can be defined for each semantic model recognition subject to indicate its respective semantic features.
[0077] As mentioned above, event accumulation is performed with different time period granularities for different spaces. Therefore, the time period for the semantic recognition unit 102 to perform semantic recognition on objects in different spaces is also different, and the time period granularity of the obtained semantic recognition results is also different. Correspondingly, the time period granularity of the encoding message obtained after the semantic encoding unit 103 encodes the semantic recognition results will also be different due to the different spaces where the objects are located. The electronic device 100 (for example, its communication unit 104, such as Figure 5As shown, the encoded message obtained by encoding the semantic recognition result of the object in each space can be sent according to the time - period granularity of each space.
[0078] For example, each semantic feature item can be encoded from the top - down of the semantic encoding tree. The semantic encoding tree is hierarchically encoded from the root node to the leaf node from top - down, which can ensure the uniqueness of encoding. In the following, hexadecimal encoding will be used as an example for description.
[0079] In one example, the event camera captures a traffic scene. The semantic encoding tree includes the following levels from top - down: the first traffic classification, the second traffic classification, the semantic model recognition subject, and the semantic feature item.
[0080] It can be seen that the levels in this example respectively correspond to the items of the semantic model in Table 1. Among them, the first traffic classification includes: traffic map, traffic signal, traffic participants, and road obstacles. The second traffic classification includes: urban streets, mountain roads, road intersections, and parking lots under the traffic map category; signal lights and traffic signs under the traffic signal category; motor vehicles, non - motor vehicles, and pedestrians under the traffic participant category; and immovable obstacles and movable obstacles under the road obstacle category. The semantic model recognition subject includes the objects under each second traffic classification, such as buildings, trees, street lights, etc. on urban streets; zebra crossings, road lane lines, etc. at road intersections; motor vehicle signal lights, pedestrian crossing signal lights, etc.; cars, fire trucks, etc.; walking pedestrians, non - motor vehicle drivers pushing their vehicles after getting off, etc. The semantic feature item includes the feature name, value length, and value of the semantic model recognition subject (i.e., using the TLV structure). Taking a walking pedestrian as an example, its semantic feature item can include height, speed, position coordinates, posture, carried items, relevance, etc.
[0081] Among them, the code length of the first traffic classification can be 2 bits, the code length of the second traffic classification can be 6 bits, the code length of the semantic model recognition subject can be 8 bits, the code length of the feature name in the semantic feature item can be 8 bits, the code length of the value length (for example, the number of bytes representing the value) can be 8 bits, and the code length of the value can be 0 to 255 bytes. The schematic diagram of the formed codeword is as Figure 6 shown. It should be understood that the code lengths of each field are not restrictive, and the semantics corresponding to different codeword values are not restrictive.
[0082] Figure 7 shows a schematic example of the semantic encoding tree. Among them, the first traffic classification is shown as traffic classification - I class in the figure, and the second traffic classification is shown as traffic classification - II class in the figure. In Figure 7In the example, some semantic feature items include the correlation between the semantic model recognition subject and other semantic model recognition subjects, and the information of this correlation can be used to optimize and adjust the semantic information of the overall space. The semantic encoding unit 103 can judge whether there is a correlation between two or more objects, and add a correlation semantic feature item to the corresponding object when it is judged that there is a correlation. As Figure 7 shown, the correlation can include, for example, movement accompaniment, space occupancy, etc.
[0083] In this way, when the receiving end performs semantic decoding, after extracting the correlation semantic feature items of multiple moving objects in the same static space, it can perform analysis and processing. According to the correlation of multiple moving objects, it can comprehensively judge whether the decoded semantic information is correct, and if there is an error, it can perform error tolerance processing, which will be described in detail later.
[0084] Table 2 below shows an example of a semantic encoding table in a traffic scenario (for example, the electronic device 100 is located in a vehicle or RSU in V2X).
[0085] Table 2
[0086]
[0087]
[0088]
[0089]
[0090] Among them, the encodings in parentheses are the encodings written in hexadecimal form. Taking 0xFFFF 0xFF in the first row as an example, it represents the position coordinates (x, y) of a building in the urban street category under the traffic map category.
[0091] In addition, the encoded message can also include a field for the vehicle networking application scenario. This field can be placed, for example, at the front of the encoded message. The vehicle networking application scenario can include: urban street blind area perception sharing scenario, mountain area high-altitude falling object perception sharing scenario, highway high-speed obstacle avoidance perception sharing scenario, tunnel exit blind area perception sharing scenario, but is not limited to this. Table 3 below shows an encoding example of the vehicle networking application scenario.
[0092] Table 3
[0093] Serial number V2X scenario Encoding 1 Blind spot perception sharing for urban streets 0xFFFF 2 High altitude falling object perception sharing in mountainous areas 0xFFFE 3 Collision avoidance perception sharing on highway expressways 0xFFFD 4 Blind spot perception sharing at tunnel exits 0xFFFC … … …
[0094] In the scenarios listed in Table 3, the advantages of the event camera can be fully demonstrated. For example, in order to save computing power, the operations of the electronic device 100 or the semantic recognition and encoding operations therein can be performed only when the vehicle is in the above scenarios. The electronic device 100 can, for example, identify the above vehicle networking application scenarios based on the position on the map.
[0095] In addition, the semantic encoding unit 103 is further configured to determine the semantic model recognition subject for generating the encoded message and the semantic feature items to be included according to the V2X application scenario. In other words, it may not be necessary to encode all the identified objects in the space, but only the objects in the area of interest or the moving objects. On the other hand, for the objects that need to be encoded, it may not be necessary to include all the semantic feature items, but only the semantic feature items of interest can be selected. The semantic encoding unit 103 can select the objects that need to be encoded and the semantic feature items to be encoded according to the V2X application scenario, thereby reducing the overhead of the encoded message.
[0096] Table 4 below shows an example of the V2X scenario semantic model mapping table, which shows examples of the semantic models corresponding to different V2X application scenarios.
[0097] Table 4
[0098]
[0099]
[0100]
[0101] In addition, when the electronic device 100 is located at a V2X vehicle or RSU, the communication unit 104 can also expand the vehicle networking message set to add the DayIII Ext Msg message body that supports semantic communication. For example, the above encoded message can be included in the DayIII Ext Msg message body.
[0102] Specifically, the third-phase expansion can be carried out based on the interaction messages defined in standards and specifications such as YD / T 3709-2020 "Message Layer Technical Requirements for LTE-Based Vehicle Networking Wireless Communication Technology", CSAE XXXX, and C-ITS XXXX "Cooperative Intelligent Transportation System Vehicle Communication System Application Layer and Application Data Interaction Standard Phase II". Figure 8The example of extending the DayIII Ext Msg message body MsgFrameIII is shown. In MsgFrameIII, four types of messages are defined: traffic map Msg_SC_Map, traffic signal Msg_SC_Signal, traffic participant Msg_SC_Participant, road obstacle Msg_SC_Obstacle, and other extensible items Msg_SC_Test.
[0103] To sum up, the electronic device 100 according to this embodiment can distinguish objects of different categories and perform semantic recognition and semantic encoding of objects of different categories by performing spatiotemporal segmentation on the pulse video data obtained by the event camera, which can greatly reduce the amount of data transmission and improve the accuracy of information transmission, reduce the processing delay of end-to-end perception communication recognition, reduce computing power requirements, and better support vehicle network application scenarios.
[0104] <Second embodiment>
[0105] Figure 9 FIG. 1 shows a functional module block diagram of an electronic device 200 according to another embodiment of the present application. Figure 9 As shown, the electronic device 200 includes: a communication unit 201, configured to receive coded messages from other electronic devices, where the coded messages are obtained by the other electronic devices by: performing spatiotemporal segmentation on the pulse video data obtained by the event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on the objects of different categories based on the pulse video data corresponding to the objects of different categories; and a decoding and semantic reorganization unit 202, configured to perform semantic decoding on the coded messages and perform semantic fusion reorganization (hereinafter also referred to as semantic reorganization) to obtain the semantics of the original scene captured by the event camera.
[0106] The communication unit 201 and the decoding and semantic reassembly unit 202 may be implemented by one or more processing circuits and at least one memory. The processing circuit may be implemented as a chip, a processor, etc., and the at least one memory may be any form of storage device such as RAM, ROM, flash memory, etc. The at least one memory is used to store computer program codes and data required for the processing circuit to perform processing. It should be understood that Figure 9 The various functional units in the electronic device shown in the figure are merely logical modules divided according to the specific functions they implement, and are not intended to limit the specific implementation method. In addition, the functions of the decoding and semantic reorganization unit 202 can also be implemented using a brain-like chip, which is not restrictive.
[0107] When the electronic device 200 is used in an autonomous driving scenario, the electronic device 200 can be set on the vehicle side. The vehicle described here can more generally refer to various user devices located on the vehicle. The electronic device 200 in this embodiment can act as a data requester and as a decoding device.
[0108] It should also be noted that the electronic device 200 can be implemented at the chip level, or it can also be implemented at the device level. For example, the electronic device 200 can work as a vehicle or user device itself, and can also include external devices such as memory, transceivers (not shown in the figure), etc. The memory can be used to store programs and related data information that need to be executed to implement various functions of the vehicle or user device. The transceiver may include one or more communication interfaces to support communication with different devices (for example, other vehicles or user devices, etc.), and the implementation form of the transceiver is not specifically limited here.
[0109] However, this is not restrictive, and the electronic device 200 may also be set on the network side, such as a base station, a roadside unit, etc.
[0110] The encoded message here can, for example, be generated and sent by the electronic device 100 in the first embodiment. For example, the encoded message can include semantic feature items of a semantic model recognition subject, where the semantic model recognition subject is the object that has undergone semantic recognition. Each semantic feature item can be encoded from top to bottom according to a semantic encoding tree. When an event camera captures a traffic scene, the semantic encoding tree can include the following layers from top to bottom: a first traffic category, a second traffic category, a semantic model recognition subject, and semantic feature items. The first traffic category includes: traffic maps, traffic signals, traffic participants, and road obstacles; the second traffic category includes: urban streets, mountain roads, road intersections, and parking lots under the traffic map category; traffic lights and traffic signs under the traffic signal category; motor vehicles, non-motor vehicles, and pedestrians under the traffic participant category; and immovable obstacles and movable obstacles under the road obstacle category. The semantic model recognition subject includes objects under each second traffic category; and the semantic feature items include the feature name, value length, and value of the semantic model recognition subject. The semantic feature items may also include correlations between the semantic model recognition subject and other semantic model recognition subjects. The encoded message may be included in the DayIII Ext Msg message body, which is an extension of the vehicle network message set to support semantic communication. The specific details have been described in detail in the first embodiment, which is also applicable to this embodiment and will not be repeated here.
[0111] The communication unit 201 receives multiple encoded messages. For example, the decoding and semantic reassembly unit 202 is configured to: align encoded messages of different categories of objects in time, where the encoded messages of different categories of objects are sent at different time period granularities; and combine the semantically decoded semantic messages according to spatial positions to obtain the complete semantics of the original scene.
[0112] Figure 10 The following illustrates an exemplary operation of the decoding and semantic reassembly unit 202. Specifically, it decodes and restores semantic information for high-speed motion entities with extremely short sampling periods, such as a time period granularity of 0.1 seconds; decodes and restores semantic information for low-speed motion entities and stationary luminous objects with short sampling periods, such as a time period granularity of 1 second; decodes and restores semantic information for stationary objects within a local stationary space with a long sampling period, such as a time period granularity of 10 seconds; and decodes and restores semantic information for stationary objects within a global stationary space with a long sampling period, such as a time period granularity of 30 seconds. For example, the decoding and semantic reassembly unit 202 can determine the time period granularity used to align encoded messages based on current needs, such as scenario requirements. For example, for low-speed motion object recognition and analysis scenarios, alignment can be performed at a 1-second granularity. In this case, the semantic information for high-speed motion entities with a 0.1-second period granularity is accumulated to 1 second (only the most recent one can be retained), while the semantic information for stationary objects with long periods of 10 seconds or 30 seconds can serve as shared background semantics for multiple low-speed motion object recognitions. For high-speed motion object recognition and analysis scenarios, alignment can be performed at a 0.1-second granularity, and so on.
[0113] For example, the decoding and semantic reorganization unit 202 may perform semantic decoding with reference to the following Table 6. It should be understood that Table 6 is only an illustrative example, which may correspond to Table 2 in the first embodiment.
[0114] Table 6
[0115]
[0116]
[0117]
[0118] In addition, the decoding and semantic reassembly unit 202 may perform scene decoding with reference to the following Table 7. It should be understood that Table 7 is only an illustrative example, which may correspond to Table 3 in the first embodiment.
[0119] Table 7
[0120]
[0121]
[0122] Next, the decoding and semantic recombination unit 202 combines the decoded semantic information according to the spatial position. For example, the decoding and semantic recombination unit 202 combines the overall semantic information of all spaces in the order of high-speed moving monomers, low-speed moving monomers, stationary luminous objects, stationary objects in a local stationary space, and stationary objects in a global stationary space according to the coordinate position information of different objects in the pulse event image frame.
[0123] The decoding and semantic recombination unit 202 can also perform semantic error tolerance processing based on the correlation information in the semantic feature items. The semantic error tolerance processing can include the correction of incorrect semantics and the recovery of lost semantics. For example, the decoding and semantic recombination unit 202 can judge whether there is a correlation between two or more objects according to the correlation information. If the judgment is yes, correlation analysis is performed. For example, the correlation can include space occupancy, movement accompaniment, etc. The decoding and semantic recombination unit 202 can also judge whether there is a semantic error or loss based on the correlation analysis. If so, semantic error tolerance processing can be performed, such as using the semantic information of other related objects to recover or correct the lost or incorrect semantics.
[0124] In addition, the electronic device 200 may further include an analysis and prediction unit 203, configured to analyze and process the semantics of the obtained original scene using a scene analysis and prediction model to obtain risk assessment and / or prediction warning information.
[0125] For example, the analysis and processing here can be intelligent analysis and processing. The analysis and prediction unit 203 can select a suitable artificial intelligence analysis model as the scene prediction analysis model according to the scene type. When the electronic device 200 is located in a vehicle in V2X, the scene type may include vehicle networking application scenarios, and the vehicle networking application scenarios may include, for example: urban street blind spot perception sharing scenarios, mountain area high-altitude falling object perception sharing scenarios, highway high-speed obstacle avoidance perception sharing scenarios, tunnel exit blind spot perception sharing scenarios.
[0126] The analysis and prediction unit 203 can determine the scene type based on the vehicle networking scene knowledge stored in the background knowledge base. The vehicle networking scene knowledge includes, for example, knowledge about urban streets, road intersections, etc. In addition, this vehicle networking scene knowledge can also be used for semantic recombination in the decoding and semantic recombination unit 202, as Figure 12 shown.
[0127] In addition, the encoded message may also include a field for vehicle networking application scenarios. The analysis and prediction unit 203 can determine the scene type based on such a field.
[0128] Figure 13A schematic diagram showing the scenario analysis prediction and early warning performed by the analysis and prediction unit 203 is presented, which shows an example of an AI analysis and prediction model and the information on the provided analysis and prediction, risk assessment, and perception sharing early warning.
[0129] Table 8 below shows a V2X scenario analysis model table of the correspondence between V2X scenarios and applicable AI analysis and prediction models. It should be understood that this table is also merely exemplary.
[0130] Table 8
[0131] Serial number V2X scenario Analysis model 1 Blind spot perception sharing for urban streets Vehicle / pedestrian / non-motor vehicle movement trajectory tracking and prediction model 2 Blind spot perception sharing for urban streets Human body posture recognition and behavior prediction analysis model 3 High altitude falling object perception sharing in mountainous areas Mountain rockfall tracking and recognition and danger prediction model 4 Collision avoidance perception sharing on highway expressways High speed parabolic tracking and recognition and danger prediction model 5 Blind spot perception sharing at tunnel exits High exposure blind spot detection model … … …
[0132] For example, after determining the scenario type, the analysis and prediction unit 203 selects the corresponding AI analysis and prediction model, and imports the restored scenario semantic information into the AI analysis and prediction model for AI analysis processing, so as to obtain valuable information such as various risk assessments and prediction early warnings. It should be understood that the scenario analysis and prediction model is not necessarily an AI analysis and prediction model, and can also be a conventional analysis and prediction model, which is not restrictive.
[0133] In addition, the communication unit 201 can also provide the obtained risk assessment and / or prediction early warning information to other devices such as surrounding vehicles. For example, when the electronic device 100 is located on the perception vehicle side and the electronic device 200 is located on the network side, by performing the above semantic recombination and AI analysis and prediction on the network side and directly providing the prediction early warning information to the vehicles in need, the requirement for the computing power of the vehicles can be reduced.
[0134] On the other hand, when the vehicle where the electronic device 200 is located can work as a data provider, the electronic device 200 can also include units similar to the spatio-temporal segmentation unit 101 and the semantic recognition unit 102 in the electronic device 100 and the analysis and prediction unit 203, and perform scenario analysis and prediction based on the result of semantic recognition.
[0135] For ease of understanding, Figure 14 A schematic diagram showing the operations of the transmitting end and the receiving end when the solution of the present application is applied to the V2X scenario is presented. As Figure 14As shown, the transmitting V2X platform can be set on the V2X-RSU side or the sensing vehicle side, and the receiving V2X platform can be set on the side of the V2X perimeter warning vehicle, that is, the vehicle side with data requirements. An event camera and various chips for execution processing, such as traditional computing chips like AI chips and CPUs, can be set on the transmitting side. Various chips for execution processing, such as traditional computing chips like AI chips and CPUs, can be set on the receiving side. The event camera on the transmitting side obtains pulse event data and provides it to the AI chip for pulse video semantic recognition. The obtained semantic recognition results are used for semantic encoding, and the encoded message obtained through semantic encoding can be subjected to V2X channel encoding. The finally obtained encoding result is sent to the receiving side through C-V2X communication. Note that various existing V2X channel encoding methods can be adopted here without limitation. The receiving side performs V2X channel decoding on the received encoding result (using the V2X channel decoding method corresponding to the V2X channel encoding method adopted by the transmitting side), and further performs semantic decoding on the obtained semantic encoding message after decoding, so as to restore the original scene semantics. The AI chip can also be used to perform scene semantic analysis and prediction.
[0136] Among them, the pulse video data obtained by the event camera can be for specific V2X application scenarios. For example, the transmitting side can determine which scenario the current scene belongs to according to various V2X scene knowledge stored in the background knowledge base, or the user can specify the scenario. The transmitting side selects the corresponding semantic recognition model by referring to the semantic recognition model selection table (for example, Table 1), and performs semantic recognition on the event data subjected to spatio-temporal segmentation using the selected semantic recognition model. Next, the transmitting side can select the semantic model recognition subject and semantic feature items by scenario for semantic encoding, and this process can be executed by referring to the scene semantic model mapping table (for example, Table 4), the scene encoding table (for example, Table 3), and the semantic encoding table (for example, Table 1). Correspondingly, the receiving side can also perform decoding of the semantic model recognition subject and semantic feature items by referring to the semantic decoding table (for example, Table 6) and the scene decoding table (for example, Table 7), then perform scene semantic recombination, and perform scene analysis and prediction to obtain possible warning information. In the scene analysis and prediction, the AI analysis and prediction model to be used can be selected by referring to the scene analysis model table (for example, Table 8).
[0137] In addition, although the above describes that the electronic device 200 includes the analysis and prediction unit 203, it is not limited to this. The electronic device 100 described in the first embodiment can also include the analysis and prediction unit 203 to perform scene analysis and prediction based on the result after semantic recognition.
[0138] To sum up, according to this embodiment, the electronic device 200 can perform semantic decoding and scene semantic reorganization on the received semantically coded messages, and can perform scene analysis and prediction based on scene semantics, thereby realizing end-to-end perception communication recognition with low processing delay, greatly reducing computing power requirements, improving the accuracy of information transmission, and better supporting Internet of Vehicles application scenarios.
[0139] <Third embodiment>
[0140] For the blind spot perception sharing scenario on urban streets, the electronic device 200 of this embodiment can realize vehicle / pedestrian / non-motor vehicle movement trajectory tracking and prediction based on the coded message provided by the electronic device 100 or the electronic device 100, such as non-motor vehicles crossing the road and running a red light, non-motor vehicles entering the motor vehicle lane, etc., as well as human posture recognition and behavior prediction analysis, such as prediction of dangerous postures of surrounding non-motor vehicle personnel (such as making and answering mobile phone calls while riding a bicycle, drinking beverages and eating while riding a bicycle, riding a bicycle with an umbrella in one hand, illegally carrying people / objects), analysis of the driver's human posture / gesture / dangerous driving behavior (such as making and answering calls, eating, bending over to pick up objects), and recognition and evidence collection of vehicle passengers' throwing gestures.
[0141] For the shared scenario of high-altitude falling object perception in mountainous areas, the electronic device 200 of this embodiment can perform vehicle high-speed obstacle avoidance, high-speed parabolic tracking and identification, and hazard prediction based on the coded message provided by the electronic device 100 or the electronic device 100. High-speed parabolic objects have a fast speed, a small target volume, and a great risk.
[0142] For the blind spot perception sharing scenario at the tunnel exit, the electronic device 200 of this embodiment can perform strong light blind spot detection in scenarios where light intensity mutation is more obvious based on the coded message provided by the electronic device 100 or the electronic device 100. Sudden changes in light brightness and darkness, backlighting, etc., such as blinding by strong light from the opposite side when meeting other vehicles, and vehicles also face high exposure scenarios after coming out of the tunnel. This scenario also involves other surrounding traffic participants (including but not limited to vehicles, pedestrians, cyclists and other targets), road abnormality information such as: road traffic events (such as traffic accidents, etc.), abnormal vehicle behavior (speeding, leaving the lane, driving in the wrong direction, irregular driving and abnormal stillness, etc.), road obstacles (such as fallen rocks, scattered objects, dead branches, etc.) and road conditions (such as accumulated water, ice, etc.).
[0143] In this embodiment, an example of semantic recognition and encoding and decoding operations in a blind spot perception sharing scenario on urban streets is given, and other scenarios can be similarly applied.
[0144] First, spatiotemporal segmentation of impulse video data is performed at the transmitting end (eg, electronic device 100 ). Figure 15 A diagram of impulse video data after spatial segmentation is shown. Figure 16 An example of time slices of each space obtained after time division is shown.
[0145] Then, corresponding semantic recognition models are determined for different categories of objects in different spaces, as shown in Table 9 below.
[0146] Table 9
[0147]
[0148]
[0149]
[0150] Then, semantic encoding is performed on the results of semantic recognition. For example, the semantic recognition results obtained through different semantic models can be semantically encoded respectively according to different time period granularities. Specifically, the static street background can be semantically encoded according to a coarse time period granularity, the different postures of pedestrians crossing the road can be semantically encoded according to a fine time period granularity, and the moving background objects such as background vehicles in non-concerned areas can be not encoded.
[0151] Among them, the corresponding scene can be encoded according to the scene encoding table. In this embodiment, the scene is urban street blind area perception sharing, and the encoding should be 0xFFFF according to Table 3. In addition, according to the scene, the semantic model recognition subject and the corresponding semantic feature items are selected according to the V2X scene semantic model mapping table (for example, Table 4) and the semantic recognition model selection table (for example, Table 9).
[0152] Next, different semantic models will be discussed with the second traffic classification as an example. For semantic model 1: urban street perception sharing scene - traffic map - urban street, the selection is as follows.
[0153]
[0154] The encoding is as follows:
[0155]
[0156] That is, the semantic encoding of the urban street is {0xFFFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFE 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFD 0xFF 0x08 0x0000 0x00000x0000 0x0000}, and the meaning of the encoding is: there are buildings, trees, and street lamps on the urban street, where the position of building 1 is (x1, y1), the position of tree 1 is (x1, y1), and the position of street lamp 1 is (x1, y1).
[0157] For semantic model 2: Urban street perception sharing scenario - traffic map - road intersection, the selection is as follows.
[0158]
[0159] The encoding is as follows:
[0160]
[0161]
[0162] That is, the semantic encoding of the road intersection is {0xFDFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFF 0x10 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFE 0x1B 0x0000 0x0000 0x0000 0x0000 0x02 0x0000 0x0000 0x0000 0x00000x02 0x0000 0x0000 0x0000 0x0000 0x02,0xFDFB 0xFF 0x20 0x0000 0x0000 0x00000x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x00000x0000 0x0000,0xBFFF 0xFF 0x01 0x02,0xBFFE 0xFF 0x01 0x00}. The encoding meaning is: The center position coordinates of the road intersection are (x, y); there is a pedestrian zebra crossing at the intersection, with the position from coordinates (x_from, y_from) to (x_to, y_to), and three object spaces occupied by (x1, y1), (x2, y3), (x3, y3); there are 2 lanes, the position of lane 1 is from coordinates (x1_from, y1_from) to (x1_to, y1_to), and the position of lane 2 is from coordinates (x2_from, y2_from) to (x2_to, y2_to).
[0163] For semantic model 3: Urban street perception sharing scenario - traffic signal - signal lamp, the selection is as follows.
[0164]
[0165] The encoding is as follows:
[0166]
[0167]
[0168] That is, the semantic encoding of the signal lamp is {0xFDFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000, 0xFFFC 0xFF 0x10 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000, 0xFFFC 0xFE 0x1B 0x0000 0x0000 0x0000 0x0000 0x02 0x0000 0x0000 0x0000 0x0000 0x02 0x0000 0x0000 0x0000 0x0000 0x02, 0xFDFB 0xFF 0x20 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000, 0xBFFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000, 0xBFFF 0xFE 0x01 0x02, 0xBFFE 0xFF 0x08 0x0000 0x0000 0x0000 0x0000, 0xBFFE 0xFE 0x01 0x00}, and the encoding meaning is: The motor vehicle signal lamp indicates to wait; the crosswalk signal lamp indicates that it is possible to pass.
[0169] For semantic model 4: Urban street perception sharing scenario - traffic participants - motor vehicle, the following is selected.
[0170]
[0171] The encoding is as follows:
[0172]
[0173] That is, the semantic encoding of a car (in motion) is {0x7FFF 0xFF 0x01 0x01, 0x7FFF 0xFF 0x18 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000}, and the encoding meaning is: The vehicle is driving at a low speed, trajectory: longitude and latitude coordinates 1 (x1, y1, t1), coordinates 2 (x2, y2, t2).
[0174]
[0175] In addition, there is another car as the recognition subject in the figure, and the coding is as follows:
[0176] That is, the semantic coding of the car (decelerating and stopping) is {0x7FFF 0xFF 0x01 0x00, 0x7FFF 0xFF 0x0C 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000}, and the coding meaning is: the vehicle decelerates and stops, with longitude and latitude coordinates (x, y), waiting.
[0177] For semantic model 5: Urban street perception sharing scenario - traffic participant - pedestrian, the selection is as follows.
[0178]
[0179] The coding is as follows:
[0180]
[0181]
[0182] That is, the pedestrian semantic encoding is {{0x7DFF 0xFF 0x01 0x00,0x7DFF 0xFE 0x01 0x00,0x7DFF0xFD 0x08 0x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x010x01,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x0000 0x01 0x0000 0x0000 0x00000x0000 0x01}{0x7DFF 0xFF 0x01 0x01,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x080x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x01 0x00,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x0000 0x01 0x0000 0x0000 0x0000 0x00000x01}{0x7DFF 0xFF 0x01 0x00,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x08 0x00000x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x01 0x01,0x7DFF 0xFA0x12 0x0000 0x0000 0x0000 0x0000 0x01 0x0000 0x0000 0x0000 0x0000 0x01}}, and the encoding meaning is: Pedestrian 1, tall, standing on one side of the road, position coordinates {(x1, y1)}, carrying a suitcase, correlation: is in a father-son relationship with (x2, y2) and in a husband-wife relationship with (x3, y3); Pedestrian 2, short, standing on one side of the road, position coordinates {(x2, y2)}, correlation: is in a moving companion relationship with (x1, y1) and in a moving companion relationship with (x3, y3); Pedestrian 3, tall, standing on one side of the road, position coordinates {(x3, y3)}, carrying a suitcase, correlation: is in a moving companion relationship with (x1, y1) and in a moving companion relationship with (x2, y2).
[0183] The receiving end (e.g., electronic device 200) receives the above encoded messages from the sending end and performs decoding and semantic recombination. Figure 17A schematic diagram of semantic decoding is shown. Among them, the received encoded message is on the left, the structure of the corresponding encoding tree is in the middle, and the decoded semantic information is on the right. It can be seen that in the encoded message of this example, it also includes the encoding of the vehicle networking application scenario, that is, OxFFFF at the upper left, representing the urban street blind spot perception sharing scenario.
[0184] For semantic model 1: Urban street perception sharing scenario - traffic map - urban street, the relevant information transmitted is as follows:
[0185]
[0186] The decoding is as follows:
[0187]
[0188] That is, the received semantic encoding is {0xFFFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFE 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFD 0xFF 0x08 0x0000 0x00000x0000 0x0000}, and the decoded semantic data is: urban street, building (x1, y1), tree (x2, y2), street lamp (x3, y3).
[0189] For semantic model 2: Urban street perception sharing scenario - traffic map - road intersection, the relevant information transmitted is as follows:
[0190]
[0191]
[0192] The decoding is as follows:
[0193]
[0194] That is, the received semantic encoding is {0xFDFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFF 0x10 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFE 0x1B 0x0000 0x0000 0x0000 0x0000 0x02 0x0000 0x0000 0x0000 0x00000x02 0x0000 0x0000 0x0000 0x0000 0x02,0xFDFB 0xFF 0x20 0x0000 0x0000 0x00000x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x00000x0000 0x0000,0xBFFF 0xFF 0x01 0x02,0xBFFE 0xFF 0x01 0x00}, and the decoded semantic data is as follows: the coordinates of the center position of the road intersection are (x, y); there is a pedestrian crosswalk at the intersection, with the position ranging from the coordinates (x_from, y_from) to (x_to, y_to), and three object spaces at (x1, y1), (x2, y3), and (x3, y3) are occupied; there are 2 lanes, with lane 1 ranging from the coordinates (x1_from, y1_from) to (x1_to, y1_to), and lane 2 ranging from the coordinates (x2_from, y2_from) to (x2_to, y2_to).
[0195] For semantic model 3: urban street perception sharing scenario - traffic signal - signal lamp, the relevant information transmitted is as follows:
[0196]
[0197] The decoding is as follows:
[0198]
[0199] That is, the received semantic encoding is {0xFDFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFF 0x10 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFE 0x1B 0x0000 0x0000 0x0000 0x0000 0x02 0x0000 0x0000 0x0000 0x00000x02 0x0000 0x0000 0x0000 0x0000 0x02,0xFDFB 0xFF 0x20 0x0000 0x0000 0x00000x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x00000x0000 0x0000,0xBFFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xBFFF 0xFE 0x010x02,0xBFFE 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xBFFE 0xFE 0x01 0x00}, and the decoded semantic data is: The motor vehicle signal light indicates to wait; The pedestrian crossing signal light indicates that it is possible to pass.
[0200] For semantic model 4: Urban street perception sharing scenario - traffic participant - motor vehicle, the relevant information transmitted is as follows:
[0201]
[0202] The decoding is as follows:
[0203]
[0204] That is, the received semantic encodings are {0x7FFF 0xFF 0x01 0x01,0x7FFF 0xFF 0x18 0x00000x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000} and {0x7FFF 0xFF 0x01 0x00,0x7FFF 0xFF 0x0C 0x0000 0x0000 0x0000 0x00000x0000}, and the decoded semantic data is: A car is moving at a low speed, with a trajectory: longitude and latitude coordinates 1 (x1, y1, t1), coordinate 2 (x2, y2, t2); and another car: the vehicle is decelerating and stopping, with longitude and latitude coordinates (x, y), waiting.
[0205] For semantic model 5: Urban street perception and sharing scenario - traffic participant - pedestrian, the relevant information transmitted is as follows:
[0206]
[0207]
[0208] The decoding is as follows:
[0209]
[0210] That is, the received semantic encodings are {0x7DFF 0xFF 0x01 0x00,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x08 0x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF0xFB 0x01 0x01,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x0000 0x01 0x00000x0000 0x0000 0x0000 0x01}……, and the decoded semantic data is: Pedestrian 1, height (tall), speed (low), position (x1, y1), posture (walking), carrying item (suitcase), correlation: has a moving accompaniment relationship with (x2, y2), has a moving accompaniment relationship with (x3, y3)…….
[0211] Next, reorganize according to the scenario semantics, and perform scenario analysis, prediction, and warning. An example of the overall architecture is as Figure 18 shown.
[0212] Specifically, the semantic messages with different time - period granularities (long - period, short - period, extremely short - period) can be aligned according to time first, and then the complete semantic information can be combined by space. Figure 19 An example of semantic recombination is shown.
[0213] Therefore, the following information after semantic recombination based on the background knowledge base can be obtained: In the urban street scene, the background objects include buildings (x1, y1), trees (x2, y2), street lamps (x3, y3), and there is a zebra crossing at the road intersection starting from the coordinate (x4_from, y4_from) to the position (x4_to, y4_to); the scene - aware vehicle is a sedan car, which has decelerated and stopped, and its position is (x5, y5)); 3 pedestrians appear at the starting position of the zebra crossing on the left - front side, and their position coordinates are {(x1, y1), (x2, y2), (x3, y3)}, standing on one side of the zebra crossing, and 2 of them are carrying suitcases.
[0214] During the transmission of the encoded message, partial information loss or error may occur. According to the embodiments of the present application, the lost or incorrect information can be restored by using the correlation information items during the semantic recombination process.
[0215] For example, in Figure 19In the example shown, if, for the semantic model 5 regarding pedestrians, the semantic codes are received: {{0x7DFF 0xFF 0x01 0x00,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x08 0x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x01 0x01,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x00 0x0000 0x0000 0x0000 0x0000 0x01}{0x7DFF 0xFF 0x01 0x00,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x08 0x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x01 0x01,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x00 0x0000 0x0000 0x0000 0x0000 0x02}}. The decoded semantic data is: Pedestrian 1, height (tall), speed (low), position (x1, y1), posture (walking), carried item (suitcase), correlation: in a moving accompaniment relationship with (x2, y2), in a moving accompaniment relationship with (x3, y3); Pedestrian 3, height (tall), speed (low), position (x3, y3), posture (walking), carried item (suitcase), correlation: in a moving accompaniment relationship with (x1, y1), in a moving accompaniment relationship with (x2, y2).
[0216] Among them, the data of Pedestrian 2 is missing. In this example, the semantic correlation between symbols can be used to help the receiving end recover the missing information. For example, the data of Pedestrian 2 can be recovered by performing correlation recognition and semantic fault tolerance processing on the data of Pedestrian 1 and Pedestrian 3. The obtained semantic data of Pedestrian 2 is: 3 pedestrians appear on the zebra crossing in the upper left front, walking at a low speed, position coordinates {(x1, y1), (x1, y1), (x3, y)}), and 2 of them carry suitcases. Therefore, through this correlation recognition and semantic fault tolerance processing, the correct rate of information transmission can be improved.
[0217] As Figure 18 shown, after the scene semantic recombination, an appropriate AI analysis model can be selected according to the scene to further obtain analysis and prediction information. For example, in the above example, the scene is the perception and sharing of blind spots in urban streets. According to the scene analysis model table (for example, Table 8 in the second embodiment), the following analysis model can be selected.
[0218] Serial number V2X scenario Analysis model 1 Blind spot perception sharing for urban streets Vehicle / pedestrian / non-motor vehicle movement trajectory tracking and prediction model 2 Blind spot perception sharing for urban streets Human body posture recognition and behavior prediction analysis model
[0219] After analyzing the semantic data using the selected analysis model, for example, the following prediction, warning, and evaluation information can be obtained. That is, pedestrian posture analysis prediction: The pedestrian standing on the left side of the crosswalk has a tendency to cross the road; pedestrian risk assessment: Low risk (low speed, 2 people dragging suitcases, and other normal postures); surrounding vehicle perception and sharing warning: Send a pedestrian crossing warning message to alert surrounding vehicles to slow down and wait.
[0220] It should be understood that the descriptions given above are merely exemplary and not restrictive.
[0221] <Fourth Embodiment>
[0222] In the process of describing the electronic device in the above embodiments, some processes or methods are clearly also disclosed. In the following, without repeating some details already discussed above, an overview of these methods is given. However, it should be noted that although these methods are disclosed in the process of describing the electronic device, these methods do not necessarily use those components described or are necessarily executed by those components. For example, the embodiments of the electronic device can be implemented partially or completely using hardware and / or firmware, while the methods for the electronic device discussed below can be completely implemented by computer-executable programs, although these methods can also use the hardware and / or firmware of the electronic device.
[0223] Figure 20 A flowchart of a method for an electronic device according to an embodiment of the present application is shown. As Figure 20 described, the method includes: performing spatio-temporal segmentation on the pulsed video data obtained by an event camera to obtain pulsed video data corresponding to different categories of objects (S11); respectively performing semantic recognition on different categories of objects based on the pulsed video data corresponding to different categories of objects (S12); and encoding the results of the semantic recognition to obtain an encoded message to be provided to other electronic devices (S13). This method can be executed, for example, at a vehicle or RSU in a vehicle-to-everything network.
[0224] For example, in step S11, the spatio-temporal segmentation can be performed as follows: The space corresponding to the pulsed video data is segmented according to a predetermined separation principle, and the predetermined separation principle can include one or more of the following: global or region of interest, movement status of the region of interest; and for each segmented space, event accumulation based on the pulsed video data is performed with a time period granularity corresponding to the space.
[0225] Step S11 may further include: for each of the segmented spaces, determining a time period granularity corresponding to the space based on the light intensity change in the space and the motion condition of the object in the space, where, when the light intensity change in the space is greater and the object moves faster, the time period granularity corresponding to the space is determined to be smaller.
[0226] For example, in step S12, for each category of objects among different categories of objects, a semantic recognition model corresponding to the category of objects may be used to perform semantic recognition.
[0227] In step S13, the object that has undergone semantic recognition may be used as the semantic model recognition subject, and an encoded message may be generated for the semantic model recognition subject. The encoded message includes semantic feature items of the semantic model recognition subject. For example, each semantic feature item may be encoded from top to bottom according to a semantic encoding tree. When the event camera captures a traffic scene, the semantic encoding tree includes the following levels from top to bottom: first traffic classification, second traffic classification, semantic model recognition subject, and semantic feature item. For example, the first traffic classification includes: traffic map, traffic signal, traffic participant, and road obstacle; the second traffic classification includes: urban street, mountain road, road intersection, and parking lot under the traffic map category, signal light and traffic sign under the traffic signal category, motor vehicle, non-motor vehicle, and pedestrian under the traffic participant category, and immovable obstacle and movable obstacle under the road obstacle category; the semantic model recognition subject includes objects under each second traffic classification; and the semantic feature item includes the feature name, value length, and value of the semantic model recognition subject.
[0228] In addition, the semantic feature item may further include the correlation between the semantic model recognition subject and other semantic model recognition subjects. For example, as Figure 21 shown, step S13 may further include the following sub-steps: determining whether there is a correlation between two or more objects (S131); when it is determined that there is a correlation, adding a correlation semantic feature item to the corresponding object (S132), and then proceeding to semantic encoding (S133), otherwise directly proceeding to S133.
[0229] In step S13, the semantic model recognition subject for which the encoded message is to be generated and the semantic feature items to be included may also be determined according to the vehicle networking application scenario. The encoded message may also include a field of the vehicle networking application scenario. The encoded message obtained by encoding the semantic recognition result of the objects in the space may be sent according to the time period granularity of each space.
[0230] When the method is applied to the vehicle networking, the method may further include expanding the vehicle networking message set to add the DayIII Ext Msg message body that supports semantic communication. The DayIII Ext Msg message body includes encoded messages.
[0231] The above method corresponds to the electronic device 100 in the first embodiment. The relevant detailed description has been given in the first and third embodiments and will not be repeated here.
[0232] Figure 22 The flowchart of a method for an electronic device according to another embodiment of the present application is shown. As Figure 22 described, the method includes: receiving an encoded message from another electronic device (S21), where the encoded message is obtained by the other electronic device as follows: performing spatio-temporal segmentation on the pulsed video data obtained by an event camera to obtain pulsed video data corresponding to different categories of objects; respectively performing semantic recognition on different categories of objects based on the pulsed video data corresponding to different categories of objects; and encoding the results of the semantic recognition; and performing semantic decoding and semantic fusion and recombination on the encoded message to obtain the semantics of the original scene captured by the event camera (S22). This method can be executed, for example, on the vehicle side in vehicle networking.
[0233] In the case of a vehicle, for example, the encoded message may be included in the DayIII Ext Msg message body, and the DayIII Ext Msg message body is an expansion of the vehicle networking message set to support semantic communication.
[0234] For example, step S22 may include: aligning the encoded messages for different categories of objects according to time, where the encoded messages for different categories of objects are sent with different time period granularities; and combining the semantically decoded semantic information according to spatial positions to obtain the complete semantics of the original scene. For example, the time period granularity for aligning the encoded messages can be determined based on current requirements.
[0235] Similarly, the encoded message includes semantic feature items for identifying the subject of the semantic model, where the subject of the semantic model identification is an object that has undergone semantic identification. Each semantic feature item is encoded from top to bottom according to the semantic encoding tree. In the case where the event camera captures a traffic scene, the semantic encoding tree from top to bottom may include the following levels: First traffic classification, Second traffic classification, Subject of semantic model identification, and Semantic feature items. The first traffic classification includes: traffic map, traffic signal, traffic participants, and road obstacles; the second traffic classification includes: urban streets, mountain roads, road intersections, and parking lots under the traffic map category, signal lights and traffic signs under the traffic signal category, motor vehicles, non-motor vehicles, and pedestrians under the traffic participant category, and immovable obstacles and movable obstacles under the road obstacle category; the subject of semantic model identification includes objects under each second traffic classification; and the semantic feature items include the feature name, value length, and value of the subject of the semantic model identification.
[0236] In addition, the semantic feature items may further include the correlation between the subject of the semantic model identification and other subjects of the semantic model identification. The correlation includes, for example, movement following or space occupancy, etc. Semantic fault tolerance processing can be performed based on the information of the correlation in the semantic feature items. Semantic fault tolerance processing includes, for example, correction of incorrect semantics and recovery of lost semantics.
[0237] Figure 23 An example of a flowchart showing the steps S22 including semantic fault tolerance processing is shown. The specific process is as follows: perform semantic decoding on the encoded message (S221); perform a correlation judgment on whether there is a correlation between two or more objects based on the result of the semantic decoding (S222), for example, judge whether there are correlation semantic feature items for two or more objects; if it is judged to be yes in step S222, proceed to step S223 for correlation analysis, for example, judge whether there is movement following, or space occupancy, etc., otherwise proceed to S226 to perform scene semantic recombination; then in step S224, judge whether there is a semantic error or loss based on the result of the correlation analysis; if it is judged that there is a semantic error or loss in step S224, proceed to step S225 for fault tolerance processing, otherwise proceed to S226 to perform scene semantic recombination.
[0238] In addition, as Figure 22 shown, the above method may further include step S23: using a scene analysis prediction model to analyze and process the semantics of the obtained original scene to obtain risk assessment and / or prediction warning information. For example, an appropriate artificial intelligence analysis model can be selected as the scene analysis prediction model according to the scene type. In the case of the vehicle network, the scene type may include vehicle network application scenarios.
[0239] For example, the scenario type can be determined based on the vehicle networking scenario knowledge stored in the background knowledge base. The encoded message can also include a field for the vehicle networking application scenario.
[0240] The above method corresponds to the electronic device 200 in the second embodiment. The relevant detailed description has been given in the second and third embodiments and will not be repeated here.
[0241] The technology of the present disclosure can be applied to various products.
[0242] For example, the electronic devices 100 and 200 can be implemented as various user devices. The user device can be implemented as a mobile terminal (such as a smart phone, a tablet personal computer (PC), a notebook PC, a portable game terminal, a portable / dongle-type mobile router, and a digital camera device), a vehicle-mounted terminal (such as a car navigation device), or a vehicle. The user device can also be implemented as a terminal that performs machine-to-machine (M2M) communication (also referred to as a machine type communication (MTC) terminal). In addition, the user device can be a wireless communication module (such as an integrated circuit module including a single wafer) installed on each of the above terminals.
[0243] [Application Examples of User Devices]
[0244] (First Application Example)
[0245] Figure 24 FIG. is a block diagram showing an example of a schematic configuration of a smart phone 900 to which the technology of the present disclosure can be applied. The smart phone 900 includes a processor 901, a memory 902, a storage device 903, an external connection interface 904, a camera device 906, a sensor 907, a microphone 908, an input device 909, a display device 910, a speaker 911, a wireless communication interface 912, one or more antenna switches 915, one or more antennas 916, a bus 917, a battery 918, and an auxiliary controller 919.
[0246] The processor 901 can be, for example, a CPU or a system on a chip (SoC), and controls the functions of the application layer and other layers of the smart phone 900. The memory 902 includes a RAM and a ROM, and stores data and programs executed by the processor 901. The storage device 903 can include a storage medium, such as a semiconductor memory and a hard disk. The external connection interface 904 is an interface for connecting an external device (such as a memory card and a universal serial bus (USB) device) to the smart phone 900.
[0247] The imaging device 906 includes an image sensor (such as a charge-coupled device (CCD) and a complementary metal-oxide semiconductor (CMOS)), and generates a captured image. The sensor 907 may include a set of sensors, such as a measurement sensor, a gyro sensor, a geomagnetic sensor, and an acceleration sensor. The microphone 908 converts the sound input to the smart phone 900 into an audio signal. The input device 909 includes, for example, a touch sensor configured to detect a touch on the screen of the display device 910, a keypad, a keyboard, a button, or a switch, and receives an operation or information input from the user. The display device 910 includes a screen (such as a liquid crystal display (LCD) and an organic light-emitting diode (OLED) display), and displays an output image of the smart phone 900. The speaker 911 converts the audio signal output from the smart phone 900 into sound.
[0248] The wireless communication interface 912 supports any cellular communication scheme (such as LTE and LTE-Advanced), and performs wireless communication. The wireless communication interface 912 generally may include, for example, a BB processor 913 and an RF circuit 914. The BB processor 913 may perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and perform various types of signal processing for wireless communication. At the same time, the RF circuit 914 may include, for example, a mixer, a filter, and an amplifier, and transmits and receives wireless signals via an antenna 916. Note that although the figure shows a case where one RF link is connected to one antenna, this is only illustrative, and also includes a case where one RF link is connected to multiple antennas through multiple phase shifters. The wireless communication interface 912 may be a chip module on which the BB processor 913 and the RF circuit 914 are integrated. As Figure 24 shown, the wireless communication interface 912 may include multiple BB processors 913 and multiple RF circuits 914. Although Figure 24 an example where the wireless communication interface 912 includes multiple BB processors 913 and multiple RF circuits 914 is shown, the wireless communication interface 912 may also include a single BB processor 913 or a single RF circuit 914.
[0249] In addition, in addition to the cellular communication scheme, the wireless communication interface 912 may support other types of wireless communication schemes, such as a short-range wireless communication scheme, a near-field communication scheme, and a wireless local area network (LAN) scheme. In this case, the wireless communication interface 912 may include a BB processor 913 and an RF circuit 914 for each wireless communication scheme.
[0250] Each of the antenna switches 915 switches the connection destination of the antenna 916 among multiple circuits (such as circuits for different wireless communication schemes) included in the wireless communication interface 912.
[0251] Each of the antennas 916 includes single or multiple antenna elements (such as the multiple antenna elements included in a MIMO antenna), and is used for the wireless communication interface 912 to transmit and receive wireless signals. As Figure 24 shown, the smart phone 900 may include multiple antennas 916. Although Figure 24 an example where the smart phone 900 includes multiple antennas 916 is shown, the smart phone 900 may also include a single antenna 916.
[0252] In addition, the smart phone 900 may include an antenna 916 for each wireless communication scheme. In this case, the antenna switch 915 may be omitted from the configuration of the smart phone 900.
[0253] The bus 917 connects the processor 901, the memory 902, the storage device 903, the external connection interface 904, the imaging device 906, the sensor 907, the microphone 908, the input device 909, the display device 910, the speaker 911, the wireless communication interface 912, and the auxiliary controller 919 to each other. The battery 918 supplies power to Figure 24 the respective blocks of the smart phone 900 shown via a feeder line, which is partially shown as a dashed line in the figure. The auxiliary controller 919 operates the minimum necessary functions of the smart phone 900, for example, in the sleep mode.
[0254] In Figure 24 the smart phone 900 shown, the communication unit 104 and the transceiver of the electronic device 100 may be implemented by the wireless communication interface 912. At least a part of the functions may also be implemented by the processor 901 or the auxiliary controller 919. For example, the processor 901 or the auxiliary controller 919 may perform spatio-temporal segmentation, semantic recognition, and semantic encoding on the pulse video data by executing the functions of the spatio-temporal segmentation unit 101, the semantic recognition unit 102, the semantic encoding unit 103, and the communication unit 104, and provide the encoded message to other electronic devices, thereby greatly reducing the data transmission volume and improving the correct rate of information transmission.
[0255] The communication unit 201 of the electronic device 200 may be implemented by the wireless communication interface 912. At least a part of the functions may also be implemented by the processor 901 or the auxiliary controller 919. For example, the processor 901 or the auxiliary controller 919 may perform decoding, semantic fusion and recombination, and scene analysis and prediction on the semantic encoded message from other electronic devices by executing the functions of the communication unit 201, the decoding and semantic recombination unit 202, and the analysis and prediction unit 203, so as to achieve end-to-end perceptual communication recognition with low processing delay, greatly reduce the computing power requirements, improve the correct rate of information transmission, and better support the vehicle networking application scenario.
[0256] (Second application example)
[0257] Figure 25 is a block diagram showing an example of a schematic configuration of an in-vehicle navigation device 920 to which the technology of the present disclosure can be applied. The in-vehicle navigation device 920 includes a processor 921, a memory 922, a Global Positioning System (GPS) module 924, a sensor 925, a data interface 926, a content player 927, a storage medium interface 928, an input device 929, a display device 930, a speaker 931, a wireless communication interface 933, one or more antenna switches 936, one or more antennas 937, and a battery 938.
[0258] The processor 921 can be, for example, a CPU or an SoC, and controls the navigation function and other functions of the in-vehicle navigation device 920. The memory 922 includes a RAM and a ROM, and stores data and programs executed by the processor 921.
[0259] The GPS module 924 uses GPS signals received from GPS satellites to measure the position of the in-vehicle navigation device 920 (such as latitude, longitude, and altitude). The sensor 925 can include a set of sensors, such as a gyro sensor, a geomagnetic sensor, and an air pressure sensor. The data interface 926 is connected to, for example, an in-vehicle network 941 via a terminal not shown, and acquires data generated by the vehicle (such as vehicle speed data).
[0260] The content player 927 reproduces content stored in a storage medium (such as a CD and a DVD) inserted into the storage medium interface 928. The input device 929 includes, for example, a touch sensor, a button, or a switch configured to detect a touch on the screen of the display device 930, and receives operations or information input from a user. The display device 930 includes a screen such as an LCD or an OLED display, and displays an image of the navigation function or the reproduced content. The speaker 931 outputs the sound of the navigation function or the reproduced content.
[0261] The wireless communication interface 933 supports any cellular communication scheme (such as LTE and LTE-Advanced), and performs wireless communication. The wireless communication interface 933 generally can include, for example, a BB processor 934 and an RF circuit 935. The BB processor 934 can perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and perform various types of signal processing for wireless communication. At the same time, the RF circuit 935 can include, for example, a mixer, a filter, and an amplifier, and transmits and receives wireless signals via the antenna 937. The wireless communication interface 933 can also be a single chip module on which the BB processor 934 and the RF circuit 935 are integrated. As Figure 25 shown, the wireless communication interface 933 can include multiple BB processors 934 and multiple RF circuits 935. AlthoughFigure 25 An example is shown in which the wireless communication interface 933 includes a plurality of BB processors 934 and a plurality of RF circuits 935, but the wireless communication interface 933 may also include a single BB processor 934 or a single RF circuit 935.
[0262] In addition to the cellular communication scheme, the wireless communication interface 933 may support other types of wireless communication schemes, such as short - range wireless communication schemes, near - field communication schemes, and wireless LAN schemes. In this case, for each wireless communication scheme, the wireless communication interface 933 may include a BB processor 934 and an RF circuit 935.
[0263] Each of the antenna switches 936 switches the connection destination of the antenna 937 among a plurality of circuits (such as circuits for different wireless communication schemes) included in the wireless communication interface 933.
[0264] Each of the antennas 937 includes one or more antenna elements (such as a plurality of antenna elements included in a MIMO antenna), and is used for the wireless communication interface 933 to transmit and receive wireless signals. As Figure 25 shown, the car navigation device 920 may include a plurality of antennas 937. Although Figure 25 an example is shown in which the car navigation device 920 includes a plurality of antennas 937, the car navigation device 920 may also include a single antenna 937.
[0265] In addition, the car navigation device 920 may include an antenna 937 for each wireless communication scheme. In this case, the antenna switch 936 may be omitted from the configuration of the car navigation device 920.
[0266] The battery 938 supplies power to each block of the car navigation device 920 shown via a feeder line, which is partially shown as a dotted line in the figure. The battery 938 accumulates the power supplied from the vehicle. Figure 25 shown, the car navigation device 920 of which the battery 938 supplies power to each block via a feeder line, which is partially shown as a dotted line in the figure. The battery 938 accumulates the power supplied from the vehicle.
[0267] In Figure 25 the car navigation device 920 shown, the communication unit 104 of the electronic device 100, the transceiver may be implemented by the wireless communication interface 933. At least a part of the functions may also be implemented by the processor 921. For example, the processor 921 can perform spatio - temporal segmentation, semantic recognition, and semantic encoding on the pulse video data by executing the functions of the spatio - temporal segmentation unit 101, the semantic recognition unit 102, the semantic encoding unit 103, and the communication unit 104, and provide the encoded message to other electronic devices, thereby greatly reducing the data transmission volume and improving the correct rate of information transmission.
[0268] The communication unit 201 of the electronic device 200 can be implemented by the wireless communication interface 933. At least a part of the functions can also be implemented by the processor 921. For example, the processor 921 can perform the functions of the communication unit 201, the decoding and semantic recombination unit 202, and the analysis and prediction unit 203 to decode and semantically fuse and recombine the semantically encoded messages from other electronic devices and perform scenario analysis and prediction, so as to achieve end-to-end perception communication recognition with low processing latency, greatly reduce the computing power requirements, improve the accuracy of information transmission, and better support the vehicle networking application scenario.
[0269] The technology of the present disclosure can also be implemented as an in-vehicle system (or vehicle) 940 including one or more blocks of an automotive navigation device 920, an in-vehicle network 941, and a vehicle module 942. The vehicle module 942 generates vehicle data (such as vehicle speed, engine speed, and fault information), and outputs the generated data to the in-vehicle network 941.
[0270] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that for those skilled in the art, all or any steps or components of the methods and apparatuses of the present disclosure can be implemented in any computing device (including a processor, a storage medium, etc.) or a network of computing devices in the form of hardware, firmware, software, or a combination thereof, which can be achieved by those skilled in the art using their basic circuit design knowledge or basic programming skills after reading the description of the present disclosure.
[0271] Moreover, the present disclosure also proposes a program product storing machine-readable instruction codes. When the instruction codes are read and executed by a machine, the methods according to the embodiments of the present disclosure can be executed.
[0272] Correspondingly, the storage medium for carrying the above program product storing machine-readable instruction codes is also included in the disclosure of the present disclosure. The storage medium includes but is not limited to floppy disks, optical discs, magneto-optical discs, memory cards, memory sticks, and the like.
[0273] In the case of implementing the present disclosure by software or firmware, a program constituting the software is installed from a storage medium or a network to a computer with a dedicated hardware structure (such as Figure 26 the general-purpose computer 2600 shown). When various programs are installed on this computer, it can perform various functions and the like.
[0274] In Figure 26Among them, the central processing unit (CPU) 2601 performs various processes according to the programs stored in the read-only memory (ROM) 2602 or the programs loaded from the storage section 2608 into the random access memory (RAM) 2603. In the RAM 2603, data required when the CPU 2601 performs various processes and so on is also stored as needed. The CPU 2601, ROM 2602, and RAM 2603 are connected to each other via a bus 2604. An input / output interface 2605 is also connected to the bus 2604.
[0275] The following components are connected to the input / output interface 2605: an input section 2606 (including a keyboard, a mouse, etc.), an output section 2607 (including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.), a storage section 2608 (including a hard disk, etc.), a communication section 2609 (including a network interface card such as a LAN card, a modem, etc.). The communication section 2609 performs communication processing via a network such as the Internet. As needed, a drive 2610 may also be connected to the input / output interface 2605. A removable medium 2611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 2610 as needed, so that a computer program read therefrom is installed into the storage section 2608 as needed.
[0276] In the case where the above series of processes are implemented by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 2611.
[0277] Those skilled in the art should understand that such a storage medium is not limited to Figure 26 the removable medium 2611 shown in which a program is stored and distributed separately from the device to provide the program to the user. Examples of the removable medium 2611 include a magnetic disk (including a floppy disk (registered trademark)), an optical disk (including a compact disc read-only memory (CD-ROM) and a digital versatile disc (DVD)), a magneto-optical disk (including a mini disc (MD) (registered trademark)), and a semiconductor memory. Alternatively, the storage medium may be the ROM 2602, a hard disk included in the storage section 2608, etc., in which a program is stored and distributed to the user together with the device containing them.
[0278] It should also be noted that in the devices, methods, and systems of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure. And, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel or independently of each other.
[0279] Finally, it should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Additionally, without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0280] Although the embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings, it should be understood that the above-described embodiments are only for illustrating the present disclosure and do not constitute a limitation to the present disclosure. For those skilled in the art, various modifications and changes can be made to the above embodiments without departing from the essence and scope of the present disclosure. Therefore, the scope of the present disclosure is only defined by the appended claims and their equivalent meanings.
[0281] The present technology can also be configured as follows.
[0282] (1) An electronic device, comprising:
[0283] At least one processor; and
[0284] At least one memory, including computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, cause the electronic device to perform:
[0285] Perform spatio-temporal segmentation on the pulsed video data obtained by an event camera to obtain pulsed video data corresponding to different classes of objects;
[0286] Based on the pulsed video data corresponding to different classes of objects respectively, perform semantic recognition on the different classes of objects; and
[0287] Encode the results of the semantic recognition to obtain an encoded message to be provided to other electronic devices.
[0288] (2) The electronic device according to (1), wherein the at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to perform the spatio-temporal segmentation as follows:
[0289] Segment the space corresponding to the pulsed video data according to a predetermined separation principle, the predetermined separation principle including one or more of the following: global or region of interest, the movement status of the region of interest; and
[0290] For each of the segmented spaces, event accumulation based on pulsed video data is performed at a time period granularity corresponding to that space.
[0291] (3) The electronic device according to (2), wherein the at least one memory and the computer program code are further configured to, via the at least one processor, cause the electronic device to determine, for each of the segmented spaces, a time period granularity corresponding to that space based on the light intensity change in that space and the motion condition of an object in that space, wherein, in a case where the light intensity change in the space is greater and the object moves faster, the time period granularity corresponding to that space is determined to be smaller.
[0292] (4) The electronic device according to (1), wherein the at least one memory and the computer program code are further configured to, via the at least one processor, cause the electronic device to:
[0293] For each object of different categories of objects, perform the semantic recognition using a semantic recognition model corresponding to the object of that category.
[0294] (5) The electronic device according to (1), wherein the at least one memory and the computer program code are further configured to, via the at least one processor, cause the electronic device to perform the semantic encoding as follows:
[0295] Take the object that has undergone semantic recognition as the semantic model recognition subject, and generate the encoding message for the semantic model recognition subject, the encoding message including semantic feature items of the semantic model recognition subject.
[0296] (6) The electronic device according to (5), wherein each semantic feature item is encoded from top to bottom according to a semantic encoding tree.
[0297] (7) The electronic device according to (6), wherein the event camera captures traffic scenes, and the semantic encoding tree from top to bottom includes the following levels: first traffic classification, second traffic classification, semantic model recognition subject, semantic feature item.
[0298] (8) The electronic device according to (7), wherein,
[0299] The first traffic classification includes: traffic map, traffic signal, traffic participant, and road obstacle;
[0300] The second traffic classification includes: urban streets, mountain roads, road intersections, and parking lots under the traffic map category; signal lights and traffic signs under the traffic signal category; motor vehicles, non-motor vehicles, and pedestrians under the traffic participant category; and immovable obstacles and movable obstacles under the road obstacle category;
[0301] The semantic model recognition subject includes objects under each second traffic classification; and
[0302] The semantic feature items include the feature name, value length, and value of the semantic model recognition subject.
[0303] (9) The electronic device according to (8), wherein the semantic feature items include the correlation between the semantic model recognition subject and other semantic model recognition subjects.
[0304] (10) The electronic device according to (7), wherein the at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to:
[0305] Determine the semantic model recognition subject for generating the encoded message and the semantic feature items to be included according to the vehicle networking application scenario.
[0306] (11) The electronic device according to (10), wherein the encoded message further includes a field of the vehicle networking application scenario.
[0307] (12) The electronic device according to (2), wherein the at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to:
[0308] Send the encoded message obtained by encoding the semantic recognition result of the objects in the space according to the time period granularity of each space.
[0309] (13) The electronic device according to (1), wherein the electronic device is located at a vehicle or a roadside unit in the vehicle networking,
[0310] The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to:
[0311] Expand the vehicle networking message set to add a DayIII Ext Msg message body that supports semantic communication.
[0312] (14) The electronic device according to (13), wherein the DayIII Ext Msg message body includes the encoded message.
[0313] (15) An electronic device, comprising:
[0314] At least one processor; and
[0315] At least one memory, including computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, cause the electronic device to perform:
[0316] Receiving an encoded message from another electronic device, the encoded message obtained by the other electronic device by: performing spatio-temporal segmentation on pulsed video data obtained by an event camera to obtain pulsed video data corresponding to different categories of objects; respectively performing semantic recognition on the different categories of objects based on the pulsed video data corresponding to the different categories of objects; and encoding the results of the semantic recognition; and
[0317] Performing semantic decoding and semantic fusion recombination on the encoded message to obtain the semantics of the original scene captured by the event camera.
[0318] (16) The electronic device according to (15), wherein the at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to:
[0319] Align the encoded messages for different categories of objects in time, wherein the encoded messages for the different categories of objects are sent with different time period granularities; and
[0320] Combine the semantic information after semantic decoding according to spatial positions to obtain the complete semantics of the original scene.
[0321] (17) The electronic device according to (15), wherein the encoded message includes semantic feature items of a semantic model recognition subject, and the semantic model recognition subject is an object on which semantic recognition has been performed.
[0322] (18) The electronic device according to (17), wherein each semantic feature item is encoded from top to bottom according to a semantic encoding tree.
[0323] (19) The electronic device according to (18), wherein the event camera captures a traffic scene, and the semantic encoding tree includes the following levels from top to bottom: first traffic classification, second traffic classification, semantic model recognition subject, semantic feature item.
[0324] (20) The electronic device according to (19), wherein,
[0325] The first traffic classification includes: traffic map, traffic signal, traffic participant, and road obstacle;
[0326] The second traffic classification includes: urban streets, mountain roads, road intersections, and parking lots under the traffic map category; traffic lights and traffic signs under the traffic signal category; motor vehicles, non-motor vehicles, and pedestrians under the traffic participant category; and immovable obstacles and movable obstacles under the road obstacle category;
[0327] The semantic model recognition subject includes objects under each second traffic classification; and
[0328] The semantic feature items include the feature name, value length, and value of the semantic model recognition subject.
[0329] (21) The electronic device according to (20), wherein the semantic feature items include the correlation between the semantic model recognition subject and other semantic model recognition subjects.
[0330] (22) The electronic device according to (21), wherein the at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to:
[0331] Perform semantic error tolerance processing based on the information of the correlation in the semantic feature items.
[0332] (23) The electronic device according to (22), wherein the semantic error tolerance processing includes correction of incorrect semantics and recovery of lost semantics.
[0333] (24) The electronic device according to (15), wherein the at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to:
[0334] Analyze and process the semantics of the obtained original scenario using a scenario analysis prediction model to obtain risk assessment and / or prediction warning information.
[0335] (25) The electronic device according to (24), wherein the at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to:
[0336] Select a suitable artificial intelligence analysis model as the scenario analysis prediction model according to the scenario type.
[0337] (26) The electronic device according to (25), wherein the electronic device is located at a vehicle in a vehicle network, and the scenario type includes vehicle network application scenarios.
[0338] (27) The electronic device according to (26), wherein the at least one memory and the computer program code are further configured to, via the at least one processor, cause the electronic device to:
[0339] Determine the scenario type based on the vehicle networking scenario knowledge stored in the background knowledge base.
[0340] (28) The electronic device according to (26), wherein the encoded message further includes a field of the vehicle networking application scenario.
[0341] (29) The electronic device according to (16), wherein the at least one memory and the computer program code are further configured to, via the at least one processor, cause the electronic device to:
[0342] Determine the time period granularity for aligning the encoded message based on the current demand.
[0343] (30) The electronic device according to (15), wherein the encoded message is included in the DayIII Ext Msg message body, and the DayIII Ext Msg message body is an extension of the vehicle networking message set to support semantic communication.
[0344] (31) A method for an electronic device, comprising:
[0345] Performing spatio-temporal segmentation on the pulsed video data obtained by the event camera to obtain pulsed video data corresponding to different categories of objects;
[0346] Semantically identifying the different categories of objects respectively based on the pulsed video data corresponding to the different categories of objects; and
[0347] Encoding the result of the semantic identification to obtain an encoded message to be provided to other electronic devices.
[0348] (32) A method for an electronic device, comprising:
[0349] Receiving an encoded message from other electronic devices, where the other electronic devices obtain the encoded message by: performing spatio-temporal segmentation on the pulsed video data obtained by the event camera to obtain pulsed video data corresponding to different categories of objects; semantically identifying the different categories of objects respectively based on the pulsed video data corresponding to the different categories of objects; and encoding the result of the semantic identification; and
[0350] Performing semantic decoding and semantic fusion recombination on the encoded message to obtain the semantics of the original scene captured by the event camera.
[0351] (33) A computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, cause the processor to execute the method according to (31) or (32).
[0352] (34) A computer program product comprising a computer program / instructions, wherein the computer program / instructions, when executed by a processor, implement the steps of the method according to (31) or (32).
Claims
1. An electronic device comprising: at least one processor; as well as at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, cause the electronic device to execute: Performing spatiotemporal segmentation on the pulse video data obtained by the event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to the objects of different categories; and The result of the semantic recognition is encoded to obtain an encoded message to be provided to other electronic devices.
2. The electronic device according to claim 1, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to perform the spatiotemporal segmentation as follows: Segmenting the space corresponding to the pulse video data according to a predetermined separation principle, wherein the predetermined separation principle includes one or more of the following: global or region of interest, and movement of the region of interest; as well as For each divided space, event accumulation based on the pulse video data is performed at a time period granularity corresponding to the space.
3. The electronic device according to claim 1, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: For each category of objects in the different categories of objects, the semantic recognition is performed using a semantic recognition model corresponding to the category of objects.
4. The electronic device according to claim 1, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to perform the semantic encoding as follows: The object that has undergone semantic recognition is used as a semantic model recognition subject, and the coded message is generated for the semantic model recognition subject, wherein the coded message includes the semantic feature items of the semantic model recognition subject.
5. An electronic device comprising: at least one processor; as well as at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, cause the electronic device to execute: Receiving a coded message from another electronic device, the coded message being obtained by the other electronic device by: performing spatiotemporal segmentation on pulse video data obtained by an event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on the objects of different categories based on the pulse video data corresponding to the objects of different categories; and encoding the results of the semantic recognition; and The encoded message is semantically decoded and semantically fused and reorganized to obtain the semantics of the original scene captured by the event camera. The electronic device according to claim 5 , wherein: The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: The obtained semantics of the original scene are analyzed and processed using a scene analysis prediction model to obtain risk assessment and / or prediction warning information.
7. A method for an electronic device, comprising: Performing spatiotemporal segmentation on the pulse video data obtained by the event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to the objects of different categories; as well as The result of the semantic recognition is encoded to obtain an encoded message to be provided to other electronic devices.
8. A method for an electronic device, comprising: receiving a coded message from another electronic device, the coded message being obtained by the other electronic device by performing spatiotemporal segmentation on impulse video data obtained by an event camera to obtain impulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to the objects of different categories; and encoding the results of semantic recognition; as well as The encoded message is semantically decoded and semantically fused and reorganized to obtain the semantics of the original scene captured by the event camera.
9. A computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method according to claim 7 or 8.
10. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the steps of the method according to claim 7 or 8 are implemented.