Electronic device, method for electronic device, and computer-readable storage medium and computer program product
Through the spatiotemporal segmentation and semantic recognition of event cameras, combined with the SCNN model, the problems of light changes and high computing power requirements in autonomous driving are solved, and efficient data transmission and recognition are achieved.
Patent Information
- Application Number
- PCT/CN2025/075368
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-04
- Filing Date
- 2025-01-27
- Publication Date
- 2025-08-07
AI Technical Summary
In autonomous driving, existing vision algorithms have problems such as a huge impact on light and dark changes, large data volume and high computing power requirements, and traditional coding methods have failed to effectively utilize the semantic features of signals to optimize transmission.
The event camera is used to perform spatiotemporal segmentation and semantic recognition of pulsed video data, optimize data transmission through semantic coding, use pulse convolutional neural network (SCNN) model to classify objects, and encode and decode in the V2X scenario of the Internet of Vehicles.
It reduces the amount of data transmission, improves the accuracy of information transmission, reduces the computing power requirement, and adapts to complex autonomous driving scenarios and vehicle networking environments.
Smart Images

Figure CN2025075368_07082025_PF_FP_ABST
Abstract
Description
Electronic device, method for electronic device, computer-readable storage medium, and computer program product
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on February 4, 2024, with application number 202410160422.9 and invention name “Electronic device, method for electronic device, computer-readable storage medium and computer program product”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of wireless communications, and in particular, to a technology for semantic communication of pulse video data, and more particularly, to an electronic device, a method for an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0003] An event camera can be simply understood as a sensor that only senses moving objects. Each pixel in an event camera has an independent photoelectric sensing module. When the brightness change at that pixel exceeds a set threshold, event data (sometimes called pulse data) is generated and output. Furthermore, because all pixels operate independently, the data output of an event camera is asynchronous and spatially sparse. This is the biggest difference between event cameras and standard cameras, and also their core innovation. The benefit of this imaging paradigm is that it significantly reduces redundant data, thereby improving the computational efficiency of post-processing algorithms.
[0004] The application of vision algorithms has become an indispensable component of the development of autonomous driving. However, current vision algorithms still have some limitations: First, cameras are susceptible to sudden changes in light intensity and backlighting; second, cameras generate a large amount of data during operation, thus requiring extremely high computing power. Event cameras offer advantages such as extremely fast response speed, reduced invalid information, lower computing power and power consumption, and high dynamic range. They can help autonomous vehicles reduce the complexity of information processing, improve driving safety, and operate normally in extremely bright and dark environments.
[0005] On the other hand, the limitations of Shannon or classical information theory manifest themselves in their exclusive consideration of signal transmission, without regard to the signal's meaning or the observer's understanding of that meaning. Semantic coding extends traditional coding to a new dimension: the semantic dimension. Semantic features (such as the communication scenario, purpose, context, and language context) can alter the probability distribution of symbols during transmission, and there may be a certain degree of correlation between symbols in the context of transmission. Utilizing semantic features as background knowledge of the sender to optimize the codeword allocation process can achieve efficient information expression. When the same data is effectively encoded, the required expected code length is significantly reduced, thereby reducing transmission and storage overhead and achieving conciseness and clarity. Summary of the Invention
[0006] A brief overview of the present disclosure is provided below to provide a basic understanding of certain aspects of the present disclosure. It should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify key or important aspects of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is simply to present certain concepts in a simplified form as a prelude to the more detailed description discussed later.
[0007] According to one aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory comprising computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, enable the electronic device to execute: performing spatiotemporal segmentation on pulse video data obtained by an event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to the objects of different categories; and encoding the results of the semantic recognition to obtain encoded messages to be provided to other electronic devices.
[0008] According to another aspect of the present disclosure, a method for an electronic device is provided, comprising: performing spatiotemporal segmentation on pulse video data obtained by an event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to objects of different categories; and encoding the results of the semantic recognition to obtain a coded message to be provided to other electronic devices.
[0009] According to one aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory comprising computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, enable the electronic device to execute: receiving coded messages from other electronic devices, the coded messages being obtained by the other electronic devices by: performing spatiotemporal segmentation on pulse video data obtained by an event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to objects of different categories, respectively; and encoding the results of the semantic recognition; and performing semantic decoding on the coded messages and performing semantic fusion reorganization to obtain the semantics of the original scene captured by the event camera.
[0010] According to another aspect of the present disclosure, a method for an electronic device is provided, comprising: receiving a coded message from another electronic device, the coded message being obtained by the other electronic device by: performing spatiotemporal segmentation on pulse video data obtained by an event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to objects of different categories, respectively; and encoding the results of the semantic recognition; and performing semantic decoding on the coded message and performing semantic fusion reorganization to obtain the semantics of the original scene captured by the event camera.
[0011] According to other aspects of the present disclosure, a computer program code and a computer program product for implementing the above-mentioned method for an electronic device, as well as a computer-readable storage medium having the computer program code for implementing the above-mentioned method for an electronic device recorded thereon are also provided.
[0012] The electronic device and method according to the embodiments of the present application distinguish objects of different categories by performing spatiotemporal segmentation on the pulse video data obtained by the event camera and perform semantic recognition and semantic encoding on objects of different categories, which can greatly reduce the amount of data transmission and improve the accuracy of information transmission.
[0013] These and other advantages of the present disclosure will become more apparent through the following detailed description of the preferred embodiments of the present disclosure in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to further illustrate the above and other advantages and features of the present disclosure, the following is a further detailed description of the specific embodiments of the present disclosure in conjunction with the accompanying drawings. The drawings, together with the detailed description below, are included in this specification and form a part of this specification. Elements with the same function and structure are represented by the same reference numerals. It should be understood that these drawings only depict typical examples of the present disclosure and should not be regarded as limiting the scope of the present disclosure. In the drawings:
[0015] FIG1 shows a functional module block diagram of an electronic device according to an embodiment of the present application;
[0016] FIG2 shows a schematic example of spatial segmentation;
[0017] FIG3 shows a schematic example of time segmentation;
[0018] FIG4 shows a schematic flow chart of performing spatiotemporal segmentation on pulse video data and classification and selection of a semantic recognition model;
[0019] FIG5 shows a functional module block diagram of an electronic device according to an embodiment of the present application;
[0020] FIG6 shows a schematic diagram of the codeword composition of each semantic feature item;
[0021] FIG7 shows a schematic example of a semantic coding tree;
[0022] FIG8 shows an example of extending the DayIII Ext Msg message body MsgFrameIII;
[0023] FIG9 shows a functional module block diagram of an electronic device according to another embodiment of the present application;
[0024] FIG10 shows an exemplary operation process of the decoding and semantic reorganization unit;
[0025] FIG11 shows a functional module block diagram of an electronic device according to another embodiment of the present application;
[0026] FIG12 shows a schematic architecture of scene semantic reconstruction;
[0027] FIG13 shows a schematic diagram of scenario analysis prediction and early warning;
[0028] FIG14 is a schematic diagram showing the operations of a transmitting end and a receiving end when the solution of the present application is applied to a V2X scenario;
[0029] FIG15 shows a graphical representation of impulse video data after spatial segmentation;
[0030] FIG16 shows an example of time slices of various spaces obtained after time segmentation;
[0031] FIG17 shows a schematic diagram of semantic decoding;
[0032] FIG18 is a schematic diagram showing the overall architecture of scene semantic reconstruction and scene analysis, prediction, and early warning;
[0033] Figure 19 shows an example of semantic reorganization;
[0034] FIG20 shows a flowchart of a method for an electronic device according to an embodiment of the present application;
[0035] FIG21 shows an example of sub-steps of step S13;
[0036] FIG22 shows a flowchart of a method for an electronic device according to another embodiment of the present application;
[0037] FIG23 shows an example of a flowchart including step S22 of semantic fault tolerance processing;
[0038] FIG24 is a block diagram showing an example of a schematic configuration of a smartphone to which the technology of the present disclosure can be applied;
[0039] FIG25 is a block diagram showing an example of a schematic configuration of a car navigation device to which the technology of the present disclosure can be applied; and
[0040] 26 is a block diagram of an exemplary structure of a general-purpose personal computer in which the method and / or apparatus and / or system according to the embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION
[0041] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings. For the sake of clarity and conciseness, not all features of an actual implementation are described in this specification. However, it should be understood that in the process of developing any such actual implementation, many implementation-specific decisions must be made in order to achieve the developer's specific goals, such as compliance with system and business-related constraints, which may vary from implementation to implementation. In addition, it should be understood that although the development work may be very complex and time-consuming, it is a routine task for those skilled in the art who benefit from the contents of this disclosure.
[0042] It is also necessary to explain here that, in order to avoid obscuring the present disclosure due to unnecessary details, the accompanying drawings only show the device structure and / or processing steps that are closely related to the solution according to the present disclosure, while other details that are not closely related to the present disclosure are omitted.
[0043] <First embodiment>
[0044] As mentioned previously, the characteristics of event cameras make them suitable for complex scenarios in autonomous driving. Furthermore, semantic communication can expand the encoding dimension, transmitting more information with less overhead. The amount of pulse video data from event cameras varies greatly. Furthermore, in V2X scenarios, poor communication environments can lead to recognition errors on the event camera's sensor side. Therefore, improving the accuracy of information transmission is a key issue.
[0045] In this embodiment, an electronic device 100 is provided for performing semantic recognition on pulse video data and semantically encoding the recognition results. The following description primarily uses autonomous driving or V2X applications as exemplary application scenarios, but this is not restrictive and is provided for convenience and clarity of description. The embodiments of this application can be applied to any scenario where data with similar characteristics can be obtained and where similar requirements exist.
[0046] Figure 1 shows a functional module block diagram of an electronic device 100 according to this embodiment. As shown in Figure 1, the electronic device 100 includes: a spatiotemporal segmentation unit 101, configured to perform spatiotemporal segmentation on the pulse video data obtained by the event camera to obtain pulse video data corresponding to objects of different categories; a semantic recognition unit 102, configured to perform semantic recognition on objects of different categories based on the pulse video data corresponding to objects of different categories; and a semantic encoding unit 103, configured to encode the result of the semantic recognition to obtain an encoded message to be provided to other electronic devices.
[0047] Among them, the spatiotemporal segmentation unit 101, the semantic recognition unit 102, and the semantic encoding unit 103 can be implemented by one or more processing circuits and at least one memory. The processing circuit can be implemented as a chip, a processor, etc., and the at least one memory can be any form of storage device such as RAM, ROM, flash memory, etc. The at least one memory is used to store computer program code and data required for the processing circuit to perform processing. In addition, it should be understood that the various functional units in the electronic device shown in Figure 1 are only logical modules divided according to the specific functions they implement, and are not used to limit the specific implementation method. In addition, the functions of the spatiotemporal segmentation unit 101, the semantic recognition unit 102, and the semantic encoding unit 103 can also be implemented using a brain-like chip, which is not restrictive.
[0048] When the electronic device 100 is used in an autonomous driving scenario, the electronic device 100 can be installed on a roadside unit (RSU) or a vehicle capable of sensing. The vehicle described here can more generally refer to various user devices located on the vehicle and capable of accessing various sensors. The electronic device 100 in this embodiment can function as a data provider and as an encoding device.
[0049] It should also be noted that the electronic device 100 can be implemented at the chip level, or it can also be implemented at the device level. For example, the electronic device 100 can work as an RSU, a vehicle or a user device itself, and may also include external devices such as a memory, a transceiver (not shown in the figure), etc. The memory can be used to store programs and related data information that need to be executed to implement various functions of the RSU, vehicle or user device. The transceiver may include one or more communication interfaces to support communication with different devices (for example, other RSUs, vehicles or user devices, etc.), and the implementation form of the transceiver is not specifically limited here.
[0050] Event cameras are different from traditional frame cameras. Frame cameras output images one by one at a fixed frame rate, ultimately forming a video stream. Event cameras, on the other hand, only record pixel brightness changes, calling these changes in light intensity events. These characteristics make event cameras particularly well-suited for many specific scenarios in autonomous driving.
[0051] These specific scenarios include, for example: scenarios with obvious sudden changes in light intensity, including sudden changes in light brightness and backlight, such as blinding from strong light on the opposite side when meeting other vehicles, and high-exposure scenarios faced by vehicles after coming out of a tunnel; scenarios with too much light or too little light, such as in a late-night environment, where the frame camera cannot recognize surrounding objects due to the extremely dim surrounding light, while the event camera can still effectively recognize surrounding objects; lateral blind spot perception scenarios: traditional frame cameras are not ideal for detecting lateral pedestrians / non-motor vehicles (ghosting scenarios, especially electric vehicles), and they do not have time to respond to pedestrians who suddenly appear in the blind spot where the view is blocked by the vehicle or obstacle in front, while the event camera can perceive danger signals more quickly; high-speed obstacle avoidance scenarios, such as when a vehicle is driving fast on the highway and encounters a tire on the road ahead, the event camera can quickly identify the tire in front and take timely obstacle avoidance actions; low-power scenarios (on-board deployment scenarios of electric vehicles), compared with traditional methods, the calculation method based on the event camera can greatly save vehicle battery consumption, thereby extending the vehicle's range.
[0052] Event camera data exhibits a certain degree of sparsity in its two-dimensional structure. For example, if a target object moves only at time t0 but remains stationary thereafter, only one event will be recorded at time t0, with no data generated thereafter. In other words, an event camera only generates asynchronous pulse signals for the changing portion, potentially generating only tens of KB of data within 10 seconds. Because event cameras, unlike traditional frame cameras, output an asynchronous event pulse video data stream, this data stream requires temporal and spatial classification. Specifically, in the temporal dimension, the order in which light intensity changes occur for objects at different levels leads to the generation of pulse video event data, resulting in a temporal distribution of the pulse video event data stream captured by the event camera. In the spatial dimension, static background objects, dynamic background objects, and moving objects in the target area within the same space experience different light intensity changes, resulting in different levels of attention. Therefore, segmentation and classification of the pulse video data from event cameras, which contains multi-layered information, is necessary from both temporal and spatial perspectives.
[0053] It is understood that in autonomous driving scenarios, the images of the surrounding environment captured by the RSU or vehicle may include various types of objects, such as stationary objects (buildings, streetlights, etc.), slow-moving objects (pedestrians, etc.), and high-speed moving objects (cars, etc.). Due to the characteristics of event cameras, the pulse video data corresponding to different types of objects have different characteristics, and therefore need to be treated differently during processing.
[0054] For example, the spatiotemporal segmentation unit 101 is configured to segment the space corresponding to the pulse video data according to a predetermined separation principle, and the predetermined separation principle may include one or more of the following: global or region of interest, movement status of the region of interest; and for each segmented space, event accumulation based on the pulse video data at a time period granularity corresponding to the space.
[0055] In other words, the spatiotemporal segmentation unit 101 can perform segmentation in the spatial dimension first and then in the temporal dimension. For example, in the spatial dimension, spatial semantic segmentation is performed according to different classifications such as static background objects, dynamic background objects, and moving objects in the target area, thereby dividing the global still space, local high-definition still space, and each moving monomer space. The moving monomer space can be further divided into high-speed moving monomer space and low-speed moving monomer space according to the speed of movement. Therefore, the pulse video data of the event camera is divided into the global still space and its objects, the local high-definition still space and its objects, and the moving monomer space and its objects, which correspond to the entire photosensitive pixel matrix of the event camera's dynamic visual sensor, the photosensitive pixel matrix of the static region of interest (ROI), and the photosensitive pixel matrix of the dynamic ROI. For ease of understanding, FIG2 shows a schematic example of spatial segmentation.
[0056] Furthermore, the spatiotemporal segmentation unit 101 can determine the time period granularity corresponding to each segmented space based on the light intensity variation and the motion of objects in that space. The greater the light intensity variation and the faster the object motion, the smaller the time period granularity corresponding to that space. Specifically, the event camera pulse data corresponding to the different spatially segmented spaces can be accumulated in the time dimension according to different time periods to form pixel matrix data, i.e., event image frames.
[0057] FIG3 shows a schematic example of time segmentation. The time period granularity of the high-speed moving single space is the smallest, while the time period granularity of the global static space is the largest. For example, long time period accumulation is suitable for the recognition of static objects with small light intensity changes and few pulse event data, short time period accumulation is suitable for the recognition of low-speed moving objects with large but slow light intensity changes and more pulse event data, and extremely short time period accumulation is suitable for the recognition of high-speed moving objects with large and fast light intensity changes and a large amount of pulse event data. Accordingly, the electronic device 100 will subsequently send coded messages for different spaces at different time granularities, that is, send the coded messages obtained by encoding the semantic recognition results of the objects in the space according to the time period granularity of each space.
[0058] After the spatiotemporal segmentation unit 101 obtains the pulse video data corresponding to objects of different categories, the semantic recognition unit 102 performs semantic recognition on the objects of different categories based on the pulse video data corresponding to the objects of different categories.
[0059] For example, the semantic recognition unit 102 can perform semantic recognition for each category of objects in different categories using a semantic recognition model corresponding to the category of objects. The semantic recognition model can be an artificial intelligence (AI) model, such as a pulse convolutional neural network (SCNN) model. The pulse neural network has the characteristics of event-driven, asynchronous operation, and extremely low power consumption, and the generation of pulse signals is very consistent with the event stream output method of the event camera based on timestamps. Therefore, in the example of this application, the SCNN model will be used as an example of the semantic recognition model for description, but it is not limited to this.
[0060] From the classification perspective of semantic recognition models, for example, semantic recognition models can include one or more of the following: static dark object recognition model, static luminous object recognition model, static oscillating object recognition model (static oscillating object recognition filtering model), low-speed motion object recognition model and high-speed motion object recognition model. A corresponding SCNN model can be established for each of these recognition models. The SCNN model structure includes a pulse coding layer, a pulse convolution layer, a pulse pooling layer, and a fully connected layer. The SCNN model fully integrates the advantages of the SNN (pulse neural network) model and the CNN (convolutional neural network) model, with fast training and recognition speeds, while also reducing computing power consumption and saving a lot of computing costs.
[0061] Figure 4 shows a schematic flow chart for the spatiotemporal segmentation of pulse video data and the classification and selection of semantic recognition models. First, the event camera obtains raw pulse video data (also called pulse event data). Then, the pulse video data is spatially segmented. For example, according to separation principles such as "global / ROI" and "ROI movement", it is segmented into global static space and its objects, local high-definition static space and its objects, and moving single-unit space and its objects. This is because the semantic recognition models can be different for static objects and moving objects.
[0062] Next, time segmentation is performed, that is, event accumulation is performed according to different time period granularities. For example, for dynamic object recognition with extremely short periods, a high-speed moving object recognition model is selected.
[0063] For slow-moving objects like pedestrians, select the slow-moving object recognition model. In addition to identifying object categories based on external contours, some scenarios require further recognition of human motion, such as standing still, walking slowly, or running quickly. In these scenarios, standard AI object classification models are inadequate and require further optimization through intensive training. As shown in the figure, select the slow-moving object posture recognition model.
[0064] For long-period static object recognition, it is possible to further distinguish whether the object itself is luminous. Static objects that emit light themselves will bring about obvious changes in light intensity, and the accumulated pulse event data is large, so that a clearer event image frame can be formed. The conventionally trained AI recognition model (static luminous object recognition model) can recognize such objects. However, for non-luminous static objects, when the event camera is stationary, there will be basically no obvious changes in light intensity, and the event image frame formed is relatively fuzzy. The AI recognition model can be further optimized through intensive training and other processes to identify such objects. As shown in the figure, a static dark object recognition model is used. For non-luminous static objects, it is also necessary to continue to distinguish whether the object will be affected by the surrounding environment such as wind, rain, and snow, causing the object to vibrate and shake, thereby bringing about obvious changes in light intensity. The AI recognition model for such static oscillating objects can also be further optimized through intensive training and other processes. As shown in the figure, a static oscillating object recognition filter model is further selected. For static objects that do not emit light and are not affected by the surrounding environment, different AI recognition models need to be used based on the motion / stationary state of the event camera itself: when the event camera itself is in motion (such as deployed on an autonomous driving vehicle), there will be obvious changes in light intensity on the surface of the static object, and the conventionally trained AI recognition model can recognize such objects; when the event camera itself is stationary (such as deployed on an RSU), there will be basically no obvious changes in light intensity on the surface of the static object, and the AI recognition model can be further optimized through intensive training and other processes to recognize such objects, as shown in the figure, by further selecting the static dark object fuzzy recognition model.
[0065] It should be understood that the classification and selection of the semantic recognition models shown in FIG4 are merely exemplary and can be appropriately increased, decreased or adjusted in actual applications.
[0066] For example, the static dark object recognition model can be used to identify static objects that do not emit light, such as surrounding street buildings, pedestrians waiting on the roadside, parked vehicles, and lane markings. Furthermore, because the motion of the event camera itself generates optical flow in the surrounding static background, different resolution recognition models can be used depending on whether the event camera itself is in motion.
[0067] The static luminous object recognition model is used to identify objects that emit light and produce changes in light intensity, such as traffic lights, flashing street lights, and neon lights in street windows.
[0068] The static oscillating object recognition model is used to identify surrounding background objects with periodic oscillating motion, such as street trees swaying in the wind, falling raindrops / snowflakes, and fluttering roadside flags.
[0069] The low-speed moving object recognition model identifies low-speed moving objects, such as pedestrians crossing the road on the pedestrian zebra crossing, vehicles traveling at low speed on the lane, non-motor vehicles on the non-motor vehicle lane, etc.; if different motion postures of objects need to be recognized, different posture recognition models need to be further defined.
[0070] The high-speed moving object recognition model is used to identify high-speed moving objects, such as high-speed projectiles, falling rocks in mountainous areas, motor vehicles, non-motor vehicles, pedestrians, or animals crossing at high speed at forks in the road.
[0071] Table 1 below shows an example of a table for selecting semantic recognition models when an event camera is shooting a traffic scene. The left side shows the semantic model, where Class I and Class II represent traffic classifications, the recognition subjects are the objects to be recognized, and the features represent the parameter items used to describe the corresponding objects. The right side shows the selection of recognition models, which are divided into three types: long cycle, short cycle, and very short cycle. The specific recognition models have been described above. The mark √ represents that the recognition subject corresponding to the row is identified using the recognition model corresponding to the column. It should be understood that Table 1 is only exemplary and not restrictive.
[0072] Table 1
[0073] The semantic recognition unit 102 uses the semantic recognition model of the corresponding category to accurately identify each object and its related features (hereinafter referred to as semantic feature items). The semantic encoding unit 103 uses the semantically recognized object as the semantic model recognition subject and generates a coded message for the semantic model recognition subject. The coded message includes the semantic feature items of the semantic model recognition subject. In other words, a coded message can be defined for each semantic model recognition subject to indicate its various semantic features.
[0074] As mentioned above, different time period granularities are used for event accumulation for different spaces. Therefore, the time period for semantic recognition by the semantic recognition unit 102 for objects in different spaces is also different, and the time period granularity of the semantic recognition results obtained is also different. Accordingly, the time period granularity of the encoded message obtained after the semantic encoding unit 103 encodes the semantic recognition result will also vary depending on the space in which the object is located. The electronic device 100 (for example, its communication unit 104, as shown in FIG5 ) can send the encoded message obtained by encoding the semantic recognition result of the object in the space according to the time period granularity of each space.
[0075] For example, each semantic feature item can be encoded from top to bottom according to a semantic coding tree. The semantic coding tree is encoded hierarchically from the root node to the leaf nodes, ensuring unique encoding. The following description uses hexadecimal encoding as an example.
[0076] In one example, an event camera shoots a traffic scene, and the semantic coding tree includes the following levels from top to bottom: a first traffic classification, a second traffic classification, a semantic model recognition subject, and a semantic feature item.
[0077] As can be seen, the layers in this example correspond to the semantic model items in Table 1. The first traffic category includes traffic maps, traffic signals, traffic participants, and road obstacles. The second traffic category includes urban streets, mountain roads, road intersections, and parking lots under the traffic map category; traffic lights and traffic signs under the traffic signal category; motor vehicles, non-motor vehicles, and pedestrians under the traffic participant category; and immovable and movable obstacles under the road obstacle category. The semantic model identification subjects include objects in each of the second traffic categories, such as buildings, trees, and streetlights on urban streets; pedestrian crossings and lane markings at road intersections; motor vehicle signals and crosswalk signals; cars and fire trucks; pedestrians; and non-motor vehicle drivers who dismount and push their vehicles. Semantic feature items include the feature name, value length, and value (i.e., using a TLV structure) of the semantic model identification subject. For example, the semantic feature items of pedestrians may include height, speed, location coordinates, posture, accompanying items, and relevance.
[0078] The code length for the first traffic classification can be 2 bits, the code length for the second traffic classification can be 6 bits, the code length for the semantic model identification subject can be 8 bits, the code length for the feature name in the semantic feature item can be 8 bits, the code length for the value length (e.g., the number of bytes representing the value) can be 8 bits, and the code length for the value can be 0 to 255 bytes. A schematic diagram of the codeword structure is shown in Figure 6. It should be understood that the code lengths of each field are not restrictive, and the semantics corresponding to different codeword values are also not restrictive.
[0079] FIG7 shows a schematic example of a semantic coding tree. The first traffic classification is shown in the figure as traffic classification-I, and the second traffic classification is shown in the figure as traffic classification-II. In the example of FIG7 , some semantic feature items include the correlation between the semantic model recognition subject and other semantic model recognition subjects. The information of this correlation can be used to optimize and adjust the semantic information of the overall space. The semantic coding unit 103 can determine whether there is a correlation between two or more objects, and add a correlation semantic feature item to the corresponding object if it is determined that there is a correlation. As shown in FIG7 , the correlation may include, for example, movement accompaniment, space occupancy, etc.
[0080] In this way, when the receiving end performs semantic decoding, it can extract the correlation semantic feature items of multiple moving objects in the same static space, and then perform analysis and processing. Based on the correlation of multiple moving objects, it can comprehensively judge whether the decoded semantic information is correct. If there are errors, fault tolerance processing can be performed, which will be described in detail later.
[0081] Table 2 below shows an example of a semantic coding table in a traffic scenario (for example, the electronic device 100 is located in a vehicle or RSU in V2X).
[0082] Table 2
[0083] The codes in brackets are hexadecimal codes. For example, 0xFFFF 0xFF in row 1 represents the location coordinates (x, y) of a building in the city street category under the traffic map category.
[0084] In addition, the encoded message may also include a field for IoV application scenarios. This field can be placed at the beginning of the encoded message, for example. IoV application scenarios may include, but are not limited to, blind spot awareness sharing scenarios on urban streets, high-altitude falling object awareness sharing scenarios in mountainous areas, high-speed obstacle avoidance awareness sharing scenarios on highways, and blind spot awareness sharing scenarios at tunnel exits. Table 3 below shows encoding examples for IoV application scenarios.
[0085] Table 3
[0086] The advantages of event cameras can be fully demonstrated in the scenarios listed in Table 3. For example, to save computing power, the electronic device 100, or the semantic recognition and encoding operations therein, can be performed only when the vehicle is in the aforementioned scenarios. For example, the electronic device 100 can identify the aforementioned connected vehicle application scenarios based on the location on a map.
[0087] In addition, the semantic coding unit 103 is also configured to determine the semantic model identification subject to generate the encoded message and the semantic feature items to be included according to the V2X application scenario. In other words, it may not be necessary to encode all identified objects in the space, but only objects in the area of interest or moving objects need to be encoded. On the other hand, for objects that need to be encoded, it may not be necessary to include all semantic feature items, but only semantic feature items of interest can be selected. The semantic coding unit 103 can select the objects to be encoded and the semantic feature items to be encoded according to the V2X application scenario, thereby reducing the overhead of the encoded message.
[0088] Table 4 below shows an example of a V2X scenario semantic model mapping table, which shows examples of semantic models corresponding to different V2X application scenarios.
[0089] Table 4
[0090] In addition, when the electronic device 100 is located in a V2X vehicle or RSU, the communication unit 104 can also expand the vehicle network message set to add a DayIII Ext Msg message body that supports semantic communication. For example, the DayIII Ext Msg message body can include the above-mentioned encoded message.
[0091] Specifically, the third phase of expansion can be based on the interactive messages defined in standards such as YD / T 3709-2020 "Technical Requirements for the Message Layer of Wireless Communication Technology for Internet of Vehicles Based on LTE", CSAE XXXX, and C-ITS XXXX "Standard for the Application Layer and Application Data Interaction of Cooperative Intelligent Transport Systems for Vehicles, Phase II". Figure 8 shows an example of an expansion of the DayIII Ext Msg message body, MsgFrameIII. MsgFrameIII defines four types of messages: traffic map Msg_SC_Map, traffic signal Msg_SC_Signal, traffic participant Msg_SC_Participant, and road obstacle Msg_SC_Obstacle, as well as an additional extensible item Msg_SC_Test.
[0092] To sum up, the electronic device 100 according to this embodiment can distinguish objects of different categories and perform semantic recognition and semantic encoding of objects of different categories by performing spatiotemporal segmentation on the pulse video data obtained by the event camera, which can greatly reduce the amount of data transmission and improve the accuracy of information transmission, reduce the processing delay of end-to-end perception communication recognition, reduce computing power requirements, and better support vehicle network application scenarios.
[0093] <Second embodiment>
[0094] Figure 9 shows a functional module block diagram of an electronic device 200 according to another embodiment of the present application. As shown in Figure 9, the electronic device 200 includes: a communication unit 201, configured to receive coded messages from other electronic devices, and the coded messages are obtained by other electronic devices by: performing spatiotemporal segmentation on the pulse video data obtained by the event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on the objects of different categories based on the pulse video data corresponding to the objects of different categories; and a decoding and semantic reorganization unit 202, configured to semantically decode the coded messages and perform semantic fusion reorganization (hereinafter also referred to as semantic reorganization) to obtain the semantics of the original scene captured by the event camera.
[0095] The communication unit 201 and the decoding and semantic reassembly unit 202 can be implemented by one or more processing circuits and at least one memory. The processing circuit can be implemented as a chip, processor, etc., and the at least one memory can be any form of storage device such as RAM, ROM, flash memory, etc. The at least one memory is used to store computer program code and data required for the processing circuit to perform processing. It should be understood that the various functional units in the electronic device shown in Figure 9 are merely logical modules divided according to the specific functions they implement, and are not used to limit the specific implementation method. In addition, the functions of the decoding and semantic reassembly unit 202 can also be implemented using a brain-like chip, which is not restrictive.
[0096] When the electronic device 200 is used in an autonomous driving scenario, the electronic device 200 can be set on the vehicle side. The vehicle described here can more generally refer to various user devices located on the vehicle. The electronic device 200 in this embodiment can act as a data requester and as a decoding device.
[0097] It should also be noted that the electronic device 200 can be implemented at the chip level, or it can also be implemented at the device level. For example, the electronic device 200 can work as a vehicle or user device itself, and can also include external devices such as memory, transceivers (not shown in the figure), etc. The memory can be used to store programs and related data information that need to be executed to implement various functions of the vehicle or user device. The transceiver may include one or more communication interfaces to support communication with different devices (for example, other vehicles or user devices, etc.), and the implementation form of the transceiver is not specifically limited here.
[0098] However, this is not restrictive, and the electronic device 200 may also be set on the network side, such as a base station, a roadside unit, etc.
[0099] The encoded message here can, for example, be generated and sent by the electronic device 100 in the first embodiment. For example, the encoded message can include semantic feature items of a semantic model recognition subject, where the semantic model recognition subject is the object that has undergone semantic recognition. Each semantic feature item can be encoded from top to bottom according to a semantic encoding tree. When an event camera captures a traffic scene, the semantic encoding tree can include the following layers from top to bottom: a first traffic category, a second traffic category, a semantic model recognition subject, and semantic feature items. The first traffic category includes: traffic maps, traffic signals, traffic participants, and road obstacles; the second traffic category includes: urban streets, mountain roads, road intersections, and parking lots under the traffic map category; traffic lights and traffic signs under the traffic signal category; motor vehicles, non-motor vehicles, and pedestrians under the traffic participant category; and immovable obstacles and movable obstacles under the road obstacle category. The semantic model recognition subject includes objects under each second traffic category; and the semantic feature items include the feature name, value length, and value of the semantic model recognition subject. The semantic feature items may also include correlations between the semantic model recognition subject and other semantic model recognition subjects. The encoded message may be included in the DayIII Ext Msg message body, which is an extension of the vehicle network message set to support semantic communication. The specific details have been described in detail in the first embodiment, which is also applicable to this embodiment and will not be repeated here.
[0100] The communication unit 201 receives multiple encoded messages. For example, the decoding and semantic reassembly unit 202 is configured to: align encoded messages of different categories of objects in time, where the encoded messages of different categories of objects are sent at different time period granularities; and combine the semantically decoded semantic messages according to spatial positions to obtain the complete semantics of the original scene.
[0101] Figure 10 illustrates an exemplary operational process of the decoding and semantic reassembly unit 202. Specifically, it decodes and restores semantic information for high-speed motion entities with extremely short sampling periods, such as a 0.1-second time period granularity; decodes and restores semantic information for low-speed motion entities and stationary luminous objects with short time periods, such as a 1-second time period granularity; decodes and restores semantic information for stationary objects within a local stationary space with a long time period, such as a 10-second time period granularity; and decodes and restores semantic information for stationary objects within a global stationary space with a long time period, such as a 30-second time period granularity. For example, the decoding and semantic reassembly unit 202 can determine the time period granularity used to align encoded messages based on current needs, such as scenario requirements. For example, for low-speed motion object recognition and analysis scenarios, alignment can be performed at a 1-second granularity. In this case, high-speed motion entity semantics with a 0.1-second period granularity are accumulated to 1 second (only the most recent one can be retained), while the semantics for stationary objects with long periods of 10 seconds or 30 seconds can serve as shared background semantics for multiple low-speed motion object recognitions. For high-speed motion object recognition and analysis scenarios, alignment can be performed at a 0.1-second granularity, and so on.
[0102] For example, the decoding and semantic reorganization unit 202 may perform semantic decoding with reference to the following Table 6. It should be understood that Table 6 is only an illustrative example, which may correspond to Table 2 in the first embodiment.
[0103] Table 6
[0104] In addition, the decoding and semantic reassembly unit 202 may perform scene decoding with reference to the following Table 7. It should be understood that Table 7 is only an illustrative example, which may correspond to Table 3 in the first embodiment.
[0105] Table 7
[0106] Next, the decoding and semantic reassembly unit 202 combines the decoded semantic information according to spatial position. For example, based on the coordinate position information of different objects in the impulse event image frame, the decoding and semantic reassembly unit 202 combines the overall semantic information of the entire space in the order of high-speed moving objects, low-speed moving objects, stationary luminous objects, stationary objects in the local stationary space, and stationary objects in the global stationary space.
[0107] The decoding and semantic reorganization unit 202 can also perform semantic error-tolerance processing based on the correlation information in the semantic feature items. Semantic error-tolerance processing can include correction of erroneous semantics and recovery of lost semantics. For example, the decoding and semantic reorganization unit 202 can determine whether there is a correlation between two or more objects based on the correlation information. If so, a correlation analysis is performed. For example, the correlation can include spatial occupancy, motion accompaniment, etc. The decoding and semantic reorganization unit 202 can also determine whether there is a semantic error or loss based on the correlation analysis. If so, semantic error-tolerance processing can be performed, such as using the semantic information of other related objects to recover or correct the lost or erroneous semantics.
[0108] In addition, the electronic device 200 may further include an analysis and prediction unit 203 configured to analyze and process the semantics of the obtained original scene using a scene analysis and prediction model to obtain risk assessment and / or prediction and warning information.
[0109] For example, the analysis process here may be intelligent analysis process, and the analysis and prediction unit 203 may select an appropriate artificial intelligence analysis model as a scenario prediction analysis model based on the scenario type. In the case where the electronic device 200 is located in a vehicle in V2X, the scenario type may include a vehicle network application scenario, which may include, for example: a shared scenario of blind spot perception on urban streets, a shared scenario of high-altitude falling object perception in mountainous areas, a shared scenario of high-speed obstacle avoidance perception on highways, and a shared scenario of blind spot perception at tunnel exits.
[0110] The analysis and prediction unit 203 can determine the scene type based on the IoV scene knowledge stored in the background knowledge base. IoV scene knowledge includes, for example, knowledge about city streets and road intersections. Furthermore, this IoV scene knowledge can also be used for semantic reorganization in the decoding and semantic reorganization unit 202, as shown in Figure 12.
[0111] In addition, the encoded message may also include a field for the vehicle network application scenario. The analysis and prediction unit 203 may determine the scenario type based on such a field.
[0112] FIG13 shows a schematic diagram of scenario analysis, prediction, and warning performed by the analysis and prediction unit 203 , which shows an example of an AI analysis and prediction model and the provided analysis and prediction, risk assessment, and perception sharing warning information.
[0113] Table 8 below shows a V2X scenario analysis model table that shows the correspondence between V2X scenarios and applicable AI analysis and prediction models. It should be understood that this table is only exemplary.
[0114] Table 8
[0115] For example, after determining the scenario type, the analysis and prediction unit 203 selects a corresponding AI analysis and prediction model and imports the restored scenario semantic information into the AI analysis and prediction model for AI analysis and processing, thereby obtaining valuable information such as various risk assessments, predictions, and warnings. It should be understood that the scenario analysis and prediction model does not necessarily have to be an AI analysis and prediction model, but can also be a conventional analysis and prediction model, and this is not restrictive.
[0116] In addition, the communication unit 201 can also provide the obtained risk assessment and / or prediction and warning information to other devices, such as surrounding vehicles. For example, if the electronic device 100 is located on the sensing vehicle side and the electronic device 200 is located on the network side, by performing the above-mentioned semantic reorganization and AI analysis and prediction on the network side, the prediction and warning information can be provided directly to the vehicle in need, which can reduce the computing power requirements of the vehicle.
[0117] On the other hand, when the vehicle in which the electronic device 200 is located is capable of working as a data provider, the electronic device 200 may also include a unit similar to the spatiotemporal segmentation unit 101 and the semantic recognition unit 102 in the electronic device 100, as well as an analysis and prediction unit 203, and perform scene analysis and prediction based on the results of semantic recognition.
[0118] For ease of understanding, Figure 14 shows a schematic diagram of the operation of the transmitter and receiver when the solution of the present application is applied to a V2X scenario. As shown in Figure 14, the transmitter V2X platform can be set on the V2X-RSU side or the sensing vehicle side, and the receiver V2X platform can be set on the V2X surrounding warning vehicle side, that is, the vehicle side with data demand. The transmitter can be equipped with an event camera and various chips for performing processing, such as AI chips and traditional computing chips such as CPUs. The receiver can be equipped with various chips for performing processing, such as AI chips and traditional computing chips such as CPUs. The event camera on the transmitter obtains pulse event data and provides it to the AI chip for pulse video semantic recognition. The obtained semantic recognition result is used for semantic encoding, and the encoded message obtained by semantic encoding can be subjected to V2X channel coding. The final obtained encoding result is sent to the receiver via C-V2X communication. Note that various existing V2X channel coding methods can be used here without limitation. The receiving end performs V2X channel decoding on the received encoding results (using the V2X channel decoding method corresponding to the V2X channel encoding method used by the transmitting end), and further performs semantic decoding on the semantically encoded message obtained after decoding to restore the original scene semantics. AI chips can also be used to perform scene semantic analysis and prediction.
[0119] The pulse video data captured by the event camera can be targeted at specific V2X application scenarios. For example, the transmitter can determine the current scenario based on various V2X scenario knowledge stored in the background knowledge base, or the user can specify the scenario. The transmitter selects a corresponding semantic recognition model based on a semantic recognition model selection table (e.g., Table 1) and uses the selected semantic recognition model to perform semantic recognition on the spatiotemporal segmented event data. Next, the transmitter selects the semantic model recognition subject and semantic feature items for semantic encoding based on the scenario. This process can be performed by referring to a scene semantic model mapping table (e.g., Table 4), a scene encoding table (e.g., Table 3), and a semantic encoding table (e.g., Table 1). Correspondingly, the receiver can also decode the semantic model recognition subject and semantic feature items based on a semantic decoding table (e.g., Table 6) and a scene decoding table (e.g., Table 7), then perform scene semantic reconstruction and perform scene analysis and prediction to obtain possible warning information. During scene analysis and prediction, the AI analysis and prediction model to be used can be selected based on a scene analysis model table (e.g., Table 8).
[0120] In addition, although the electronic device 200 is described above as including the analysis and prediction unit 203, it is not limited to this. The electronic device 100 described in the first embodiment may also include the analysis and prediction unit 203 to perform scene analysis and prediction based on the results of semantic recognition.
[0121] To sum up, according to this embodiment, the electronic device 200 can perform semantic decoding and scene semantic reorganization on the received semantically coded messages, and can perform scene analysis and prediction based on scene semantics, thereby realizing end-to-end perception communication recognition with low processing delay, greatly reducing computing power requirements, improving the accuracy of information transmission, and better supporting Internet of Vehicles application scenarios.
[0122] <Third embodiment>
[0123] For the blind spot perception sharing scenario on urban streets, the electronic device 200 of this embodiment can realize vehicle / pedestrian / non-motor vehicle movement trajectory tracking and prediction based on the coded message provided by the electronic device 100 or the electronic device 100, such as non-motor vehicles crossing the road and running a red light, non-motor vehicles entering the motor vehicle lane, etc., as well as human posture recognition and behavior prediction analysis, such as prediction of dangerous postures of surrounding non-motor vehicle personnel (such as making and answering mobile phone calls while riding a bicycle, drinking beverages and eating while riding a bicycle, riding a bicycle with an umbrella in one hand, illegally carrying people / objects), analysis of the driver's human posture / gesture / dangerous driving behavior (such as making and answering calls, eating, bending over to pick up objects), and recognition and evidence collection of vehicle passengers' throwing gestures.
[0124] For the shared scenario of high-altitude falling object perception in mountainous areas, the electronic device 200 of this embodiment can perform vehicle high-speed obstacle avoidance, high-speed parabolic tracking and identification, and hazard prediction based on the coded message provided by the electronic device 100 or the electronic device 100. High-speed parabolic objects have a fast speed, a small target volume, and a great risk.
[0125] For the blind spot perception sharing scenario at the tunnel exit, the electronic device 200 of this embodiment can perform strong light blind spot detection in scenarios where light intensity mutation is more obvious based on the coded message provided by the electronic device 100 or the electronic device 100. Sudden changes in light brightness and darkness, backlighting, etc., such as blinding by strong light from the opposite side when meeting other vehicles, and vehicles also face high exposure scenarios after coming out of the tunnel. This scenario also involves other surrounding traffic participants (including but not limited to vehicles, pedestrians, cyclists and other targets), road abnormality information such as: road traffic events (such as traffic accidents, etc.), abnormal vehicle behavior (speeding, leaving the lane, driving in the wrong direction, irregular driving and abnormal stillness, etc.), road obstacles (such as fallen rocks, scattered objects, dead branches, etc.) and road conditions (such as accumulated water, ice, etc.).
[0126] In this embodiment, an example of semantic recognition and encoding and decoding operations in a blind spot perception sharing scenario on urban streets is given, and other scenarios can be similarly applied.
[0127] First, the impulse video data is temporally and spatially segmented at the transmitting end (eg, electronic device 100). FIG15 shows a diagram of the impulse video data after spatial segmentation. FIG16 shows an example of time slices of each space obtained after temporal segmentation.
[0128] Then, corresponding semantic recognition models are determined for objects of different categories in different spaces, as shown in Table 9 below.
[0129] Table 9
[0130] The semantic recognition results are then semantically encoded. For example, semantic recognition results obtained by different semantic models can be semantically encoded at different time period granularities. Specifically, the static street background can be semantically encoded at a coarse-grained time period, while the different postures of pedestrians crossing the road can be semantically encoded at a fine-grained time period. Moving background objects such as background vehicles in non-focus areas can be left unencoded.
[0131] The corresponding scene can be coded according to the scene coding table. In this embodiment, the scene is urban street blind spot perception sharing, and the code should be 0xFFFF according to Table 3. In addition, the semantic model recognition subject and corresponding semantic feature items are selected according to the V2X scene semantic model mapping table (for example, Table 4) and the semantic recognition model selection table (for example, Table 9).
[0132] The following discussion of different semantic models is based on the second traffic classification. For Semantic Model 1: City Street Perception Sharing Scenario - Traffic Map - City Street, the selections are as follows.
[0133] The code is as follows:
[0134] That is, the semantic encoding of a city street is {0xFFFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000, 0xFFFE 0xFF 0x08 0x0000 0x0000 0x0000 0x0000, 0xFFFD 0xFF 0x08 0x0000 0x0000 0x0000 0x0000}, which means that a city street has buildings, trees, and street lights, where building 1 is at (x1, y1), tree 1 is at (x1, y1), and street light 1 is at (x1, y1).
[0135] For Semantic Model 2: Urban Street Perception Shared Scenario - Traffic Map - Road Intersection, select as follows.
[0136] The code is as follows:
[0137] That is, the semantic encoding of the road intersection is {0xFDFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFF 0x10 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFE 0x1B 0x0000 0x0000 0x0000 0x0000 0x0000 0x02 0x0000 0x0000 0x0000 0x0000 0x02,0xFDFB 0xFF 0x20 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xBFFF 0xFF 0x01 0x02,0xBFFE 0xFF 0x01 0x00}, the encoding means: the coordinates of the center of the road intersection are (x, y); there is a pedestrian crossing at the intersection, the position is from the coordinates (x_from, y_from) to (x_to, y_to), and the three objects (x1, y1), (x2, y3), and (x3, y3) occupy the space; there are two lanes, the position of lane 1 is from the coordinates (x1_from, y1_from) to (x1_to, y1_to), and the position of lane 2 is from the coordinates (x2_from, y2_from) to (x2_to, y2_to).
[0138] For semantic model 3: Urban Street Perception Shared Scenario - Traffic Signal - Traffic Light, select as follows.
[0139] The code is as follows:
[0140] That is, the semaphore semantics are encoded as {0xFDFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFF 0x10 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFE 0x1B 0x0000 0x0000 0x0000 0x0000 0x0000 0x02 0x0000 0x0000 0x0000 0x0000 0x02,0xFDFB 0xFF 0x20 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xBFFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xBFFF 0xFE 0x01 0x02,0xBFFE 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xBFFE 0xFE 0x01 0x00}, the coding means: the motor vehicle signal light indicates please wait; the pedestrian crossing signal light indicates you can go.
[0141] For Semantic Model 4: Urban Street Perception Shared Scenario - Traffic Participants - Motor Vehicles, the selections are as follows.
[0142] The code is as follows:
[0143] That is, the semantic code of the car (moving) is {0x7FFF 0xFF 0x01 0x01, 0x7FFF 0xFF 0x18 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000}, which means: the vehicle is moving at a low speed, and the trajectory is: longitude and latitude coordinates 1 (x1, y1, t1), and coordinates 2 (x2, y2, t2).
[0144] In addition, there is another car in the picture as the identification subject, which is coded as follows:
[0145] That is, the semantic code of the car (slowing down and stopping) is {0x7FFF 0xFF 0x01 0x00, 0x7FFF 0xFF 0x0C 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000}, which means: the vehicle slows down and stops, the longitude and latitude coordinates are (x, y), and it is waiting.
[0146] For Semantic Model 5: Urban Street Perception Shared Scenario - Traffic Participants - Pedestrians, the selections are as follows.
[0147] The code is as follows:
[0148] That is, the pedestrian semantics is encoded as {{0x7DFF 0xFF 0x01 0x00,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x08 0x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x01 0x01,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x0000 0x01 0x0000 0x0000 0x0000 0x01}{0x7DFF 0xFF 0x01 0x01,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x08 0x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x01 0x00,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x0000 0x01 0x0000 0x0000 0x0000 0x0000 0x01}{0x7DFF 0xFF 0x01 0x00,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x08 0x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x01 0x01,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x0000 0x01 0x0000 0x0000 0x0000 0x0000 0x01}}, the encoding meaning is: Pedestrian 1, tall, standing on one side of the road, position coordinates {(x1, y1)}, carrying a suitcase, correlation: and (x2, y2) are in a parent-child relationship, and (x3, y3) are in a husband-wife relationship; Pedestrian 2, short, standing on one side of the road, position coordinates {(x2, y2)}, correlation: and (x1, y1) are in a moving companion relationship, and (x3, y3) are in a moving companion relationship; Pedestrian 3, tall, standing on one side of the road, position coordinates {(x3, y3)}, carrying a suitcase, correlation: and (x1, y1) are in a moving companion relationship, and (x2, y2) are in a moving companion relationship.
[0149] The receiving end (for example, the electronic device 200) receives the above-mentioned coded messages from the sending end, and performs decoding and semantic reorganization. Figure 17 shows a schematic diagram of semantic decoding. Among them, the left side is the received coded message, the middle is the structure of the corresponding coding tree, and the right side is the decoded semantic information. It can be seen that the coded message of this example also includes the encoding of the Internet of Vehicles application scenario, that is, the OxFFFF at the top left, which represents the blind spot perception sharing scenario of urban streets.
[0150] For semantic model 1: City Street Perception Sharing Scenario - Traffic Map - City Street, the following relevant information is transmitted:
[0151] Decoded as follows:
[0152] That is, the received semantic code is {0xFFFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000, 0xFFFE 0xFF 0x08 0x0000 0x0000 0x0000 0x0000, 0xFFFD 0xFF 0x08 0x0000 0x0000 0x0000 0x0000}, and the decoded semantic data is: city streets, buildings (x1, y1), trees (x2, y2), and street lights (x3, y3).
[0153] For semantic model 2: Urban Street Perception Sharing Scenario - Traffic Map - Road Intersection, the following relevant information is transmitted:
[0154] Decoded as follows:
[0155] That is, the received semantic code is {0xFDFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFF 0x10 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFE 0x1B 0x0000 0x0000 0x0000 0x0000 0x0000 0x02 0x0000 0x0000 0x0000 0x0000 0x02,0xFDFB 0xFF 0x20 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xBFFF 0xFF 0x01 0x02,0xBFFE 0xFF 0x01 0x00}, the decoded semantic data is: the coordinates of the center of the road intersection are (x, y); there is a pedestrian crossing at the intersection, the position is from the coordinates (x_from, y_from) to (x_to, y_to), (x1, y1), (x2, y3), (x3, y3) three objects occupy the space; there are 2 lanes, lane 1 is located from the coordinates (x1_from, y1_from) to (x1_to, y1_to), lane 2 is located from the coordinates (x2_from, y2_from) to (x2_to, y2_to).
[0156] For semantic model 3: Urban Street Perception Sharing Scenario - Traffic Signals - Traffic Lights, the following relevant information is transmitted:
[0157] Decoded as follows:
[0158] That is, the received semantic code is {0xFDFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFF 0x10 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xFFFC 0xFE 0x1B 0x0000 0x0000 0x0000 0x0000 0x0000 0x02 0x0000 0x0000 0x0000 0x0000 0x02,0xFDFB 0xFF 0x20 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000,0xBFFF 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xBFFF 0xFE 0x01 0x02,0xBFFE 0xFF 0x08 0x0000 0x0000 0x0000 0x0000,0xBFFE 0xFE 0x01 0x00}, the decoded semantic data is: the motor vehicle signal light indicates please wait; the pedestrian crossing signal light indicates you can go.
[0159] For semantic model 4: Urban Street Perception Sharing Scenario - Traffic Participants - Motor Vehicles, the following relevant information is transmitted:
[0160] Decoded as follows:
[0161] That is, the received semantic encoding is {0x7FFF 0xFF 0x01 0x01,0x7FFF 0xFF 0x18 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000} and {0x7FFF 0xFF 0x01 0x00,0x7FFF 0xFF 0x0C 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000 0x0000}, the decoded semantic data is: a car is traveling at a low speed, trajectory: latitude and longitude coordinates 1 (x1, y1, t1), coordinates 2 (x2, y2, t2); and another car: the vehicle slows down and stops, latitude and longitude coordinates (x, y), waiting.
[0162] For semantic model 5: Urban Street Perception Shared Scenario - Traffic Participants - Pedestrians, the following relevant information is transmitted:
[0163] Decoded as follows:
[0164] That is, the received semantic code is {0x7DFF 0xFF 0x01 0x00,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x08 0x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x01 0x01,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x0000 0x01 0x0000 0x0000 0x0000 0x01}…, the decoded semantic data is: pedestrian 1, height (high), speed (low), position (x1, y1), posture (walking), items carried (luggage), correlation: and (x2, y2) are moving accompanying relationships, and (x3, y3) are moving accompanying relationships….
[0165] Next, it is reorganized according to the scene semantics, and scene analysis, prediction and early warning are performed. An example of the overall architecture is shown in Figure 18.
[0166] Specifically, semantic messages of different time period granularities (long period, short period, very short period) can be aligned according to time first, and then the complete semantic information can be combined according to space. Figure 19 shows an example of semantic reorganization.
[0167] Therefore, the following information can be obtained after semantic reorganization based on the background knowledge base: urban street scene, background objects include buildings (x1, y1), trees (x2, y2), and street lights (x3, y3), and there is a pedestrian zebra crossing at the road intersection starting from the coordinates (x4_from, y4_from) to the position (x4_to, y4_to); the scene-perceived vehicle is a car, which has slowed down and stopped, and the position is (x5, y5); 3 pedestrians appear at the starting position of the zebra crossing in the left front, with position coordinates {(x1, y1), (x2, y2), (x3, y3)}, standing on one side of the zebra crossing, two of whom are carrying suitcases.
[0168] During the transmission of the coded message, some information may be lost or erroneous. According to the embodiment of the present application, the lost or erroneous information can be restored by using the correlation information item during the semantic reorganization process.
[0169] For example, in the example shown in FIG19 , if the semantic code is received for semantic model 5 related to pedestrians: {{0x7DFF 0xFF 0x01 0x00,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x08 0x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x01 0x01,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x0000 0x00 0x0000 0x0000 0x0000 0x0000 0x0000 0x01}{0x7DFF 0xFF 0x01 0x00,0x7DFF 0xFE 0x01 0x00,0x7DFF 0xFD 0x08 0x0000 0x0000 0x0000 0x0000,0x7DFF 0xFC 0x01 0x00,0x7DFF 0xFB 0x01 0x01,0x7DFF 0xFA 0x12 0x0000 0x0000 0x0000 0x0000 0x00 0x0000 0x0000 0x0000 0x0000 0x02}}. The decoded semantic data are: pedestrian 1, height (high), speed (low), position (x1, y1), posture (walking), items carried (luggage suitcase), correlation: and (x2, y2) are moving accompanying relations, and (x3, y3) are moving accompanying relations; pedestrian 3, height (high), speed (low), position (x3, y3), posture (walking), items carried (luggage suitcase), correlation: and (x1, y1) are moving accompanying relations, and (x2, y2) are moving accompanying relations.
[0170] The data for pedestrian 2 is lost. In this example, the semantic correlation between symbols can be used to help the receiver recover the lost information. For example, the data for pedestrian 1 and pedestrian 3 can be used to perform correlation identification and semantic fault tolerance processing to recover the data for pedestrian 2. The obtained semantic data for pedestrian 2 is as follows: three pedestrians appear on the zebra crossing in front of the left, walking at a low speed, with location coordinates {(x1, y1), (x1, y1), (x3, y3)}, two of whom are carrying suitcases. Therefore, through this correlation identification and semantic fault tolerance processing, the accuracy of information transmission can be improved.
[0171] As shown in Figure 18, after the scene semantics are reorganized, an appropriate AI analysis model can be selected based on the scene to further obtain analysis and prediction information. For example, in the above example, the scene is urban street blind spot perception sharing. According to the scene analysis model table (for example, Table 8 in the second embodiment), the following analysis model can be selected.
[0172] After analyzing semantic data using the selected analysis model, for example, the following prediction, warning, and assessment information can be obtained. For example, pedestrian posture analysis predicts that a pedestrian standing on the left side of the zebra crossing is likely to cross the road; pedestrian risk assessment indicates low risk (low speed, two people carrying luggage, and otherwise normal posture); and shared warnings based on surrounding vehicle perception: A pedestrian crossing warning is issued, alerting surrounding vehicles to slow down and wait.
[0173] It should be understood that the above description is only illustrative, and not restrictive.
[0174] <Fourth embodiment>
[0175] In the process of describing the electronic device in the above embodiments, it is obvious that some processes or methods are also disclosed. Below, an overview of these methods is given without repeating some of the details already discussed above, but it should be noted that although these methods are disclosed in the process of describing the electronic device, these methods do not necessarily use the components described or are not necessarily performed by those components. For example, the embodiments of the electronic device can be partially or completely implemented using hardware and / or firmware, and the methods for the electronic device discussed below can be completely implemented by a computer-executable program, although these methods can also use the hardware and / or firmware of the electronic device.
[0176] FIG20 shows a flowchart of a method for an electronic device according to an embodiment of the present application. As shown in FIG20 , the method includes: performing spatiotemporal segmentation on pulse video data obtained by an event camera to obtain pulse video data corresponding to objects of different categories (S11); performing semantic recognition on objects of different categories based on the pulse video data corresponding to the objects of different categories (S12); and encoding the results of the semantic recognition to obtain an encoded message to be provided to other electronic devices (S13). This method can be performed, for example, on a vehicle or RSU in a connected vehicle network.
[0177] For example, in step S11, spatiotemporal segmentation can be performed as follows: the space corresponding to the pulse video data is segmented according to a predetermined separation principle, and the predetermined separation principle may include one or more of the following: global or region of interest, movement status of the region of interest; and for each segmented space, event accumulation based on the pulse video data is performed at a time period granularity corresponding to the space.
[0178] Step S11 may also include: for each divided space, determining the time period granularity corresponding to the space based on the light intensity change in the space and the movement status of the object in the space, wherein the greater the light intensity change in the space and the faster the object moves, the smaller the time period granularity corresponding to the space is determined to be.
[0179] For example, in step S12 , semantic recognition may be performed on objects of each category among objects of different categories using a semantic recognition model corresponding to the category of objects.
[0180] In step S13, the semantically recognized object can be used as a semantic model recognition subject, and an encoded message is generated for the semantic model recognition subject, the encoded message including the semantic feature items of the semantic model recognition subject. For example, each semantic feature item can be encoded from top to bottom according to a semantic encoding tree. When an event camera is shooting a traffic scene, the semantic encoding tree includes the following layers from top to bottom: a first traffic classification, a second traffic classification, a semantic model recognition subject, and semantic feature items. For example, the first traffic classification includes: traffic maps, traffic signals, traffic participants, and road obstacles; the second traffic classification includes: urban streets, mountain roads, road intersections, and parking lots under the traffic map category; traffic lights and traffic signs under the traffic signal category; motor vehicles, non-motor vehicles, and pedestrians under the traffic participant category; and immovable obstacles and movable obstacles under the road obstacle category; the semantic model recognition subject includes objects under each second traffic classification; and the semantic feature items include the feature name, value length, and value of the semantic model recognition subject.
[0181] In addition, semantic feature items may also include correlations between the semantic model recognition subject and other semantic model recognition subjects. For example, as shown in Figure 21, step S13 may also include the following sub-steps: determining whether there is a correlation between two or more objects (S131); if a correlation is determined to exist, adding correlation semantic feature items to the corresponding objects (S132), and then proceeding to semantic encoding (S133); otherwise, proceeding directly to S133.
[0182] In step S13, the semantic model identification subject and the semantic feature items to be included in the generated encoded message can also be determined based on the IoV application scenario. The encoded message can also include a field related to the IoV application scenario. The encoded message obtained by encoding the semantic recognition results of objects in each space can be sent at the time period granularity of each space.
[0183] When the method is applied to the Internet of Vehicles (IoV), the method may further include extending the IoV message set to add a DayIII Ext Msg message body that supports semantic communication, wherein the DayIII Ext Msg message body includes an encoded message.
[0184] The above method corresponds to the electronic device 100 in the first embodiment, and the relevant detailed descriptions have been given in the first embodiment and the third embodiment, which will not be repeated here.
[0185] FIG22 shows a flowchart of a method for an electronic device according to another embodiment of the present application. As shown in FIG22 , the method includes: receiving a coded message from another electronic device (S21), the coded message being obtained by the other electronic device by: performing spatiotemporal segmentation on the pulse video data obtained by the event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to the objects of different categories; and encoding the results of the semantic recognition; and performing semantic decoding on the coded message and performing semantic fusion and reorganization to obtain the semantics of the original scene captured by the event camera (S22). This method can be executed on the vehicle side of a connected vehicle, for example.
[0186] In the case of a vehicle, for example, the encoded message may be included in a DayIII Ext Msg message body, which is an extension to the vehicle networking message set to support semantic communication.
[0187] For example, step S22 may include: aligning encoded messages for different categories of objects in time, where the encoded messages for different categories of objects are sent at different time period granularities; and combining semantically decoded semantic information according to spatial position to obtain the complete semantics of the original scene. For example, the time period granularity used to align the encoded messages may be determined based on current needs.
[0188] Similarly, the encoded message includes semantic feature items of the semantic model recognition subject, and the semantic model recognition subject is an object that has been semantically recognized. Each semantic feature item is encoded from top to bottom according to the semantic coding tree. In the case where the event camera is shooting a traffic scene, the semantic coding tree may include the following levels from top to bottom: a first traffic classification, a second traffic classification, a semantic model recognition subject, and a semantic feature item. The first traffic classification includes: traffic maps, traffic signals, traffic participants, and road obstacles; the second traffic classification includes: urban streets, mountain roads, road intersections and parking lots under the traffic map category, traffic lights and traffic signs under the traffic signal category, motor vehicles, non-motor vehicles and pedestrians under the traffic participant category, and immovable obstacles and movable obstacles under the road obstacle category; the semantic model recognition subject includes objects under each second traffic classification; and the semantic feature item includes the feature name, value length and value of the semantic model recognition subject.
[0189] Furthermore, semantic feature items can also include correlations between semantic model recognition subjects and other semantic model recognition subjects. Correlations, for example, include tracking motion or spatial occupancy. Semantic error tolerance can be performed based on the correlation information in the semantic feature items. Semantic error tolerance includes, for example, correcting erroneous semantics and recovering lost semantics.
[0190] FIG23 shows an example of a flowchart of step S22 including semantic error tolerance processing. The specific process is as follows: semantically decode the encoded message ( S221 ); perform a correlation determination based on the semantic decoding result to determine whether two or more objects are correlated ( S222 ), for example, determining whether there are correlation semantic feature items for the two or more objects; if the determination in step S222 is yes, proceed to step S223 to perform a correlation analysis, for example, to determine whether there is motion tracking or spatial occupancy; otherwise, proceed to step S226 to perform scene semantic reconstruction; then, in step S224, determine whether there are semantic errors or loss based on the correlation analysis results; if it is determined in step S224 that there are semantic errors or loss, proceed to step S225 to perform error tolerance processing; otherwise, proceed to step S226 to perform scene semantic reconstruction.
[0191] Furthermore, as shown in FIG22 , the method may further include step S23: using a scenario analysis and prediction model to analyze and process the semantics of the obtained original scenario to obtain risk assessment and / or prediction and warning information. For example, an appropriate artificial intelligence analysis model may be selected as the scenario analysis and prediction model based on the scenario type. In the case of a vehicle network, the scenario type may include a vehicle network application scenario.
[0192] For example, the scenario type can be determined based on the Internet of Vehicles scenario knowledge stored in the background knowledge base. The encoded message can also include a field for the Internet of Vehicles application scenario.
[0193] The above method corresponds to the electronic device 200 in the second embodiment, and the relevant detailed descriptions have been given in the second and third embodiments and will not be repeated here.
[0194] The technology of the present disclosure can be applied to various products.
[0195] For example, the electronic devices 100 and 200 may be implemented as various user devices. The user device may be implemented as a mobile terminal (such as a smartphone, a tablet personal computer (PC), a notebook PC, a portable game terminal, a portable / dongle-type mobile router, and a digital camera), an in-vehicle terminal (such as a car navigation device), or a vehicle. The user device may also be implemented as a terminal that performs machine-to-machine (M2M) communication (also known as a machine-type communication (MTC) terminal). In addition, the user device may be a wireless communication module (such as an integrated circuit module including a single chip) installed on each of the above-mentioned terminals.
[0196] [Application examples on user devices]
[0197] (First application example)
[0198] 24 is a block diagram showing an example of a schematic configuration of a smartphone 900 to which the technology of the present disclosure can be applied. The smartphone 900 includes a processor 901, a memory 902, a storage device 903, an external connection interface 904, a camera 906, a sensor 907, a microphone 908, an input device 909, a display device 910, a speaker 911, a wireless communication interface 912, one or more antenna switches 915, one or more antennas 916, a bus 917, a battery 918, and an auxiliary controller 919.
[0199] The processor 901 may be, for example, a CPU or a system on a chip (SoC), and controls the functions of the application layer and other layers of the smartphone 900. The memory 902 includes RAM and ROM, and stores data and programs executed by the processor 901. The storage device 903 may include storage media such as semiconductor memories and hard disks. The external connection interface 904 is an interface for connecting external devices (such as memory cards and universal serial bus (USB) devices) to the smartphone 900.
[0200] The camera 906 includes an image sensor such as a charge coupled device (CCD) and a complementary metal oxide semiconductor (CMOS) and generates a captured image. The sensor 907 may include a group of sensors such as a measurement sensor, a gyroscope sensor, a geomagnetic sensor, and an acceleration sensor. The microphone 908 converts the sound input to the smartphone 900 into an audio signal. The input device 909 includes, for example, a touch sensor, a keypad, a keyboard, a button, or a switch configured to detect a touch on the screen of the display device 910, and receives an operation or information input from the user. The display device 910 includes a screen such as a liquid crystal display (LCD) and an organic light emitting diode (OLED) display and displays an output image of the smartphone 900. The speaker 911 converts the audio signal output from the smartphone 900 into sound.
[0201] The wireless communication interface 912 supports any cellular communication scheme (such as LTE and LTE-Advanced) and performs wireless communications. The wireless communication interface 912 may typically include, for example, a BB processor 913 and an RF circuit 914. The BB processor 913 may perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and may also perform various types of signal processing for wireless communications. Meanwhile, the RF circuit 914 may include, for example, mixers, filters, and amplifiers, and transmit and receive wireless signals via an antenna 916. Note that while the figure shows a scenario where one RF link is connected to one antenna, this is merely illustrative, and also encompasses scenarios where one RF link is connected to multiple antennas via multiple phase shifters. The wireless communication interface 912 may be a chip module on which the BB processor 913 and RF circuit 914 are integrated. As shown in FIG. 24 , the wireless communication interface 912 may include multiple BB processors 913 and multiple RF circuits 914. While FIG. 24 illustrates an example in which the wireless communication interface 912 includes multiple BB processors 913 and multiple RF circuits 914, the wireless communication interface 912 may also include a single BB processor 913 or a single RF circuit 914.
[0202] In addition, in addition to the cellular communication scheme, the wireless communication interface 912 can support other types of wireless communication schemes, such as a short-range wireless communication scheme, a near-field communication scheme, and a wireless local area network (LAN) scheme. In this case, the wireless communication interface 912 may include a BB processor 913 and an RF circuit 914 for each wireless communication scheme.
[0203] Each of the antenna switches 915 switches a connection destination of the antenna 916 between a plurality of circuits (eg, circuits for different wireless communication schemes) included in the wireless communication interface 912 .
[0204] Each of the antennas 916 includes a single or multiple antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for transmitting and receiving wireless signals via the wireless communication interface 912. As shown in FIG24, the smartphone 900 may include multiple antennas 916. Although FIG24 shows an example in which the smartphone 900 includes multiple antennas 916, the smartphone 900 may also include a single antenna 916.
[0205] In addition, the smartphone 900 may include an antenna 916 for each wireless communication scheme. In this case, the antenna switch 915 may be omitted from the configuration of the smartphone 900.
[0206] The bus 917 connects the processor 901, the memory 902, the storage device 903, the external connection interface 904, the camera 906, the sensor 907, the microphone 908, the input device 909, the display device 910, the speaker 911, the wireless communication interface 912, and the auxiliary controller 919. The battery 918 supplies power to the various blocks of the smartphone 900 shown in FIG. 24 via feeders, which are partially shown as dotted lines in the figure. The auxiliary controller 919 operates the minimum necessary functions of the smartphone 900, for example, in sleep mode.
[0207] In the smartphone 900 shown in FIG24 , the communication unit 104 and transceiver of the electronic device 100 may be implemented by the wireless communication interface 912. At least a portion of the functionality may also be implemented by the processor 901 or the auxiliary controller 919. For example, the processor 901 or the auxiliary controller 919 may perform spatiotemporal segmentation, semantic recognition, and semantic encoding on the impulse video data by executing the functions of the spatiotemporal segmentation unit 101, the semantic recognition unit 102, the semantic encoding unit 103, and the communication unit 104, and provide the encoded message to other electronic devices, thereby significantly reducing the amount of data transmission and improving the accuracy of information transmission.
[0208] The communication unit 201 of the electronic device 200 can be implemented by the wireless communication interface 912. At least part of the functions can also be implemented by the processor 901 or the auxiliary controller 919. For example, the processor 901 or the auxiliary controller 919 can decode and semantically fuse and reorganize semantically encoded messages from other electronic devices and perform scene analysis and prediction by executing the functions of the communication unit 201, the decoding and semantic reorganization unit 202, and the analysis and prediction unit 203, thereby achieving end-to-end perception, communication and recognition with low processing latency, greatly reducing computing power requirements, improving the accuracy of information transmission, and better supporting vehicle network application scenarios.
[0209] (Second application example)
[0210] 25 is a block diagram showing an example of a schematic configuration of a car navigation device 920 to which the technology of the present disclosure can be applied. The car navigation device 920 includes a processor 921, a memory 922, a global positioning system (GPS) module 924, a sensor 925, a data interface 926, a content player 927, a storage medium interface 928, an input device 929, a display device 930, a speaker 931, a wireless communication interface 933, one or more antenna switches 936, one or more antennas 937, and a battery 938.
[0211] The processor 921 may be, for example, a CPU or an SoC, and controls a navigation function and other functions of the car navigation apparatus 920. The memory 922 includes a RAM and a ROM, and stores data and programs executed by the processor 921.
[0212] The GPS module 924 measures the position (such as latitude, longitude, and altitude) of the car navigation device 920 using GPS signals received from GPS satellites. The sensor 925 may include a group of sensors such as a gyroscope sensor, a geomagnetic sensor, and an air pressure sensor. The data interface 926 is connected to, for example, the in-vehicle network 941 via an unillustrated terminal and acquires data generated by the vehicle (such as vehicle speed data).
[0213] The content player 927 reproduces content stored in a storage medium (such as a CD or DVD) inserted into the storage medium interface 928. The input device 929 includes, for example, a touch sensor, button, or switch configured to detect a touch on the screen of the display device 930, and receives an operation or information input from the user. The display device 930 includes a screen such as an LCD or OLED display and displays an image of a navigation function or reproduced content. The speaker 931 outputs the sound of the navigation function or the reproduced content.
[0214] The wireless communication interface 933 supports any cellular communication scheme (such as LTE and LTE-Advanced) and performs wireless communication. The wireless communication interface 933 may generally include, for example, a BB processor 934 and an RF circuit 935. The BB processor 934 may perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and perform various types of signal processing for wireless communication. Meanwhile, the RF circuit 935 may include, for example, a mixer, a filter, and an amplifier, and transmit and receive wireless signals via an antenna 937. The wireless communication interface 933 may also be a chip module on which the BB processor 934 and the RF circuit 935 are integrated. As shown in Figure 25, the wireless communication interface 933 may include multiple BB processors 934 and multiple RF circuits 935. Although Figure 25 shows an example in which the wireless communication interface 933 includes multiple BB processors 934 and multiple RF circuits 935, the wireless communication interface 933 may also include a single BB processor 934 or a single RF circuit 935.
[0215] In addition, in addition to the cellular communication scheme, the wireless communication interface 933 can support other types of wireless communication schemes, such as a short-range wireless communication scheme, a near field communication scheme, and a wireless LAN scheme. In this case, for each wireless communication scheme, the wireless communication interface 933 can include a BB processor 934 and an RF circuit 935.
[0216] Each of the antenna switches 936 switches a connection destination of the antenna 937 between a plurality of circuits included in the wireless communication interface 933 , such as circuits for different wireless communication schemes.
[0217] Each of the antennas 937 includes a single or multiple antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for transmitting and receiving wireless signals via the wireless communication interface 933. As shown in FIG25, the car navigation device 920 may include multiple antennas 937. Although FIG25 shows an example in which the car navigation device 920 includes multiple antennas 937, the car navigation device 920 may also include a single antenna 937.
[0218] Furthermore, the car navigation device 920 may include an antenna 937 for each wireless communication scheme. In this case, the antenna switch 936 may be omitted from the configuration of the car navigation device 920.
[0219] The battery 938 supplies power to the respective blocks of the car navigation device 920 shown in Fig. 25 via a feeder line, which is partially shown as a dotted line in the figure. The battery 938 accumulates the power supplied from the vehicle.
[0220] In the car navigation device 920 shown in FIG25 , the communication unit 104 and transceiver of the electronic device 100 can be implemented by the wireless communication interface 933. At least a portion of the functionality can also be implemented by the processor 921. For example, the processor 921 can perform spatiotemporal segmentation, semantic recognition, and semantic encoding on the impulse video data by executing the functions of the spatiotemporal segmentation unit 101, the semantic recognition unit 102, the semantic encoding unit 103, and the communication unit 104, and provide the encoded message to other electronic devices, thereby significantly reducing the amount of data transmission and improving the accuracy of information transmission.
[0221] The communication unit 201 of the electronic device 200 can be implemented by the wireless communication interface 933. At least a portion of its functionality can also be implemented by the processor 921. For example, the processor 921 can perform decoding and semantic fusion reorganization of semantically encoded messages from other electronic devices, as well as scenario analysis and prediction, by executing the functions of the communication unit 201, the decoding and semantic reorganization unit 202, and the analysis and prediction unit 203. This can achieve end-to-end perception, communication, and recognition with low processing latency, significantly reducing computing power requirements, improving the accuracy of information transmission, and better supporting Internet of Vehicles application scenarios.
[0222] The technology of the present disclosure can also be implemented as an in-vehicle system (or vehicle) 940 including a car navigation device 920, an in-vehicle network 941, and one or more blocks of a vehicle module 942. The vehicle module 942 generates vehicle data (such as vehicle speed, engine speed, and fault information) and outputs the generated data to the in-vehicle network 941.
[0223] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be pointed out that for those skilled in the art, it is understandable that all or any steps or components of the methods and devices of the present disclosure can be implemented in any computing device (including a processor, storage medium, etc.) or a network of computing devices in the form of hardware, firmware, software, or a combination thereof. This can be achieved by those skilled in the art using their basic circuit design knowledge or basic programming skills after reading the description of the present disclosure.
[0224] Furthermore, the present disclosure also provides a program product storing machine-readable instruction codes. When the instruction codes are read and executed by a machine, the method according to the embodiment of the present disclosure can be executed.
[0225] Accordingly, the storage medium for carrying the program product storing the machine-readable instruction code is also included in the disclosure of the present invention, including but not limited to a floppy disk, an optical disk, a magneto-optical disk, a memory card, a memory stick, and the like.
[0226] When the present disclosure is implemented through software or firmware, the programs constituting the software are installed from a storage medium or a network to a computer with a dedicated hardware structure (such as the general-purpose computer 2600 shown in Figure 26). When various programs are installed on the computer, it can perform various functions, etc.
[0227] 26 , a central processing unit (CPU) 2601 executes various processes according to a program stored in a read-only memory (ROM) 2602 or a program loaded from a storage section 2608 to a random access memory (RAM) 2603. In the RAM 2603, data required when the CPU 2601 executes various processes, etc., is also stored as needed. The CPU 2601, the ROM 2602, and the RAM 2603 are connected to each other via a bus 2604. An input / output interface 2605 is also connected to the bus 2604.
[0228] The following components are connected to the input / output interface 2605: an input section 2606 (including a keyboard, a mouse, etc.), an output section 2607 (including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and speakers, etc.), a storage section 2608 (including a hard disk, etc.), and a communication section 2609 (including a network interface card such as a LAN card, a modem, etc.). The communication section 2609 performs communication processing via a network such as the Internet. A drive 2610 may also be connected to the input / output interface 2605 as needed. A removable medium 2611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed in the drive 2610 as needed, so that a computer program read therefrom is installed in the storage section 2608 as needed.
[0229] In the case where the above-described series of processing is realized by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 2611 .
[0230] It should be understood by those skilled in the art that such storage media is not limited to the removable medium 2611 shown in FIG. 26 , which stores the program and is distributed separately from the device to provide the program to the user. Examples of the removable medium 2611 include magnetic disks (including floppy disks (registered trademark)), optical disks (including compact disk read-only memories (CD-ROMs) and digital versatile disks (DVDs)), magneto-optical disks (including minidiscs (MDs) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be a ROM 2602, a hard disk included in the storage portion 2608, or the like, in which the program is stored and distributed to the user together with the device containing them.
[0231] It should also be noted that in the apparatus, method, and system of the present disclosure, each component or step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure. Furthermore, the steps of performing the above series of processes can naturally be performed in chronological order according to the order of description, but do not necessarily need to be performed in chronological order. Certain steps can be performed in parallel or independently of each other.
[0232] Finally, it should be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. Furthermore, in the absence of further limitations, an element defined by the phrase "comprises a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0233] Although the embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, it should be understood that the embodiments described above are merely illustrative of the present disclosure and are not intended to limit the present disclosure. Those skilled in the art will appreciate that various modifications and variations can be made to the above embodiments without departing from the spirit and scope of the present disclosure. Therefore, the scope of the present disclosure is solely defined by the appended claims and their equivalents.
Claims
1. An electronic device comprising: at least one processor; as well as at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, cause the electronic device to execute: Performing spatiotemporal segmentation on the pulse video data obtained by the event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to the objects of different categories; and The result of the semantic recognition is encoded to obtain an encoded message to be provided to other electronic devices.
2. The electronic device according to claim 1, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to perform the spatiotemporal segmentation as follows: Segmenting the space corresponding to the pulse video data according to a predetermined separation principle, wherein the predetermined separation principle includes one or more of the following: global or region of interest, and movement of the region of interest; as well as For each divided space, event accumulation based on the pulse video data is performed at a time period granularity corresponding to the space.
3. The electronic device according to claim 2, wherein The at least one memory and the computer program code are also configured to enable the electronic device to determine, through the at least one processor, for each divided space, a time period granularity corresponding to the space based on the light intensity change in the space and the movement status of the object in the space, wherein the greater the light intensity change in the space and the faster the object moves, the smaller the time period granularity corresponding to the space is determined to be.
4. The electronic device according to claim 1, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: For each category of objects in the different categories of objects, the semantic recognition is performed using a semantic recognition model corresponding to the category of objects.
5. The electronic device according to claim 1, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to perform the semantic encoding as follows: The object that has undergone semantic recognition is used as a semantic model recognition subject, and the coded message is generated for the semantic model recognition subject, wherein the coded message includes the semantic feature items of the semantic model recognition subject. The electronic device according to claim 5 , wherein: Each semantic feature item is encoded from top to bottom according to the semantic coding tree.
7. The electronic device according to claim 6, wherein: The event camera shoots traffic scenes, and the semantic coding tree includes the following levels from top to bottom: first traffic classification, second traffic classification, semantic model recognition subject, and semantic feature item.
8. The electronic device according to claim 7, wherein: The first traffic classification includes: traffic maps, traffic signals, traffic participants, and road obstacles; The second traffic classification includes: urban streets, mountain roads, road intersections and parking lots under the traffic map category, signal lights and traffic signs under the traffic signal category, motor vehicles, non-motor vehicles and pedestrians under the traffic participant category, and immovable obstacles and movable obstacles under the road obstacle category; The semantic model recognition subject includes objects under each second traffic classification; and The semantic feature item includes the feature name, value length and value of the semantic model identification subject.
9. The electronic device according to claim 8, wherein The semantic feature item includes the correlation between the semantic model recognition subject and other semantic model recognition subjects.
10. The electronic device according to claim 7, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: According to the application scenario of the Internet of Vehicles, the semantic model identification subject to generate the encoded message and the semantic feature items to be included are determined.
11. The electronic device according to claim 10, wherein: The encoded message also includes a field for the Internet of Vehicles application scenario.
12. The electronic device according to claim 2, wherein, The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: The encoded message obtained by encoding the semantic recognition result of the object in each space is sent according to the time period granularity of the space.
13. The electronic device according to claim 1, wherein The electronic device is located in a vehicle or a roadside unit in the vehicle network. The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: Expand the Internet of Vehicles message set to add the DayIII Ext Msg message body to support semantic communication.
14. The electronic device according to claim 13, wherein: The DayIII Ext Msg message body includes the encoded message.
15. An electronic device comprising: at least one processor; as well as at least one memory including computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, cause the electronic device to execute: Receiving a coded message from another electronic device, the coded message being obtained by the other electronic device by: performing spatiotemporal segmentation on pulse video data obtained by an event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on the objects of different categories based on the pulse video data corresponding to the objects of different categories; and encoding the results of the semantic recognition; and The encoded message is semantically decoded and semantically fused and reorganized to obtain the semantics of the original scene captured by the event camera.
16. The electronic device according to claim 15, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: aligning coded messages for different categories of objects in time, wherein the coded messages for the different categories of objects are sent at different time period granularities; and The semantic information after semantic decoding is combined according to spatial positions to obtain the complete semantics of the original scene.
17. The electronic device according to claim 15, wherein: The coded message includes semantic feature items of a semantic model recognition subject, where the semantic model recognition subject is an object that has been semantically recognized.
18. The electronic device according to claim 17, wherein: Each semantic feature item is encoded from top to bottom according to the semantic coding tree.
19. The electronic device according to claim 18, wherein The event camera shoots traffic scenes, and the semantic coding tree includes the following levels from top to bottom: first traffic classification, second traffic classification, semantic model recognition subject, and semantic feature item.
20. The electronic device according to claim 19, wherein The first traffic classification includes: traffic maps, traffic signals, traffic participants, and road obstacles; The second traffic classification includes: urban streets, mountain roads, road intersections and parking lots under the traffic map category, signal lights and traffic signs under the traffic signal category, motor vehicles, non-motor vehicles and pedestrians under the traffic participant category, and immovable obstacles and movable obstacles under the road obstacle category; The semantic model recognition subject includes objects under each second traffic classification; and The semantic feature item includes the feature name, value length and value of the semantic model identification subject.
21. The electronic device according to claim 20, wherein The semantic feature item includes the correlation between the semantic model recognition subject and other semantic model recognition subjects.
22. The electronic device according to claim 21, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: Semantic fault tolerance processing is performed based on the correlation information in the semantic feature items.
23. The electronic device according to claim 22, wherein: The semantic fault tolerance processing includes correction of erroneous semantics and recovery of lost semantics.
24. The electronic device according to claim 15, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: The obtained semantics of the original scene are analyzed and processed using a scene analysis prediction model to obtain risk assessment and / or prediction warning information.
25. The electronic device according to claim 24, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: Select a suitable artificial intelligence analysis model as the scenario analysis prediction model according to the scenario type.
26. The electronic device according to claim 25, wherein The electronic device is located in a vehicle in the Internet of Vehicles, and the scenario type includes an Internet of Vehicles application scenario.
27. The electronic device according to claim 26, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: The scenario type is determined based on the Internet of Vehicles scenario knowledge stored in the background knowledge base.
28. The electronic device according to claim 26, wherein The encoded message also includes a field for the Internet of Vehicles application scenario.
29. The electronic device according to claim 16, wherein The at least one memory and the computer program code are further configured to, through the at least one processor, cause the electronic device to: A time period granularity for aligning the encoded messages is determined based on current requirements.
30. The electronic device according to claim 15, wherein The encoded message is included in a DayIII Ext Msg message body, which is an extension of the Internet of Vehicles message set to support semantic communication.
31. A method for an electronic device, comprising: Performing spatiotemporal segmentation on the pulse video data obtained by the event camera to obtain pulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to the objects of different categories; as well as The result of the semantic recognition is encoded to obtain an encoded message to be provided to other electronic devices.
32. A method for an electronic device, comprising: receiving a coded message from another electronic device, the coded message being obtained by the other electronic device by performing spatiotemporal segmentation on impulse video data obtained by an event camera to obtain impulse video data corresponding to objects of different categories; performing semantic recognition on objects of different categories based on the pulse video data corresponding to the objects of different categories; and encoding the results of semantic recognition; as well as The encoded message is semantically decoded and semantically fused and reorganized to obtain the semantics of the original scene captured by the event camera.
33. A computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method according to claim 31 or 32.
34. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the steps of the method according to claim 31 or 32 are implemented.
Citation Information
Patent Citations
Video data processing method and device, electronic equipment and storage medium
CN109922372A
Event-based adaptation of coding parameters for video image encoding
CN112534816A
Low, small and slow target identification and positioning method and system based on mixed vision
CN115035470A
Neuromorphic visual target identification method and device
CN117370858A
Frame rate adjusting method, device, equipment and system
CN117425064A