Dynamic differential transforms for segmentation

By generating differential images and processing image features using transformation operations, the inconsistency problem of image segmentation mask under different image signal processor settings is solved, and more stable segmentation mask generation is achieved, and the quality of image processing is improved.

CN120303697APending Publication Date: 2025-07-11QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380079321.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-03
Filing Date
2023-09-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the image segmentation mask has inconsistent semantics or instance representations under different image signal processor settings, resulting in time inconsistency problems and flicker artifacts.

Method used

By generating a differential image and processing the features of the differential image and the previous image using a transformation operation, a combined feature representation is generated in combination with the current image features, thereby generating a consistent segmentation mask.

Benefits of technology

Reduce or eliminate time inconsistencies between segmentation masks and improve the quality of image processing operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303697A_ABST
    Figure CN120303697A_ABST
Patent Text Reader

Abstract

Techniques and systems are provided for generating one or more partition masks. For example, a method can include generating a differential image based on a difference between a current image and a previous image. The method can also include processing the differential image and a feature representing the previous image using a transformation operation to generate a transformed feature representation of the previous image. The method can include combining the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image. The method can also include generating a segmentation mask for the current image based on the combined feature representation of the current image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to processing image data to perform segmentation (e.g., semantic segmentation, instance segmentation, etc.). For example, aspects of the present disclosure include systems and techniques for performing segmentation using differential images or different images (e.g., based on the difference between an input image for a current time frame and an input image for a previous time frame). Background Art

[0002] An increasing number of devices or systems (e.g., autonomous vehicles such as autonomous and semi-autonomous vehicles, drones or unmanned aerial vehicles (UAVs), mobile robots; mobile devices such as mobile phones, extended reality (XR) devices, and other suitable devices or systems) include multiple sensors for collecting information about the environment, and a processing system for processing the information for various purposes such as for route planning, navigation, collision avoidance, etc.). An example of such a system is an advanced driver assistance system (ADAS) for an autonomous or semi-autonomous vehicle.

[0003] A device or system may perform segmentation on sensor data (e.g., one or more images) to generate a segmentation output (e.g., a segmentation mask or a segmentation map). Based on the segmentation, objects may be identified and labeled with corresponding classifications of specific objects (e.g., a person, a car, a background, etc.) within the image or video. This labeling may be performed on a per-pixel basis. The segmentation mask may be a representation of the labels of the image or view. The segmentation output may then be used to perform one or more operations such as image processing (e.g., blurring a portion of the image). Maintaining the consistency of the segmentation output over time (referred to as temporal consistency) may be difficult, resulting in visual defects in the output. Summary of the Invention

[0004] The following presents a simplified summary of one or more aspects related to the present disclosure. Accordingly, the following summary should not be considered an exhaustive overview of all contemplated aspects, nor should it be considered to identify key or critical elements of all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts in a simplified form related to one or more aspects of the mechanisms disclosed herein prior to the detailed description presented below.

[0005] According to at least one example, a processor-implemented method is provided. The method includes: generating a difference image based on a difference between a current image and a previous image; using a transformation operation to process the difference image and features representing the previous image to generate a transformed feature representation of the previous image; combining the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image; and generating a segmentation mask for the current image based on the combined feature representation of the current image.

[0006] In another example, an apparatus is provided. The apparatus includes: at least one memory; and at least one processor coupled to the at least one memory and configured to: generate a difference image based on a difference between a current image and a previous image; use a transformation operation to process the difference image and features representing the previous image to generate a transformed feature representation of the previous image; combine the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image; and generate a segmentation mask for the current image based on the combined feature representation of the current image.

[0007] In another example, a non-transitory computer-readable medium storing instructions is provided. The instructions, when executed by one or more processors, cause the one or more processors to: generate a difference image based on a difference between a current image and a previous image; use a transformation operation to process the difference image and features representing the previous image to generate a transformed feature representation of the previous image; combine the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image; and generate a segmentation mask for the current image based on the combined feature representation of the current image.

[0008] In another example, an apparatus is provided. The apparatus includes: means for generating a difference image based on a difference between a current image and a previous image; means for using a transformation operation to process the difference image and features representing the previous image to generate a transformed feature representation of the previous image; means for combining the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image; and means for generating a segmentation mask for the current image based on the combined feature representation of the current image.

[0009] In some aspects, the device is, is part of, and / or includes the following: a computing device or component of a vehicle or vehicle (e.g., an autonomous vehicle), a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a camera, a wearable device (e.g., a network-connected watch, etc.), a personal computer, a laptop computer, a server computer, or other device. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the device described above may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more accelerometers, any combination thereof, and / or other sensors).

[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood in reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.

[0011] The foregoing and other features and aspects will become more apparent when reference is made to the following specification, claims, and appended drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Exemplary aspects of the present application are described in detail below with reference to the following drawings:

[0013] Figure 1A and Figure 1B are block diagrams of vehicles suitable for implementing various aspects in accordance with aspects of the present disclosure.

[0014] Figure 1C are block diagrams of components of a vehicle suitable for implementing various aspects in accordance with aspects of the present disclosure;

[0015] Figure 1D illustrates an exemplary implementation of a system on a chip (SOC) according to some examples;

[0016] Figure 2A is a component block diagram of components of an exemplary vehicle management system according to various aspects;

[0017] Figure 2B is a component block diagram of components of another exemplary vehicle management system according to various aspects;

[0018] Figures 3A to 3DAnd Figure 4 is a diagram illustrating an example of a neural network according to some aspects;

[0019] Figure 5 are an image and corresponding segmentation masks illustrating examples of inconsistent segmentation results of the same image under different Image Signal Processor (ISP) settings according to aspects of the present disclosure;

[0020] Figure 6 are segmentation masks illustrating examples of inconsistent segmentation results between adjacent images according to aspects of the present disclosure;

[0021] Figure 7 is a diagram illustrating an example of a machine learning system for generating a segmentation mask from an image according to aspects of the present disclosure;

[0022] Figure 8 is a diagram illustrating an example of the concatenation of features of a previous image and features of a current image according to aspects of the present disclosure;

[0023] Figure 9 is a diagram illustrating an example of a machine learning system including a transformation operation for generating a segmentation mask from an image according to aspects of the present disclosure;

[0024] Figure 10 is a diagram illustrating an example of a machine learning system including a transformation operation for generating a segmentation mask from an image using a difference image according to aspects of the present disclosure;

[0025] Figure 11 is a diagram illustrating an example of a convolution operation that varies based on the value of a difference image according to aspects of the present disclosure;

[0026] Figure 12 is a flowchart illustrating an example of a process for processing one or more images according to aspects of the present disclosure; and

[0027] Figure 13 illustrates an example computing device architecture of an example computing device that can implement the various techniques described herein. Detailed Description

[0028] Certain aspects of the present disclosure are provided below. Some of these aspects may be applied independently, and some of them may be applied in combination, which will be apparent to those skilled in the art. In the following description, specific details are set forth for purposes of explanation to provide a thorough understanding of the aspects of the present application. However, it is obvious that the aspects may be implemented without these specific details. The accompanying drawings and description are not intended to be restrictive.

[0029] The following description provides only example aspects and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the example aspects will provide those skilled in the art with a description that can be used to implement the example aspects. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0030] In some cases, such as before performing one or more operations on an image or video (e.g., autonomous or semi-autonomous driving operations, applying effects to an image, etc.), an image or video frame may be processed to identify one or more objects present within the image or video frame. For example, adding a virtual background to a video conference may include identifying the objects in the foreground (e.g., people), and modifying all parts of the video frame except for the pixels under the objects. In some cases, objects in an image may be identified by using one or more neural networks or other machine learning (ML) models to assign segmentation classes (e.g., human class, car class, background class, etc.) to each pixel in the frame and then grouping adjacent pixels that share a segmentation class to form objects of the segmentation class (e.g., people, cars, background, etc.). This technique may be referred to as per-pixel segmentation. The per-pixel labels may be referred to as segmentation masks (also referred to herein as segmentation maps).

[0031] An example of one type of segmentation is semantic segmentation that treats multiple objects of the same class as a single entity or instance (e.g., all detected people within an image are treated as the "person" class). Another type of segmentation treats multiple objects of the same class as different entities or instances (e.g., the first person detected in an image is the first instance of the "person" class, and the second person detected in the image is the second instance of the "person" class).

[0032] In some cases, per-pixel segmentation may include inputting an image into an ML model such as (but not limited to) a convolutional neural network (CNN). The ML model may process the image to output a segmentation mask or segmentation map for the image. The segmentation mask may include segmentation class information for each pixel in the frame. In some cases, the segmentation mask may be configured to retain information only for pixels corresponding to one or more classes (e.g., for pixels classified as people), thereby isolating the selected classified pixels from other classified pixels (e.g., isolating person pixels from background pixels).

[0033] Segmentation can be important for different devices or applications that include one or more cameras of mobile devices, vehicles, extended reality (XR) devices, Internet of Things (IoT) devices, and the like. Current segmentation solutions (e.g., deployed on devices) may face problems of inconsistent semantic or instance representations based on different camera settings (e.g., image signal processing (ISP) settings) and inconsistent semantic or instance representations over time (referred to as temporal inconsistency). For example, current ML systems that generate segmentation masks may produce segmentation masks with flickering artifacts due to inconsistent predictions between images or frames.

[0034] Systems, apparatuses, electronic devices, methods (also referred to as processes), and computer-readable media (collectively referred to as "systems and technologies" herein) are described herein that provide a machine learning system that utilizes a differential image to generate a segmentation mask for an image. For example, a transformation operation of a machine learning system can use a differential image to transform a previous image such that the previous image is pixel-aligned with the current image (e.g., to represent an object in the previous image in a pose similar to the pose of the object in the current image). In some examples, a computing device can generate a differential image based on the difference between a current image and a previous image. The computing device can use a transformation operation of the machine learning system (e.g., a convolution operation performed using at least one convolutional filter, a transformer operation performed using at least one transformer block, or other transformation operations) to process the differential image and features representing the previous image to generate a transformed feature representation of the previous image. The computing device can combine the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image. The computing device can then generate a segmentation mask for the current image based on the combined feature representation of the current image.

[0035] Aspects of the present application will be described with reference to the accompanying drawings.

[0036] The systems and technologies described herein can be implemented by any type of system or device. An illustrative example of a system that can be used to implement the systems and technologies described herein is a vehicle (e.g., an autonomous or semi-autonomous vehicle) or a system or component of the vehicle (e.g., an ADAS or other system or component). Other examples of systems or devices that can be used to perform the techniques described herein can include mobile devices (e.g., mobile phones or so-called "smartphones" or other mobile devices), XR devices (e.g., VR devices, AR devices, MR devices, etc.), cameras, wearable devices (e.g., network-connected watches, etc.), and / or other types of systems or devices.

[0037] Figure 1A and Figure 1B is a diagram illustrating an example vehicle 100 that can implement the systems and technologies described herein. Refer toFigure 1A and Figure 1B ,Vehicle 100 may include a control unit 140 and a plurality of sensors 102 to 138, including a satellite geolocation system receiver (e.g., a sensor) 108, occupancy sensors 112, 116, 118, 126, 128, tire pressure sensors 114, 120, cameras 122, 136, microphones 124, 134, impact sensors 130, radar 132, and light detection and ranging (LIDAR) 138. The plurality of sensors 102 to 138 disposed in or on the vehicle can be used for various purposes, such as autonomous and semi-autonomous navigation and control, collision avoidance, position determination, etc., and to provide sensor data about objects and people in or on the vehicle 100. The sensors 102 to 138 can include one or more sensors from a variety of sensors capable of detecting various information useful for navigation and collision avoidance. Each of the sensors 102 to 138 can communicate with the control unit 140 and with each other either wired or wirelessly. Specifically, the sensors can include one or more cameras 122, 136 or other optical or optoelectronic sensors. The sensors can also include other types of object detection and ranging sensors, such as radar 132, lidar 138, IR sensors, and ultrasonic sensors. The sensors can also include tire pressure sensors 114, 120, humidity sensors, temperature sensors, satellite geolocation sensors 108, accelerometers, vibration sensors, gyroscopes, gravimeters, impact sensors 130, dynamometers, pressure gauges, strain sensors, fluid sensors, chemical sensors, gas content analyzers, pH sensors, radiation sensors, Geiger counters, neutron detectors, biomaterial sensors, microphones 124, 134, occupancy sensors 112, 116, 118, 126, 128, proximity sensors, and other sensors.

[0038] The vehicle control unit 140 can be configured with processor-executable instructions to perform various aspects using information received from various sensors, particularly cameras 122, 136, radar 132, and lidar 138. In some aspects, the control unit 140 can use distance and relative positioning information (e.g., relative azimuth angle) obtainable from the radar 132 and / or lidar 138 sensors to supplement the processing of camera images. The control unit 140 can be further configured to use information about other vehicles determined using various aspects to control the steering, braking, and speed of the vehicle 100 when operating in autonomous or semi-autonomous mode.

[0039] Figure 1C is a block diagram of components of a system 150 that illustrates components and support systems suitable for implementing various aspects. Refer to Figure 1A 、 1Band 1C, the vehicle 100 may include a control unit 140, which may include various circuits and devices for controlling the operation of the vehicle 100. In Figure 1C the example illustrated in, the control unit 140 includes a processor 164, a memory 166, an input module 168, an output module 170, and a radio module 172. The control unit 140 may be coupled to and configured to control the drive control assembly 154, the navigation assembly 156, and one or more sensors 158 of the vehicle 100.

[0040] The control unit 140 may include a processor 164, which may be configured with processor-executable instructions to control the maneuvering, navigation, and / or other operations of the vehicle 100, including operations in various aspects. The processor 164 may be coupled to the memory 166. The control unit 140 may include an input module 168, an output module 170, and a radio module 172.

[0041] The radio module 172 may be configured for wireless communication. The radio module 172 may exchange signals 182 (e.g., command signals for controlling maneuvering, signals from navigation facilities, etc.) with a network node 180 and may provide the signals 182 to the processor 164 and / or the navigation assembly 156. In some aspects, the radio module 172 may enable the vehicle 100 to communicate with a wireless communication device 190 via a wireless communication link 92. The wireless communication link 92 may be a bi-directional or uni-directional communication link and may use one or more communication protocols.

[0042] The input module 168 may receive sensor data from one or more vehicle sensors 158 and electronic signals from other components (including the drive control assembly 154 and the navigation assembly 156). The output module 170 may be used to communicate with or activate various components of the vehicle 100, including the drive control assembly 154, the navigation assembly 156, and the sensors 158.

[0043] The control unit 140 may be coupled to the drive control assembly 154 to control the physical elements of the vehicle 100 related to the maneuvering and navigation of the vehicle, such as engines, motors, throttles, steering elements, other control elements, braking or decelerating elements, and so on. The drive control assembly 154 may also include components for controlling other devices of the vehicle, and these other devices include environmental controls (e.g., air conditioning and heating), external and / or internal lighting, internal and / or external information displays (which may include display screens or other devices for displaying information), safety devices (e.g., tactile devices, audible alarms, etc.), and other similar devices.

[0044] The control unit 140 may be coupled to the navigation component 156 and may receive data from the navigation component 156. The control unit 140 may be configured to use such data to determine the current location and orientation of the vehicle 100 and an appropriate route toward a destination. In various aspects, the navigation component 156 may include or be coupled to a GNSS receiver system (e.g., one or more Global Positioning System (GPS) receivers) that enables the vehicle 100 to determine its current location using Global Navigation Satellite System (GNSS) signals. Alternatively or additionally, the navigation component 156 may include a radio navigation receiver for receiving navigation beacons or other signals from radio nodes such as Wi-Fi access points, cellular network sites, radio stations, remote computing devices, other vehicles, etc. Under the control of the drive control component 154, the processor 164 may control the vehicle 100 to navigate and maneuver. The processor 164 and / or the navigation component 156 may be configured to communicate with a server 184 on a network 186 (e.g., the Internet) using wireless signals 182 exchanged over a cellular data network via a network node 180 to receive commands for controlling maneuvers, receiving data useful in navigation, providing real-time location reports, and evaluating other data.

[0045] The control unit 140 may be coupled to one or more sensors 158. The sensors 158 may include the sensors 102 to 138 as described and may be configured to provide various data to the processor 164.

[0046] Although the control unit 140 is described as including separate components, in some aspects, some or all of the components (e.g., the processor 164, the memory 166, the input module 168, the output module 170, and the radio module 172) may be integrated in a single device or module (such as a System-on-Chip (SOC) processing device). Such an SOC processing device may be configured for use in a vehicle and may be configured (such as configured with processor-executable instructions executed in the processor 164) to perform operations in various aspects when installed in the vehicle.

[0047] Figure 1D An example implementation of a System-on-Chip (SOC) 105 is illustrated. The SOC 105 may include a central processing unit (CPU) 110 or a multi-core CPU configured to perform one or more of the functions described herein. In some aspects, the SOC 105 may be based on the ARM instruction set. In some cases, the CPU 110 may be associated with Figure 1Cis similar to the processor 164. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device (e.g., a neural network with weights), latencies, frequency bin information, task information, and other information can be stored in a memory block associated with the neural processing unit (NPU) 125, stored in a memory block associated with the CPU 110, stored in a memory block associated with the graphics processing unit (GPU) 115, stored in a memory block associated with the digital signal processor (DSP) 106, stored in the memory block 185, and / or can be distributed across multiple blocks. Instructions executed at the CPU 110 can be loaded from a program memory associated with the CPU 110 or can be loaded from the memory block 185.

[0048] The SOC 105 may also include additional processing blocks customized for specific functions, such as the GPU 115, the DSP 106, the NPU 125, the connection block 135, and the multimedia processor 145. In some cases, the connection block 135 may include a fifth-generation new radio (5G NR) connection, a fourth-generation long-term evolution (4G LTE) connection, Wi-Fi TM connection, a universal serial bus (USB) connection, Bluetooth TM connection, etc. In some examples, the multimedia processor 145 may, for example, detect and recognize gestures or perform other functions, such as generating a segmentation mask according to the systems and techniques described herein. In some aspects, the NPU 125 is implemented in the CPU 110, the DSP 106, and / or the GPU 115. The SOC 105 may also include a sensor processor 155, one or more image signal processors (ISPs) 175, and / or a navigation module 195. In some cases, the navigation module 195 may include a global positioning system (GPS) or a global navigation satellite system (GNSS). In some cases, the navigation module 195 may be similar to Figure 1C the navigation component 156. In some examples, the sensor processor 155 may receive inputs from, for example, one or more sensors 158. In some cases, the connection block 135 may be similar to Figure 1C the radio module 172.

[0049] Figure 2A illustrates examples of vehicle applications, subsystems, computing elements, or units within a vehicle management system 200 that can be utilized within a vehicle (such as Figure 1A the vehicle 100). Refer to Figures 1A to 2A, in some aspects, various vehicle applications, computing elements or units within the vehicle management system 200 can be implemented within a system of interconnected computing devices (e.g., subsystems) that communicate data and commands to each other. In other aspects, the vehicle management system 200 can be implemented as multiple vehicle applications executing within a single computing device, such as separate threads, processes, algorithms, or computing elements. However, the use of the term vehicle application when describing various aspects is not intended to imply or require that the corresponding functionality be implemented within a single autonomous (or semi-autonomous) vehicle management system computing device, although this is a potential implementation aspect. Instead, the use of the term vehicle application is intended to include subsystems with independent processors, computing elements (e.g., threads, algorithms, subroutines, etc.) running in one or more computing devices, and combinations of subsystems and computing elements.

[0050] In various aspects, vehicle applications executed in the vehicle management system 200 can include (but are not limited to) a radar perception vehicle application 202, a camera perception vehicle application 204, a positioning engine vehicle application 206, a map fusion and arbitration vehicle application 208, a route vehicle planning application 210, a sensor fusion and road world model (RWM) management vehicle application 212, a motion planning and control vehicle application 214, and a behavior planning and prediction vehicle application 216. The vehicle applications 202 - 216 are merely examples of some of the vehicle applications in an example configuration of the vehicle management system 200. In other configurations consistent with various aspects, other vehicle applications can be included, such as additional vehicle applications for other perception sensors (e.g., lidar perception layers, etc.), additional vehicle applications for planning and / or control, additional vehicle applications for modeling, etc., and / or some of the vehicle applications 202 - 216 can be excluded from the vehicle management system 200. Each of the vehicle applications 202 to 216 can exchange data, computation results, and commands.

[0051] The vehicle management system 200 can receive and process data from sensors (e.g., radar, lidar, cameras, inertial measurement unit (IMU), etc.), navigation systems (e.g., GPS receivers, IMUs, etc.), vehicle networks (e.g., controller area network (CAN) bus), and databases in memory (e.g., digital map data). The vehicle management system 200 can output vehicle control commands or signals to a drive-by-wire (DBW) system / control unit 220. The DBW system / control unit 220 is a system, subsystem, or computing device that directly interfaces with the vehicle steering, throttle, and brake controllers. Figure 2AThe configurations of the vehicle management system 200 and the DBW system / control unit 220 illustrated in are merely example configurations, and other configurations of the vehicle management system and other vehicle components may be used in various aspects. In some examples, Figure 2A the configurations of the vehicle management system 200 and the DBW system / control unit 220 illustrated in may be used in vehicles configured for autonomous or semi-autonomous operation, while different configurations may be used in non-autonomous vehicles.

[0052] The radar perception vehicle application 202 may receive data from one or more detection and ranging sensors (such as radar (e.g., radar 132) and / or lidar (e.g., lidar 138)), and process the data to identify and determine the positions of other vehicles and objects near the vehicle 100. The radar perception vehicle application 202 may include using neural network processing and artificial intelligence methods to identify objects and vehicles, and transmitting such information to the sensor fusion and RWM management vehicle application 212.

[0053] The camera perception vehicle application 204 may receive data from one or more cameras (such as cameras (e.g., cameras 122, 136)), and process the data to identify and determine the positions of other vehicles and objects near the vehicle 100. The camera perception vehicle application 204 may include using neural network processing and artificial intelligence methods to identify objects and vehicles. The camera perception vehicle application 204 may transmit such information to the sensor fusion and RWM management vehicle application 212.

[0054] The positioning engine vehicle application 206 may receive data from various sensors and process the data to determine the positioning of the vehicle (e.g., vehicle 100). The various sensors may include but are not limited to GPS sensors, IMUs, and / or other sensors connected via a bus (such as a CAN bus). The positioning engine vehicle application 206 may also utilize inputs from one or more cameras (such as cameras (e.g., cameras 122, 136)) and / or any other available sensors (such as radar, lidar, etc.).

[0055] The map fusion and arbitration vehicle application 208 can access data within a high-definition (HD) map database, receive the output received from the localization engine vehicle application 206, and process the data to further determine the location of the vehicle 100 within the map, such as the position within a traffic lane, the location within a street map, etc. The HD map database can be stored in a memory (e.g., memory 166). For example, the map fusion and arbitration vehicle application 208 can convert the latitude and longitude information from the GPS to a position within the ground road map included in the HD map database. GPS position localization includes errors, so the map fusion and arbitration vehicle application 208 can be used to determine the best-guess position of the vehicle 100 within the road based on the arbitration between the GPS coordinates and the HD map data. For example, although the GPS coordinates can place the vehicle 100 near the middle of a two-lane road in the HD map, the map fusion and arbitration vehicle application 208 can determine that the vehicle 100 is most likely aligned with the driving lane consistent with the driving direction according to the driving direction. The map fusion and arbitration vehicle application 208 can transfer the map-based position information to the sensor fusion and RWM management vehicle application 212.

[0056] The route planning vehicle application 210 can utilize the HD map and the input from an operator or dispatcher to plan the route for the vehicle 100 to follow to a specific destination. The route planning vehicle application 210 can transfer the map-based position information to the sensor fusion and RWM management vehicle application 212. However, the use of a priori maps is not necessary for other vehicle applications (such as the sensor fusion and RWM management vehicle application 212, etc.). For example, other stacks can operate and / or control the vehicle based only on perception data without the provided map, with the concept of constructing lane, boundary, and local maps as the perception data is received.

[0057] The sensor fusion and RWM management vehicle application 212 can receive data and outputs generated by one or more of the radar sensing vehicle application 202, the camera sensing vehicle application 204, the map fusion and arbitration vehicle application 208, and the route planning vehicle application 210. The sensor fusion and RWM management vehicle application 212 can use some or all of such inputs to estimate or refine the position and state of the vehicle 100 relative to the road, other vehicles on the road, and / or other objects near the vehicle 100. For example, the sensor fusion and RWM management vehicle application 212 can combine image data from the camera sensing vehicle application 204 with the arbitrated map position information from the map fusion and arbitration vehicle application 208 to refine the determined position of the vehicle within the traffic lane. As another example, the sensor fusion and RWM management vehicle application 212 can combine object recognition and image data from the camera sensing vehicle application 204 with object detection and ranging data from the radar sensing vehicle application 202 to determine and refine the relative positions of other vehicles and objects near the vehicle. As another example, the sensor fusion and RWM management vehicle application 212 can receive information about the positioning and driving direction of other vehicles from vehicle-to-vehicle (V2V) communication (such as via the CAN bus) and combine this information with the information from the radar sensing vehicle application 202 and the camera sensing vehicle application 204 to refine the positions and movements of the other vehicles. The sensor fusion and RWM management vehicle application 212 can output the refined position and state information of the vehicle 100 and the refined position and state information of other vehicles and objects near the vehicle to the motion planning and control vehicle application 214 and / or the behavior planning and prediction vehicle application 216.

[0058] As a further example, the sensor fusion and RWM management vehicle application 212 can use dynamic traffic control instructions to direct the vehicle 100 to change speed, lane, driving direction, or other navigation elements, and combine this information with other received information to determine refined position and state information. The sensor fusion and RWM management vehicle application 212 can output the refined position and state information of the vehicle 100 and the refined position and state information of other vehicles and objects near the vehicle 100 via wireless communication (such as through a cellular vehicle-to-everything (C-V2X) connection, other wireless connections, etc.) to the motion planning and control vehicle application 214, the behavior planning and prediction vehicle application 216, and / or devices remote from the vehicle 100, such as data servers, other vehicles, etc.

[0059] In some examples, the sensor fusion and RWM management vehicle application 212 can monitor sensed data from various sensors (such as sensed data from the radar sensing vehicle application 202, the camera sensing vehicle application 204, other sensing vehicle applications, etc.) and / or data from one or more sensors themselves to analyze conditions in the vehicle sensor data. The sensor fusion and RWM management vehicle application 212 can be configured to detect conditions in the sensor data, such as sensor measurements being at, above, or below a threshold, certain types of sensor measurements occurring, etc. The sensor fusion and RWM management vehicle application 212 can output the sensor data as part of the refined location and status information of the vehicle 100 provided to the behavior planning and prediction vehicle application 216 and / or provided to devices (such as data servers, other vehicles, etc.) remote from the vehicle 100 via wireless communication (such as via a C-V2X connection, other wireless connections, etc.).

[0060] The refined location and status information can include vehicle descriptors associated with the vehicle 100 and the vehicle owner and / or operator, such as: vehicle specifications (e.g., size, weight, color, on-board sensor types, etc.); vehicle location, speed, acceleration, direction of travel, attitude, orientation, destination, fuel / power level, and other status information; vehicle emergency status (e.g., the vehicle is an emergency vehicle or a private individual is in an emergency situation); vehicle restrictions (e.g., heavy / wide load, turning restrictions, high occupancy vehicle (HOV) authorization, etc.); vehicle capabilities (e.g., all-wheel drive, four-wheel drive, snow tires, chains, supported connection types, on-board sensor operating status, on-board sensor resolution level, etc.); equipment problems (e.g., low tire pressure, weak brakes, sensor failures, etc.); owner / operator travel preferences (e.g., preferred lanes, roads, routes, and / or destinations, preference to avoid tolls or highways, preference for the fastest route, etc.); permission to provide sensor data to a data proxy server (e.g., 184); and / or owner / operator identification information.

[0061] The behavior planning and prediction vehicle application 216 of the autonomous vehicle system 200 can use the refined location and status information of the vehicle 100 and the location and status information of other vehicles and objects output from the sensor fusion and RWM management vehicle application 212 to predict the future behavior of other vehicles and / or objects. For example, the behavior planning and prediction vehicle application 216 can use such information to predict the future relative location of these other vehicles based on its own vehicle's location and speed and the location and speed of other vehicles near the vehicle. Such predictions can take into account information from the HD map and route planning to anticipate changes in relative vehicle location as the host vehicle and other vehicles travel along the road. The behavior planning and prediction vehicle application 216 can output other vehicle and object behavior and location predictions to the motion planning and control vehicle application 214.

[0062] In addition, the behavior planning and prediction vehicle application 216 can use object behavior combined with location prediction to plan and generate control signals for controlling the motion of the vehicle 100. For example, based on the route planning information, the refined location in the road information, and the relative location and motion of other vehicles, the behavior planning and prediction vehicle application 216 can determine that the vehicle 100 needs to change lanes and accelerate, such as to maintain or achieve a minimum distance from other vehicles and / or to prepare for a turn or an exit. As a result, the behavior planning and prediction vehicle application 216 can calculate or otherwise determine the steering angle of the wheels and the change to the throttle setting, and this steering angle and change to the throttle setting will be commanded, along with various such parameters necessary to effect such lane change and acceleration, to the motion planning and control vehicle application 214 and the DBW system / control unit 220. One such parameter can be the calculated steering wheel command angle.

[0063] The motion planning and control vehicle application 214 can receive the data and information output from the sensor fusion and RWM management vehicle application 212 and the other vehicle and object behavior and location predictions from the behavior planning and prediction vehicle application 216, and use this information to plan and generate control signals for controlling the motion of the vehicle 100, and to verify that such control signals meet the safety requirements of the vehicle 100. For example, based on the route planning information, the refined location in the road information, and the relative location and motion of other vehicles, the motion planning and control vehicle application 214 can verify various control commands or instructions and pass them on to the DBW system / control unit 220.

[0064] The DBW system / control unit 220 can receive these commands or instructions from the motion planning and control vehicle application 214 and translate such information into mechanical control signals for controlling the wheel angles, brakes, and throttle of the vehicle 100. For example, the DBW system / control unit 220 can respond to the calculated steering wheel command angle by sending a corresponding control signal to the steering wheel controller.

[0065] In various aspects, the vehicle management system 200 can include functionality that performs safety checks or supervision of various commands, plans, or other decisions of various vehicle applications that may affect vehicle and occupant safety. Such safety check or supervision functionality can be implemented within a dedicated vehicle application or distributed among various vehicle applications and included as part of that functionality. In some aspects, multiple safety parameters can be stored in the memory, and the safety check or supervision functionality can compare the determined values (e.g., relative spacing to nearby vehicles, distance from the road centerline, etc.) with the corresponding safety parameters and issue a warning or command in the case of a violation or impending violation of the safety parameters. For example, the safety or supervision functionality in the behavior planning and prediction vehicle application 216 (or a separate vehicle application) can determine the current or future spacing between another vehicle (refined by the sensor fusion and RWM management vehicle application 212) and the vehicle 100 (e.g., based on the world model refined by the sensor fusion and RWM management vehicle application 212), compare the spacing with the safety spacing parameter stored in the memory, and issue an instruction to the motion planning and control vehicle application 214 to accelerate, decelerate, or turn in the case of a current or predicted spacing violation of the safety spacing parameter. As another example, the safety or supervision functionality in the motion planning and control vehicle application 214 (or a separate vehicle application) can compare the determined or commanded steering wheel command angle with the safety wheel angle limit or parameter and issue an override command and / or alarm in response to the commanded angle exceeding the safety wheel angle limit.

[0066] Some safety parameters stored in the memory can be static (i.e., do not change over time), such as the maximum vehicle speed. Other safety parameters stored in the memory can be dynamic as these parameters are continuously or periodically determined or updated based on vehicle state information and / or environmental conditions. Non-limiting examples of safety parameters include: maximum safe speed, maximum braking pressure, maximum acceleration, and safety steering wheel angle limits, all of which can vary depending on road and weather conditions.

[0067] Figure 2B Illustrates examples of vehicle applications, subsystems, computing elements, or units within the vehicle management system 250 that can be utilized within the vehicle 100. Refer toFigures 1A to 2B , in some aspects, the vehicle applications 202, 204, 206, 208, 210, 212, and 216 of the vehicle management system 200 may be similar to those vehicle applications described in the reference Figure 2A , and the vehicle management system 250 may operate similarly to the vehicle management system 200, except that the vehicle management system 250 may transfer various data or instructions to the vehicle safety and collision avoidance system 252 instead of the DBW system / control unit 220. For example, Figure 2B the configuration of the vehicle management system stack 250 and the vehicle safety and collision avoidance system 252 illustrated in

[0068] In various aspects, the behavior planning and prediction vehicle application 216 and / or the sensor fusion and RWM management vehicle application 212 may output data to the vehicle safety and collision avoidance system 252. For example, the sensor fusion and RWM management vehicle application 212 may output sensor data as part of the refined position and status information of the vehicle 100 provided to the vehicle safety and collision avoidance system 252. The vehicle safety and collision avoidance system 252 may use the refined position and status information of the vehicle 100 to make safety determinations regarding the vehicle 100 and / or the occupants of the vehicle 100. As another example, the behavior planning and prediction vehicle application 216 may output behavior models and / or predictions related to the movement of other vehicles to the vehicle safety and collision avoidance system 252. The vehicle safety and collision avoidance system 252 may use the behavior models and / or predictions related to the movement of other vehicles to make safety determinations regarding the vehicle 100 and / or the occupants of the vehicle 100.

[0069] In various aspects, the vehicle safety and collision avoidance system 252 may include functionality for performing safety checks or supervision of various commands, plans, or other decisions for various vehicle applications that may affect vehicle and occupant safety and the actions of a human driver. In some aspects, various safety parameters may be stored in a memory, and the vehicle safety and collision avoidance system 252 may compare a determined value (e.g., relative spacing to a nearby vehicle, distance from a road centerline, etc.) to a corresponding safety parameter and issue a warning or command if the safety parameter is violated or will be violated. For example, the vehicle safety and collision avoidance system 252 may determine a current or future spacing (e.g., based on a world model refined by sensor fusion and RWM managed vehicle application 212) between another vehicle (as refined by sensor fusion and RWM managed vehicle application 212) and the vehicle, compare the spacing to a safety spacing parameter stored in the memory, and issue an instruction to the driver to accelerate, decelerate, or turn if the current or predicted spacing violates the safety spacing parameter. As another example, the vehicle safety and collision avoidance system 252 may compare a change in the steering wheel angle of a human driver to a safety wheel angle limit or parameter and issue an override command and / or alarm in response to the steering wheel angle exceeding the safety wheel angle limit.

[0070] As indicated above, segmentation may be performed on the image data to generate a segmentation mask for the image data. In some cases, one or more machine learning techniques may be used to perform the segmentation, such as using one or more neural networks. A neural network is an example of a machine learning system, and a neural network may include an input layer, one or more hidden layers, and an output layer. Data is provided to input nodes of the input layer, processing is performed by hidden nodes of one or more hidden layers, and an output is produced by output nodes of the output layer. A deep learning network typically includes multiple hidden layers. Each layer of a neural network may include a feature map or activation map, which may include artificial neurons (or nodes). The feature map may include filters, kernels, etc. A node may include one or more weights for indicating the importance of a node in one or more of the layers. In some cases, a deep learning network may have a series of many hidden layers, where early layers are used to determine simple and low-level characteristics of the input, and later layers build a hierarchy of more complex and abstract characteristics.

[0071] Deep learning architectures can learn hierarchical structures of features. For example, if presented with visual data, the first layer can learn to identify relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to identify spectral power in specific frequencies. The second layer takes the output of the first layer as input and can learn to identify combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to identify common visual objects or spoken phrases.

[0072] When applied to problems with a natural hierarchical structure, deep learning architectures can perform particularly well. For example, the classification of motorized vehicles can benefit from first learning to identify wheels, windshields, and other features. These features can be combined in different ways at higher levels to identify cars, trucks, and airplanes.

[0073] Neural networks can be designed with a variety of connection patterns. In a feedforward network, information flows from lower layers to higher layers, where each neuron in a given layer communicates with neurons in the higher layer. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In a recurrent connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can help identify patterns that span more than one block of input data that is sequentially delivered to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when the identification of high-level concepts can assist in discerning specific low-level features of the input. The connections between layers of a neural network can be fully connected or locally connected. Various examples of neural network architectures are described below with respect to Figures 3A to 4 description.

[0074] Neural networks can be designed with a variety of connection patterns. In a feedforward network, information flows from lower layers to higher layers, where each neuron in a given layer communicates with neurons in the higher layer. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In a recurrent connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can help identify patterns that span more than one block of input data that is sequentially delivered to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when the identification of high-level concepts can assist in discerning specific low-level features of the input.

[0075] The connections between the layers of a neural network can be fully connected or locally connected. Figure 3A An example of a fully connected neural network 302 is illustrated. In the fully connected neural network 302, the neurons in the first layer can convey their outputs to each neuron in the second layer, such that each neuron in the second layer will receive inputs from each neuron in the first layer. Figure 3B An example of a locally connected neural network 304 is illustrated. In the locally connected neural network 304, the neurons in the first layer can be connected to a limited number of neurons in the second layer. More generally, the locally connected layer of the locally connected neural network 304 can be configured such that each neuron in the layer will have the same or a similar connection pattern, but the connection strengths can have different values (e.g., value 310, value 312, value 314, and value 316). The locally connected pattern of connections can give rise to spatially distinct receptive fields in higher layers, as the higher layer neurons in a given region can receive inputs that are tuned through training to the characteristics of a restricted portion of the total input to the network.

[0076] One example of a locally connected neural network is a convolutional neural network. Figure 3C An example of a convolutional neural network 306 is illustrated. The convolutional neural network 306 can be configured such that the connection strengths associated with the inputs (e.g., input 308) for each neuron in the second layer are shared. Convolutional neural networks can be well-suited for problems where the spatial location of the input is meaningful. In accordance with aspects of the present disclosure, the convolutional neural network 306 can be used to perform one or more aspects of video compression and / or decompression.

[0077] One type of convolutional neural network is a deep convolutional network (DCN). Figure 3D A detailed example of a DCN 300 designed to identify visual features from an image 326 input from an image capture device 330 (such as an in-vehicle camera) is illustrated. The DCN 300 of the current example can be trained to identify traffic signs and the numbers provided on the traffic signs. The DCN 300 can be trained for other tasks, such as identifying lane markings or identifying traffic signals.

[0078] The DCN 300 can be trained using supervised learning. During training, an image (such as image 326 of a speed limit sign) can be presented to the DCN 300, and then a forward pass can be computed to produce an output 322. The DCN 300 can include a feature extraction part and a classification part. When receiving the image 326, the convolutional layer 332 can apply a convolutional kernel (not shown) to the image 326 to generate a first set of feature maps 318. As an example, the convolutional kernel of the convolutional layer 332 can be a 5x5 kernel that generates 28x28 feature maps. In this example, since four different feature maps are generated in the first set of feature maps 318, four different convolutional kernels are applied to the image 326 at the convolutional layer 332. The convolutional kernel can also be referred to as a filter or a convolutional filter.

[0079] The first set of feature maps 318 can be subsampled by a max pooling layer (not shown) to generate a second set of feature maps 320. The max pooling layer reduces the size of the first set of feature maps 318. That is, the size of the second set of feature maps 320 (such as 14x14) is smaller than the size of the first set of feature maps 318 (such as 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 320 can be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).

[0080] In Figure 3D the example, the second set of feature maps 320 is convolved to generate a first feature vector 324. Additionally, the first feature vector 324 is further convolved to generate a second feature vector 328. Each feature of the second feature vector 328 can include numbers corresponding to possible features of the image 326, such as "sign", "60", and "100". A softmax function (not shown) can convert the numbers in the second feature vector 328 into probabilities. Thus, the output 322 of the DCN 300 is the probability that the image 326 includes one or more features.

[0081] In this example, the probabilities of "sign" and "60" in the output 322 are higher than the probabilities of other numbers (such as "30", "40", "50", "70", "80", "90", and "100") in the output 322. Before training, the output 322 produced by the DCN 300 may be incorrect. Therefore, the error between the output 322 and the target output can be computed. The target output is the ground truth of the image 326 (e.g., "sign" and "60"). Then the weights of the DCN 300 can be adjusted such that the output 322 of the DCN 300 is closer to the target output.

[0082] To adjust the weights, the learning algorithm can compute the gradient vector of the weights. The gradient can indicate the amount by which the error will increase or decrease when the weights are adjusted. At the top layer, the gradient can directly correspond to the value of the weights connecting the activated neurons in the penultimate layer and the neurons in the output layer. In the lower layers, the gradient can depend on the values of the weights and the error gradients computed in the higher layers. The weights can then be adjusted to reduce the error. This way of adjusting the weights can be called "backpropagation" because it involves a "backward pass" through the neural network.

[0083] In practice, the error gradient of the weights can be computed on a small number of examples such that the computed gradient is close to the true error gradient. This approximation method can be called stochastic gradient descent. Stochastic gradient descent can be repeated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level. After learning, new images can be presented to the DCN, and the forward pass through the network can produce an output 322 that can be considered as an inference or prediction of the DCN.

[0084] A deep belief network (DBN) is a probabilistic model that includes multiple layers of hidden nodes. A DBN can be used to extract hierarchical representations of a training data set. A DBN can be obtained by stacking layers of restricted Boltzmann machines (RBMs). An RBM is a type of artificial neural network that can learn a probability distribution from a set of inputs. Since an RBM can learn a probability distribution without information about the class to which each input should be classified, an RBM is typically used for unsupervised learning. Using a hybrid paradigm of unsupervised and supervised, the bottom RBM of the DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of the inputs from the previous layer and the target classes) and can be used as a classifier.

[0085] A deep convolutional network (DCN) is a network of convolutional networks configured with additional pooling and normalization layers. The DCN has achieved state-of-the-art performance on many tasks. The DCN can be trained using supervised learning, where both the input targets and the output targets are known for many paradigms and are used to modify the weights of the network by using the gradient descent method.

[0086] The DCN can be a feed-forward network. Additionally, as described above, the connections from the neurons in the first layer of the DCN to a set of neurons in the next higher layer are shared across the neurons in the first layer. The feed-forward and shared connections of the DCN can be used for fast processing. For example, the computational burden of the DCN may be much smaller than that of a similarly sized neural network that includes recurrent or feedback connections.

[0087] The processing of each layer of the convolutional network can be considered as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then the convolutional network trained on this input can be considered three-dimensional, where two spatial dimensions are along the axes of the image, and the third dimension captures color information. The output of the convolutional connection can be considered to form a feature map in subsequent layers, where each element in this feature map (e.g., feature map 320) receives inputs from a certain range of neurons in the previous layer (e.g., feature map 318) and from each of the multiple channels. The values in the feature map can be further processed with a non-linearity (such as rectification, max(0,x)). The values from adjacent neurons can be further pooled, which corresponds to downsampling and can provide additional local invariance and dimensionality reduction.

[0088] Figure 4 FIG. is a block diagram illustrating an example of a deep convolutional network 450. Based on connections and weight sharing, the deep convolutional network 450 can include multiple different types of layers. As Figure 4 shown, the deep convolutional network 450 includes convolutional blocks 454A, 454B. Each of the convolutional blocks 454A, 454B can be configured with a convolutional layer (CONV) 456, a normalization layer (LNorm) 458, and a max pooling layer (MAX POOL) 460.

[0089] The convolutional layer 456 can include one or more convolutional filters, which can be applied to the input data 452 to generate a feature map. Although only two convolutional blocks 454A, 454B are shown, the present disclosure is not limited thereto, but any number of convolutional blocks (e.g., convolutional blocks 454A, 454B) can be included in the deep convolutional network 450 according to design preferences. The normalization layer 458 can normalize the output of the convolutional filter. For example, the normalization layer 458 can provide whitening or lateral inhibition. The max pooling layer 460 can provide spatially downsampled aggregation to achieve local invariance and dimensionality reduction.

[0090] For example, the parallel filter bank of the deep convolutional network can be loaded onto the CPU 110 or GPU 115 of the SOC 105 to achieve high performance and low power consumption. In an alternative aspect, the parallel filter bank can be loaded onto the DSP 106 or ISP 175 of the SOC 105. Additionally, the deep convolutional network 450 can access other processing blocks that may exist on the SOC 105, such as the sensor processor 155 and navigation module 195 dedicated to sensors and navigation, respectively.

[0091] The deep convolutional network 450 may also include one or more fully connected layers, such as layer 462A (labeled "FC1") and layer 462B (labeled "FC2"). The deep convolutional network 450 may also include a logistic regression (LR) layer 464. Between each layer 456, 458, 460, 462A, 462B, 464 of the deep convolutional network 450 are weights (not shown) to be updated. The output of each of these layers (e.g., 456, 458, 460, 462A, 462B, 464) can serve as the input to a subsequent one of these layers (e.g., 456, 458, 460, 462A, 462B, 464) of the deep convolutional network 450 to learn a hierarchical feature representation from the input data 452 (e.g., image, audio, video, sensor data, and / or other input data) provided at the initial convolutional block 454A. The output of the deep convolutional network 450 is a classification score 466 for the input data 452. The classification score 466 can be a set of probabilities, where each probability is the probability that the input data includes a feature from a set of features.

[0092] As previously noted, segmentation may be important for many use cases including extended reality (XR) applications (e.g., AR, VR, MR, etc.), autonomous driving, cameras of mobile devices, IoT devices or systems, etc. Current segmentation solutions (e.g., deployed on devices) may provide inconsistent semantic representations. Figure 5 An example of inconsistent segmentation results for the same image under different image signal processor (ISP) settings is illustrated. As shown, a first image 502 with a first ISP setting produces a segmentation mask 504. However, based on processing a second image 506 (e.g., of the same scene) with a second ISP setting, a segmentation mask 508 can be generated with pixel classifications for a first portion 503 and a second portion 505 that are different compared to similar portions in the segmentation mask 504.

[0093] Figure 6 An example of inconsistent segmentation results over time between adjacent images is illustrated, which is referred to as temporal inconsistency between segmentation masks. As shown, a segmentation mask 604 is generated for the current image that includes pixel classifications that are inconsistent compared to a segmentation mask 602 generated for a previous image (an image in a video or other image sequence that is before the current image). The result of temporally inconsistent segmentation masks (due to inconsistent predictions between images or frames) is a flickering artifact.

[0094] One possible way to address temporal inconsistency using a neural network is to use optical flow to warp features between images or frames. However, optical flow can be challenging due to the large amount of computation on the device making the neural network very slow.

[0095] In some cases, techniques for solving temporal inconsistency are to aggregate features from one or more previous images and use the aggregated features to process the current image to generate a segmentation mask. Figure 7 FIG. 4 is an illustration of an example of a machine learning system 700 configured to perform such a method. As shown, a previous image 702 (at time T) is processed by a machine learning model 704 (shown at time instance T) to generate features 706 representing the previous image 702. A current image 703 (at time T+1, which is the next time step after time T) is also processed by the machine learning model 704 (shown at time instance T+1) to generate features 707 representing the current image 703. The features 706 representing the previous image 702 are combined (e.g., concatenated) with the features 707 representing the current image 703 to generate combined features for the current image 703.

[0096] The features 706 representing the previous image 702 (and / or a combined representation based on combining the features 706 with features (not shown) generated for the previous image at T-1) can be processed by a machine learning operation 708 to generate a segmentation mask 710 for the previous image 702. In some aspects, the machine learning operation 708 can be a convolutional operation, such as (but not limited to) a two-dimensional (2D) convolutional operation using a 1x1 convolutional filter (Conv2d1x1). The combined features (of features 706 and 707) generated for the current image 703 can be processed by the machine learning operation 708 to generate a segmentation mask 711 for the current image 703.

[0097] Using features from previous images helps to provide consistent results. However, concatenating two features that are not pixel-aligned can lead to additional inconsistencies in the segmentation mask. Figure 8 FIG. 8 is an illustration of an example of misaligned concatenated features 813. For example, features 806 can be generated (e.g., by the machine learning model 704) based on a previous image (e.g., at time instance T), and features 807 can be generated (e.g., by the machine learning model 704) based on a current image (an image that appears after the previous image in a video or other image sequence; e.g., at time instance T+1). The features 806 can be combined with the features 807 (e.g., by a concatenation operation referred to as "concat") to generate concatenated features (also referred to as combined features) 813. As shown, the pose of the person represented in the features 806 is not aligned with the pose of the person represented in the features 807, resulting in the concatenated features 813 being misaligned.

[0098] In some cases, blocks (e.g., transform operation blocks) can be added to a machine learning system to transform features generated for a previous image such that features corresponding to an object in the previous image are aligned (and thus pixel - aligned) with features corresponding to the same object in the previous image. Figure 9 is an illustration of an example of a machine learning system 700 that includes a transform operation 912 added to Figure 7 a machine learning system 900 for generating a segmentation mask from an image. Adding such a block can improve performance, but such a machine learning system 900 can be modified to provide an understanding of the localization of objects in the current image.

[0099] As noted above, the systems and techniques described herein provide a machine learning system that utilizes a differential image to generate a segmentation mask for an image. Figure 10 is an illustration of an example of a machine learning system 1000 that includes a transform operation 1012 (which may correspond to transform operation 912). The transform operation 1012 uses a differential image 1014 to transform a previous image 1002 (e.g., from time instance T) such that features 1006 representing the previous image 1002 are pixel - aligned with features 1006 representing the current image 1003 (e.g., from time instance T + 1 or later). Thus, the differential image 1014 can be used by the transform operation 1012 to transform the features 1006 to the next time step (e.g., time instance T + 1).

[0100] As Figure 10 shown, the machine learning system 1000 performs a difference operation to determine the difference between a previous image 1002 (at time T) and a current image 1003 (at time T + 1, which can be the next time step after time T). The result of the difference operation between the previous image 1002 and the current image 1003 is a differential image 1014 (also referred to as a difference image). In some cases, the difference operation can include determining the difference between each pixel of the current image 1003 and each corresponding pixel of the previous image 1002 (at the common position within the image frame), thereby producing a difference value for each pixel position in the differential image 1014. In some aspects, the differential image can be multiplied by one or more segmentation masks from one or more previous outputs of the machine learning system, which can produce one or more “masked” differential images (e.g., a batch of masked differential images).

[0101] The previous image 1002 is processed by a machine learning model 1004 (shown at time instance T) to generate features 1006 representing the previous image 1002. In some cases, the machine learning model 1004 may identify certain features in the input image. In some examples, the machine learning model 1004 may include one or more layers (e.g., hidden layers such as convolutional layers, normalization layers, pooling layers, and / or other layers) or transformer blocks that can generate feature maps for identifying certain features. In some cases, the machine learning model 1004 includes an encoder-decoder neural network architecture. Exemplary examples of the machine learning model 1004 include Figure 3A a fully connected neural network 302, Figure 3B a locally connected neural network 304, Figure 3C a convolutional neural network 306, Figure 3D a deep convolutional network (DCN) 300, and / or other types of ML models.

[0102] The current image 1003 is also processed by the machine learning model 1004 (shown at time instance T + 1) to generate features 1007 representing the current image 1003. In some cases, the machine learning system 1000 may generate a difference image by determining the difference between the intermediate features generated by the machine learning model 1004 for the previous image 1002 and the intermediate features generated by the machine learning model 1004 for the current image 1003. For example, the intermediate features may be output by one or more intermediate layers (e.g., hidden layers before the final layer) of the machine learning model 1004.

[0103] The features 1006 representing the previous image 1002 are combined (e.g., concatenated using a concatenation operation) with the pixels or features of the difference image 1014 (or masked difference image) to generate combined features 1015. Subsequently, the combined features 1015 are processed using a transformation operation 1012 to generate transformed features 1016 for the previous image 1002. The transformed features 1016 for the previous image 1002 are combined (e.g., concatenated) with the features 1007 representing the current image 1003 to generate combined features (also referred to as concatenated features) 1017. The resulting combined features 1017 are processed by a machine learning operation 1008 (e.g., a Conv2d 1x1 operation) to generate a segmentation mask 1011 for the current image 1003. The features 1006 representing the previous image 1002 (and / or a combined representation based on combining the features 1006 with the transformed features generated for the previous image at T - 1) may also be processed by the machine learning operation 1008 to generate a segmentation mask 1010 for the previous image 1002.

[0104] In some aspects, the transform operation 1012 can be a convolution operation performed using a convolution filter or kernel (such as a 2D 3x3 convolution filter or kernel). In some cases, the transform operation 1012 can be a deformable convolution operation performed using a deformable convolution filter or kernel (such as using DCN-v2n). In some cases, the transform operation 1012 can be a transformer block. In some examples, the key of the transformer is the previous feature, and the query is the difference image 1014 (or masked difference image).

[0105] In some aspects, the parameters of the transform operation 1012 (e.g., the weights of the convolution filter or kernel, the weight offsets of the deformable convolution such as DCN-v2, etc.) are fixed in different iterations of the machine learning system 1000 that generates the segmentation mask for the input image. For example, the weights of the convolution filter of the transform operation 1012 can be kept fixed (or constant) when transforming different difference images.

[0106] In some aspects, the parameters of the transform operation 1012 (e.g., the weights of the convolution filter or kernel, the weight offsets of the deformable convolution such as DCN-v2, etc.) can vary or be modified for each iteration of the machine learning system 1000 based on the difference image 1014 (when processing a new image to generate the segmentation mask for the new image). Figure 11 is an illustration of an example of a system 1100 that can change the transform operation 1112 (such as a convolution operation) based on the value of the difference image 1114, which can be similar to the difference image 1014 of Figure 10 the.

[0107] As Figure 11As shown in the illustrative example, the difference image 1114 is input to the non-maximum suppression engine 1122. The difference image is shown as having a height (H) and a width (W) that can be of any suitable size. In some examples, the non-maximum suppression engine 1122 can perform a max pooling operation (e.g., using one or more max pooling layers). In other examples, the non-maximum suppression engine 1122 can perform other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. The max pooling operation can include downsampling for dimensionality reduction, which can help the difference image information provide a better transformation of the features of the previous image. In some cases, the max pooling can be performed by applying a max pooling filter (e.g., having a size of 2x2) with a certain step amount (e.g., equal to the dimension of the filter, such as a step amount of 2) to the difference image 1114. The output from the max pooling filter can include the maximum value in each sub-region that the filter convolves around. Using a 2x2 filter as an illustrative example, each unit in the pooling layer can summarize the region of 2×2 nodes in the previous layer (where each node is a value in the difference image 1114). For example, four values (nodes) in the difference image 1114 can be analyzed by the 2×2 max pooling filter at each iteration of the filter, and the maximum value among the four values is output as the "max" value. As noted above, in some examples, an L2 norm pooling filter can also be used. The L2 norm pooling filter includes calculating the square root of the sum of the squares of the values in a 2×2 region (or other suitable region) of the difference image 1114 (instead of calculating the maximum value as in max pooling), and using the calculated value as the output.

[0108] The output from the non-maximum suppression engine 1122 includes an array of values of reduced dimension having a height H s and a width W s The array is flattened from a two-dimensional (2D) representation (having a height and width H s x W s to a one-dimensional (1D) representation (e.g., having H s *W svector or tensor of dimensions). The 1D representation is input to the transform adaptation engine 1124, which is configured to generate or determine values for the transform operation 1112 (such as transform operation 1012). In some aspects, the transform adaptation engine 1124 can include a multi-layer perceptron (MLP) network, a fully connected layer, and / or other deep neural networks. The transform adaptation engine 1124 processes the 1D representation of the differential image 1114 to generate a 1D set of parameter values (such as weights) of size K*K. For example, the transform adaptation engine 1124 can include one or more convolutional filters (and / or other types of machine learning operations) that process the 1D representation of the differential image 1114 to generate K*K parameter values. The 1D K*K parameter values can then be reshaped to generate an array 1126 of parameter values having dimensions of height K × width K (K×K). The K×K array 1126 can be used as a convolutional filter or kernel for the transform operation 1112. Thus, the transform operation 1112 (using the K×K array 1126 as a filter or kernel) is determined based on the pixels or feature values of the differential image 1114. Using this technique, the transform operation 1112 can be adapted based on each specific differential image determined based on at least two images (such as adjacent images or video frames in a video).

[0109] As described above with respect to Figure 10 the transform operation 1112 can be used to generate transformed features for the previous image used to generate the differential image 1114. For example, the transform operation 1112 can use the differential image 1114 to transform the features representing the previous image to the next time step (corresponding to the time step of the current image used to generate the differential image 1114) such that the features representing the previous image are pixel-aligned with the features representing the current image.

[0110] Using the differential image-based systems and techniques described herein can reduce or eliminate temporal inconsistencies between segmentation masks and thus can improve the quality of image processing operations.

[0111] Figure 12 is a flowchart of a process 1200 for processing one or more images illustrative of aspects of the present disclosure. In some examples, the process 1200 can be performed by a computing device or by a component or system of a computing device (such as a chipset, such as Figure 1D the SOC 105). The computing device can implement a machine learning system (such as Figure 10 the machine learning system 1000) to perform the differential image-based techniques described herein. The computing device can include a vehicle (such as Figure 1Aa vehicle 100), or a computing system or component of a vehicle, a mobile device such as a mobile phone, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a network-connected wearable device (e.g., a network-connected watch), or other computing devices. In some cases, the computing device may include Figure 13 a computing system 1300. The operations of process 1200 may be implemented as software components executed and run on one or more processors (e.g., Figure 13 the processor 1310 of or other processors). The sending and receiving of signals by the computing device in process 1200 may be enabled, for example, by one or more antennas and / or one or more transceivers (e.g., wireless transceivers).

[0112] At block 1202, the computing device (or its component) may generate a difference image based on the difference between the current image and the previous image. In some examples, the current image may include Figure 10 an image 1003 (at time T + 1), the previous image may include an image 1002 (at time T), and the difference image may include a difference image 1014.

[0113] In some aspects, the computing device (or its component) may use a machine learning model to process the previous image to generate intermediate features representing the previous image (e.g., features output by one or more intermediate hidden layers of the machine learning model 1004 at Figure 10 time T). The computing device (or its component) may use a machine learning model to process the current image to generate intermediate features representing the current image (e.g., features output by one or more intermediate hidden layers of the machine learning model 1004 at Figure 10 time T + 1). In some cases, the computing device (or its component) may further determine the difference between the intermediate features representing the previous image and the intermediate features representing the current image. The computing device (or its component) may then generate a difference image based on the difference between the intermediate features representing the previous image and the intermediate features representing the current image.

[0114] At block 1204, the computing device (or its component) may use a transformation operation to process the difference image and the features representing the previous image to generate a transformed feature representation of the previous image. In some examples, the features representing the previous image may include Figure 10 features 1006, the transformation operation may include a transformation operation 1012, and the transformed feature representation of the previous image may include transformed features 1016. In some aspects, the computing device (or its component) may use a machine learning model to process the previous image to generate features representing the previous image. In some cases, the machine learning model may include Figure 10The machine learning model 1004 (at time T) of the machine learning system 1000.

[0115] In some aspects, a computing device (or its components) may combine a differential image with features representing a previous image to generate a combined feature representation of the previous image (e.g., as shown in Figure 10 ). In some cases, to process the differential image and the features representing the previous image (using a transformation operation) to generate a transformed feature representation of the previous image, the computing device (or its components) may use a transformation operation to process the combined feature representation of the previous image to generate a transformed feature representation of the previous image (e.g., as shown in Figure 10 ).

[0116] In some examples, the transformation operation includes a convolution operation performed using at least one convolutional filter. In some cases, the weights of the at least one convolutional filter are fixed. In some cases, the weights of the at least one convolutional filter are modified based on the differential image (e.g., by the system 1100 described in Figure 11 ). In some aspects, the at least one convolutional filter includes a deformable convolution. In some aspects, at least one weight offset of the deformable convolution is modified based on the differential image. In some examples, the transformation operation includes a transformer operation performed using at least one transformer block.

[0117] At block 1206, the computing device (or its components) may combine the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image. In some aspects, the features representing the current image may include Figure 10 the features 1007, and the combined feature representation of the current image may include the combined feature 1017. In some cases, the computing device (or its components) may use a machine learning model to process the current image to generate features representing the current image. In some examples, the machine learning model may include Figure 10 the machine learning model 1004 (at time T+1) of the machine learning system 1000.

[0118] At block 1208, the computing device (or its components) may generate a segmentation mask for the current image based on the combined feature representation of the current image. In some aspects, the segmentation mask for the current image may include the Figure 10 segmentation mask 1011 generated for the image 1003. Referring to Figure 10 As an example, a machine learning operation 1008 (e.g., a Conv2d 1x1 operation) may process the combined feature 1017 to generate a segmentation mask 1011 for the current image 1003.

[0119] Figure 13FIG. is an example diagram illustrating a system for implementing certain aspects of the present disclosure. Specifically, Figure 13 illustrates an example of a computing system 1300, which can be any computing device that constitutes an internal computing system, a remote computing system, a camera, or any component thereof, where the components of the system communicate with each other using connection 1305. Connection 1305 can be a physical connection using a bus or a direct connection to the processor 1310, such as in a chipset architecture. Connection 1305 can also be a virtual connection, a networking connection, or a logical connection.

[0120] In some aspects, the computing system 1300 is a distributed system, where the functions described in the present disclosure can be distributed within one data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the system components represent many such components that each perform some or all of the functions for which the component is described. In some aspects, the components can be physical or virtual devices.

[0121] The example system 1300 includes at least one processing unit (CPU or processor) 1310 and connection 1305, which communicatively couples various system components including system memory 1315 (such as read-only memory (ROM) 1320 and random access memory (RAM) 1325) to the processor 1310. The computing system 1300 can include a cache 1312 that is directly connected to, in close proximity to, or integrated as part of the processor 1310 and is a high-speed memory.

[0122] The processor 1310 can include any general-purpose processor and hardware services or software services (such as services 1332, 1334, and 1336 stored in the storage device 1330 and configured to control the processor 1310), as well as a dedicated processor in which software instructions are incorporated into the actual processor design. The processor 1310 can be substantially a complete stand-alone computing system that includes multiple cores or processors, buses, memory controllers, caches, etc. The multi-core processor can be symmetric or asymmetric.

[0123] To enable user interaction, the computing system 1300 includes an input device 1345 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. The computing system 1300 can also include an output device 1335 that can be one or more of a plurality of output mechanisms. In some cases, a multi-mode system can enable a user to provide multiple types of input / output to communicate with the computing system 1300.

[0124] The computing system 1300 may include a communication interface 1340, which may generally govern and manage user input and system output. The communication interface may perform or facilitate the reception and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including using audio jack / plug, microphone jack / plug, Universal Serial Bus (USB) port / plug, Apple TM Lightning TM port / plug, Ethernet port / plug, fiber optic port / plug, dedicated wired port / plug, 3G, 4G, 5G, and / or other cellular data network wireless signaling, Bluetooth TM wireless signaling, Bluetooth TM Low Energy (BLE) wireless signaling, iBeacon TM wireless signaling, Radio Frequency Identification (RFID) wireless signaling, Near Field Communication (NFC) wireless signaling, Dedicated Short Range Communication (DSRC) wireless signaling, 802.11 Wi-Fi wireless signaling, Wireless Local Area Network (WLAN) signaling, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), Infrared (IR) communication wireless signaling, Public Switched Telephone Network (PSTN) signaling, Integrated Services Digital Network (ISDN) signaling, ad hoc network signaling, radio wave signaling, microwave signaling, infrared signaling, visible light signaling, ultraviolet light signaling, wireless signaling along the electromagnetic spectrum, or some combination thereof. The communication interface 1340 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 1300 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There are no restrictions on operating on any particular hardware arrangement, and thus the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0125] The storage device 1330 can be a non-volatile and / or non-transitory and / or computer-readable memory device, and can be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as cassette tapes, flash memory cards, solid state memory devices, digital versatile discs, cartridges, floppy disks, hard disks, magnetic tapes, magnetic strips, any other magnetic storage medium, flash memory, memristor memory, any other solid state memory, compact disc read-only memory (CD-ROM) optical discs, rewritable compact discs (CDs), digital video discs (DVDs), Blu-ray discs (BDDs), holographic optical discs, another optical medium, secure digital (SD) cards, micro secure digital (microSD) cards, Memory cards, smart card chips, EMV chips, subscriber identity module (SIM) cards, mini / micro / nano / pico SIM cards, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (e.g., level 1 (L1) cache, level 2 (L2) cache, level 3 (L3) cache, level 4 (L4) cache, level 5 (L5) cache, other (L#) cache), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge and / or combinations thereof.

[0126] The storage device 1330 may include software services, servers, services, etc. When the code defining such software is executed by the processor 1310, the code causes the system to perform functions. In some aspects, the hardware services that perform specific functions may include software components stored in a computer-readable medium connected to the necessary hardware components (such as the processor 1310, the connection 1305, the output device 1335, etc.) to perform functions. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium may include non-transitory media, in which data can be stored and which does not include carrier waves and / or transient electrical signals propagated wirelessly or through a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memories, memories, or memory devices. The computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. By passing and / or receiving information, data, arguments, parameters, or memory contents, a code segment can be coupled to another code segment or a hardware circuit. Information, arguments, parameters, data, etc. can be passed, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0127] Specific details are provided in the above description to provide a thorough understanding of the various aspects and examples provided herein, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the exemplary aspects of the present application have been described in detail herein, it is to be understood that the various inventive concepts can be implemented and employed in other various ways, and the appended claims are not to be construed as including these variations unless limited by the prior art. The various features and aspects of the above applications can be used separately or jointly. Additionally, without departing from the broader scope of the specification, the aspects can be utilized in any number of environments and applications beyond those described herein. Therefore, the specification and the drawings should be considered illustrative rather than restrictive. For purposes of illustration, the methods are described in a particular order. It should be appreciated that in alternative aspects, the methods can be performed in a different order than that described.

[0128] For the sake of clarity in explanation, in some instances, the present technology may be presented as including separate functional blocks that include devices, device components, steps, or routines in a method embodied in software or a combination of hardware and software. Additional components other than those shown and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the aspects.

[0129] Furthermore, those skilled in the art should understand that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each particular application, but such specific implementation decisions should not be construed as causing a departure from the scope of the present disclosure.

[0130] The various aspects or examples may be described above as a process or method, which is depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. The process is terminated when the operations of the process are completed, but the process may have additional steps not included in the figures. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or the main function.

[0131] The processes and methods according to the above examples may be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or a group of functions. Portions of the computer resources used may be accessed via a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code. Examples of computer-readable media that may be used to store instructions, the information used, and / or the information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.

[0132] In some aspects, a computer-readable storage device, medium, and memory may include a wire or wireless signal containing a bitstream, etc. However, when mentioned, non-transitory computer-readable storage media explicitly exclude media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.

[0133] Those skilled in the art will understand that information and signals can be represented using any of a variety of different technologies and methods. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description may, in some cases, be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof, depending in part on the specific application, in part on the desired design, in part on the corresponding technology, etc.

[0134] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein can be implemented or executed using hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof, and can be in any form factor. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) for performing the necessary tasks can be stored in a computer-readable or machine-readable medium. The processor can execute the necessary tasks. Examples of form factors include: laptop devices, smart phones, mobile phones, tablet devices, or other personal computers with a small form factor, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functions described herein can also be embodied in peripheral devices or plug-in cards. By further example, such functions can also be implemented on a circuit board in different chips or different processes executed on a single device.

[0135] Instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in this disclosure.

[0136] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device such as a cellular phone, or an integrated circuit device with multiple uses, including applications in wireless communication devices such as cellular phones and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially realized by a computer-readable data storage medium including program code including instructions that, when executed, perform one or more of the above-described methods, algorithms, and / or operations. The computer-readable data storage medium may form part of a computer program product, which may include packaging material. The computer-readable medium may include a memory or data storage medium, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques may be at least partially realized by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0137] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term “processor” may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.

[0138] One of ordinary skill in the art will appreciate that, without departing from the scope of this specification, the less-than (“<”) and greater-than (“>”) symbols or terms used herein may be replaced, respectively, with the less-than-or-equal-to (“≤”) and greater-than-or-equal-to (“≥”) symbols.

[0139] In cases where a component is described as “configured to” perform certain operations, such configuration can be implemented, for example, by designing an electronic circuit or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.

[0140] The phrase “coupled to” or “communicatively coupled to” means that any component is directly or indirectly physically connected to another component, and / or any component is directly or indirectly in communication with another component (e.g., connected to that other component via a wired or wireless connection and / or other suitable communication interface).

[0141] Claim language or other language that recites “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language that recites “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language that recites “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language that recites “at least one of A and B” or “at least one of A or B” can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.

[0142] Claim language or other language that recites “at least one processor, the at least one processor being configured to”, “at least one processor being configured to”, etc. indicates that one processor or multiple processors (in any combination) can perform the associated operations. For example, claim language that recites “at least one processor, the at least one processor being configured to: X, Y, and Z” means that a single processor can be used to perform operations X, Y, and Z; or multiple processors are each assigned tasks for a particular subset of operations X, Y, and Z such that the multiple processors together perform X, Y, and Z; or a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language that recites “at least one processor, the at least one processor being configured to: X, Y, and Z” can mean that any single processor can perform at least a subset of operations X, Y, and Z.

[0143] Exemplary aspects of the present disclosure include:

[0144] Aspect 1. A processor-implemented method for generating one or more segmentation masks, the processor-implemented method comprising: generating a difference image based on a difference between a current image and a previous image; using a transformation operation to process the difference image and features representing the previous image to generate a transformed feature representation of the previous image; combining the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image; and generating a segmentation mask for the current image based on the combined feature representation of the current image.

[0145] Aspect 2. The processor-implemented method according to aspect 1, the processor-implemented method further comprising: using a machine learning model to process the previous image to generate the features representing the previous image.

[0146] Aspect 3. The processor-implemented method according to aspect 2, the processor-implemented method further comprising: using the machine learning model to process the current image to generate the features representing the current image.

[0147] Aspect 4. The processor-implemented method according to any one of aspects 2 or 3, wherein generating the difference image based on the difference between the current image and the previous image comprises: using the machine learning model to process the previous image to generate intermediate features representing the previous image; using the machine learning model to process the current image to generate intermediate features representing the current image; determining a difference between the intermediate features representing the previous image and the intermediate features representing the current image; and generating the difference image based on the difference between the intermediate features representing the previous image and the intermediate features representing the current image.

[0148] Aspect 5. The processor-implemented method according to any one of aspects 1 to 4, the processor-implemented method further comprising: combining the difference image with the features representing the previous image to generate a combined feature representation of the previous image.

[0149] Aspect 6. The processor-implemented method according to aspect 5, wherein using the transformation operation to process the difference image and the features representing the previous image to generate the transformed feature representation of the previous image comprises: using the transformation operation to process the combined feature representation of the previous image to generate the transformed feature representation of the previous image.

[0150] Aspect 7. The processor-implemented method according to any one of aspects 1 to 6, wherein the transformation operation comprises a convolution operation performed using at least one convolution filter.

[0151] Aspect 8. The processor-implemented method according to aspect 7, wherein the weights of the at least one convolutional filter are fixed.

[0152] Aspect 9. The processor-implemented method according to aspect 7, wherein the weights of the at least one convolutional filter are modified based on the difference image.

[0153] Aspect 10. The processor-implemented method according to any one of aspects 7 to 9, wherein the at least one convolutional filter includes deformable convolution.

[0154] Aspect 11. The processor-implemented method according to aspect 10, wherein at least one weight offset of the deformable convolution is modified based on the difference image.

[0155] Aspect 12. The processor-implemented method according to any one of aspects 1 to 6, wherein the transformation operation includes a transformer operation performed using at least one transformer block.

[0156] Aspect 13. An apparatus for generating one or more segmentation masks, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: generate a difference image based on the difference between the current image and the previous image; use a transformation operation to process the difference image and the features representing the previous image to generate a transformed feature representation of the previous image; combine the transformed feature representation of the previous image with the features representing the current image to generate a combined feature representation of the current image; and generate a segmentation mask for the current image based on the combined feature representation of the current image.

[0157] Aspect 14. The apparatus according to aspect 13, wherein the at least one processor is configured to: process the previous image using a machine learning model to generate the features representing the previous image.

[0158] Aspect 15. The apparatus according to aspect 14, wherein the at least one processor is configured to: process the current image using the machine learning model to generate the features representing the current image.

[0159] Aspect 16. The apparatus according to any one of Aspects 14 or 15, wherein in order to generate the difference image based on the difference between the current image and the previous image, the at least one processor is configured to: process the previous image using the machine learning model to generate intermediate features representing the previous image; process the current image using the machine learning model to generate intermediate features representing the current image; determine the difference between the intermediate features representing the previous image and the intermediate features representing the current image; and generate the difference image based on the difference between the intermediate features representing the previous image and the intermediate features representing the current image.

[0160] Aspect 17. The apparatus according to any one of Aspects 13 to 16, wherein the at least one processor is configured to: combine the difference image with the features representing the previous image to generate a combined feature representation of the previous image.

[0161] Aspect 18. The apparatus according to Aspect 17, wherein in order to process the difference image and the features representing the previous image to generate the transformed feature representation of the previous image, the at least one processor is configured to: process the combined feature representation of the previous image using the transformation operation to generate the transformed feature representation of the previous image.

[0162] Aspect 19. The apparatus according to any one of Aspects 13 to 18, wherein the transformation operation includes a convolution operation performed using at least one convolution filter.

[0163] Aspect 20. The apparatus according to Aspect 19, wherein the weights of the at least one convolution filter are fixed.

[0164] Aspect 21. The apparatus according to Aspect 19, wherein the weights of the at least one convolution filter are modified based on the difference image.

[0165] Aspect 22. The apparatus according to any one of Aspects 19 to 21, wherein the at least one convolution filter includes deformable convolution.

[0166] Aspect 23. The apparatus according to Aspect 22, wherein at least one weight offset of the deformable convolution is modified based on the difference image.

[0167] Aspect 24. The apparatus according to any one of Aspects 13 to 18, wherein the transformation operation includes a transformer operation performed using at least one transformer block.

[0168] Aspect 25. A non-transitory computer-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform the operations according to any one of Aspects 1 to 24.

[0169] Aspect 26. An apparatus comprising one or more components for performing the operations according to any one of Aspects 1 to 24.

Claims

1. A processor-implemented method for generating one or more segmentation masks, the processor-implemented method comprising: Generating a difference image based on a difference between a current image and a previous image; Processing the difference image and features representing the previous image using a transformation operation to generate a transformed feature representation of the previous image; Combining the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image; and Generating a segmentation mask for the current image based on the combined feature representation of the current image.

2. The processor-implemented method according to claim 1, the processor-implemented method further comprising: Processing the previous image using a machine learning model to generate the features representing the previous image.

3. The processor-implemented method according to claim 2, the processor-implemented method further comprising: Processing the current image using the machine learning model to generate the features representing the current image.

4. The processor-implemented method according to claim 2, wherein generating the difference image based on the difference between the current image and the previous image comprises: Processing the previous image using the machine learning model to generate intermediate features representing the previous image; Processing the current image using the machine learning model to generate intermediate features representing the current image; Determining a difference between the intermediate features representing the previous image and the intermediate features representing the current image; and Generating the difference image based on the difference between the intermediate features representing the previous image and the intermediate features representing the current image.

5. The processor-implemented method according to claim 1, the processor-implemented method further comprising: Combining the difference image with the features representing the previous image to generate a combined feature representation of the previous image.

6. The processor-implemented method according to claim 5, wherein processing the difference image and the features representing the previous image using the transformation operation to generate the transformed feature representation of the previous image comprises: Processing the combined feature representation of the previous image using the transformation operation to generate the transformed feature representation of the previous image.

7. The processor-implemented method according to claim 1, wherein the transformation operation comprises a convolution operation performed using at least one convolution filter.

8. The processor-implemented method according to claim 7, wherein the weights of the at least one convolution filter are fixed.

9. The processor-implemented method according to claim 7, wherein the weights of the at least one convolution filter are modified based on the difference image.

10. The processor-implemented method according to claim 7, wherein the at least one convolution filter comprises a deformable convolution.

11. The processor-implemented method according to claim 10, wherein at least one weight offset of the deformable convolution is modified based on the difference image.

12. The processor-implemented method according to claim 1, wherein the transformation operation includes a transformer operation performed using at least one transformer block.

13. An apparatus for generating one or more segmentation masks, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: generate a difference image based on a difference between a current image and a previous image; process the difference image and features representing the previous image using a transformation operation to generate a transformed feature representation of the previous image; combine the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image; and generate a segmentation mask for the current image based on the combined feature representation of the current image.

14. The apparatus according to claim 13, wherein the at least one processor is configured to: process the previous image using a machine learning model to generate the features representing the previous image.

15. The apparatus according to claim 14, wherein the at least one processor is configured to: process the current image using the machine learning model to generate the features representing the current image.

16. The apparatus according to claim 14, wherein in order to generate the difference image based on the difference between the current image and the previous image, the at least one processor is configured to: process the previous image using the machine learning model to generate intermediate features representing the previous image; process the current image using the machine learning model to generate intermediate features representing the current image; determine a difference between the intermediate features representing the previous image and the intermediate features representing the current image; and generate the difference image based on the difference between the intermediate features representing the previous image and the intermediate features representing the current image.

17. The apparatus according to claim 13, wherein the at least one processor is configured to: combine the difference image with the features representing the previous image to generate a combined feature representation of the previous image.

18. The apparatus according to claim 17, wherein in order to process the difference image and the features representing the previous image to generate the transformed feature representation of the previous image, the at least one processor is configured to: process the combined feature representation of the previous image using the transformation operation to generate the transformed feature representation of the previous image.

19. The apparatus according to claim 13, wherein the transformation operation includes a convolution operation performed using at least one convolutional filter.

20. The apparatus according to claim 19, wherein the weights of the at least one convolutional filter are fixed.

21. The apparatus according to claim 19, wherein the weights of the at least one convolutional filter are modified based on the difference image.

22. The apparatus according to claim 19, wherein the at least one convolutional filter includes deformable convolution.

23. The apparatus according to claim 22, wherein at least one weight offset of the deformable convolution is modified based on the difference image.

24. The apparatus according to claim 13, wherein the transformation operation includes a transformer operation performed using at least one transformer block.

25. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: generate a difference image based on a difference between a current image and a previous image; process the difference image and features representing the previous image using a transformation operation to generate a transformed feature representation of the previous image; combine the transformed feature representation of the previous image with features representing the current image to generate a combined feature representation of the current image; and generate a segmentation mask for the current image based on the combined feature representation of the current image.

26. The non-transitory computer-readable medium according to claim 25, wherein the instructions, when executed by the one or more processors, cause the one or more processors to: process the previous image using a machine learning model to generate the features representing the previous image; and process the current image using the machine learning model to generate the features representing the current image.

27. The non-transitory computer-readable medium according to claim 26, wherein, in order to generate the difference image based on the difference between the current image and the previous image, the instructions, when executed by the one or more processors, cause the one or more processors to: process the previous image using the machine learning model to generate intermediate features representing the previous image; process the current image using the machine learning model to generate intermediate features representing the current image; determine a difference between the intermediate features representing the previous image and the intermediate features representing the current image; and generate the difference image based on the difference between the intermediate features representing the previous image and the intermediate features representing the current image.

28. The non-transitory computer-readable medium according to claim 25, wherein the instructions, when executed by the one or more processors, cause the one or more processors to: combine the difference image with the features representing the previous image to generate a combined feature representation of the previous image.

29. The non-transitory computer-readable medium according to claim 28, wherein, in order to process the difference image and the features representing the previous image to generate the transformed feature representation of the previous image, the instructions, when executed by the one or more processors, cause the one or more processors to: process the combined feature representation of the previous image using the transformation operation to generate the transformed feature representation of the previous image.

30. The non-transitory computer-readable medium according to claim 25, wherein the transformation operation includes a convolution operation performed using at least one convolution filter.