A method and apparatus for identifying a vehicle turn signal
By using spatial and temporal encoders to process images in autonomous vehicles, the characteristic and timing signals of turn signals are generated, solving the problem of inaccurate turn signal recognition and improving the safety of autonomous driving and the efficiency of communication between vehicles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING VOYAGER TECH CO LTD
- Filing Date
- 2024-12-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies struggle to accurately identify turn signal status in autonomous vehicles, impacting effective communication between vehicles and decision-making by autonomous driving systems.
The target image is processed using a spatial encoder and a temporal encoder to generate the spatial and temporal features of the turn signal. The timing signal of the turn signal is determined by the recognition system to accurately identify the vehicle's light status information.
It improves safety during autonomous driving by accurately recognizing turn signal status, enhancing communication between vehicles and the autonomous driving system's understanding capabilities.
Smart Images

Figure CN122135333A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to a method, apparatus, device, computer-readable storage medium, and computer program product for recognizing vehicle turn signals. Background Technology
[0002] In the development of autonomous driving technology, vehicle turn signal recognition technology plays a crucial role. It is not only essential for effective communication between vehicles, but also key to the autonomous driving system's ability to understand and respond to the driving intentions of other vehicles. Through high-precision cameras and advanced image processing algorithms, autonomous vehicles can capture and analyze the turn signal signals of surrounding vehicles in real time. This technology can accurately interpret the flashing patterns of turn signals, whether it's a left turn, right turn, or an emergency braking warning, all of which can be quickly identified and incorporated into the decision-making process of the autonomous driving system. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for recognizing vehicle turn signals is provided. The method includes: processing a target image associated with a vehicle using a spatial encoder in a recognition system to determine a first spatial feature associated with a first turn signal of the vehicle and a second spatial feature associated with a second turn signal of the vehicle; generating a first temporal feature corresponding to the first turn signal and a second temporal feature associated with the second turn signal using a time encoder in the recognition system based on the first spatial feature, the second spatial feature, and time information corresponding to the target image; and determining a first temporal signal of the first turn signal and a second temporal signal of the second turn signal based on the first and second temporal features to determine the vehicle's light status information.
[0004] In a second aspect of this disclosure, an apparatus for identifying vehicle turn signals is provided. The apparatus includes: a processing module configured to process a target image associated with a vehicle using a spatial encoder in an identification system to determine a first spatial feature associated with a first turn signal and a second spatial feature associated with a second turn signal; a first generation module configured to generate a first temporal feature corresponding to the first turn signal and a second temporal feature associated with the second turn signal using a time encoder in the identification system based on the first spatial feature, the second spatial feature, and time information corresponding to the target image; and a determining module configured to determine a first temporal signal of the first turn signal and a second temporal signal of the second turn signal based on the first and second temporal features to determine vehicle light status information.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method of the first aspect.
[0008] It should be understood that the content described in this summary section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram of an example identification system that can be implemented in accordance with embodiments of the present disclosure is shown;
[0011] Figure 2 A schematic diagram illustrating an example process for recognizing vehicle turn signals according to some embodiments of the present disclosure is shown;
[0012] Figure 3 A schematic structural block diagram of an example device for identifying vehicle turn signals according to some embodiments of the present disclosure is shown; and
[0013] Figure 4 A block diagram of an apparatus capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0017] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0018] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0019] As briefly mentioned earlier, vehicle turn signal recognition technology plays a crucial role in the development of autonomous driving technology. It is not only essential for effective communication between vehicles, but also key to the autonomous driving system's ability to understand and respond to the driving intentions of other vehicles. Through high-precision cameras and advanced image processing algorithms, autonomous vehicles can capture and analyze the turn signal signals of surrounding vehicles in real time.
[0020] Embodiments of this disclosure propose a vehicle turn signal recognition scheme. According to various embodiments of this disclosure, a spatial encoder in the recognition system processes a target image associated with a vehicle to determine a first spatial feature associated with a first turn signal and a second spatial feature associated with a second turn signal; a time encoder in the recognition system generates a first temporal feature corresponding to the first turn signal and a second temporal feature associated with the second turn signal based on the first spatial feature, the second spatial feature, and time information corresponding to the target image; and based on the first and second temporal features, a first temporal signal for the first turn signal and a second temporal signal for the second turn signal are determined to determine the vehicle's light status information.
[0021] In this way, the embodiments of this disclosure can more accurately identify the vehicle's light status information, thereby increasing safety during autonomous driving.
[0022] Example Environment
[0023] Figure 1 A schematic diagram of an example identification system 100 that can be implemented according to embodiments of the present disclosure is shown. The identification system 110 can be deployed in an autonomous vehicle or in a server that communicates with the autonomous vehicle. As shown, the identification system 100 may include a spatial encoder 110 and a temporal encoder 120. The spatial encoder 110 and the temporal encoder 120 are implemented based on a Transformer model.
[0024] exist Figure 1 In the recognition system 100, the recognition system 100 can acquire a target image 130 associated with the vehicle. The target image 130 can be a set of images captured by a camera device deployed on the autonomous vehicle, or it can be a preset set of images. The target image 130 includes at least the vehicle's turn signals.
[0025] After acquiring the target image 130, the recognition system 110 can call the spatial encoder 110 to process the target image 130 to obtain the spatial features of the vehicle's turn signals. Further, the recognition system 110 can process the spatial features of the turn signals using the time encoder 120 to obtain the temporal features of the turn signals. Finally, the recognition system 110 can determine the vehicle's turn signal status information based on the temporal features.
[0026] In some embodiments, the spatial encoder 110 may be implemented as a combination of a Tiny-Vit and a Multilayer Perceptron (MLP). Tiny-Vit is a small transformer model that transfers knowledge from a large model to a small model through knowledge distillation to take advantage of large-scale datasets.
[0027] In some embodiments, the temporal encoder 120 may be implemented, for example, as a stack of multiple Transformer structures. The Transformer structure includes at least one self-attention layer. The self-attention layer is the core component of the Transformer structure, which allows the model to pay attention to each position in the sequence when processing the sequence, thereby better understanding the contextual information of the sequence.
[0028] It should be understood that the structure and function of environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0029] Example process
[0030] Figure 2 A flowchart of an example process 200 for a travel service according to some embodiments of the present disclosure is shown. Process 200 may be implemented at identification system 100. References are made below. Figure 1 Describe the process 200.
[0031] like Figure 2 As shown, in block 210, the recognition system 100 uses a spatial encoder to process a target image associated with a vehicle to determine a first spatial feature associated with the vehicle's first turn signal and a second spatial feature associated with the vehicle's second turn signal.
[0032] As an example, such as Figure 1 As shown, the recognition system 100 can acquire a target image 130 associated with the vehicle. The target image 130 can be a set of images captured by a camera device deployed in the vehicle, or it can be a preset set of images. The target image 130 includes the vehicle's turn signal information. In some scenarios, the target image 130 can be a set of photos of the rear of the vehicle, or it can be video content including the rear view of the vehicle.
[0033] After acquiring the target image 130, the recognition system 100 can call the spatial encoder 110 to process the target image 130 to determine the first spatial feature and the second spatial feature associated with the first turn signal and the second turn signal, respectively.
[0034] Continue to refer to Figure 2 In box 220, the recognition system 100 uses a time encoder to generate a first temporal feature corresponding to the first turn signal and a second temporal feature associated with the second turn signal based on the first spatial feature, the second spatial feature and the time information corresponding to the target image.
[0035] As an example, such as Figure 1As shown, after obtaining the first spatial feature and the second spatial feature, the recognition system 100 can input the first spatial feature, the second spatial feature, and the time information corresponding to the target image into the time encoder 120 to generate the first temporal feature and the second temporal feature. The time information corresponding to the target image may be, for example, the position of the target image in a set of reference images.
[0036] In some embodiments, the spatial encoder and / or time encoder are implemented based on a converter model. For example, such as... Figure 1 As shown, the spatial encoder 110 can be implemented as a combination of a small vision transformer 115 and multiple multilayer perceptrons. The temporal encoder 120 can be implemented as a superposition of multiple transformer structures. The spatial and temporal encoders with a full Transformer structure enable more consistent configuration of the optimizer and learning rate, thereby achieving optimal optimization results.
[0037] Since turn signal recognition depends on the input order of images, and the time encoder 120 is implemented as a Transformer structure that processes input data in parallel, in some embodiments, the recognition system 100 can update the first spatial feature and the second spatial feature based on time information. As an example, the recognition system 100 can perform position encoding on the input target image 130 to determine the time information of the target image 130. The recognition system 100 can update the first spatial feature and the second spatial feature based on the time information.
[0038] Furthermore, the recognition system 100 can construct a first feature sequence corresponding to the first turn signal and a second feature sequence corresponding to the second turn signal based on the updated first spatial features and the updated second spatial features, respectively. As an example, the recognition system 100 can construct a first feature sequence corresponding to the first turn signal and a second feature sequence corresponding to the second turn signal within the target time period based on the updated first spatial features and second spatial features.
[0039] Finally, the recognition system 100 can use a time encoder to process the first feature sequence and the second feature sequence respectively to generate a first temporal feature and a second temporal feature. As an example, the recognition system 100 can invoke the time encoder 120 to process the first feature sequence and the second feature sequence to generate a first temporal feature and a second temporal feature associated with a target time period. The target time period can indicate an acquisition cycle of a set of images to which the target image 130 belongs. For example, if a vehicle's camera captures a 5-second video, then the target time period could be 5 seconds.
[0040] In some embodiments, the identification system 100 may determine multiple feature components corresponding to multiple dimensions of the first spatial feature and the second spatial feature based on time information. Further, the identification system 100 may update the first spatial feature and the second spatial feature based on the multiple feature components.
[0041] As an example, the recognition system 100 can perform position encoding on the T-frame images based on the time information of the target image 130 to determine the order of each image in the T-frame images. The position encoding PE(t,d) (i.e., the feature component) can be calculated using a combination of sine and cosine functions. The position encoding can be calculated, for example, based on the following formula:
[0042]
[0043] Where t is the temporal index of each image, t∈{1,2,…T}. d is the dimensional index, d∈{1,2,…D}, where D is the dimension of the spatial features corresponding to each frame of images. k is the index of half the dimension (i.e., the result of d divided by 2, rounded down).
[0044] After calculating the position code for each frame of the image, the recognition system 100 can match the target position code with the target spatial feature x. t The features of the target space are then updated by adding them together. The specific formula is as follows:
[0045] x′t=x t +PE(t,;) (3)
[0046] Where PE(t,;) is the target location code corresponding to the target spatial features, x′ t This refers to the updated target space features.
[0047] In some embodiments, time information indicates the number of the target image in the time dimension. As an example, a set of images includes a first image, a target image, and a second image. The first image, target image, and second image are arranged in chronological order. Then, the first image is numbered 1, the target image is numbered 2, and the second image is numbered 3.
[0048] Continue to refer to Figure 2 In box 230, the recognition system 100 determines the first timing signal of the first turn signal and the second timing signal of the second turn signal based on the first timing feature and the second timing feature, so as to determine the vehicle's light status information.
[0049] As an example, such as Figure 1As shown, after the recognition system 100 generates a left-turn timing feature (i.e., a first timing feature) 122 and a right-turn timing feature (i.e., a second timing feature) 124 through the time encoder 120, the recognition system 100 can input the left-turn timing feature 122 and the right-turn timing feature 124 to the multilayer perceptrons 140-1 and 140-2 respectively to determine the first timing signal and the second timing signal. The first timing signal can be, for example, the first operating state (flashing or off) of the left-turn light during the target time period, and the second timing signal can be, for example, the second operating state (flashing or off) of the right-turn light during the target time period. Further, the recognition system 100 can determine the vehicle's light status information based on the first timing signal and the second timing signal. The vehicle's light status information can, for example, indicate the vehicle's driving state (e.g., left-turn state, right-turn state, braking state, etc.).
[0050] In some embodiments, the vehicle's light status information indicates one of the following states: a first state, corresponding to a first timing signal being a flashing signal and a second timing signal being an off signal; a second state, corresponding to a first timing signal being an off signal and a second timing signal being a flashing signal; a third state, corresponding to both the first and second timing signals being flashing signals; and a fourth state, corresponding to both the first and second timing signals being off signals.
[0051] As an example, the first state corresponds to the left turn signal flashing and the right turn signal being off, meaning the first state is a left turn. The second state corresponds to the left turn signal being off and the right turn signal flashing, meaning the second state is a right turn. The third state corresponds to both the left and right turn signals flashing, meaning the third state is a hazard warning state. The fourth state corresponds to both the left and right turn signals being off, meaning the fourth state is a stopped state.
[0052] In some embodiments, the light status information is first light status information, and the recognition system 100 can also process the target image to generate second light status information of the vehicle's brake lights. As an example, such as... Figure 1 As shown, the recognition system 100 can process the target image using the spatial encoder 110 to obtain brake signal features 170. Further, the recognition system 100 can generate second lamp status information for the vehicle's brake lights based on the brake signal features. The second lamp status information indicates that the left and right turn signals are in a constantly illuminated state.
[0053] The training process of the recognition system 100 will be further described below. In some embodiments, the recognition system 100 is trained by a suitable training device (e.g., a server) using sample images.
[0054] In some embodiments, similar to the reasoning process described above, the server can use the recognition system 100 to be trained to generate recognition information of the sample image. The recognition information at least indicates the predicted light state information of the turn signal in the sample image, which is determined based on the spatial features output by the spatial encoder.
[0055] Furthermore, the server can determine the training loss of the recognition system, at least based on a comparison of the predicted turn signal status information and the reference turn signal status information of the sample image, to train the recognition system. Specifically, the predicted turn signal status information and the reference turn signal status information can indicate whether the turn signal is illuminated.
[0056] As an example, such as Figure 1 As shown, the server can use the spatial features output by the spatial encoder 110 in the recognition system 100 to generate recognition information for the sample image. In some embodiments, the recognition information may include the timing light status information described above, such as whether the turn signals are flashing.
[0057] Additionally, the recognition information may also include the predicted state of the turn signal in the current image. In some embodiments, during the training phase, the recognition system 100 may also determine the first predicted state of the left / right turn signal in the sample image based on the spatial features of the left / right turn signal in the sample image, such as whether it is illuminated.
[0058] Furthermore, the server can determine a first loss value 150 associated with the first turn signal and a second loss value 155 associated with the second turn signal based on the first predicted light state and the first reference light state information, respectively, thereby training the spatial encoder 110. As an example, the first reference light state information can indicate the first actual light state of the left / right turn signal in the sample image, such as whether it is actually lit.
[0059] Alternatively, during the training phase, the recognition system 100 may determine a second predicted light state corresponding to the left / right turn signal in the sample image set, such as whether it is flashing, based on the temporal characteristics of the left / right turn signal in the sample image set.
[0060] Furthermore, the server can determine a third loss value 160 associated with the first turn signal and a fourth loss value 165 associated with the second turn signal, respectively, based on the second predicted light state and the second reference light state information, thereby training the temporal encoder 120. As an example, the second reference light state information can indicate the second actual light state of the left / right turn signal determined based on the sample image set, for example, whether it is actually flashing.
[0061] In some embodiments, the server can perform data augmentation on the sample image. Specifically, the server can acquire a captured reference image. Further, the server can apply image processing procedures to the reference image to obtain the sample image. The image processing procedures are used to modify the content of the reference image.
[0062] As an example, the server can acquire reference images captured by cameras deployed on a vehicle and perform data augmentation on these images. Data augmentation includes adjusting the size and attributes of the reference image (e.g., saturation, brightness, hue, etc.) and horizontally flipping the reference image. This effectively increases the number of sample images, making the model trained on these sample images more robust.
[0063] In some embodiments, the server can replace the content of at least one region in a reference image with preset content to obtain a sample image. For example, the server can occlude a portion of the reference image. For instance, the reference image is an image of the rear of a vehicle. The server can use a preset mask to cover vehicle components such as the rear bumper, rear windshield, and tires in the rear image of the vehicle. This improves the robustness of the recognition system across multiple scenarios.
[0064] In some embodiments, the server can crop at least one edge region of a reference image to obtain a sample image. For example, the reference image is a rear view of a vehicle. The server can crop out portions of the rear view of the vehicle that are not associated with the vehicle's turn signals (e.g., lane lines, traffic lights, etc.) to obtain the sample image. This improves the training efficiency of the recognition system 100.
[0065] In this way, the embodiments of this disclosure can more accurately identify the vehicle's light status information, thereby increasing safety during autonomous driving.
[0066] Example devices and equipment
[0067] Figure 3 A schematic structural block diagram of a device 300 for identifying vehicle turn signals according to certain embodiments of the present disclosure is shown. The device 300 may be implemented as or included in the identification system 100. The various modules / components in the device 300 may be implemented by hardware, software, firmware, or any combination thereof.
[0068] As shown in the figure, the device 300 includes a processing module 310 configured to process a target image associated with a vehicle using a spatial encoder in the recognition system to determine a first spatial feature associated with a first turn signal of the vehicle and a second spatial feature associated with a second turn signal of the vehicle; a first generation module 320 configured to generate a first temporal feature corresponding to the first turn signal and a second temporal feature associated with the second turn signal using a time encoder in the recognition system based on the first spatial feature, the second spatial feature and time information corresponding to the target image; and a determination module 330 configured to determine a first temporal signal of the first turn signal and a second temporal signal of the second turn signal based on the first temporal feature and the second temporal feature to determine the vehicle's light status information.
[0069] In some embodiments, the first generation module 320 is further configured to: update the first spatial feature and the second spatial feature based on time information; construct a first feature sequence corresponding to the first turn signal and a second feature sequence corresponding to the second turn signal based on the updated first spatial feature and the updated second spatial feature; and process the first feature sequence and the second feature sequence using a time encoder to generate the first temporal feature and the second temporal feature.
[0070] In some embodiments, the generation module 320 is further configured to: determine multiple feature components corresponding to multiple dimensions of the first spatial feature and the second spatial feature based on time information; and update the first spatial feature and the second spatial feature based on the multiple feature components.
[0071] In some embodiments, time information indicates the number of the target image in the time dimension.
[0072] In some embodiments, the recognition system is trained based on the following process: generating recognition information of sample images using the recognition system, the recognition information indicating at least the predicted state information of the turn signals in the sample images, the predicted state information being determined based on spatial features output by a spatial encoder; and determining the training loss of the recognition system based at least on a comparison of the predicted state information and reference state information of the sample images, to train the recognition system, wherein the predicted state information and the reference state information indicate whether the turn signals are illuminated.
[0073] In some embodiments, the sample image is determined based on the following process: acquiring a captured reference image; and applying an image processing procedure to the reference image to obtain the sample image, wherein the image processing procedure is used to modify the content of the reference image.
[0074] In some embodiments, applying image processing procedures to a reference image to obtain a sample image includes: replacing the content of at least one region in the reference image with preset content; or cropping at least one edge region of the reference image.
[0075] In some embodiments, the lamp status information indicates one of the following states: a first state, which corresponds to a first timing signal being a flashing signal and a second timing signal being an off signal; a second state, which corresponds to a first timing signal being an off signal and a second timing signal being a flashing signal; a third state, which corresponds to both the first and second timing signals being flashing signals; and a fourth state, which corresponds to both the first and second timing signals being off signals.
[0076] In some embodiments, the spatial encoder and / or time encoder are implemented based on a converter model.
[0077] In some embodiments, the lamp status information is first lamp status information, and the device 300 further includes a second generation module configured to process the target image using a recognition system to generate second lamp status information of the vehicle's brake lights.
[0078] Figure 4 A block diagram illustrating a computing device 400 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 4 The computing device 400 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 4 The computing device 400 shown can be used to implement Figure 1 The identification system 100.
[0079] like Figure 4 As shown, computing device 400 is in the form of a general-purpose computing device. Components of computing device 400 may include, but are not limited to, one or more processors or processing units 410, memory 420, storage devices 430, one or more communication units 440, one or more input devices 450, and one or more output devices 460. Processing unit 410 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 420. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 400.
[0080] Computing device 400 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to computing device 400, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 420 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 430 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 400.
[0081] The computing device 400 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 4 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 420 may include computer program product 425 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0082] The communication unit 440 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 400 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 400 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.
[0083] Input device 450 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 460 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 400 can also communicate as needed with one or more external devices (not shown) via communication unit 440. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 400, or with any device (e.g., network card, modem, etc.) that enables computing device 400 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).
[0084] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0085] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0086] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0087] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0089] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for recognizing vehicle turn signals, comprising: The spatial encoder in the recognition system is used to process the target image associated with the vehicle to determine a first spatial feature associated with the vehicle's first turn signal and a second spatial feature associated with the vehicle's second turn signal; The time encoder in the recognition system generates a first temporal feature corresponding to the first turn signal and a second temporal feature associated with the second turn signal based on the first spatial feature, the second spatial feature, and the time information corresponding to the target image. as well as Based on the first timing feature and the second timing feature, a first timing signal of the first turn signal and a second timing signal of the second turn signal are determined to determine the vehicle's light status information.
2. The method according to claim 1, wherein generating a first timing feature corresponding to the first turn signal and a second timing feature associated with the second turn signal comprises: Based on the time information, update the first spatial feature and the second spatial feature; Based on the updated first spatial features and the updated second spatial features, a first feature sequence corresponding to the first turn signal and a second feature sequence corresponding to the second turn signal are constructed respectively. as well as The first feature sequence and the second feature sequence are processed by the time encoder to generate the first temporal feature and the second temporal feature.
3. The method according to claim 2, wherein updating the first spatial feature and the second spatial feature based on the time information comprises: Based on the time information, multiple feature components corresponding to multiple dimensions of the first spatial feature and the second spatial feature are determined; as well as Based on the multiple feature components, the first spatial feature and the second spatial feature are updated.
4. The method according to claim 1, wherein the time information indicates the number of the target image in the time dimension.
5. The method of claim 1, wherein the recognition system is trained based on the following process: The recognition system generates recognition information for sample images, the recognition information indicating at least the predicted turn signal state information in the sample images, the predicted turn signal state information being determined based on the spatial features output by the spatial encoder; and The training loss of the recognition system is determined based at least on a comparison between the predicted light state information and the reference light state information of the sample image, in order to train the recognition system, wherein the predicted light state information and the reference light state information indicate whether the turn signal is lit.
6. The method of claim 5, wherein the sample image is determined based on the following process: Acquire reference images for shooting; and An image processing procedure is applied to the reference image to obtain the sample image, wherein the image processing procedure is used to change the content of the reference image.
7. The method of claim 6, wherein applying an image processing procedure to the reference image to obtain the sample image comprises: Replace the content of at least one region in the reference image with preset content; or At least one edge region of the reference image is cropped.
8. The method of claim 1, wherein the vehicle's light status information indicates one of the following states: The first state corresponds to: the first timing signal being a flashing signal and the second timing signal being an off signal; The second state corresponds to: the first timing signal being an off signal and the second timing signal being a flashing signal; The third state corresponds to the situation where both the first timing signal and the second timing signal are flashing signals. The fourth state corresponds to the situation where both the first timing signal and the second timing signal are off signals.
9. The method of claim 1, wherein the spatial encoder and / or the time encoder is implemented based on a converter model.
10. The method according to claim 1, wherein the lamp status information is first lamp status information, further comprising: The target image is processed using a recognition system to generate the second lamp status information of the vehicle's brake lights.
11. A device for identifying vehicle turn signals, comprising: The processing module is configured to process a target image associated with a vehicle using a spatial encoder in the recognition system to determine a first spatial feature associated with a first turn signal of the vehicle and a second spatial feature associated with a second turn signal of the vehicle. The first generation module is configured to use the time encoder in the recognition system to generate a first temporal feature corresponding to the first turn signal and a second temporal feature associated with the second turn signal based on the first spatial feature, the second spatial feature and the time information corresponding to the target image; as well as The determining module is configured to determine a first timing signal of the first turn signal and a second timing signal of the second turn signal based on the first timing feature and the second timing feature, so as to determine the vehicle's light status information.
12. An electronic device, comprising: At least one processing unit; as well as At least one memory is coupled to at least one processing unit and stores instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.
14. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.