State determination method and device, electronic equipment, medium and vehicle
By acquiring temporal information of target image sets around the vehicle, and using a pre-trained bird's-eye view model to extract speed changes and size information, the problem of inaccurate judgment of the state of objects with speed jumps in existing technologies is solved, achieving higher accuracy of state information and user experience.
Patent Information
- Application Number
- CN202410605917.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies, when using intelligent vehicle systems to determine the state of objects around a vehicle, cannot accurately determine the state of objects with sudden changes in speed based solely on speed information, resulting in a poor user experience.
By acquiring the temporal information of the target image set, and utilizing the feature extraction module and decoding network of the pre-trained bird's-eye view model, the velocity change information and velocity magnitude information of the target are extracted and determined, thereby accurately judging the state of the target.
It improves the accuracy of target status information, avoids inaccurate status information caused by sudden changes in target speed, and enhances user experience.
Smart Images

Figure CN120963709A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of response time, and more particularly to a state determination method, apparatus, electronic device, medium, and vehicle. Background Technology
[0002] With the continuous development of internet technology and the rapid popularization of mobile terminals, intelligent vehicle systems have become a focus of attention for major manufacturers. When users use intelligent vehicle systems to judge the state of objects around the vehicle, current technology usually only predicts the state of surrounding objects based on their speed information, and cannot accurately judge the state of objects with sudden changes in speed, thus affecting the user experience. Summary of the Invention
[0003] This disclosure provides a state determination method, apparatus, electronic device, medium, and vehicle, which proposes a method for determining target state information based on target speed change information and speed magnitude information, thereby improving the accuracy of target state information and avoiding inaccurate state information determined due to sudden changes in target speed.
[0004] A first aspect of this disclosure proposes a state determination method, which includes: acquiring a target image set with temporal information; extracting feature vectors from the target image set using a feature extraction module of a pre-trained bird's-eye view model; determining velocity change information and velocity magnitude information of the target in the target image set based on the temporal information and feature vectors, using a decoding network of the pre-trained bird's-eye view model, wherein the velocity change information is used to indicate the motion state of the target; and determining the state information of the target based on the velocity change information and velocity magnitude information.
[0005] In some embodiments of this disclosure, the feature extraction module of the pre-trained model extracts the feature vector of the target image set by: using a convolutional neural network to extract a first feature from the target image set; and using a spatial transformation algorithm to project the first feature onto a preset coordinate system to obtain the feature vector.
[0006] In some embodiments of this disclosure, determining the velocity change information and velocity magnitude information of a target in a target image set based on temporal information and feature vectors, using a pre-trained bird's-eye view model decoding network and a multilayer perception network, includes: determining a first detection parameter set of the target image based on temporal information and feature vectors, using a pre-trained bird's-eye view model decoding network and a multilayer perception network, the first detection parameter set including at least: target attribute information, velocity change information, and tracking result information; or, determining the tracking result and a second detection parameter set of the target in the target image based on temporal information and feature vectors, using a pre-trained bird's-eye view model decoding network and a multilayer perception network, the second detection parameter set including at least: velocity change information and velocity magnitude information; and determining the velocity change information and velocity magnitude information of the target based on the first detection parameter set, or based on the tracking result and the second detection parameter set.
[0007] In some embodiments of this disclosure, determining the first detection parameter set of a target image based on temporal information and feature vectors, using a decoding network of a pre-trained bird's-eye view model and a multilayer perceptron includes: determining the decoding features in the target image set using the decoding network based on the feature vectors; and determining the first detection parameter set of the target image using the multilayer perceptron based on the decoding features and temporal information.
[0008] In some embodiments of this disclosure, determining the tracking result and second detection parameter set of the target image based on temporal information and feature vectors, using a decoding network and a multilayer perceptron of a pre-trained bird's-eye view model, includes: determining the second detection parameter set of the target image set using a multilayer perceptron based on feature vectors and temporal information; and determining the tracking result information of the target image set using a decoding network based on feature vectors and temporal information.
[0009] In some embodiments of this disclosure, the method further includes: obtaining a reference parameter set of the target, the reference parameter set including at least: reference velocity change information and reference velocity magnitude information; determining reference state information of the target based on the reference parameter set; and comparing the difference between the reference state information and the state information with a preset threshold to determine the accuracy of the state information.
[0010] A second aspect of this disclosure provides a state determination apparatus, comprising: an acquisition module for acquiring a set of target images with temporal information; an extraction module for extracting feature vectors from the target image set using a feature extraction module of a pre-trained bird's-eye view model; a processing module for determining velocity change information and velocity magnitude information of a target in the target image set based on the temporal information and feature vectors, using a decoding network of the pre-trained bird's-eye view model, wherein the velocity change information is used to indicate the motion state of the target; and a determination module for determining the state information of the target based on the velocity change information and velocity magnitude information.
[0011] A third aspect of this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the methods described in the first aspect of this disclosure.
[0012] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to cause a computer to perform the method described in the first aspect of this disclosure.
[0013] A fifth aspect of this disclosure provides a vehicle including the apparatus described in the second aspect of the preceding embodiment or the electronic equipment described in the third aspect of the preceding embodiment.
[0014] In summary, the state determination method proposed in this disclosure includes: acquiring a target image set with temporal information; extracting feature vectors from the target image set using the feature extraction module of a pre-trained bird's-eye view model; determining the velocity change information and velocity magnitude information of the target in the target image set based on the temporal information and feature vectors, using the decoding network of the pre-trained bird's-eye view model; and determining the target's state information based on the velocity change information and velocity magnitude information. This method improves the accuracy of state information by determining the target's velocity change information and velocity magnitude information, and avoids inaccurate state information determined due to sudden changes in target velocity.
[0015] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0017] Figure 1 This is a flowchart of a state determination method according to an embodiment of the present disclosure;
[0018] Figure 2 This is a flowchart of another state determination method according to an embodiment of the present disclosure;
[0019] Figure 3 This is a flowchart of another state determination method according to an embodiment of the present disclosure;
[0020] Figure 4 This is a schematic diagram of the structure of a state determination device according to an embodiment of the present disclosure;
[0021] Figure 5 This is a block diagram illustrating an electronic device for implementing the state determination method of this disclosure, according to an exemplary embodiment. Detailed Implementation
[0022] Embodiments of this disclosure are described in detail below, with examples of embodiments shown in the accompanying drawings, wherein the same or similar reference numerals identify the same or similar originals or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0023] With the continuous development of internet technology and the rapid popularization of mobile terminals, intelligent vehicle systems have become a focus of attention for major manufacturers. When users use intelligent vehicle systems to judge the state of objects around the vehicle, current technology usually only predicts the state of surrounding objects based on their speed information, and cannot accurately judge the state of objects with sudden changes in speed, thus affecting the user experience.
[0024] This disclosure aims to propose a method for determining target state information based on target velocity change information and velocity magnitude information, with the goal of improving the accuracy of target state information and avoiding inaccurate state information determined due to sudden changes in target velocity.
[0025] The method proposed in this disclosure can be widely applied to code development and maintenance in areas such as vehicle driving, vehicle-assisted driving, autonomous driving, and vehicle electronic control. The application scenarios are not limited in the embodiments of this disclosure.
[0026] The method for determining the state provided in this application will be described in detail below with reference to the accompanying drawings.
[0027] Figure 1 This is a flowchart of a state determination method according to an embodiment of this disclosure. Figure 1 The illustrated embodiment shows that the state determination method includes:
[0028] Step 101: Obtain the target image set with temporal information.
[0029] In some embodiments, vehicle sensors, cameras, etc., can be used to collect road conditions around the vehicle body to obtain a target image set with time-series information. This disclosure does not limit the method of obtaining the target image set.
[0030] In some embodiments, the target image set may be images of pedestrians, vehicles, roadblocks, and other objects acquired around the user's vehicle. This disclosure does not limit the specific targets in the target image set.
[0031] In some embodiments, the timing information may be time-labeled images when acquiring the target image set.
[0032] In some embodiments, the target image set can be determined by acquiring a video segment and performing methods such as cropping or frame extraction on the video. In this case, the time frames in the video can be determined as the temporal information of the images in the target image set.
[0033] Step 102: Use the feature extraction module of the pre-trained bird's-eye view model to extract the feature vector of the target image set.
[0034] In some embodiments, the feature extraction module of a pre-trained bird's-eye view (BEV) model can be used to convert feature information in the icon image set into vector information in a predetermined coordinate system, thereby laying the foundation for determining the velocity change information and state information of the target.
[0035] Specifically, the first feature in the target image set can be extracted by using the convolutional neural network in the model; the first feature can be projected onto a preset coordinate system using a spatial transformation algorithm to obtain the feature vector.
[0036] In some embodiments, feature vectors can be used to indicate velocity change information and velocity magnitude information of targets in a target image set.
[0037] Step 103: Based on the temporal information and feature vectors, use the decoding network of the pre-trained bird's-eye view model to determine the velocity change information and velocity magnitude information of the target in the target image set.
[0038] In some embodiments, a first set of detection parameters for a target image can be determined using a decoding network and a multilayer perceptual network of a pre-trained bird's-eye view model based on temporal information and feature vectors. The first set of detection parameters includes at least: target attribute information, velocity change information, and tracking result information. Then, based on the first set of detection parameters, the velocity change information and velocity magnitude information of the target can be determined.
[0039] Specifically, based on the feature vectors, the decoding features in the target image set can be determined using a decoding network; and based on the decoding features and temporal information, the first detection parameter set of the target image can be determined using a multilayer perceptron.
[0040] In some embodiments, the tracking result and a second detection parameter set of the target image can be determined based on temporal information and feature vectors, using the decoding network and multilayer perceptron of a pre-trained bird's-eye view model. The second detection parameter set includes at least: velocity change information and velocity magnitude information; and then the velocity change information and velocity magnitude information of the target can be determined based on the tracking result and the second detection parameter set.
[0041] Specifically, based on the feature vectors, the decoding features in the target image set can be determined using a decoding network; and based on the decoding features and temporal information, the first detection parameter set of the target image can be determined using a multilayer perceptron.
[0042] In some embodiments, velocity change information is used to indicate the motion state of a target.
[0043] Step 104: Determine the target's state information based on the velocity change information and velocity magnitude information.
[0044] In some embodiments, the velocity change information and velocity magnitude information can be directly determined as the target's state information, or the velocity change information and velocity magnitude information can be weighted and fused, and the fused result can be determined as the target's state information. This disclosure does not limit this.
[0045] In summary, the state determination method proposed in this disclosure includes: acquiring a target image set with temporal information; extracting feature vectors from the target image set using the feature extraction module of a pre-trained bird's-eye view model; determining the velocity change information and velocity magnitude information of the target in the target image set based on the temporal information and feature vectors, using the decoding network of the pre-trained bird's-eye view model; and determining the target's state information based on the velocity change information and velocity magnitude information. This method improves the accuracy of state information by determining the target's velocity change information and velocity magnitude information, and avoids inaccurate state information determined due to sudden changes in target velocity.
[0046] Figure 2 This is a flowchart of a state determination method according to an embodiment of the present disclosure. Figure 1 The illustrated embodiment is further explained as follows: Figure 2 The illustrated embodiment shows that the state determination method includes:
[0047] Step 201: Obtain the target image set with temporal information.
[0048] In this disclosure, the principle of step 201 is similar to Figure 1 Step 101 in the illustrated embodiment is the same and can be referred to accordingly. Figure 1 The relevant descriptions will not be repeated here.
[0049] Step 202: Use the convolutional neural network in the feature extraction module to extract the first feature from the target image set.
[0050] In some embodiments, the first feature may be used to indicate texture features, edge features, etc. in the target image set. This disclosure does not limit the specific features of the first feature. Any feature that can be used to determine speed magnitude information and speed change information can be the first feature.
[0051] In some embodiments, a convolutional neural network can be used to perform convolution operations on the data in the target image set, thereby extracting features from the target image set.
[0052] Step 203: Using the spatial transformation algorithm in the feature extraction module, the first feature is projected onto a preset coordinate system to obtain the feature vector.
[0053] In some embodiments, a spatial transformation algorithm can be used to project the first feature onto a preset coordinate system, and a BEV encoder can be used to encode the first feature to obtain a feature vector.
[0054] Spatial conversion algorithms include, for example, BEV conversion algorithms based on pure vision or BEV conversion algorithms based on multimodal fusion, which are not limited in this disclosure.
[0055] In some embodiments, a spatial transformation algorithm is used to transform the first feature in the target image set into the same coordinate system, which facilitates the subsequent acquisition of velocity change information and velocity magnitude information.
[0056] Step 204: Based on the feature vectors, use the decoding network to determine the decoding features in the target image set.
[0057] In some embodiments, the decoding network is used to decode the feature vectors to determine the decoded features as the decoded features of the target image.
[0058] Step 205: Based on the decoded features and temporal information, a first set of detection parameters for the target image is determined using a multilayer perceptron.
[0059] In some embodiments, the first set of detection parameters includes at least: target attribute information, velocity change information, and tracking result information.
[0060] Among them, target attribute information is used to indicate the category of each target in the target image data, such as indicating that the target is a pedestrian, vehicle, etc. Target attribute information can also be used to indicate the attributes of each target in the target image, such as the target's width, height, etc.; velocity change information can be used to indicate the target's motion state, such as indicating that the target is in a deceleration state, acceleration state, etc.; tracking result information is used to indicate the target's motion trajectory.
[0061] In some embodiments, the decoded features can be extracted using a multilayer perceptual network (MLP) of a pre-trained bird's-eye view model. At the same time, the temporal change state of the decoded features can be determined based on the temporal information, thereby determining the first set of detection parameters for the target image.
[0062] Step 206: Determine the target's velocity change information and velocity magnitude information based on the first detection parameter set.
[0063] In some embodiments, the speed magnitude information can be determined based on the tracking result information and speed change information, and the target attribute information can be used to correspond each target with the speed change information and speed magnitude information, thereby determining the target's speed change information and speed magnitude information.
[0064] Step 207: Determine the target's state information based on the velocity change information and velocity magnitude information.
[0065] In this disclosure, the principle of step 207 is similar to Figure 1 Step 103 in the illustrated embodiment is the same and can be referred to accordingly. Figure 1 The relevant descriptions will not be repeated here.
[0066] Step 208: Obtain the reference parameter set of the target.
[0067] In some embodiments, the reference parameter set includes at least: reference velocity change information and reference velocity magnitude information.
[0068] In some embodiments, the reference parameter set may be the ground truth values of objects of the same type as the target used by the pre-trained bird's-eye view model during the training process.
[0069] In some embodiments, the reference parameter set of the target may be stored in the database of the device executing the method, so that the reference parameter set can be obtained by database query. This disclosure does not limit the method of obtaining the reference parameter set.
[0070] Step 209: Determine the reference state information of the target based on the reference parameter set.
[0071] In some embodiments, reference state information of the target is determined based on a set of reference parameters. The method for determining the reference state information is the same as the method for determining the state information described above, and can be referred to the relevant descriptions of steps 201-207, which will not be repeated here.
[0072] Step 210: Compare the difference between the reference state information and the state information with a preset threshold to determine the accuracy of the state information.
[0073] In some embodiments, when the difference between the reference state information and the state information is greater than a preset threshold, it can be determined that the state information is inaccurate.
[0074] In some embodiments, when the difference between the reference status information and the status information is less than or equal to a preset threshold, it can be determined that the status information is accurate.
[0075] In some embodiments, steps 208 to 210 are optional.
[0076] In summary, the state determination method provided by the embodiments of this disclosure includes: acquiring a target image set with temporal information; extracting a first feature from the target image set using a convolutional neural network in the feature extraction module; projecting the first feature onto a preset coordinate system using a spatial transformation algorithm in the feature extraction module to obtain a feature vector; determining decoded features in the target image set using a decoding network based on the feature vector; determining a first detection parameter set of the target image using a multilayer perceptron based on the decoded features and temporal information; determining velocity change information and velocity magnitude information of the target based on the first detection parameter set; determining the state information of the target based on the velocity change information and velocity magnitude information; acquiring a reference parameter set of the target; determining reference state information of the target based on the reference parameter set; and comparing the difference between the reference state information and the state information with a preset threshold to determine the accuracy of the state information. This method transforms the first feature in the wooden plaque image set to the same coordinate system, which facilitates the determination of the target's velocity change information and velocity magnitude information, thereby improving the determination rate of state information. At the same time, by utilizing the velocity change information and velocity magnitude information, the accuracy of the target's state information is improved, and the inaccuracy of the determined state information due to sudden changes in the target's velocity can be avoided.
[0077] Figure 3 This is a flowchart of a state determination method according to an embodiment of the present disclosure. Figure 1 The illustrated embodiment is further explained as follows: Figure 3 The illustrated embodiment shows that the state determination method includes:
[0078] Step 301: Obtain the target image set with temporal information.
[0079] Step 302: Use the convolutional neural network in the feature extraction module to extract the first feature from the target image set.
[0080] Step 303: Using the spatial transformation algorithm in the feature extraction module, the first feature is projected onto a preset coordinate system to obtain the feature vector.
[0081] In this disclosure, the principles of steps 301-303 are the same as those of... Figure 2 Steps 201-203 in the illustrated embodiment are the same and can be referred to accordingly. Figure 2 The relevant descriptions will not be repeated here.
[0082] Step 304: Based on the feature vector and temporal information, a second detection parameter set for the target image set is determined using a multilayer perceptron.
[0083] In some embodiments, the second set of detection parameters includes at least: velocity change information and velocity magnitude information.
[0084] Among them, velocity change information is used to indicate the target's motion state, such as indicating whether the target is in a deceleration or acceleration state; velocity magnitude information is used to indicate the magnitude of the velocity.
[0085] In some embodiments, feature vectors can be extracted using a multilayer perceptual network of a pre-trained bird's-eye view model. Simultaneously, based on temporal information, the temporal changes of the feature vectors can be determined, thereby determining the second set of detection parameters for the target image.
[0086] In some embodiments, when determining velocity change information, feature vectors can also be extracted using a multilayer perceptual network of a pre-trained bird's-eye view model to obtain the velocity uncertainty of the target, thereby determining the accuracy of the velocity magnitude information.
[0087] The following loss function can be used for training the bird's-eye view model:
[0088]
[0089] Where L represents the loss function and σ represents the pre-set confidence level.
[0090] Step 305: Based on the feature vector and temporal information, use the decoding network to determine the tracking result information of the target image set.
[0091] In some embodiments, the decoding network is used to decode the feature vectors to combine the decoded features with temporal information to determine the target's trajectory (i.e., tracking result information).
[0092] In some embodiments, tracking result information is used to indicate the trajectory of the target.
[0093] Step 306: Based on the tracking results and the second detection parameter set, determine the velocity change information and velocity magnitude information of the target.
[0094] In some embodiments, the tracking results correspond velocity change information and velocity magnitude information to the target, thereby determining the target's velocity change information and velocity magnitude information.
[0095] Step 307: Determine the target's state information based on the velocity change information and velocity magnitude information.
[0096] Step 308: Obtain the reference parameter set of the target.
[0097] Step 309: Determine the reference state information based on the reference parameter set.
[0098] Step 310: Compare the difference between the reference state information and the state information with a preset threshold to determine the accuracy of the state information.
[0099] In this disclosure, the principles of steps 307-310 are the same as those of... Figure 2 Steps 207-210 in the illustrated embodiment are the same and can be referred to accordingly. Figure 2 The relevant descriptions will not be repeated here.
[0100] In summary, the state determination method provided by this disclosure includes: acquiring a target image set with temporal information; extracting a first feature from the target image set using a convolutional neural network in a feature extraction module; projecting the first feature onto a preset coordinate system using a spatial transformation algorithm in the feature extraction module to obtain a feature vector; determining a second detection parameter set for the target image set using a multilayer perceptron based on the feature vector and the temporal information; determining tracking result information of the target image set using a decoding network based on the feature vector and the temporal information; determining velocity change information and velocity magnitude information of the target based on the tracking result and the second detection parameter set; determining state information of the target based on the velocity change information and velocity magnitude information; acquiring a reference parameter set for the target; determining reference state information of the target based on the reference parameter set; and comparing the difference between the reference state and the state information with a preset threshold to determine the accuracy of the state information. This method transforms the first feature in the wooden plaque image set to the same coordinate system, which facilitates the determination of the target's velocity change information and velocity magnitude information, thereby improving the determination rate of state information. At the same time, by utilizing the velocity change information and velocity magnitude information, the accuracy of the target's state information is improved, and the inaccuracy of the determined state information due to sudden changes in the target's velocity can be avoided.
[0101] Therefore, this disclosure has the following beneficial effects:
[0102] 1. This method extracts features from the target image set, thereby determining the target's velocity change and magnitude information, accurately judging the target's state information, and improving the accuracy of the state information. It avoids situations where the determined state information is inaccurate due to sudden changes in the target's velocity.
[0103] Corresponding to the methods provided in the above embodiments, this disclosure also provides a state determination device. Since the device provided in this disclosure corresponds to the methods provided in the above embodiments, the implementation of the methods is also applicable to the device provided in this embodiment, and will not be described in detail in this embodiment.
[0104] Figure 4 This is a schematic diagram of the structure of a state determination device 400 according to an embodiment of this disclosure. Figure 4 As shown, the state determination device includes:
[0105] The acquisition module 410 is used to acquire a target image set with temporal information;
[0106] The extraction module 420 is used to extract feature vectors of the target image set using the feature extraction module of the pre-trained bird's-eye view model.
[0107] Processing module 430 is used to determine the velocity change information and velocity magnitude information of the target in the target image set based on the temporal information and feature vectors and using the decoding network of the pre-trained bird's-eye view model. The velocity change information is used to indicate the motion state of the target.
[0108] The determination module 440 is used to determine the target's state information based on the velocity change information and velocity magnitude information.
[0109] In some embodiments, the extraction module 420 is further configured to: extract a first feature from the target image set using a convolutional neural network in the feature extraction module; and project the first feature onto a preset coordinate system using a spatial transformation algorithm in the feature extraction module to obtain the feature vector.
[0110] In some embodiments, the processing module 430 is further configured to determine a first set of detection parameters for the target image based on temporal information and feature vectors, using a decoding network and a multilayer perceptual network of a pre-trained bird's-eye view model. The first set of detection parameters includes at least: target attribute information, velocity change information and tracking result information.
[0111] Alternatively, based on temporal information and feature vectors, the tracking result and second detection parameter set of the target in the target image are determined using the decoding network and multilayer perceptual network of the pre-trained bird's-eye view model. The second detection parameter set includes at least: velocity change information and velocity magnitude information. Based on the first detection parameter set, or based on the tracking result and the second detection parameter set, the velocity change information and velocity magnitude information of the target are determined.
[0112] In some embodiments, the processing module 430 is further configured to: determine the decoding features in the target image set using a decoding network based on the feature vectors; and determine the first detection parameter set of the target image using a multilayer perceptron based on the decoding features and temporal information.
[0113] In some embodiments, the processing module 430 is further configured to: determine a second detection parameter set of the target image set using a multilayer perceptron based on feature vectors and temporal information; and determine tracking result information of the target image set using a decoding network based on feature vectors and temporal information.
[0114] In some embodiments, the determining module 440 is further configured to: obtain a reference parameter set of the target, the reference parameter set including at least: reference velocity change information and reference velocity magnitude information; determine reference state information of the target based on the reference parameter set; and compare the difference between the reference state information and the state information with a preset threshold to determine the accuracy of the state information.
[0115] In summary, the state determination device includes: an acquisition module for acquiring a set of target images with temporal information; an extraction module for extracting feature vectors from the target image set using a feature extraction module of a pre-trained bird's-eye view model; a processing module for determining the velocity change information and velocity magnitude information of the target in the target image set based on the temporal information and feature vectors using a decoding network of the pre-trained bird's-eye view model; and a determination module for determining the state information of the target based on the velocity change information and velocity magnitude information. The device proposed in this disclosure improves the accuracy of the state information by determining the target's velocity change information and velocity magnitude information, thus avoiding inaccurate state information determined due to sudden changes in the target's velocity.
[0116] The methods and apparatus provided in the embodiments of this application have been described above. To implement the functions of the methods provided in the embodiments of this application, the electronic device may include a hardware structure and software modules, and may implement the above functions in the form of a hardware structure, software modules, or a hardware structure plus software modules. One of the above functions may be executed in the form of a hardware structure, software modules, or a hardware structure plus software modules.
[0117] Figure 5 This is a block diagram illustrating an electronic device 500 for implementing the above-described state determination method according to an exemplary embodiment.
[0118] For example, electronic device 500 can be a mobile phone, computer, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0119] Reference Figure 5 The electronic device 500 may include one or more of the following components: processing component 502, memory 504, power supply component 506, multimedia component 508, audio component 510, input / output (I / O) interface 512, sensor component 514, and communication component 516.
[0120] Processing component 502 typically controls the overall operation of electronic device 500, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 502 may include one or more processors 520 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 502 may include one or more modules to facilitate interaction between processing component 502 and other components. For example, processing component 502 may include a multimedia module to facilitate interaction between multimedia component 508 and processing component 502.
[0121] Memory 504 is configured to store various types of data to support the operation of electronic device 500. Examples of such data include instructions for any application or method operating on electronic device 500, contact data, phonebook data, messages, pictures, videos, etc. Memory 504 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0122] Power supply component 506 provides power to various components of electronic device 500. Power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 500.
[0123] Multimedia component 508 includes a screen that provides an output interface between electronic device 500 and user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 508 includes a front-facing camera and / or a rear-facing camera. When electronic device 500 is in an operating mode, such as a shooting mode or video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0124] Audio component 510 is configured to output and / or input audio signals. For example, audio component 510 includes a microphone (MIC) configured to receive external audio signals when electronic device 500 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 504 or transmitted via communication component 516. In some embodiments, audio component 510 also includes a speaker for outputting audio signals.
[0125] I / O interface 512 provides an interface between processing component 502 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0126] Sensor assembly 514 includes one or more sensors for providing state assessments of various aspects of electronic device 500. For example, sensor assembly 514 may detect the on / off state of electronic device 500, the relative positioning of components such as the display and keypad of electronic device 500, changes in position of electronic device 500 or a component of electronic device 500, the presence or absence of user contact with electronic device 500, orientation or acceleration / deceleration of electronic device 500, and temperature changes of electronic device 500. Sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 514 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0127] Communication component 516 is configured to facilitate wired or wireless communication between electronic device 500 and other devices. Electronic device 500 can access wireless networks according to communication standards, such as WiFi, 2G or 3G, 5G LTE, 5G NR (NewRadio), or combinations thereof. In one exemplary embodiment, communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 516 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented according to radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0128] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0129] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, which can be executed by a processor 520 of an electronic device 500 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0130] Embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the state determination method described in the above embodiments of this disclosure.
[0131] Embodiments of this disclosure also propose a vehicle including a status determination device as described above or an electronic device as described above.
[0132] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0133] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0134] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.
[0135] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a system according to a computer, a system including a processing module, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (control method), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic device, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0136] It should be understood that various parts of the embodiments of this disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0137] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0138] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a single processing module, or each unit can exist physically separately, or two or more units can be integrated into a single module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The aforementioned storage medium can be a read-only memory, a hard disk, or an optical disk, etc.
[0139] Although embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A method for determining a state, characterized in that, The method includes: Acquire a target image set with temporal information; The feature vector of the target image set is extracted using the feature extraction module of the pre-trained bird's-eye view model; Based on the time series information and the feature vector, the velocity change information and velocity magnitude information of the targets in the target image set are determined using the decoding network of the pre-trained bird's-eye view model; Based on the speed change information and the speed magnitude information, the state information of the target is determined.
2. The method according to claim 1, characterized in that, The feature extraction module utilizing the pre-trained model extracts the feature vectors of the target image set, including: The first feature in the target image set is extracted using the convolutional neural network in the feature extraction module. The first feature is projected onto a preset coordinate system using the spatial transformation algorithm in the feature extraction module to obtain the feature vector.
3. The method according to claim 1, characterized in that, The step of determining the velocity change information and velocity magnitude information of the targets in the target image set using the decoding network of the pre-trained bird's-eye view model based on the temporal information and the feature vector includes: Based on the time series information and the feature vector, the first detection parameter set of the target image is determined using the decoding network and multilayer perceptron of the pre-trained bird's-eye view model. The first detection parameter set includes at least: target attribute information, velocity change information and tracking result information. Alternatively, based on the temporal information and the feature vector, the tracking result of the target in the target image and the second detection parameter set are determined using the decoding network and multilayer perceptron of the pre-trained bird's-eye view model. The second detection parameter set includes at least: velocity change information and velocity magnitude information. Based on the first set of detection parameters, or based on the tracking result and the second set of detection parameters, determine the velocity change information and velocity magnitude information of the target.
4. The method according to claim 3, characterized in that, The step of determining the first detection parameter set of the target image based on the temporal information and the feature vector, using the decoding network and multilayer perceptron of the pre-trained bird's-eye view model, includes: Based on the feature vector, the decoding network is used to determine the decoding features in the target image set; Based on the decoding features and the temporal information, the first detection parameter set of the target image is determined using the multilayer perceptron.
5. The method according to claim 3, characterized in that, The step of determining the tracking result and the second detection parameter set of the target image based on the temporal information and the feature vector, using the decoding network and multilayer perceptron of the pre-trained bird's-eye view model, includes: Based on the feature vector and the temporal information, a second detection parameter set for the target image set is determined using a multilayer perceptron. Based on the feature vector and the temporal information, the tracking result information of the target image set is determined using the decoding network.
6. The method according to claim 1, characterized in that, The method further includes: Obtain a reference parameter set for the target, the reference parameter set including at least: reference velocity change information and reference velocity magnitude information; Based on the reference parameter set, determine the reference state information of the target; The accuracy of the status information is determined by comparing the difference between the reference status information and the status information with a preset threshold.
7. A state determination device, characterized in that, The device includes: The acquisition module is used to acquire a set of target images with temporal information; The extraction module is used to extract feature vectors from the target image set using the feature extraction module of the pre-trained bird's-eye view model; The processing module is used to determine the velocity change information and velocity magnitude information of the target in the target image set based on the time series information and the feature vector, using the decoding network of the pre-trained bird's-eye view model. The velocity change information is used to indicate the motion state of the target. The determination module is used to determine the state information of the target based on the speed change information and the speed magnitude information.
8. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
10. A vehicle, characterized in that, Includes the state determination device as described in claim 7 or the electronic device as described in claim 8.