Method and device for determining position, electronic equipment and storage medium
By using the historical time queue and the position information collected by the camera in extended real-life devices, the accuracy of the handle position information when the camera is in the camera's camera blind spot is solved, and more accurate and stable position prediction is achieved.
Patent Information
- Application Number
- CN202311459843.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-06
AI Technical Summary
When using extended reality devices, the position information of the handle is difficult to accurately track when it is within the camera's camera blind spot, resulting in errors and deviations in position prediction.
The position of the target object is determined by storing the position information of the most recent n historical moments of the target object in the historical time queue, and when the target object is outside the camera blind spot, the position information collected by the camera is used as the current position information.
It effectively prevents displacement information deviations caused by long-term integral sensor information, and improves the accuracy and stability of position prediction.
Smart Images

Figure CN119941840A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a method, device, electronic device and storage medium for determining a position. Background Art
[0002] When using an extended reality device, a handle is usually used to control the extended reality device. The position of the handle is important information for controlling the content displayed by the extended reality device. The position of the handle is usually tracked based on the integration of an inertial measurement unit (IMU). Summary of the invention
[0003] The present disclosure provides a method, device, electronic device and storage medium for determining a position.
[0004] The present disclosure adopts the following technical solutions.
[0005] In some embodiments, the present disclosure provides a method for determining a location, comprising:
[0006] Determine n historical moments of the target object in the historical time queue and whether the target moment is located in the blind spot of the camera; the historical time queue is used to store the historical position information of the target object at the latest n historical moments before the target moment, where n is a preset positive integer not less than 2;
[0007] In response to the target object being located outside the camera blind spot at n historical moments and at a target moment, the position information of the target object at the target moment collected by the camera is used as the position information of the target object at the target moment.
[0008] In some embodiments, the present disclosure provides a device for determining a position, comprising:
[0009] A control unit, used to determine n historical moments of the target object in a historical time queue and whether the target moment is located in a blind spot of the camera; the historical time queue is used to store historical position information of the target object at the latest n historical moments before the target moment, where n is a preset positive integer not less than 2;
[0010] The control unit is also used for: in response to the target object being located outside the camera blind spot at n historical moments and the target moment, using the position information of the target object at the target moment collected by the camera as the position information of the target object at the target moment.
[0011] In some embodiments, the present disclosure provides an electronic device, comprising: at least one memory and at least one processor;
[0012] The memory is used to store program codes, and the processor is used to call the program codes stored in the memory to execute the above method.
[0013] In some embodiments, the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and when the program code is executed by a processor, the processor is prompted to perform the above method.
[0014] The method for determining position provided by the embodiment of the present disclosure can prevent displacement information deviation. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0016] Figure 1 is a schematic diagram of using an extended reality device according to an embodiment of the present disclosure.
[0017] Figure 2 Detailed description of the invention The invention is a flowchart of a method for determining a position according to an embodiment of the present invention.
[0018] Figure 3 It is a schematic diagram of the historical time queue in different states of an embodiment of the present disclosure.
[0019] Figure 4 Detailed description of the invention The invention is a flowchart of a method for determining a position according to an embodiment of the present invention.
[0020] Figure 5 It is an operation flow chart of the position estimation model of the embodiment of the present disclosure.
[0021] Figure 6 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0023] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0024] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0026] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0027] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in sequence and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0030] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modification of “one” mentioned in the present disclosure is illustrative rather than restrictive, and those skilled in the art should understand that it should be understood as “one or more” unless otherwise clearly indicated in the context.
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] The solution provided by the embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0034] The extended reality (XR) technology in one or more embodiments of the present disclosure may be mixed reality technology, augmented reality technology, or virtual reality technology. Extended reality technology can combine reality and virtuality through computers to provide users with an extended reality space that allows human-computer interaction. In the extended reality space, users can use extended reality devices such as helmet-mounted displays (HMDs) to engage in social interaction, entertainment, learning, work, telecommuting, and UGC (User Generated Content) creation.
[0035] refer to Figure 1 Users can enter the extended reality space through extended reality devices such as head-mounted glasses, and control their own virtual characters (Avatar) in the extended reality space to engage in social interaction, entertainment, learning, remote work, etc. with virtual characters controlled by other users.
[0036] In one embodiment, in the extended reality space, the user can implement relevant interactive operations through a controller, and the controller may be a handle. For example, the user performs relevant operation controls by operating buttons on the handle.
[0037] The extended reality devices described in the embodiments of the present disclosure may include but are not limited to the following types:
[0038] The computer-side extended reality device uses the computer to perform related calculations and data output of the extended reality function, and the external computer-side extended reality device uses the data output by the computer to achieve the extended reality effect.
[0039] The mobile expansion device supports the setting of a mobile terminal (such as a smart phone) in various ways (such as a head-mounted display with a dedicated card slot). Through a wired or wireless connection with the mobile terminal, the mobile terminal performs relevant calculations of the extended reality function and outputs data to the mobile extended reality device, such as watching extended reality videos through the mobile terminal's APP.
[0040] The all-in-one extended reality device has a processor for performing related calculations for extended reality functions, and thus has independent extended reality input and output functions. It does not need to be connected to a computer or mobile terminal and has a high degree of freedom in use.
[0041] Of course, the form of the extended reality device is not limited to this, and it can be further miniaturized or enlarged as needed.
[0042] The extended reality device is equipped with a posture detection sensor (such as a nine-axis sensor) for detecting posture changes of the extended reality device in real time. If the user wears the extended reality device, then when the user's head posture changes, the real-time posture of the head will be transmitted to the processor to calculate the user's gaze point in the extended reality space environment, and the image within the user's gaze range (i.e., virtual field of view) in the three-dimensional model of the extended reality space environment is calculated based on the gaze point and displayed on the display screen, giving people an immersive experience as if they were watching in the real environment.
[0043] The controller tracking scheme can generally be divided into optical tracking (such as using a camera to shoot) and inertial measurement unit integration-based methods. Optical tracking can provide high-precision position results, but when the controller moves to the camera's blind spot, vision-based optical tracking cannot give accurate position predictions. In addition, extended reality devices usually run on embedded terminals, which face resource constraints.
[0044] In the related technology, sensors are used for integration, and the speed, acceleration and other information are integrated over time to obtain the motion trajectory of the handle in the camera blind spot. In this way, over time, the IMU will cause errors to accumulate due to temperature, noise and other reasons, causing the final predicted position to shift continuously in a certain direction, and the prediction result is completely unusable.
[0045] The time series is modeled based on a time series network (for example: Recurrent Neural Network, RNN, Long Short-Term Memory, LSTM, Gate Recurrent Unit, GRU), and the position prediction of the handle in the camera blind spot is transformed into a time series prediction problem. However, this method must maintain the hidden layer information of the neural network model. In order to maintain the hidden layer information of the neural network model, position prediction is also required when the handle is outside the camera blind spot. The neural network model must keep running both when the handle is in the camera blind spot and outside the camera blind spot.
[0046] In some embodiments of the present disclosure, a method for determining a position is proposed, such as Figure 1 As shown, including:
[0047] S11, determining the n historical moments of the target object in the historical time queue and whether the target moment is located in the blind spot of the camera.
[0048] In some embodiments, the execution end of the method proposed in the present disclosure can be any extended reality device in the present disclosure, and the extended reality device can include a head-mounted display device and a matching handle (or other accessories that need to be positioned, such as a handle, a leg ring, etc.), taking the handle as an example, the handle can be divided into a left-hand handle and a right-hand handle, and the target object can be the left-hand handle or the right-hand handle of the extended reality device. The method can be used to determine the position of the left-hand handle or the right-hand handle. The positions of the left and right handles usually do not have a strong correlation, so the positions of the left and right handles are determined separately. In some embodiments, a historical time queue is provided, and the historical time queue is used to store the historical position information of the target object at the latest n historical moments before the target moment, n is a preset positive integer not less than 2, and n can be, for example, a positive integer not less than 10, 20, 30, 40 or 50, for example, 60; the historical position information is the position information of the target object at the historical moment, and in some embodiments, the position information can be periodically obtained, and the latest n historical moments can be the historical moments of n cycles before the target moment, for example, 30 cycles can be set within 1 second, and n can be 60, then the historical time queue can be the historical moment within 2 seconds before the target moment. In some embodiments, the historical moment and the target moment may be the frame acquisition moment of the image captured by the camera. The camera may be located on a head-mounted display device of the extended reality device. The target object (handle) has a light spot for positioning. The camera captures the light spot on the handle to obtain the 6-DOF information (position and angle, etc.) of the target object. The camera may capture images periodically, for example, 30 frames of images per second, and the position and posture change information of the target object may be obtained when the camera captures the image. Therefore, each historical moment and the target moment are also the frame acquisition moments of the captured image. Each frame of the image captured by the camera corresponds to the position and posture change information and position information of a target object. In some embodiments, when the position information of the handle depends on the position information of the head-mounted display device, the historical time queue also contains the position information such as the position, speed, angle, acceleration and / or angular velocity information of the head-mounted display device at the historical moment. In some embodiments, the target object described in the present disclosure being located in the blind spot of the camera includes that the camera is unable to capture the target object. In other embodiments, the target object being located in the blind spot of the camera may also include: although the camera can capture the target object, the captured image cannot be used to determine the position information of the target object (for example, due to insufficient clarity, insufficient number of light spots, etc.). In this case, the target object is also considered to be located in the blind spot.
[0049] S12: In response to the target object being located outside the camera blind area at n historical moments and at the target moment, using the position information of the target object at the target moment collected by the camera as the position information of the target object at the target moment.
[0050] In some embodiments, when the target object is outside the camera blind spot at n historical moments and the target moment, it indicates that the target object is already in a position that can be photographed by the camera and has been in the position that can be photographed by the camera for a certain period of time. At this time, the position information of the target object can be determined by the image taken by the camera at the target moment. At this time, there is no need to determine the position information of the target object at the target moment based on the historical time queue and posture change information. For example, in some embodiments, a neural network model is used to determine the position information of the target object at the target moment based on the historical time queue and posture change information. However, when the target object is outside the camera blind spot at n historical moments and the target moment, the neural network model will not work.
[0051] The method proposed in some embodiments of the present disclosure, when the target object is located outside the camera blind spot at n historical moments and the target moment, uses the position information collected by the camera at the target moment as the position information of the target object at the target moment, thereby avoiding the displacement deviation problem caused by long-term integration of sensor information.
[0052] In some embodiments of the present disclosure, the method further includes: obtaining position and posture change information of the target object at the target time. In response to the target object being located in a camera blind spot at at least one of the n historical moments and the target moment, determining the position information of the target object at the target moment according to the historical time queue and the position and posture change information.
[0053] In some embodiments, obtaining the posture change information may be performed before or after step S11. Taking the method used in an extended reality device as an example, both the handle and the head-mounted display device may have posture sensors, such as one or more of an acceleration sensor and an angular velocity sensor. Sensor information (such as one or more of acceleration information and angular velocity information) obtained from the posture sensor may be used as the posture change information, or feature extraction (such as integration, where the integration time is from the target moment to the last historical moment) may be performed on the sensor information to obtain the posture change information. Specifically, the posture change information may include the speed, angle, acceleration and / or angular velocity of the target object. In some embodiments, in the extended reality device, the head-mounted display device may be used as the origin of the space, and the position of the handle may be represented as a position dependent on the head-mounted display device. At this time, because the posture change information of the handle depends on the head-mounted display device, the posture change information of the handle needs to contain posture data such as the position, speed, angle, acceleration and / or angular velocity of the head-mounted display device. In some embodiments, the camera periodically captures images, and the moment of capturing images is the frame acquisition moment. The target moment may be the frame acquisition moment of the last captured image, that is, the moment of the last camera capturing an image from the current moment (including the current moment). Therefore, the posture change information is: the posture change information of the target object during the period from the last captured image to the last captured image.
[0054] In some embodiments, when the target object is located in a camera blind spot at a historical moment or a target moment, the position information of the target object at the target moment is predicted using n historical position information and the position change information at the target moment in the historical time queue. For example, the n historical position information and the position change information at the target moment can be input into a neural network model to output the position information. The output position information at the target moment can be the relative position of the target moment relative to the last historical moment before the target moment, because it is difficult to determine the absolute position in the camera blind spot, and it is simpler to determine the relative position relative to the last historical moment before the target moment, which is more suitable for devices with less computing resources. Kalman filtering can be performed on the determined position information to improve accuracy.
[0055] In some embodiments of the present disclosure, a historical time sequence is set to save the location information of the target object at n historical moments. When the target moment and the historical moment are both outside the camera blind spot, the location information of the target object is determined using the optical tracking of the camera, without using the neural network model to calculate the location information of the target object. When the target object is located in the camera blind spot at the target moment or n historical moments, the location information at the target moment is predicted using the historical time sequence.
[0056] In some embodiments of the present disclosure, after determining that the target object is located in the camera blind spot at least one of the n historical moments and the target moment, a position estimation model is used to determine the position information of the target object at the target moment according to the historical time queue and the posture change information; after determining that the target object is outside the camera blind spot at n historical moments and the target moment, the position estimation model is maintained or set to a non-working state.
[0057] In some embodiments, traditional LSTM, GRU and other timing networks need to use hidden layer data to calculate the position information of the target object in the camera blind spot. Therefore, no matter where the target object is located, the timing network needs to always calculate the position information of the target object. In the present disclosure, there is no need to always use a position prediction model (such as a neural network model) to predict the position information of the target object. When the target object is not in the camera blind spot at both the target moment and the historical moment, there is no need to use the position prediction model, the position prediction model does not work, and the historical time queue and posture change information do not need to be input into the position prediction model. The position prediction model is only allowed to work when the target object is in the camera blind spot. This can reduce resource consumption. In the present disclosure, when the target object is in the camera blind spot, the predicted position information refers to the historical position information of n historical moments, rather than relying solely on the posture change information, thereby preventing position information offset.
[0058] In some embodiments, after the location information of the target object is determined, the historical time sequence can be updated, and the location information of the target object at the target moment can be placed into the historical time sequence according to the first-in-first-out principle, and the location information of the target object at the earliest historical moment in the historical time queue can be removed.
[0059] In order to better illustrate the embodiments of the present disclosure, a specific embodiment is proposed below. In this embodiment, the execution end of the method is an extended reality device. First, the handle of the extended reality device and the posture sensor information of the head-mounted display device at the current moment are obtained, and then feature extraction is performed to obtain posture change information (the speed and angular velocity of the left and right handles, the position and angular velocity of the head-mounted display, etc.). The head-mounted camera shoots 30 frames of images in 1 second. The historical time queue includes: the position information of the handle corresponding to the 60 historical moments of the camera shooting images in the previous 2 seconds, that is, the position information corresponding to the 60 frames of images in the previous 2 seconds. Taking the target moment as the current moment as an example, if the handle is located in the area that the camera of the head-mounted display device can shoot at the current moment and the 60 historical moments within the previous 2 seconds (that is, outside the camera blind area, such as Figure 3(a), the blank box indicates that the handle is outside the camera blind spot at the moment the camera captures the image, and the box pointed by the arrow is the position information that needs to be determined corresponding to the moment when the most recent frame of image was captured), then the image of the handle captured by the camera is used to determine the position information of the handle, without the need to use a neural network model for prediction, and then the historical time queue is updated according to the first-in-first-out principle. Specifically, there is a positioning light spot on the handle, and the position information of the handle is determined by the positioning light spot. If at the current moment and in the 60 historical moments within the previous 2 seconds, the handle is located in an area that the head-mounted display device camera cannot capture (i.e., the camera blind spot, such as Figure 3 (b) and 3(c) The black box indicates that the handle is in the camera blind spot when the camera takes the image. The posture change information and the historical time queue are input into the neural network model, and the predicted handle position information is output. In order to improve the prediction accuracy of the model, it can be optimized to output the result optimized by Kalman filtering and update the historical time queue. The output position information is the relative position of the current moment relative to the most recent historical moment. In the disclosed embodiment, the camera collects images at a periodicity of 30 frames per second, and predicts the position information corresponding to the most recent frame of the image through the position information corresponding to the historical 60 frames of images, thereby avoiding the predicted trajectory deviation problem caused by being in the camera blind spot for a long time, and the neural network model is not always in operation. The neural network model does not need to work outside the camera blind spot, thereby reducing resource usage and being more suitable for use on embedded devices.
[0060] In some embodiments of the present disclosure, after determining that the target object is located in the camera blind spot at at least one of n historical moments and the target moment, and before determining the position information of the target object at the target moment based on the historical time queue and the posture change information, it also includes: in response to the target object entering the camera blind spot for the first time and the number of historical position information does not reach n, ending the method.
[0061] In some embodiments, Figure 4As shown, after obtaining the posture change information, it can be determined whether the target object is located in the camera blind spot at the historical moment and the target moment. If it is not located in the camera blind spot, the position information of the target object at the target moment will be determined based on the captured image, and then the historical time series will be updated with the determined position information. If the target object is not located in the camera blind spot at the historical moment or the target moment, it will be determined whether the target object is the first time that the device executing the method (such as an extended reality device) enters the camera blind spot after this startup. If it is not the first time to enter the camera blind spot, it means that enough position information of the target object has been accumulated, and the position information of the target object at the target moment is determined based on the historical time queue and the posture change information. If it is the first time to enter the camera blind spot, it can be determined whether the number of historical position information stored in the historical time queue reaches n; if it reaches n, the position information of the target object at the target moment is determined based on the historical time queue and the posture change information, or, if it does not reach n, the method ends. In some embodiments, such as Figure 4 As shown, if the target object does not enter the camera blind spot for the first time after the device is turned on this time, it may not have accumulated enough historical location information. At this time, it is first determined whether there is n historical location information in the historical time queue, because it may not have acquired n location information after the device is turned on this time, resulting in the historical time queue not storing n historical location information. If n historical location information has not been stored, there is not enough information to ensure the accuracy of the predicted location information, and the method in the embodiment of the present disclosure is not executed at this time. If there is enough historical location information, the location information of the target object at the target time is predicted using, for example, a neural network model.
[0062] In some embodiments of the present disclosure, determining the position information of the target object at the target moment according to the historical time queue and the position change information includes: determining the relative position of the target object at the target moment relative to the previous historical moment according to the historical time queue and the position change information, and using the relative position as the position information of the target object at the target moment. Alternatively, determining the relative position of the target object at the target moment relative to the previous historical moment according to the historical time queue and the position change information, and determining the position information of the target object at the target moment according to the relative position and the historical position information of the target object at the previous historical moment
[0063] In some embodiments, when the target object is located in a camera blind spot, it is relatively difficult to determine the absolute position of the target object, and it will consume more resources and time, and the timeliness is poor. Therefore, the output position information is the relative position of the target object relative to the previous historical moment before the target moment and the closest to the target moment in time. The previous historical moment refers to the previous historical moment before the target moment. In other embodiments, the determined relative position can be added to the historical position information of the previous historical moment to obtain the position information of the target moment, and the obtained position information can be an absolute position (for example, the absolute position is expressed by spatial coordinates).
[0064] In some embodiments of the present disclosure, determining the position information of the target object at the target time according to the historical time queue and the posture change information includes:
[0065] Inputting the n historical position information in the historical time queue and the posture change information into a position estimation model to obtain an initial predicted position;
[0066] Perform at least one iteration on the initial predicted position; determine the probability that the stage position predicted by the position estimation model at each iteration stage falls into each error plane, wherein each error plane distribution has a corresponding error expectation;
[0067] According to the probability that the stage position predicted in each iterative stage falls into each error plane and the error expectation corresponding to the error plane, the initial predicted position is corrected to obtain a corrected predicted position;
[0068] The corrected predicted position is used as the position information of the target object at the target time.
[0069] In some embodiments, Figure 5 As shown in the figure, when determining the position information of the target object at the target time, the historical time queue and posture change information are input into the position estimation model to obtain the initial predicted position P coarse, which is a rough position, which is used as the starting point of the subsequent iteration stage. It is different from the real position information, so it is iterated once or multiple times. The input of the first iteration stage is the initial predicted position, and the input of the subsequent iteration stage is the stage position output by the previous iteration stage. The iteration stage is used to perform a displacement (residual expectation) on the input position, and the sum of the displacement and the position input in the iteration stage is used as the stage position output by the iteration stage. The output of the iteration stage is the position information of the target object predicted by the iteration stage at the target time, that is, the stage position. The stage position is erroneous. The calculation of the error of the stage position is converted into the calculation of the error plane in the present disclosure. The error plane is a two-dimensional or three-dimensional plane of the prediction error. The error plane corresponds to the error expectation. After determining the error plane in which the stage position falls, the error expectation corresponding to the error plane in which it falls can also be used to determine the residual expectation of the stage position. Therefore, in the embodiment of the present disclosure, if a preset number of N iteration stages are performed, there are residual expectations of N iteration stages. These residual expectations are used to correct the initial predicted position. The corrected predicted position obtained is the position information of the target object at the target time determined by the position estimation model.
[0070] In some embodiments of the present disclosure, the initial predicted position is corrected according to the probability that the stage position predicted in each iterative stage falls into each error plane and the error expectation corresponding to the error plane, including: correcting the initial predicted position using the following formula 1:
[0071] Formula 1
[0072] In formula 1, P true To correct the predicted position, P coarse is the initial predicted position, N is the number of iteration stages, λ i is the weight of the i-th iteration stage, P is the number of error planes, x j is the expected error of the jth error plane, P j is the probability that the stage position of the i-th iteration stage falls into the j-th error plane.
[0073] In some embodiments, as shown in Formula 1, at any iteration stage, the phase position output by the iteration stage falls into each error plane P j The probability of the corresponding error expectation xj is then multiplied by the corresponding probability and the summed up to obtain the residual expectation of the iteration stage. Each iteration stage has a corresponding weight λ i , the residual expectations of each iterative stage are weighted summed up to obtain the deviation of the entire position estimation model. The initial predicted position is corrected with this deviation to obtain the corrected predicted position.
[0074] In some embodiments of the present disclosure, before step S11, a step of pre-training the position estimation model is also included. Pre-training is the process of training the model using a large amount of data with known results so that the deviation between the position predicted by the model and the actual position is minimized. Usually, this process usually requires multiple iteration stages, and the model parameters are continuously adjusted to minimize the loss function of the position estimation model.
[0075] In some embodiments of the present disclosure, the position estimation model is pre-trained in the following manner:
[0076] Input the training data into the position estimation model to obtain the initial estimated position;
[0077] Perform at least one pre-training iteration phase on the initial predicted position, wherein, in the pre-training iteration phase, an error plane prediction model is determined according to the residual of the real position of the training data and the input position of the pre-training iteration phase, the input of the error plane prediction model is the position, and the output of the error plane prediction model is the error expectation corresponding to the input position; the residual expectation of the position estimation model in the pre-training iteration phase is determined according to the error plane prediction model and the probability distribution function of the pre-training iteration phase;
[0078] Determine the loss function of the position estimation model according to the residual expectation of each iterative stage of pre-training;
[0079] The parameters of the location estimation model are adjusted to reduce the loss function of the location estimation model.
[0080] In some embodiments, during the pre-training process, the training data may include: historical position information and position change information acquired in advance, and the real position corresponding to the training data is known. In the pre-training stage, the following steps 1 to 5 are specifically performed.
[0081] Step 1: Input the training data into the position estimation model to obtain the initial estimated position, which is the same process as obtaining the initial predicted position.
[0082] Step 2: Use the difference between the true position and the input position of the pre-training iteration phase as the residual, and use the distribution of the residual to determine the error plane prediction model, which can tell what the error expectation is at different positions.
[0083] Specifically, the iterative stage of pre-training is to repeat steps 2 and 3. The input position of the first iterative stage of pre-training is the initial estimated position. A residual expectation is determined in the iterative stage of pre-training. The sum of the input position of an iterative stage of pre-training plus the residual expectation of the iterative stage will be used as the input position of the next iterative stage of pre-training. Through this step, the calculation problem of the residual expectation of the iterative stage is converted into the prediction problem of the error plane.
[0084] Step 3: Determine the residual expectation of the position estimation model in the pre-training iterative stage according to the error plane prediction model and the probability distribution function of the pre-training iterative stage.
[0085] Specifically, the input of the probability distribution function is the position, and the output of the probability distribution function is the probability of being at the input position in the pre-training iteration phase. Specifically, the residual expectation in the pre-training iteration phase is calculated using the following formula 2:
[0086] Formula 2 E(Res)=∫[E(x)*P(x)]dx
[0087] In Formula 2, E(Res) is the residual expectation of the calculated pre-training iteration phase, x is the position, E(x) is the error expectation of the error plane where position x is located, the asterisk in Formula 2 represents multiplication, P(x) is the probability of being at position x in the calculated pre-training iteration phase, and the integration interval is the entire space. Formula 2 means that for each position x, the error expectation E(x) of the error plane is multiplied by the probability distribution P(x) at that position, and then the entire space is integrated. This is equivalent to multiplying the error expectation of each position by its corresponding probability, and adding up the products at all positions, and the result is the residual expectation of the iteration phase.
[0088] Step 4: Calculate the loss function.
[0089] Specifically, the loss function of the position estimation model is determined according to the residual expectation of each iterative stage of pre-training, including:
[0090] The loss function of the position estimation model is calculated using the following formula 3:
[0091] Formula 3
[0092] In formula 3, L all represents the loss function of the position estimation model, N represents the number of pre-training iterations, that is, the number of times steps 2 and 3 are repeated. N can be set to 10, and λ i Represents the weight of the residual expectation of the i-th iteration stage of pre-training, E(Res) i Represents the residual expectation of the i-th iteration stage of pre-training. By multiplying the residual expectation of each iteration stage of pre-training by the corresponding weight and summing them up, the importance and contribution of different iteration stages can be comprehensively considered to construct a comprehensive loss function. The weight parameter λ iThe relative importance of residual expectation for each iteration stage can be controlled and can be adjusted according to specific needs. The weights and / or number of iteration stages determined when pre-training the position estimation model can be the same as the weights and / or number of iteration stages when using the position estimation model (i.e., determining the position information of the target object at the target time). The weight of the previous iteration stage is greater than the weight of the later iteration stage. In some embodiments, λ i =0.5×λ (i-1) It means that the weights decrease successively, which is consistent with the learning from coarse to fine.
[0093] Step 5: Adjust the parameters of the location estimation model to minimize the loss function.
[0094] Specifically, after adjusting the position estimation model, repeating steps 1 to 4 to determine the adjusted loss function, and adjusting the parameters to minimize the loss function simultaneously, can optimize the model parameters so that the model can achieve good performance in multiple stages and comprehensively consider the task requirements of each stage. This can improve the overall performance and generalization ability of the model and adapt to the training process of different levels and goals.
[0095] In some embodiments of the present disclosure, the proposed method can be conveniently deployed on embedded end-side hardware. In order to reduce the occupation of memory and computing resources, the time series queues with fixed lengths of limited lengths in and outside the blind spots are optimized.
[0096] The method proposed in this disclosure allows the position estimation model to make continuous position predictions at any time, and has a high tolerance for frame loss. Traditional handle blind spot prediction methods usually make position predictions at fixed time intervals, while this solution makes predictions instantly when needed, achieving more real-time and continuous prediction capabilities. This highly flexible reasoning mechanism enables the network to better adapt to different operating scenarios and input changes, providing more accurate blind spot prediction results.
[0097] The method proposed in the present disclosure does not require the position estimation model to perform position prediction when the target object is outside the camera blind spot for a long time, thereby effectively reducing power consumption. In traditional solutions, even if the handle is always outside the camera blind spot, the neural network will continue to perform position prediction, wasting a lot of computing resources and energy. The present disclosure intelligently determines whether the target is in the camera blind spot and only performs prediction when necessary, thereby minimizing power consumption and extending the battery life of embedded devices.
[0098] Traditional blind spot prediction schemes require the storage of intermediate data of the hidden layer. In the method proposed in the present invention, when the target object is outside the camera blind spot at the historical moment and the current moment, the position estimation model does not work, and there is no need to store the position information predicted by the position estimation model at this time, thereby reducing the amount of data storage. In the method proposed in the present invention, the historical time queue and posture change information can be stored in the memory to prevent the impact on data accuracy (when the neural network model is quantized using heterogeneous hardware such as an embedded neural network processor, floating-point data will be converted into fixed-point data, resulting in a decrease in accuracy).
[0099] The method proposed in the embodiments of the present disclosure can be combined with other algorithms. In particular, the method can be combined with time convolution, time series network, extended Kalman filter, discrete Kalman filter, unscented Kalman filter, Lagrangian Kalman filter and other algorithms.
[0100] The present disclosure also provides a device for determining a position, comprising:
[0101] A control unit is used to determine the n historical moments of the target object in the historical time queue and whether the target moment is located in the blind spot of the camera; the historical time queue is used to store the historical position information of the target object at the latest n historical moments before the target moment, where n is a preset positive integer not less than 2;
[0102] The control unit is also used to: in response to the target object being located outside the camera blind spot at n historical moments and the target moment, use the position information of the target object at the target moment collected by the camera as the position information of the target object at the target moment.
[0103] In some embodiments, the method further includes: an acquisition unit configured to acquire position and posture change information of the target object at a target moment.
[0104] The control unit is also used to: in response to the target object being located in the camera blind spot at at least one of the n historical moments and the target moment, determine the position information of the target object at the target moment according to the historical time queue and the posture change information.
[0105] After determining that the target object is located in the camera blind spot at at least one of n historical moments and the target moment, and before determining the position information of the target object at the target moment based on the historical time queue and the posture change information, the control unit is also used to: end the method in response to the target object entering the camera blind spot for the first time and the number of historical position information does not reach n.
[0106] In some embodiments, determining the position information of the target object at the target time according to the historical time queue and the posture change information includes:
[0107] Determine the relative position of the target object at the target moment relative to the previous historical moment according to the historical time queue and the posture change information, and use the relative position as the position information of the target object at the target moment; or
[0108] According to the historical time queue and the posture change information, the relative position of the target object at the target moment relative to the previous historical moment is determined, and the position information of the target object at the target moment is determined according to the relative position and the historical position information of the target object at the previous historical moment.
[0109] In some embodiments, after determining that the target object is located in the camera blind spot at at least one of the n historical moments and the target moment, the control unit is used to use a position estimation model to determine the position information of the target object at the target moment according to the historical time queue and the posture change information;
[0110] After determining that the target object is located outside the camera blind area at n historical moments and the target moment, the control unit is further used to: maintain or set the position estimation model to be in a non-working state.
[0111] In some embodiments, determining the position information of the target object at the target time according to the historical time queue and the posture change information includes:
[0112] Inputting the n historical position information in the historical time queue and the posture change information into a position estimation model to obtain an initial predicted position;
[0113] performing at least one iteration on the initial predicted position;
[0114] Determining the probability that the stage position predicted by the position estimation model at each iteration stage falls into each error plane, wherein each error plane distribution has a corresponding error expectation;
[0115] According to the probability that the stage position predicted in each iterative stage falls into each error plane and the error expectation corresponding to the error plane, the initial predicted position is corrected to obtain a corrected predicted position;
[0116] The corrected predicted position is used as the position information of the target object at the target time.
[0117] In some embodiments, the initial predicted position is corrected according to the probability that the stage position predicted in each iterative stage falls into each error plane and the error expectation corresponding to the error plane, including: correcting the initial predicted position using the following formula 1:
[0118] Formula 1
[0119] In formula 1, P true To correct the predicted position, P coarse is the initial predicted position, N is the number of iteration stages, λ i is the weight of the i-th iteration stage, P is the number of error planes, x j is the expected error of the jth error plane, P j is the probability that the stage position of the i-th iteration stage falls into the j-th error plane.
[0120] In some embodiments, the position estimation model is pre-trained in the following manner:
[0121] Input the training data into the position estimation model to obtain the initial estimated position;
[0122] Perform at least one pre-training iteration phase on the initial predicted position, wherein, in the pre-training iteration phase, an error plane prediction model is determined based on the residual between the real position of the training data and the input position of the pre-training iteration phase, the input of the error plane prediction model is the position, and the output of the error plane prediction model is the error expectation corresponding to the input position; the residual expectation of the position estimation model in the pre-training iteration phase is determined based on the error plane prediction model and the probability distribution function of the pre-training iteration phase, the input of the probability distribution function is the position, and the output of the probability distribution function is the probability of being located at the input position in the pre-training iteration phase;
[0123] Determine the loss function of the position estimation model according to the residual expectation of each iterative stage of pre-training;
[0124] The parameters of the location estimation model are adjusted to reduce the loss function of the location estimation model.
[0125] In some embodiments, determining the residual expectation of the position estimation model in the pre-training iteration phase according to the error plane prediction model and the probability distribution function of the pre-training iteration phase includes:
[0126] The residual expectation of the iterative stage of pre-training is calculated using the following formula 2:
[0127] Formula 2 E(Res)=∫[E(x)*P(x)]dx
[0128] In Formula 2, E(Res) is the residual expectation of the calculated pre-training iteration stage, x is the position, E(x) is the error expectation of the error plane where the position x is located, P(x) is the probability of being at the position x in the calculated pre-training iteration stage, and the integration interval is the entire space.
[0129] In some embodiments, the loss function of the position estimation model is determined according to the residual expectation of each iteration stage of pre-training, including:
[0130] The loss function of the position estimation model is calculated using the following formula 3:
[0131] Formula 3
[0132] In formula 3, L all represents the loss function of the location estimation model, N represents the number of pre-training iterations, and λ i Represents the weight of the residual expectation of the i-th iteration stage of pre-training, E(Res) i represents the residual expectation of the i-th iteration stage of pre-training.
[0133] In some embodiments, the weight of an earlier iteration stage during the iteration is greater than the weight of a later iteration stage during the iteration.
[0134] For the embodiments of the device, since they basically correspond to the method embodiments, the relevant parts refer to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separation modules may or may not be separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present embodiment. Those of ordinary skill in the art can understand and implement without paying creative work.
[0135] The method and device of the present disclosure are described above based on the embodiments and application examples. In addition, the present disclosure also provides an electronic device and a computer-readable storage medium, which are described below.
[0136] Reference below Figure 6 , which shows a schematic diagram of the structure of an electronic device (such as a terminal device or a server) 800 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in the figure is only an example and should not bring any limitation to the functions and scope of use of the embodiment of the present disclosure.
[0137] The electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0138] Typically, the following devices may be connected to the I / O interface 805: input devices 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 808 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although the electronic device 800 with various devices is shown in the figure, it should be understood that it is not required to implement or have all the devices shown. More or fewer devices may be implemented or have alternatively.
[0139] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0140] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0141] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0142] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0143] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method of the present disclosure.
[0144] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0145] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0146] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, constitute a limitation on the unit itself.
[0147] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0148] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0149] According to one or more embodiments of the present disclosure, a method for determining a position is provided, including:
[0150] Determine n historical moments of the target object in the historical time queue and whether the target moment is located in the blind spot of the camera; the historical time queue is used to store the historical position information of the target object at the latest n historical moments before the target moment, where n is a preset positive integer not less than 2;
[0151] In response to the target object being located outside the camera blind spot at n historical moments and at a target moment, the position information of the target object at the target moment collected by the camera is used as the position information of the target object at the target moment.
[0152] According to one or more embodiments of the present disclosure, a method for determining a position is provided, further comprising:
[0153] Acquire the position change information of the target object at the target moment; in response to the target object being located in the camera blind spot at at least one of the n historical moments and the target moment, determine the position information of the target object at the target moment according to the historical time queue and the position change information.
[0154] According to one or more embodiments of the present disclosure, a method for determining a position is provided, which, after determining that the target object is located in the camera blind spot at at least one of n historical moments and the target moment, and before determining the position information of the target object at the target moment according to the historical time queue and the posture change information, further includes:
[0155] In response to the target object entering the camera blind spot for the first time and the number of the historical position information does not reach n, the method is terminated.
[0156] According to one or more embodiments of the present disclosure, a method for determining a position is provided, wherein the position information of the target object at the target time is determined according to the historical time queue and the posture change information, including:
[0157] Determine the relative position of the target object at the target moment relative to the previous historical moment according to the historical time queue and the posture change information, and use the relative position as the position information of the target object at the target moment; or
[0158] According to the historical time queue and the posture change information, the relative position of the target object at the target moment relative to the previous historical moment is determined, and the position information of the target object at the target moment is determined according to the relative position and the historical position information of the target object at the previous historical moment.
[0159] According to one or more embodiments of the present disclosure, a method for determining a position is provided, wherein after determining that the target object is located in the camera blind spot at at least one of the n historical moments and the target moment, a position estimation model is used to determine the position information of the target object at the target moment according to the historical time queue and the posture change information;
[0160] After determining that the target object is located outside the camera blind area at n historical moments and the target moment, the position estimation model is maintained or set to be in a non-working state.
[0161] According to one or more embodiments of the present disclosure, a method for determining a position is provided, wherein the position information of the target object at a target time is determined according to the historical time queue and the posture change information, including:
[0162] Inputting the n historical position information in the historical time queue and the posture change information into a position estimation model to obtain an initial predicted position;
[0163] performing at least one iteration on the initial predicted position;
[0164] Determining the probability that the stage position predicted by the position estimation model at each iteration stage falls into each error plane, wherein each error plane distribution has a corresponding error expectation;
[0165] According to the probability that the stage position predicted in each iterative stage falls into each error plane and the error expectation corresponding to the error plane, the initial predicted position is corrected to obtain a corrected predicted position;
[0166] The corrected predicted position is used as the position information of the target object at the target time.
[0167] According to one or more embodiments of the present disclosure, a method for determining a position is provided, wherein an initial predicted position is corrected according to the probability that the stage position predicted in each iterative stage falls into each error plane and the error expectation corresponding to the error plane, including: correcting the initial predicted position using the following formula 1:
[0168] Formula 1
[0169] In formula 1, P true To correct the predicted position, P coarse is the initial predicted position, N is the number of iteration stages, λ i is the weight of the i-th iteration stage, P is the number of error planes, x j is the expected error of the jth error plane, P j is the probability that the stage position of the i-th iteration stage falls into the j-th error plane.
[0170] According to one or more embodiments of the present disclosure, a method for determining a position is provided, wherein the position estimation model is pre-trained in the following manner:
[0171] Input the training data into the position estimation model to obtain the initial estimated position;
[0172] Perform at least one pre-training iteration phase on the initial predicted position, wherein, in the pre-training iteration phase, an error plane prediction model is determined based on the residual between the real position of the training data and the input position of the pre-training iteration phase, the input of the error plane prediction model is the position, and the output of the error plane prediction model is the error expectation corresponding to the input position; the residual expectation of the position estimation model in the pre-training iteration phase is determined based on the error plane prediction model and the probability distribution function of the pre-training iteration phase, the input of the probability distribution function is the position, and the output of the probability distribution function is the probability of being located at the input position in the pre-training iteration phase;
[0173] Determine the loss function of the position estimation model according to the residual expectation of each iterative stage of pre-training;
[0174] The parameters of the location estimation model are adjusted to reduce the loss function of the location estimation model.
[0175] According to one or more embodiments of the present disclosure, a method for determining a position is provided, wherein a residual expectation of a position estimation model in a pre-training iterative stage is determined according to an error plane prediction model and a probability distribution function in a pre-training iterative stage, comprising:
[0176] The residual expectation of the iterative stage of pre-training is calculated using the following formula 2:
[0177] Formula 2 E(Res)=∫[E(x)*P(x)]dx
[0178] In Formula 2, E(Res) is the residual expectation of the calculated pre-training iteration stage, x is the position, E(x) is the error expectation of the error plane where the position x is located, P(x) is the probability of being at the position x in the calculated pre-training iteration stage, and the integration interval is the entire space.
[0179] According to one or more embodiments of the present disclosure, a method for determining a position is provided, wherein a loss function of a position estimation model is determined according to residual expectations of each iterative stage of pre-training, including:
[0180] The loss function of the position estimation model is calculated using the following formula 3:
[0181] Formula 3
[0182] In formula 3, L all represents the loss function of the location estimation model, N represents the number of pre-training iterations, and λ i Represents the weight of the residual expectation of the i-th iteration stage of pre-training, E(Res) i represents the residual expectation of the i-th iteration stage of pre-training.
[0183] According to one or more embodiments of the present disclosure, a method for determining a position is provided, wherein a weight in a previous iteration stage during iteration is greater than a weight in a later iteration stage during iteration.
[0184] According to one or more embodiments of the present disclosure, there is provided a device for determining a position, including:
[0185] A control unit, used to determine n historical moments of the target object in a historical time queue and whether the target moment is located in a blind spot of the camera; the historical time queue is used to store historical position information of the target object at the latest n historical moments before the target moment, where n is a preset positive integer not less than 2;
[0186] The control unit is also used for: in response to the target object being located outside the camera blind spot at n historical moments and the target moment, using the position information of the target object at the target moment collected by the camera as the position information of the target object at the target moment.
[0187] According to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one memory and at least one processor;
[0188] The at least one memory is used to store program codes, and the at least one processor is used to call the program codes stored in the at least one memory to execute any one of the above methods.
[0189] According to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium is used to store program code, and when the program code is executed by a processor, the processor is prompted to execute the above method.
[0190] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0191] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0192] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A method for determining a position, characterized in that: include: Determine n historical moments of the target object in the historical time queue and whether the target moment is located in the blind spot of the camera; the historical time queue is used to store the historical position information of the target object at the latest n historical moments before the target moment, where n is a preset positive integer not less than 2; In response to the target object being located outside the camera blind spot at n historical moments and at a target moment, the position information of the target object at the target moment collected by the camera is used as the position information of the target object at the target moment.
2. The method according to claim 1, characterized in that Also includes: Obtain the position change information of the target object at the target time; In response to the target object being located in the camera blind spot at at least one of the n historical moments and the target moment, the position information of the target object at the target moment is determined according to the historical time queue and the posture change information.
3. The method according to claim 2, characterized in that After determining that the target object is located in the camera blind spot at at least one of the n historical moments and the target moment, and before determining the position information of the target object at the target moment according to the historical time queue and the posture change information, the method further includes: In response to the target object entering the camera blind spot for the first time and the number of the historical position information does not reach n, the method is terminated.
4. The method according to claim 2, characterized in that: Determining the position information of the target object at the target time according to the historical time queue and the posture change information includes: Determine the relative position of the target object at the target moment relative to the previous historical moment according to the historical time queue and the posture change information, and use the relative position as the position information of the target object at the target moment; or According to the historical time queue and the posture change information, the relative position of the target object at the target moment relative to the previous historical moment is determined, and the position information of the target object at the target moment is determined according to the relative position and the historical position information of the target object at the previous historical moment.
5. The method according to claim 2, characterized in that After determining that the target object is located in the camera blind spot at at least one of the n historical moments and the target moment, using a position estimation model to determine the position information of the target object at the target moment according to the historical time queue and the posture change information; After determining that the target object is located outside the camera blind area at n historical moments and the target moment, the position estimation model is maintained or set to be in a non-working state.
6. The method according to claim 2, characterized in that Determining the position information of the target object at a target time according to the historical time queue and the posture change information includes: Inputting the n historical position information in the historical time queue and the posture change information into a position estimation model to obtain an initial predicted position; performing at least one iteration on the initial predicted position; Determining the probability that the stage position predicted by the position estimation model at each iteration stage falls into each error plane, wherein each error plane distribution has a corresponding error expectation; According to the probability that the stage position predicted in each iterative stage falls into each error plane and the error expectation corresponding to the error plane, the initial predicted position is corrected to obtain a corrected predicted position; The corrected predicted position is used as the position information of the target object at the target time.
7. The method according to claim 6, characterized in that According to the probability that the stage position predicted in each iterative stage falls into each error plane and the error expectation corresponding to the error plane, the initial predicted position is corrected, including: correcting the initial predicted position using the following formula 1: Formula 1 In formula 1, P true To correct the predicted position, P coarse is the initial predicted position, N is the number of iteration stages, λ i is the weight of the i-th iteration stage, P is the number of error planes, x j is the expected error of the jth error plane, P j is the probability that the stage position of the i-th iteration stage falls into the j-th error plane.
8. The method according to claim 6, characterized in that The position estimation model is pre-trained in the following manner: Input the training data into the position estimation model to obtain the initial estimated position; Perform at least one pre-training iteration phase on the initial predicted position, wherein, in the pre-training iteration phase, an error plane prediction model is determined based on the residual between the actual position of the training data and the input position of the pre-training iteration phase, the input of the error plane prediction model is the position, and the output of the error plane prediction model is the error expectation corresponding to the input position, and the residual expectation of the position estimation model in the pre-training iteration phase is determined based on the error plane prediction model and the probability distribution function of the pre-training iteration phase, the input of the probability distribution function is the position, and the output of the probability distribution function is the probability of being located at the input position in the pre-training iteration phase; Determine the loss function of the position estimation model according to the residual expectation of each iterative stage of pre-training; The parameters of the location estimation model are adjusted to reduce the loss function of the location estimation model.
9. The method according to claim 8, characterized in that The residual expectation of the position estimation model in the pre-training iteration phase is determined according to the error plane prediction model and the probability distribution function of the pre-training iteration phase, including: The residual expectation of the iterative stage of pre-training is calculated using the following formula 2: Formula 2 E(Res)=∫[E(x)*P(x)]dx In Formula 2, E(Res) is the residual expectation of the calculated pre-training iteration stage, x is the position, E(x) is the error expectation of the error plane where the position x is located, P(x) is the probability of being at the position x in the calculated pre-training iteration stage, and the integration interval is the entire space.
10. The method according to claim 8, characterized in that The loss function of the position estimation model is determined based on the residual expectation of each iterative stage of pre-training, including: The loss function of the position estimation model is calculated using the following formula 3: Formula 3 In formula 3, L all represents the loss function of the location estimation model, N represents the number of pre-training iterations, and λ i Represents the weight of the residual expectation of the i-th iteration stage of pre-training, E(Res) i represents the residual expectation of the i-th iteration stage of pre-training.
11. The method according to claim 7, characterized in that The weight of the previous iteration stage is greater than the weight of the later iteration stage.
12. A device for determining a position, characterized in that: include: A control unit, used to determine n historical moments of the target object in a historical time queue and whether the target moment is located in a blind spot of the camera; the historical time queue is used to store historical position information of the target object at the latest n historical moments before the target moment, where n is a preset positive integer not less than 2; The control unit is also used for: in response to the target object being located outside the camera blind spot at n historical moments and the target moment, using the position information of the target object at the target moment collected by the camera as the position information of the target object at the target moment.
13. An electronic device comprising: at least one memory and at least one processor; The at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, the computer-readable storage medium being used to store program code, wherein when the program code is executed by a processor, the processor is prompted to execute the method according to any one of claims 1 to 11.