Method and device for determining position, electronic equipment and storage medium

By using historical time queues and pose change information for iterative positioning in extended real-life devices, the problem of difficult to accurately determine the position of the control component is solved, and highly accurate position estimation is achieved.

CN120163867APending Publication Date: 2025-06-17BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311723077.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When using extended reality devices, the position of the control components is difficult to accurately determine, resulting in offset and inaccurate display content.

Method used

By inputting the position change information of the historical time queue and the target object at the target time into the position estimation model, the positioning accuracy is improved in at least two iteration stages, and gradually approaching the real position of the target object.

Benefits of technology

It effectively prevents the data offset from increasing continuously, improves the accuracy and reliability of position estimation, and ensures the position accuracy of the control components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163867A_ABST
    Figure CN120163867A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for determining a position, electronic equipment and a storage medium. The method for determining the position comprises the steps that a historical time queue and pose change information of a target object at a target moment are input into a position estimation model to obtain an initial prediction position, the historical time queue is used for storing historical position information of the target object at recent n historical moments before the target moment, and n is a preset positive integer not smaller than 2; executing at least two iteration stages on the initial prediction position to obtain position information of the target object at the target moment; wherein the positioning precision of any iteration stage is higher than the positioning precision of the previous iteration stage. According to the invention, data offset can be prevented from increasing continuously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, an electronic device, and a storage medium for determining a position. Background Art

[0002] When using extended reality devices such as virtual reality, mixed reality, and augmented reality, control components such as a handle are usually used to control the extended reality device. The position of the control component is an important basis for controlling the content displayed by the extended reality device. Summary of the Invention

[0003] The present disclosure provides a method, an apparatus, an electronic device, and a storage medium for determining a position.

[0004] The present disclosure adopts the following technical solutions.

[0005] In some embodiments, the present disclosure provides a method for determining a position, including:

[0006] Inputting a historical time queue and pose change information of a target object at a target moment into a position estimation model to obtain an initial predicted position, where the historical time queue is used to store historical position information of the target object at the n most recent historical moments before the target moment, and n is a positive integer not less than 2 preset;

[0007] Performing at least two iterative stages on the initial predicted position to obtain position information of the target object at the target moment;

[0008] Wherein, the positioning accuracy of any iterative stage is higher than that of the previous iterative stage.

[0009] In some embodiments, the present disclosure provides a device for determining a position, including:

[0010] A control unit configured to input a historical time queue and pose change information of a target object at a target moment into a position estimation model to obtain an initial predicted position, where the historical time queue is used to store historical position information of the target object at the n most recent historical moments before the target moment, and n is a positive integer not less than 2 preset;

[0011] The control unit is further configured to: perform at least two iterative stages according to the initial predicted position to obtain position information of the target object at the target moment;

[0012] Wherein, the positioning accuracy of any of the iterative stages is higher than that of the previous iterative stage.

[0013] In some embodiments, the present disclosure provides an electronic device, including: at least one memory and at least one processor;

[0014] Among them, the memory is used to store program code, and the processor is used to call the program code stored in the memory to execute the above method.

[0015] In some embodiments, the present disclosure provides a computer-readable storage medium for storing program code, which, when run by a processor, causes the processor to execute the above method.

[0016] The method for determining a position provided by the embodiments of the present disclosure inputs the historical time queue and the pose change information of the target object at the target moment into a position estimation model to obtain an initial predicted position. Among them, the historical time queue is used to store the historical position information of the target object at the n most recent historical moments before the target moment, and n is a positive integer not less than 2 preset; at least two iterative stages are performed on the initial predicted position to obtain the position information of the target object at the target moment; among them, the positioning accuracy of any iterative stage is higher than that of the previous iterative stage. The present disclosure can prevent the continuous increase of data deviation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.

[0018] Figure 1 is a schematic diagram of using an extended reality device in an embodiment of the present disclosure.

[0019] Figure 2 is a flowchart of the method for determining a position in an embodiment of the present disclosure.

[0020] Figure 3 is a flowchart of the method for determining a position in an embodiment of the present disclosure.

[0021] Figure 4 is a schematic diagram of a dilated convolutional neural network model in an embodiment of the present disclosure.

[0022] Figure 5 is a schematic diagram of the processing of the convolutional layer of the dilated convolutional neural network in an embodiment of the present disclosure.

[0023] Figure 6 is a schematic diagram of the structure of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] It should be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0025] For example, when responding to a user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0026] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0027] It should be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0028] It should be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0029] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0030] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed sequentially and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0031] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0032] It should be noted that the concepts such as "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0033] It should be noted that the modification of "one" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0034] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes, and are not used to limit the scope of these messages or information.

[0035] The following will describe in detail the solutions provided by the embodiments of this disclosure in conjunction with the accompanying drawings.

[0036] The Extended Reality (XR) technology in one or more embodiments of this disclosure can be mixed reality technology, augmented reality technology, or virtual reality technology. The extended reality technology can combine the real and the virtual through a computer to provide an extended reality space for human-computer interaction. In the extended reality space, users can, for example, through extended reality devices such as head-mounted displays (HMDs), carry out social interactions, entertainment, learning, work, remote work, creation of UGC (User Generated Content), etc.

[0037] Referring to Figure 1 , users can enter the extended reality space through extended reality devices such as head-mounted glasses, and control their virtual characters (Avatars) in the extended reality space to carry out social interactions, entertainment, learning, remote work, etc. with virtual characters controlled by other users.

[0038] In one embodiment, in the extended reality space, users can achieve relevant interaction operations through a controller, and the controller can be a handle. For example, users can perform relevant operation controls by operating the buttons of the handle.

[0039] The extended reality devices described in the embodiments of the present disclosure may include, but are not limited to, the following types:

[0040] Computer-side extended reality devices perform relevant calculations and data output for extended reality functions using a computer. An external computer-side extended reality device uses the data output by the computer to achieve the effect of extended reality.

[0041] Mobile extended reality devices support setting a mobile terminal (such as a smartphone) in various ways (such as a head-mounted display with a dedicated card slot). Through a wired or wireless connection with the mobile terminal, the mobile terminal performs relevant calculations for extended reality functions and outputs data to the mobile extended reality device. For example, watch extended reality videos through an APP on the mobile terminal.

[0042] All-in-one extended reality devices have a processor for performing relevant calculations for extended reality functions, and thus have independent extended reality input and output functions. They do not need to be connected to a computer or a mobile terminal, and have a high degree of freedom of use.

[0043] Of course, the implementation form of the extended reality device is not limited to this, and it can be further miniaturized or enlarged according to needs.

[0044] A sensor for attitude detection (such as a nine-axis sensor) is provided in the extended reality device to detect the attitude change of the extended reality device in real time. If the user wears the extended reality device, when the user's head attitude changes, the real-time attitude of the head will be transmitted to the processor, and based on this, the fixation point of the user's line of sight in the extended reality space environment is calculated. According to the fixation point, the image in the user's viewing range (i.e., the virtual field of view) in the three-dimensional model of the extended reality space environment is calculated and displayed on the display screen, giving people an immersive experience as if they were watching in a real environment.

[0045] The handle tracking scheme can generally be divided into optical tracking (such as using a camera to take pictures) and the method based on the integration of inertial measurement units. Optical tracking can provide high-precision position results, but when the handle moves into the blind area of the camera's shooting range, the vision-based optical tracking cannot give an accurate position prediction. In addition, extended reality devices usually run on embedded terminals, which face resource constraints.

[0046] In the related art, the method of integrating sensors is adopted. By integrating information such as speed and acceleration over time, the movement trajectory of the handle in the blind area of the camera is obtained. As time goes by, the IMU will cause errors to accumulate continuously due to reasons such as temperature and noise, making the finally predicted position shift continuously in a certain direction, and the prediction result may be completely unusable.

[0047] Modeling time series based on a temporal network (e.g., Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Gate Recurrent Unit (GRU)) transforms the position prediction of the handle in the camera blind spot into a time series prediction problem. However, this method must preserve the hidden layer information of the neural network model. To preserve the hidden layer information of the neural network model, position prediction is also required when the handle is outside the camera blind spot, and the neural network model must remain operational both when the handle is in the camera blind spot and outside the camera blind spot.

[0048] As Figure 2 shown, Figure 2 is a flowchart of the method for determining a position according to an embodiment of the present disclosure, including the following steps.

[0049] S11. Input the historical time queue and the pose change information of the target object at the target moment into the position estimation model to obtain an initial predicted position.

[0050] In some embodiments, the execution end of the method proposed in the present disclosure can be any extended reality device in the present disclosure. The extended reality device can include a head-mounted display device and a supporting control device (such as a handle, a bracelet, a leg ring, and other accessories that need to be positioned). Taking the control device as a handle, the handle can be divided into a left-hand handle and a right-hand handle. The target object can be the left-hand handle or the right-hand handle of the extended reality device. This method can be used to determine the positions of the left-hand handle and the right-hand handle. Generally, there is no strong correlation between the positions of the left and right handles. Therefore, the positions of the left-hand handle and the right-hand handle are determined separately. In some embodiments, the historical time queue is used to store the historical position information of the target object at the most recent n historical moments before the target moment, where n is a preset positive integer not less than 2. n can be, for example, a positive integer not less than 10, 20, 30, 40, or 50, such as 60. The historical position information is the position information of the target object at the historical moment. In some embodiments, the position information can be obtained periodically. The most recent n historical moments can be the historical moments of n cycles before the target moment. For example, 30 cycles can be set within 1 second, and n can be 60. Then the historical time queue can be the historical moments within 2 seconds before the target moment. In some embodiments, the historical moment and the target moment can be the frame acquisition moments of the images captured by the camera. The camera can be located on the head-mounted display device of the extended reality device. There is a light spot for positioning on the target object (such as a handle). The camera captures the light spot on the handle to obtain the 6-degree-of-freedom information (position, angle, etc.) of the target object. The camera can capture images periodically. For example, 30 images are captured per second. The pose change information of the target object can be obtained when the camera captures the image to determine the corresponding position information at the time of capture. Therefore, each historical moment and the target moment are also the frame acquisition moments of the captured images. Each frame of the image captured by the camera corresponds to a pose change information of the target object and a position information determined based on the pose change information. The target moment can be the moment when the camera captures a frame of image most recently. In some embodiments, when the position information of the handle depends on the position information of the head-mounted display device, the historical time queue also contains the position, speed, angle, acceleration, and / or angular velocity information, etc., of the head-mounted display device at the historical moment. In some embodiments, the historical time queue and the pose change information are input into the position estimation model. The position estimation model can be a neural network model. After calculation, the position estimation model will predict an initial predicted position at the target moment. The initial predicted position can be represented by spatial coordinates.

[0051] S12. Perform at least two iterative stages on the initial predicted position to obtain the position information of the target object at the target moment.

[0052] In some embodiments, the number of iterative stages can be 2, 3, or more than 3, which is not limited herein. The accuracy of the initial predicted position is relatively low, so it cannot be directly output as the position information at the target time. Instead, the accuracy needs to be improved through an iterative approach. The initial predicted position serves as the input for one iterative stage. Each iterative stage outputs a predicted stage position of the target object at the target time and can also output the error corresponding to the stage position. The stage position output by the last iterative stage. In this embodiment, the positioning accuracy of any iterative stage is higher than that of the previous iterative stage. Therefore, a coarse-to-fine positioning is achieved, gradually approaching the true position of the target object at the target time. Specifically, the positioning accuracy can refer to the accuracy of the stage position output by the iterative stage. For example, the deviation between the stage position calculated in the first iterative stage and the true position is no greater than 1 cm, and the second iterative stage can ensure that the deviation between the stage position and the true position is no greater than 0.2 cm based on the first iteration. In some embodiments, at each iterative stage, the position information (i.e., the stage position) of the target object determined in the previous iterative stage is corrected by a correction value, and the accuracy of the correction value for each iterative stage gradually increases. For example, if the accuracy of the correction value determined in one iterative stage is 1 cm, then the accuracy of the correction value determined in the next iterative stage is less than 1 cm, such as 0.2 cm. In this way, the accuracy of the position information of the target object determined at each iteration is higher.

[0053] In some embodiments of the present disclosure, by using a historical time queue, the historical position information of the target object before is comprehensively considered, and a coarse-to-fine method is adopted to determine the position information of the target object at the target time. Each iteration can get closer to the true position of the target object at the target time. By combining the two, it is possible to prevent the continuous increase of data offset and ensure the accuracy and reliability of the calculated position information of the target object at the target time.

[0054] In some embodiments, after determining the position information of the target object, the historical time series can be updated. According to the first-in, first-out principle, the position information of the target object at the target time is put into the historical time series, and the position information of the target object at the earliest historical time in the historical time queue is removed.

[0055] In some embodiments of the present disclosure, the following operations are performed at each iterative stage:

[0056] Determine the error plane corresponding to the current iterative stage. Each error plane has an error expectation. The higher the positioning accuracy of the current iterative stage, the smaller the interval of the error expectations of the error planes corresponding to the current iterative stage;

[0057] Determine the probabilities that the input position falls into the respective error planes corresponding to the current iteration stage, where the input position in the first iteration stage is the initial prediction position, and the input positions in other iteration stages are the stage positions output in the previous iteration stage;

[0058] According to the probabilities that the input position falls into the respective error planes and the error expectations corresponding to the error planes, correct the input position to obtain a corrected position;

[0059] Input the corrected position into the position estimation model to obtain the stage position of the current iteration stage.

[0060] In some embodiments, Figure 3The flowchart of the method proposed in some embodiments of the present disclosure is shown. First, obtain the historical time queue and the pose change information of the target object at the target moment, input them into the position estimation model, and then obtain the initial predicted position. Then, the iterative stage will be executed. In some embodiments, a preset number of error levels are preset. Each error plane is divided into different error levels. The error expectations of the error planes within one error level are arranged in an arithmetic progression at equal intervals. The intervals (i.e., the common differences) of the error expectations of the error planes in different error levels are different. One iterative stage corresponds to the error planes within one error level. The error plane (which can also be called the loss plane) can be a plane or a surface of the prediction error. The error plane corresponds to an error expectation, that is, the error plane is a plane used to characterize the error expectation. Through the error plane, the calculation of the error expectation is converted into the calculation of the position. If a position only falls into one error plane, the error expectation at this position is the error expectation corresponding to the error plane. However, the error expectation of a position in space may be multiple. Its essence is a probability distribution problem, that is, a position in space may fall into multiple error planes, and the possibility of falling into each error plane is described by probability. That is, the error expectation corresponding to a position is described by a probability distribution. At a position, its corresponding error expectation may be value one, value two, etc. The probabilities that its error expectation is value one, value two, etc. can be the probabilities of falling into the error planes corresponding to value one, value two, etc. The error plane is preset. Known data can be used for pre-training in advance, so as to estimate the distribution probability of the error expectations at different positions, and thus the probability that this position falls into each error plane can be known. For example, the error levels can include 10 cm, 1 cm, 0.2 cm, which means that the error expectations of the error planes within this error level are arranged at intervals of 10 cm, 1 cm, and 0.2 cm to form an arithmetic progression. One iterative stage corresponds to the error planes within one error level. The error expectations of the error planes within one error level form an arithmetic progression with the interval of the error expectation as the common difference. For example, the error expectations corresponding to the error planes with an error level of 1 cm can be -4, -3, -2, -1, 0, 1, 2, etc. First, determine the error level corresponding to the current iterative stage, so as to determine the error plane corresponding to this error level. The error plane is a plane used to evaluate the error. For example, if the position information falls into a certain error plane, the error of this position information is the error expectation corresponding to this error plane, thus converting the process of determining the error into the process of determining the error plane. When determining the error plane corresponding to the current iterative stage, it can be to select the error plane whose absolute value of the error expectation is less than a preset multiple of the interval of the error plane. For example, the interval is 1 cm and the multiple is 4, then determine the error plane whose absolute value of the error expectation is less than 1×4. In the iterative process, a pose optimizer can be used for calculation. In the first iterative stage, input the initial predicted position into the pose optimizer, and then the probability of falling into the error plane of the current error level can be obtained.Then, the error expectations corresponding to each error plane are multiplied by their corresponding probabilities and accumulated to obtain a correction value. The correction value is added to the input position to obtain the corrected position, and the corrected position is input into the position estimation model for calculation to obtain the stage position. The stage position output in the last iteration stage is used as the position information of the target object at the target moment.

[0061] For example, assume that the initial predicted position input in the first iteration stage is 40 cm, while the actual position is 42.45 cm, and the error level corresponding to the first iteration stage is 1 cm. Then the pose optimizer determines that the probabilities of falling into the error planes with error expectations of -4 cm, -3 cm, -2 cm, -1 cm, 0 cm, 1 cm, 2 cm, 3 cm, and 4 cm are 0.1, 0, 0, 0, 0, 0, 0, 0.9, and 0 respectively. Then the correction value is -4×0.1 + 3×0.9 = 2.3 cm, and the corrected position is 40 + 2.3 = 42.3 cm. Since the error level is 1 cm, the 0.3 after the decimal point is doubtful. After inputting 42.3 cm into the position estimation model, assume that the result output by the position estimation model is still 42.3 cm. Then this value is the stage position of the first iteration stage. Assume there are two error levels, then this 42.3 cm will be used as the input position of the second iteration stage. Taking the error level of the second iteration stage as 0.2 cm as an example, after the pose optimizer calculates the stage position, it is determined that the probabilities of falling into the error planes with error expectations of -0.2 cm, 0 cm, and 0.2 cm are 0, 0.5, and 0.5 respectively. Then the correction value of the second iteration stage is 0×0.5 + 0.2×0.5 = 0.1 cm, and the corrected position of the second iteration stage is 42.4 + 0.1 = 42.4 cm. It can be seen that it is already very close to the actual position of 42.45 cm. After inputting 42.4 cm into the position estimation model, the stage position of the second iteration stage will be obtained. If there are only two error levels, then the stage position obtained in the second iteration stage is used as the position information of the target object output finally at the target moment.

[0062] In some embodiments of the present disclosure, the idea of cascade and step-by-step improvement is used to predict the residual of the previous prediction. In the iterative stage of the residual, simple regression is not used, but multiple error planes are generated based on the predicted input position, and the corresponding probability distribution is predicted. By using iteration and error planes, the accuracy of position positioning can be gradually improved. This method makes full use of the prediction ability of the model and the law of error distribution, and can obtain more accurate position estimation results. The process of continuously refining the error plane enables the true position of the target to be gradually approximated, thereby improving the accuracy of position positioning.

[0063] In some embodiments of the present disclosure, the position estimation model needs to be trained before use, and the training method can be the same as that of existing neural networks. During training, data with known true position information is used, and the historical position information and pose change information for training are input. The initial predicted position is output through the position estimation model. Then, iteration is performed. An error plane is generated based on the error between the input position and the true position information, and this error plane represents the error distribution law of the predicted position. Using the error plane, the input position is corrected through regression or optimization to approximate the true position. During the training process, the parameters of the position estimation model and the pose optimizer can be continuously adjusted to reduce the loss function. When the number of iterations reaches a certain number or the loss function is less than a preset value during the training process, the training ends.

[0064] In some embodiments of the present disclosure, different bits in the numerical value of the position information for determining the position information of the target object at the target moment are used in different iteration stages. In some embodiments, the numerical value of the position information, for example, has a tens digit, a units digit, a first digit after the decimal point, and a second digit after the decimal point. For different bits in the numerical value of the position information, different iteration stages are used for determination. In this way, the bits in the numerical value of the position information determined in the previous iteration stage are higher than those determined in the subsequent iteration stage. For example, the tens digit is determined in the first iteration stage, the units digit is determined in the second iteration stage, the first digit after the decimal point is determined in the third iteration stage, and the second digit after the decimal point is determined in the fourth iteration stage. In this way, the determined position information gradually approaches the true position information through continuous iteration. In some embodiments, the accuracy of the determined bit can be gradually improved by controlling the interval of the error expectation of the error plane corresponding to each iteration stage.

[0065] In some embodiments, the absolute value of the error expectation of the error plane determined in the current iteration stage (the current iteration stage can be any iteration stage) is less than the interval of the error plane in the previous iteration stage. For example, if the interval in the first iteration stage is 1 cm, then the error expectation of the error plane in the second iteration stage is within -1 cm to 1 cm. This can improve the positioning accuracy for each iteration.

[0066] Before inputting the historical time queue and the pose change information of the target object at the target moment into the position estimation model to obtain the initial predicted position, some embodiments of the present disclosure further include:

[0067] Determine whether the target object is located in the blind area of the camera at n historical moments and the target moment in the historical time queue;

[0068] In response to the target object being located outside the blind area of the camera at n historical moments and the target moment, use the position information of the target object collected by the camera at the target moment as the position information of the target object at the target moment;

[0069] In response to the target object being located in the camera blind area at at least one of the n historical moments and the target moment, perform the step of inputting the historical time queue and the pose change information into the position estimation model to obtain the initial predicted position.

[0070] In some embodiments, it may be to first obtain a historical time queue and pose change information. Then determine whether the historical moment and the target moment are located in the imaging blind area of the camera. Taking this method applied to an extended reality device as an example, both the control device and the head-mounted display device may be equipped with pose sensors, such as one or more of an acceleration sensor and an angular velocity sensor, etc. It may be to obtain the sensor information of the pose sensor (such as one or more of acceleration information and angular velocity information) as the pose change information, or it may be to perform feature extraction on the sensor information (such as integration, and the integration duration is from the target moment to the previous historical moment) to obtain the pose change information. Specifically, the pose change information may include the following of the target object: speed, angle, acceleration, and / or angular velocity. In some embodiments, in an extended reality device, the head-mounted display device may be used as the origin of the space, and the position of the control device may represent a position dependent on the head-mounted display device. At this time, since the pose change information of the target object depends on the head-mounted display device, the pose change information of the target object needs to contain pose data such as the position, speed, angle, acceleration, and / or angular velocity of the head-mounted display device. In some embodiments, the camera periodically captures images, and the moment of capturing the image is the frame acquisition moment. The target moment may be the frame acquisition moment of the previous image capture, that is, the moment of the nearest previous camera image capture starting from the current moment (including the current moment). Therefore, the pose change information is: the pose change information of the target object during the period from the most recent image capture to the most recent previous image capture. The camera may be a camera on the head-mounted display device, and the target object may be the control device. The target object being located in the imaging blind area of the camera includes: the camera cannot capture the target object, and may also include: although the camera can capture the target object, the captured image cannot be used to determine the position information of the target object (for example, due to insufficient clarity, insufficient number of light spots, etc.), and at this time, it is also regarded as the target object being located in the imaging blind area. When the target object is outside the imaging blind area at n historical moments and the target moment, it indicates that at this time, the target object has been in a position where it can be captured by the camera and has been in this position for a certain period of time. At this time, the position information of the target object can be determined through the image captured by the camera at the target moment. At this time, it is not necessary to determine the position information of the target object at the target moment according to the historical time queue and the pose change information. For example, the position estimation model is a neural network model, and when the target object is outside the imaging blind area at n historical moments and the target moment, the position estimation model will not work, which can avoid the displacement deviation problem caused by the long-term integration of sensor information. When the target object is in the imaging blind area at n historical moments and the target moment, the position information of the target object at the target moment is determined through the position estimation model, and the position information can be filtered by Kalman filter to improve the accuracy.

[0071] In some embodiments of the present disclosure, a historical time series is set up to save the position information of the target object at n historical moments. When the target object is outside the camera blind area at the target moment and historical moments, the optical tracking of the camera is used to determine the position information of the target object, instead of continuously calculating the position information of the target object using a neural network model. When the target object is in the camera blind area at the target moment or one of the n historical moments, the historical time queue and pose change information are used to predict the position information at the target moment.

[0072] In some embodiments of the present disclosure, after determining that the target object is in the camera blind area at at least one of the n historical moments and the target moment, an initial predicted position is obtained using a position estimation model; after determining that the target object is outside the camera blind area at all of the n historical moments and the target moment, the position estimation model is maintained or set to a non-operating state.

[0073] In some embodiments, for traditional time series networks such as LSTM and GRU, since the hidden layer data needs to be used to calculate the position information in the camera blind area, the time series network needs to continuously calculate the position information of the target object regardless of where the target object is located. However, in the present disclosure, it is not necessary to continuously use the position estimation model to predict the position information of the target object. When the target object is not in the camera blind area at the target moment and historical moments, there is no need to use the position prediction model, and the position prediction model does not work. There is no need to input the historical time queue and pose change information into the position prediction model. The position prediction model only works when the target object is in the camera blind area, thus reducing resource consumption. Moreover, in the present disclosure, when the target object is in the camera blind area, the predicted position information refers to the historical position information of n historical moments, rather than relying solely on pose change information, thereby preventing the deviation of position information.

[0074] In some embodiments of the present disclosure, based on the initial predicted position, at least two iterative stages are performed to obtain the position information of the target object at the target moment, including: determining the relative position of the target object at the target moment relative to the previous historical moment, and using the relative position as the position information of the target object at the target moment; or determining the relative position of the target object at the target moment relative to the previous historical moment, and determining the position information of the target object at the target moment based on the relative position and the historical position information of the target object at the previous historical moment.

[0075] In some embodiments, the position information of the target object determined at the target moment can be a relative position, that is, the relative position relative to the historical position information of the previous historical moment of the target moment, or an absolute position, such as a position represented by spatial coordinates.

[0076] In some embodiments of the present disclosure, the position estimation model is a dilated convolutional neural network model, and the dilation coefficient of the dilated convolutional neural network model is not less than 2.

[0077] In some embodiments, the time convolutional network processes the time series representation as follows. For the historical sequences x1 to x t The outputs are y1 to y t , and for the output value of the model at time t, it depends on the values at time t and historical times.. The difference between the time convolutional network and the convolutional neural network is that the time convolutional network cannot use subsequent data, and it represents a model with historical time constraints. To make full use of information over a long time, the number of convolutional layers and the number of hidden layers must be increased accordingly. And the formats of the convolutional layers and the hidden layer units will bring problems such as: the training of the model becomes complex, the gradient dispersion of the model transmission, and the overfitting of the model. Therefore, in some embodiments of the present disclosure, a dilated convolution (DilatedConvolution) neural network model is used. Its essential problem is that a deep neural network needs to use a very large filter size or filter layers to increase the receptive field of the model. And dilated convolution can make the filter obtain a larger receptive field by processing the input of the model at equal intervals. The dilation coefficient is generally set as an exponential multiple of 2, and the filter is represented as a sequence F=(f1,f2…f K ), where f i is the filter of the i-th layer, K is the total number of layers, and the dilatation rate of the dilated convolution at the input X=(x1,x2,..x T ) equal to d is: Intuitive representation of dilated convolution Figure 4 As shown. It can be seen that for the input X, when the dilation coefficient is 2, the amount of data to be processed is halved for each layer passed through. Through the dilated convolutional neural network, the receptive field can be expanded and the computational complexity can be reduced.

[0078] In some embodiments of the present disclosure, the following operations are performed on each convolutional layer of the position estimation model: the layer input data of the current convolutional layer is subjected to two preset processes to obtain the layer output data, or after the layer input data is subjected to at least one preset process, the data obtained by performing a 1×1 convolution on the layer input data is added to obtain the layer output data; wherein, the layer input data of the first layer of the position estimation model is the historical time queue and the pose change information, and the layer output data of the current convolutional layer is the layer input data of the next convolutional layer; the preset process includes: after the layer input data is processed by dilated convolution with weight parameter normalization, it is non-linearly processed using an activation function, and then dropout unit processing is performed.

[0079] In some embodiments, compared with the absolute coordinate trajectory that directly outputs the trajectory, using the information of the sensor to predict a relative coordinate value can make the learning mode of the network simpler. Therefore, a residual block structure H(x) = F(x) + W s x is also constructed in the present disclosure, as Figure 5 shown, which shows the operations performed in each layer of the dilated convolutional neural network. Taking the layer input data of the current convolutional layer as x, if the convolutional layer is the first layer, then x is the historical time series and pose change information, and if the current convolutional layer is other layers, then x is the layer output data of the previous layer. To avoid overfitting, dilated convolutional processing with weight parameter normalization is performed on x. For the processed data, the LeakyReLU activation function is used for non-linear processing, and the output of the negative half-axis can be taken into account. Dropout processing is used during passing to increase robustness. These steps are performed once or multiple times, Figure 5 and are performed 2 times in the present disclosure, and F(x) is used to represent the final output. Meanwhile, optionally, in the short-circuit connection, when the feature dimension and the output dimension do not match, a 1×1 convolution can be added to perform dimension transformation, and a 1×1 convolution operation is performed on the layer input data X, denoted as W s X, and finally the layer output data of the current convolutional layer H(x) = F(x) + W s x is obtained.

[0080] In some embodiments of the present disclosure, taking the target object as the handle as an example, a temporal convolutional network with dilated convolution and residual structure is used to predict the position information of the target object in the camera blind area. By introducing dilated convolution and residual structure, this method can more accurately predict the position information of the target object in the blind area. Secondly, the present disclosure decouples the left and right hand handles, enabling the user to independently control the actions of each hand. Through the decoupling process, the user can operate the handle more freely, improving the flexibility and diversity of interaction. In addition, the present disclosure also designs a method for determining the position that is from coarse to fine and has good numerical stability. This method adopts a strategy of gradually refining during the prediction process, first making a rough prediction, and then gradually refining the prediction result. This method can not only improve the accuracy of the prediction, but also maintain the numerical stability, avoiding mutations and jitters in the prediction result.

[0081] In some embodiments of the present disclosure, a device for determining the position is also proposed, including:

[0082] A control unit, configured to input the historical time queue and the pose change information of the target object at the target moment into the position estimation model to obtain an initial predicted position, where the historical time queue is used to store the historical position information of the target object at the n most recent historical moments before the target moment, and n is a preset positive integer not less than 2;

[0083] The control unit is further configured to: according to the initial predicted position, perform at least two iterative stages to obtain the position information of the target object at the target time;

[0084] Wherein, the positioning accuracy of any one of the iterative stages is higher than that of the previous iterative stage.

[0085] In some embodiments, the following operations are performed in each of the iterative stages:

[0086] Determine the error plane corresponding to the current iterative stage. Each of the error planes has an error expectation. The higher the positioning accuracy of the current iterative stage, the smaller the interval of the error expectations of the error planes corresponding to the current iterative stage;

[0087] Determine the probability that the input position falls into each of the error planes corresponding to the current iterative stage. The input position of the first iterative stage is the initial predicted position, and the input positions of other iterative stages are the stage positions output by the previous iterative stage;

[0088] According to the probability that the input position falls into each of the error planes and the error expectation corresponding to the error plane, correct the input position to obtain a corrected position;

[0089] Input the corrected position into the position estimation model to obtain the stage position of the current iterative stage.

[0090] In some embodiments, each of the error planes is divided into different error levels. The error expectations of the error planes within one error level are arranged at equal intervals, the intervals of the error expectations of the error planes in different error levels are different, and one iterative stage corresponds to the error planes within one error level.

[0091] In some embodiments, different iterative stages are used to determine different bits in the numerical value of the position information of the target object at the target time.

[0092] In some embodiments, a determination unit is further included, configured to determine whether the target object is located in the camera's blind area at n historical times and the target time in the historical time queue before inputting the historical time queue and the pose change information of the target object at the target time into the position estimation model to obtain the initial predicted position;

[0093] The control unit is configured to, in response to the target object being located outside the camera's blind area at n historical times and the target time, use the position information of the target object collected by the camera at the target time as the position information of the target object at the target time;

[0094] A control unit, configured to, in response to the target object being located in the camera blind area at at least one of the n historical moments and the target moment, perform the step of inputting the historical time queue and the pose change information into a position estimation model to obtain an initial predicted position.

[0095] In some embodiments, the control unit is configured to, after determining that the target object is located in the camera blind area at at least one of the n historical moments and the target moment, use the position estimation model to obtain the initial predicted position;

[0096] The control unit is configured to, after determining that the target object is located outside the camera blind area at all of the n historical moments and the target moment, keep or set the position estimation model in a non-operating state.

[0097] In some embodiments, according to the initial predicted position, at least two iterative stages are performed to obtain the position information of the target object at the target moment, including:

[0098] Determine the relative position of the target object at the target moment relative to the previous historical moment, and use the relative position as the position information of the target object at the target moment; or,

[0099] Determine the relative position of the target object at the target moment relative to the previous historical moment, and determine the position information of the target object at the target moment according to the relative position and the historical position information of the target object at the previous historical moment.

[0100] In some embodiments, the position estimation model is a dilated convolutional neural network model, and the dilation coefficient of the dilated convolutional neural network model is not less than 2.

[0101] In some embodiments, the following operations are performed on each convolutional layer of the position estimation model:

[0102] Perform two preset processes on the layer input data of the current convolutional layer to obtain layer output data, or, after performing at least one preset process on the layer input data, add the data obtained by performing 1×1 convolution on the layer input data to obtain layer output data;

[0103] Wherein, the layer input data of the first layer of the position estimation model is the historical time queue and the pose change information, and the layer output data of the current convolutional layer is the layer input data of the next convolutional layer;

[0104] The preset process includes: performing dilated convolutional processing on the layer input data after normalizing the weight parameters, performing non-linear processing using an activation function, and then performing dropout unit processing.

[0105] For the embodiments of the apparatus, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The apparatus embodiments described above are merely illustrative, where the modules described as separation modules may or may not be separated. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0106] As described above, the method and apparatus of the present disclosure have been described based on the embodiments and application examples. In addition, the present disclosure also provides an electronic device and a computer-readable storage medium. These electronic devices and computer-readable storage media will be described below.

[0107] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device (such as a terminal device or a server) 800 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in the figure is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0108] The electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 802 or the programs loaded from the storage device 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.

[0109] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 can allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows the electronic device 800 having various devices, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices can be alternatively implemented or included.

[0110] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.

[0111] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0112] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0113] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.

[0114] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods of the present disclosure as described above.

[0115] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the “C” language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0117] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0118] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, the types of hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0119] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0120] According to one or more embodiments of the present disclosure, a method for determining a location is provided, including:

[0121] Input the historical time queue and the pose change information of the target object at the target moment into the position estimation model to obtain an initial predicted position, where the historical time queue is used to store the historical position information of the target object at the n most recent historical moments before the target moment, and n is a preset positive integer not less than 2;

[0122] Perform at least two iterative stages on the initial predicted position to obtain the position information of the target object at the target moment;

[0123] Among them, the positioning accuracy of any iterative stage is higher than that of the previous iterative stage.

[0124] According to one or more embodiments of the present disclosure, a method for determining a position is provided. In each of the iterative stages, the following operations are performed:

[0125] Determine the error plane corresponding to the current iterative stage, where each error plane has an error expectation, and the smaller the interval of the error expectations of the error planes corresponding to the current iterative stage, the higher the positioning accuracy of the current iterative stage;

[0126] Determine the probability that the input position falls into each of the error planes corresponding to the current iterative stage, where the input position of the first iterative stage is the initial predicted position, and the input positions of other iterative stages are the stage positions output by the previous iterative stage;

[0127] According to the probability that the input position falls into each of the error planes and the error expectation corresponding to the error plane, correct the input position to obtain a corrected position;

[0128] Input the corrected position into the position estimation model to obtain the stage position of the current iterative stage.

[0129] According to one or more embodiments of the present disclosure, a method for determining a position is provided. Each error plane is divided into different error levels, the error expectations of the error planes within one error level are arranged at equal intervals, the intervals of the error expectations of the error planes in different error levels are different, and one iterative stage corresponds to the error planes within one error level;

[0130] And / or,

[0131] Different iterative stages are used to determine different bits in the values of the position information of the target object at the target moment.

[0132] According to one or more embodiments of the present disclosure, a method for determining a position is provided. Before inputting the historical time queue and the pose change information of the target object at the target moment into the position estimation model to obtain an initial predicted position, it further includes:

[0133] Determine n historical moments of the target object in the historical time queue and whether the target moment is within the blind area of the camera's shooting;

[0134] In response to the target object being outside the blind area of the camera at all of the n historical moments and the target moment, use the position information of the target object at the target moment collected by the camera as the position information of the target object at the target moment;

[0135] In response to the target object being within the blind area of the camera at at least one of the n historical moments and the target moment, perform the step of inputting the historical time queue and the pose change information into the position estimation model to obtain an initial predicted position.

[0136] According to one or more embodiments of the present disclosure, a method for determining a position is provided. After determining that the target object is within the blind area of the camera at at least one of the n historical moments and the target moment, use the position estimation model to obtain the initial predicted position;

[0137] After determining that the target object is outside the blind area of the camera at all of the n historical moments and the target moment, keep or set the position estimation model in a non-working state.

[0138] According to one or more embodiments of the present disclosure, a method for determining a position is provided. According to the initial predicted position, perform at least two iterative stages to obtain the position information of the target object at the target moment, including:

[0139] Determine the relative position of the target object at the target moment relative to the previous historical moment, and use the relative position as the position information of the target object at the target moment; or,

[0140] Determine the relative position of the target object at the target moment relative to the previous historical moment, and determine the position information of the target object at the target moment according to the relative position and the historical position information of the target object at the previous historical moment.

[0141] According to one or more embodiments of the present disclosure, a method for determining a position is provided. The position estimation model is a dilated convolutional neural network model, and the dilation coefficient of the dilated convolutional neural network model is not less than 2.

[0142] According to one or more embodiments of the present disclosure, a method for determining a position is provided. Perform the following operations on each convolutional layer of the position estimation model:

[0143] The layer output data is obtained by performing two preset processes on the layer input data of the current convolutional layer, or by performing at least one preset process on the layer input data and then adding the data obtained by performing a 1×1 convolution on the layer input data;

[0144] Among them, the layer input data of the first layer of the position estimation model is the historical time queue and the pose change information, and the layer output data of the current convolutional layer is the layer input data of the next convolutional layer;

[0145] The preset process includes: performing dilated convolution processing on the layer input data after normalizing the weight parameters, performing non-linear processing using an activation function, and then performing dropout unit processing.

[0146] According to one or more embodiments of the present disclosure, there is provided a device for determining a position, including:

[0147] A control unit, configured to input a historical time queue and pose change information of a target object at a target moment into a position estimation model to obtain an initial predicted position, where the historical time queue is used to store historical position information of the target object at the n most recent historical moments before the target moment, and n is a preset positive integer not less than 2;

[0148] The control unit is further configured to: perform at least two iterative stages based on the initial predicted position to obtain the position information of the target object at the target moment;

[0149] Among them, the positioning accuracy of any one of the iterative stages is higher than that of the previous iterative stage.

[0150] According to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one memory and at least one processor;

[0151] Among them, the at least one memory is used to store program codes, and the at least one processor is used to call the program codes stored in the at least one memory to execute the method described in any one of the above.

[0152] According to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium, which is used to store program codes, and when the program codes are run by a processor, the processor is prompted to execute the above method.

[0153] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0154] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0155] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms for implementing the claims.

Claims

1. A method for determining a position, characterized in that, Including: Inputting the historical time queue and the pose change information of the target object at the target moment into a position estimation model to obtain an initial predicted position, where the historical time queue is used to store the historical position information of the target object at the n most recent historical moments before the target moment, and n is a preset positive integer not less than 2; Performing at least two iterative stages on the initial predicted position to obtain the position information of the target object at the target moment; Wherein, the positioning accuracy of any iterative stage is higher than that of the previous iterative stage.

2. The method according to claim 1, characterized in that, In each of the iterative stages, the following operations are performed: Determining the error plane corresponding to the current iterative stage, wherein each error plane has an error expectation, and the smaller the interval of the error expectations of the error planes corresponding to the current iterative stage, the higher the positioning accuracy of the current iterative stage; Determining the probability that the input position falls into each of the error planes corresponding to the current iterative stage, wherein the input position of the first iterative stage is the initial predicted position, and the input positions of other iterative stages are the stage positions output by the previous iterative stage; Correcting the input position according to the probability that the input position falls into each of the error planes and the error expectation corresponding to the error plane to obtain a corrected position; Inputting the corrected position into the position estimation model to obtain the stage position of the current iterative stage.

3. The method according to claim 1, characterized in that, Each error plane is divided into different error levels, the error expectations of the error planes within one error level are arranged at equal intervals, the intervals of the error expectations of the error planes in different error levels are different, and one iterative stage corresponds to the error planes within one error level; And / or Different iterative stages are used to determine different bits in the value of the position information of the target object at the target moment.

4. The method according to claim 1, characterized in that, Before inputting the historical time queue and the pose change information of the target object at the target moment into the position estimation model to obtain the initial predicted position, it further includes: Determining whether the target object is in the camera's blind area at the n historical moments and the target moment in the historical time queue; In response to the target object being outside the camera's blind area at the n historical moments and the target moment, using the position information of the target object at the target moment collected by the camera as the position information of the target object at the target moment; In response to the target object being in the camera's blind area at at least one of the n historical moments and the target moment, performing the step of inputting the historical time queue and the pose change information into the position estimation model to obtain the initial predicted position.

5. The method according to claim 4, characterized in that, After determining that the target object is in the camera's blind area at at least one of the n historical moments and the target moment, using the position estimation model to obtain the initial predicted position; After determining that the target object is outside the camera's blind area at the n historical moments and the target moment, keeping or setting the position estimation model in a non-working state.

6. The method according to claim 1, characterized in that, According to the initial predicted position, performing at least two iterative stages to obtain the position information of the target object at the target moment, including: Determine the relative position of the target object at the target moment relative to the previous historical moment, and use the relative position as the position information of the target object at the target moment; or, Determine the relative position of the target object at the target moment relative to the previous historical moment, and determine the position information of the target object at the target moment according to the relative position and the historical position information of the target object at the previous historical moment.

7. The method according to claim 1, characterized in that, The position estimation model is a dilated convolutional neural network model, and the dilation coefficient of the dilated convolutional neural network model is not less than 2.

8. The method according to claim 7, characterized in that, Perform the following operations on each convolutional layer of the position estimation model: Perform two preset processes on the layer input data of the current convolutional layer to obtain the layer output data, or, after performing at least one preset process on the layer input data, add the data obtained by performing 1×1 convolution on the layer input data to obtain the layer output data; Among them, the layer input data of the first layer of the position estimation model is the historical time queue and the pose change information, and the layer output data of the current convolutional layer is the layer input data of the next convolutional layer; The preset process includes: performing dilated convolutional processing on the layer input data after normalizing the weight parameters, performing non-linear processing using an activation function, and then performing dropout unit processing.

9. A device for determining a position, characterized in that, Includes: A control unit for inputting the historical time queue and the pose change information of the target object at the target moment into the position estimation model to obtain an initial predicted position, where the historical time queue is used to store the historical position information of the target object at the n most recent historical moments before the target moment, and n is a preset positive integer not less than 2; The control unit is further configured to: perform at least two iterative stages according to the initial predicted position to obtain the position information of the target object at the target moment; Among them, the positioning accuracy of any one of the iterative stages is higher than that of the previous iterative stage.

10. An electronic device, comprising: At least one memory and at least one processor; Among them, the at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the method according to any one of claims 1 to 8.

11. A computer-readable storage medium for storing program code, which, when run by a processor, causes the processor to execute the method according to any one of claims 1 to 8.