Stationary state detection method, device, storage medium, and program product
By acquiring the three-dimensional position information of visual markers associated with the object to be detected, and using optical positioning technology to detect the movement of the visual markers, the problem of high hardware cost in existing technologies is solved, and high accuracy and robustness of static state detection are achieved.
Patent Information
- Application Number
- CN202411934126.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-12-25
AI Technical Summary
In existing technologies, static state detection requires the introduction of additional hardware such as accelerometers or IMUs, which increases costs, especially when there are multiple objects to be detected, exceeding the customer's budget.
By acquiring the three-dimensional position information of multiple visual markers associated with the object under test at various measurement times, optical positioning technology is used to detect whether the visual markers have moved within a preset measurement time, and the static state of the object under test is comprehensively judged, avoiding the introduction of additional hardware.
It achieves highly accurate and robust static state detection, reduces hardware costs, and improves the stability and reliability of detection.
Smart Images

Figure CN119958503B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer software technology, and more particularly to a static state detection method, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] Stillness detection typically refers to detecting whether an object or device is stationary or motionless at a given moment. This detection is crucial for many applications. For example, in virtual photography, which utilizes Virtual Reality (VR), Augmented Reality (AR), and real-time computer graphics rendering technologies to combine a virtual environment with real-world footage, creating an effect that appears to be filmed in the real world. However, to ensure consistency between the virtual scene and the actual footage, detecting whether the camera is stationary is essential, as changes in the camera's position affect the relative position of the virtual environment, causing it to appear to jump around. Similarly, in automated production control, many production lines require ensuring that machines or workpieces are stationary during certain processes to guarantee the accurate execution of subsequent operations (such as assembly, inspection, and processing). On automated assembly lines, detecting whether workpieces are completely fixed and stationary prevents assembly errors or damage caused by shaking or displacement.
[0003] The static state detection method in related technologies requires the introduction of additional hardware, such as placing accelerometers, gyroscopes, inertial measurement units (IMUs) on the object to be detected to help confirm whether the object is stationary. This increases the detection cost, especially when there are multiple objects to be detected, the hardware expenditure may exceed the customer's budget. Summary of the Invention
[0004] In view of the above, this specification provides a method for detecting a stationary state, an electronic device, a computer-readable storage medium, and a computer program product through one or more embodiments.
[0005] To achieve the above objectives, one or more embodiments of this specification provide the following technical solutions:
[0006] According to a first aspect of one or more embodiments of this specification, a method for detecting a stationary state is provided, the method comprising:
[0007] The three-dimensional position information of multiple visual markers associated with the object to be detected at various measurement times is obtained; the three-dimensional position information is obtained based on optical positioning of the multiple visual markers.
[0008] For each of the aforementioned visual markers, using multiple three-dimensional position information of the visual markers within a preset measurement time, the system detects whether the visual markers have moved within the preset measurement time, and obtains the detection result.
[0009] Based on the detection results of the multiple visual markers at the preset measurement time, the static state detection result of the object to be detected is determined.
[0010] According to a second aspect of the embodiments of this specification, an electronic device is provided, comprising:
[0011] processor;
[0012] Memory used to store processor-executable instructions;
[0013] Wherein, when the processor executes the executable instructions, it is used to implement the method described in the first aspect.
[0014] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in the first aspect.
[0015] According to a fourth aspect of the embodiments of this specification, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0016] The technical solutions provided in the embodiments of this specification may include the following beneficial effects:
[0017] In the embodiments of this specification, the three-dimensional position information of multiple visual markers associated with the object to be detected at various measurement times can be obtained, thereby providing a data basis for subsequent static state determination; then, by utilizing the multiple three-dimensional position information of each visual marker within a preset measurement time, it is possible to accurately detect whether each visual marker has moved, and improve the stability and robustness of detection through multi-time data analysis; finally, based on the detection results of all visual markers within the preset measurement time, it is possible to comprehensively determine whether the object to be detected is in a static state, and ensure the accuracy and reliability of the detection results through the comprehensive judgment of multiple visual markers.
[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of an optical positioning process provided in an exemplary embodiment.
[0020] Figure 2This is a flowchart of a static state detection method provided in an exemplary embodiment.
[0021] Figure 3 This is a flowchart of another static state detection method provided in an exemplary embodiment.
[0022] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment.
[0023] Figure 5 This is a block diagram of a stationary state detection device provided in an exemplary embodiment. Detailed Implementation
[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0025] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.
[0026] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0027] Here, we will first explain the relevant terms mentioned in the embodiments of this specification:
[0028] (1) Optical positioning: This is a technology that uses optical sensors (such as cameras, sensor arrays, lasers, etc.) to acquire information about the position or spatial location of an object. It mainly relies on the acquisition and analysis of optical signals for positioning and is widely used in many fields, such as virtual reality (VR), augmented reality (AR), motion capture, and robot navigation.
[0029] Optical positioning typically relies on computer vision technology for object recognition and tracking. A camera captures an image of a specific area, identifies feature points or visual markers (such as color, shape, and pattern) of objects in the scene, and then calculates the positions of these feature points or visual markers in the image to determine the spatial coordinates of the objects.
[0030] To obtain the three-dimensional coordinates of an object, optical positioning systems typically use multiple cameras or sensors to acquire image information from different angles, and then calculate the object's specific position in space using the principles of triangulation. Alternatively, distance and depth can be estimated using a single camera and a reference object of known size (such as a calibration plate).
[0031] In many applications, such as virtual photography or motion capture, optical positioning systems track objects using specific cameras by placing visual markers on them. These visual markers can be reflectors, LEDs (Light Emitting Diodes), infrared markers, etc. The camera captures changes in these markers and determines the object's position and orientation based on the marker's location and movement.
[0032] (2) Rigid body: This is an important concept in physics and mechanics, referring to an object whose relative positions between its internal points remain unchanged under the action of external forces. In other words, the shape and size of a rigid body do not change during motion.
[0033] (3) The global coordinate system, also known as the world coordinate system, is a fixed and unified reference coordinate system. In this coordinate system, the position and orientation of all objects are represented based on a fixed reference point (usually the origin) and fixed axes (usually the X, Y, and Z axes).
[0034] (4) A local coordinate system, also called an object coordinate system or a relative coordinate system, is a coordinate system set relative to a specific object or reference point. The origin and axes of a local coordinate system are defined relative to that object, so they will change as the object moves.
[0035] The global and local coordinate systems are connected through coordinate transformations, such as translation and rotation. For example, in a 3D graphic, all objects in the scene might be positioned based on a global coordinate system, while each object (such as a cube) has its own local coordinate system. Suppose the origin of the cube's local coordinate system is the cube's center, and its faces are parallel to the coordinate axes. In this case, the position of each vertex of the cube is defined relative to its local coordinate system. However, in the scene, these local coordinates are transformed into global coordinates based on the cube's position and rotation in the global coordinate system.
[0036] Based on the problems in related technologies, this specification provides a static state detection method that can detect static states of objects in different scenarios without introducing additional hardware such as an IMU. It can use the three-dimensional position information of multiple visual markers associated with the object to be detected within a preset measurement time to make a comprehensive judgment, thereby accurately determining whether the object to be detected is in a static state.
[0037] The static state detection method provided in the embodiments of this specification can be executed by an electronic device, including but not limited to physical servers, virtual servers, server clusters, smartphones / mobile phones, tablet computers, personal digital assistants (PDAs), laptop computers, desktop computers, media content players, video game consoles / systems, virtual reality systems, augmented reality systems, wearable devices (e.g., watches, glasses, gloves, headwear, or any other type of device).
[0038] In some embodiments, multiple visual markers can be set on the object to be detected, or see [link to relevant documentation]. Figure 1 Multiple visual markers 30 are set on a rigid body 20, and the rigid body 20 is further placed on the object to be detected 10. The object to be detected 10 can be any object in different scenarios that requires static state detection. For example, in a virtual shooting scenario, the object to be detected 10 can be at least one of a camera, a photographed object, or a photographed person in the virtual shooting scenario; in an automated production scenario, the object to be detected 10 can be a machine or workpiece in certain processes; and in an autonomous driving and traffic safety scenario, the object to be detected 10 includes vehicles and stationary objects on the road (such as accident scenes, obstacles, illegally parked vehicles, etc.). The visual markers 30 can be reflective markers (such as infrared reflective points) or graphic markers (such as QR codes or geometrically distinctive marks such as circles or triangles). The three-dimensional geometric relationships of the multiple visual markers 30 are known and fixed; for example, their relative distances and spatial positions remain unchanged.
[0039] Therefore, optical positioning of the object 10 to be detected can be performed based on multiple visual markers 30, thereby obtaining the three-dimensional position information of each visual marker 30 at each measurement time and the pose information of the object 10 to be detected. Specifically, at each measurement time, please refer to... Figure 1Multiple optical sensors 40 (such as infrared cameras or ordinary RGB cameras) capture images of multiple visual markers 30 from different perspectives, obtaining the 2D pixel coordinates of each visual marker 30 in the image plane. Utilizing multi-view geometry principles, the 3D position information of each visual marker 30 is calculated using triangulation based on the calibration parameters (intrinsic and extrinsic parameters) of the optical sensors 40 and the 2D coordinates of the multiple optical sensors 40's perspectives. Based on the acquired 3D position information of the multiple visual markers 30, the 3D position and rotation information (i.e., pose information) of the center of the rigid body 20 are solved using the Perspective-n-Point (PnP) algorithm or other rigid body transformation estimation methods, based on the known geometric relationships of the rigid body 20. If the visual markers 30 are fixed on the object 10 to be detected, then the pose of the rigid body 20 is the pose of the object 10 to be detected.
[0040] Please see Figure 2 The diagram illustrates a flowchart of a static state detection method, which can be executed by an electronic device. The method includes:
[0041] In S101, the three-dimensional position information of multiple visual markers associated with the object to be detected at each measurement time is obtained; the three-dimensional position information is obtained based on optical positioning of multiple visual markers.
[0042] As described above, optical positioning technology can be used to obtain the three-dimensional position information of multiple visual markers associated with the object to be detected at various measurement times, thereby providing a data basis for subsequent static state judgment.
[0043] In S102, for each visual marker, multiple three-dimensional position information of the visual marker within a preset measurement time is used to detect whether the visual marker has moved within the preset measurement time, and the detection result is obtained.
[0044] In this step, by utilizing multiple three-dimensional positional information of each visual marker within a preset measurement time, the minute movements of each visual marker can be accurately detected, avoiding overall judgment errors. Through multi-time data analysis, the stability and robustness of the detection are improved.
[0045] In S103, the detection result of the static state of the object to be detected is determined based on the detection results of multiple visual markers at preset measurement times.
[0046] In this step, based on the detection results of all visual markers within a preset measurement time, the overall static state of the object under test can be comprehensively determined. By comprehensively judging multiple visual markers, the accuracy and reliability of the detection results are ensured, effectively addressing data anomalies caused by occlusion or noise of some visual markers, thus improving robustness.
[0047] In some embodiments, before performing motion detection using multiple three-dimensional positional information of various visual markers within a preset measurement time, at least one of the following preprocessing procedures can be performed to further improve detection accuracy.
[0048] In the first possible implementation, the aforementioned three-dimensional position information includes the global three-dimensional position information of the visual markers in the global coordinate system. At each measurement moment, by optically positioning multiple visual markers, the electronic device can acquire the global three-dimensional position information of the multiple visual markers associated with the object to be detected in the global coordinate system, as well as the pose information of the object to be detected. Specifically, the pose information of the object to be detected at each measurement moment describes the position and orientation of the object to be detected in the global coordinate system at that measurement moment, and can comprehensively reflect the state of the object to be detected in three-dimensional space. For example, assuming that multiple visual markers are set on a rigid body, the local coordinate system in which the multiple visual markers are located can be a coordinate system with the center of the rigid body as the origin. The pose information of the object to be detected can be determined by the relationship between the local three-dimensional position information of the multiple visual markers in the local coordinate system and the global three-dimensional position information in the global coordinate system. Conversely, if the pose information of the object to be detected and the global three-dimensional position information of the multiple visual markers in the global coordinate system are known, the local three-dimensional position information of the multiple visual markers in the local coordinate system can be calculated in reverse.
[0049] As mentioned earlier, the three-dimensional geometric relationships of multiple visual markers are known and fixed; for example, their relative distances and spatial positions remain unchanged. In other words, within the local coordinate system of the multiple visual markers, the local three-dimensional position information of each visual marker is fixed and can be pre-stored. Therefore, for each visual marker, the electronic device can convert the global three-dimensional position information of the visual marker at the same measurement time into local three-dimensional position information in the local coordinate system based on the pose information of the object to be detected at the same measurement time. If the distance between the converted local three-dimensional position information and the pre-stored local three-dimensional position information is greater than a preset distance threshold (e.g., greater than 0.5cm, 1cm, or 1.1cm, etc.), it indicates that the visual marker data at that measurement time is abnormal (e.g., noise, error, or occlusion), and the data needs to be discarded. In this case, the electronic device discards the global three-dimensional position information of that visual marker at that measurement time.
[0050] If the distance between the local 3D position information and the pre-stored local 3D position information is less than or equal to a preset distance threshold, the global 3D position information of the visual marker at that measurement time can be stored in a buffer. This buffer can then be used to retrieve multiple 3D position information within a preset measurement duration for motion detection. Each visual marker has a corresponding buffer used to cache the global 3D position information that has passed the above verification process.
[0051] For example, for n visual markers, at time t1, after the above processing, it is found that the distance between the converted local 3D position information of k visual markers and the pre-stored local 3D position information is greater than a preset threshold. Then, the 3D position information of these k visual markers at time t1 is discarded, where k≤n.
[0052] In this embodiment, considering that the three-dimensional position information of the visual marker may be noisy due to the following reasons during the actual optical positioning process: (1) sensor accuracy limitations, leading to positioning errors; (2) partial occlusion of the visual marker, leading to inaccurate recognition; (3) environmental interference, such as changes in light. By combining the global coordinate system and the local coordinate system, the global three-dimensional position information of the visual marker is transformed into the local coordinate system, and the transformation result is compared with the pre-stored local information, abnormal data can be quickly identified and eliminated, thereby avoiding the influence of erroneous data on the detection results and effectively improving the accuracy and robustness of static state detection.
[0053] In the second possible implementation, for each visual marker, attribute information of the visual marker at each measurement moment within a preset measurement time can be obtained based on the optical positioning process. If the attribute information at any measurement moment does not meet the preset attribute conditions, the electronic device can discard the three-dimensional position information of the visual marker at that measurement moment. In this embodiment, by filtering the attribute information of the visual marker obtained during the optical positioning process, abnormal or unreliable three-dimensional position information of the visual marker can be effectively filtered out, thereby improving the accuracy and robustness of the static state detection of the object to be detected.
[0054] As an example, if n visual markers are set on a rigid body, and the three-dimensional position information and attribute information corresponding to the n visual markers are obtained through an optical positioning process at a certain measurement moment, if the attribute information of a certain visual marker at the measurement moment does not meet the preset attribute conditions, the three-dimensional position information of the visual marker at the measurement moment is deleted; if the attribute information of other visual markers at the measurement moment meets the preset attribute conditions, the three-dimensional position information of other visual markers at the measurement moment is retained.
[0055] It is understood that this embodiment does not impose any restrictions on the attribute information of the visual markers, and specific settings can be made according to the actual application scenario. This embodiment does not impose any restrictions on this.
[0056] In one example, the attribute information of a visual tag includes the confidence level of its 3D position information. Some optical positioning systems can provide a confidence score (e.g., a value between 0 and 1) when acquiring the 3D position information of a visual tag, indicating the reliability of the data. A preset attribute condition could be that the confidence level must be higher than a certain threshold (e.g., 0.8). If the confidence level of a visual tag at a measurement moment is 0.6, which is lower than the preset threshold of 0.8, the data is considered unreliable and discarded.
[0057] In another example, the attribute information of a visual marker includes its sharpness or visibility. Some optical positioning systems determine the sharpness or visibility of a visual marker (e.g., if the marker is occluded or blurred due to insufficient light). Preset attribute conditions can be that the visual marker must meet certain sharpness or visibility standards. If, at a certain measurement moment, the visual marker is partially occluded and can only be partially recognized by the optical positioning system, resulting in a sharpness below the preset standard, then the 3D position information at that measurement moment is discarded.
[0058] In another example, the attribute information of a visual marker includes its size or outline shape, which some optical positioning systems provide. Preset attribute conditions may include requiring at least one of the visual marker's size and shape to be within a preset range. If the outline detection result of the visual marker at a measurement moment shows deformation (e.g., becoming too small, too large, or irregular), the three-dimensional position information at that measurement moment is discarded.
[0059] In the third possible implementation, where the detection of the stationary state of certain objects requires high real-time performance, the electronic device can compare the measurement time of the last 3D position information within a preset measurement period with the current time. If the time difference between the last measurement time and the current time is less than a preset time difference, the step of detecting whether the visual marker has moved within the preset measurement period is executed. By determining the time difference between the "measurement time of the last 3D position information" (i.e., the measurement time most recent to the current time) and the "current time," the data is ensured to be up-to-date, avoiding the impact of delayed processing of outdated data on the timeliness of stationary state detection and meeting the needs of high real-time application scenarios. If the time difference exceeds a preset threshold, it indicates that the data is outdated, and the electronic device can directly skip the movement detection step, thereby saving computing resources. The electronic device can also instruct the re-acquisition of the latest 3D position information for movement detection.
[0060] In this embodiment, by introducing a real-time judgment mechanism for time difference, the static state detection process can be made more efficient and accurate in scenarios with high real-time requirements. Provided the time difference condition is met, the electronic device can quickly execute the "visual marker movement detection" step, promptly outputting the static state result of the object to be detected, ensuring that the detection process remains synchronized with changes in the external environment, and improving real-time response capabilities.
[0061] In the fourth possible implementation, sufficient 3D position information ensures the accuracy of motion detection. For each visual marker, the electronic device can compare the number of multiple 3D position information points within a preset measurement time with the expected number corresponding to the preset measurement time. If the ratio between the number of multiple 3D position information points within the preset measurement time and the expected number corresponding to the preset measurement time is greater than a preset ratio, the step of detecting whether the visual marker has moved within the preset measurement time is executed. This embodiment ensures that the data volume is sufficiently dense and suitable for motion detection by comparing the difference between the actual number of valid 3D position information points within the preset time and the expected number, avoiding misjudgments or unreliable results due to insufficient data. Furthermore, a sufficient number of 3D position information points can filter out random errors or outliers during data acquisition, making the detection results more accurate and stable. If the actual number of valid 3D position information points within the preset time is insufficient, the electronic device can skip subsequent motion detection steps to avoid performing meaningless calculations and save computing resources. The electronic device can also instruct the acquisition of more 3D position information points at more measurement times for motion detection.
[0062] In this embodiment, by introducing a ratio judgment mechanism between the number of three-dimensional position information and the expected number, it can be ensured that the amount of data used for motion detection is sufficient, thereby improving the accuracy and robustness of the detection results.
[0063] In the fifth possible implementation, since the 3D position information of the visual markers is obtained through optical positioning, the optical sensor may experience slight fluctuations in the 2D position data during detection due to factors such as ambient light interference, hardware precision limitations, or minor jitter of the visual markers. These fluctuations are further amplified during triangulation, leading to high-frequency noise or fluctuations in the 3D position information. Therefore, the electronic device can perform low-pass filtering on multiple 3D position information of each visual marker within a preset measurement time. Low-pass filtering can effectively suppress these high-frequency noises and smooth out fluctuations in the data. After low-pass filtering, the 3D position information of the visual markers is more stable, reducing misjudgments caused by noise. The filtered data is closer to the true state, helping the device to accurately determine whether the visual markers and the object to be detected are stationary.
[0064] This embodiment effectively reduces fluctuations in three-dimensional position information caused by minute jitters during optical positioning by performing low-pass filtering on multiple three-dimensional position information of visual markers within a preset measurement time, thereby improving the accuracy and stability of static state detection.
[0065] It should be noted that the three-dimensional position information in the second to fifth implementation methods mentioned above is all global three-dimensional position information in the global coordinate system.
[0066] It is understandable that in real-world scenarios, one of the five implementation methods mentioned above can be selected for preprocessing, or any combination of two or more methods can be used. This embodiment does not impose any restrictions on this.
[0067] In some embodiments, after performing at least one of the preprocessing steps described above, the electronic device can perform motion detection based on multiple valid three-dimensional position information of each visual marker within a preset measurement time. For example, the electronic device can statistically process the multiple three-dimensional position information of the visual marker within the preset measurement time to obtain reference three-dimensional position information. Then, based on the differences between the various three-dimensional position information of the visual marker within the preset measurement time and the reference three-dimensional position information, it can determine whether the visual marker has moved within the preset measurement time. In this embodiment, through statistical processing (e.g., calculating the mean, weighted average, etc.), a stable reference position of the visual marker within the measurement time can be extracted, effectively eliminating errors caused by noise, interference, or transient anomalies at a single measurement point. By comparing the differences between the three-dimensional position information and the reference three-dimensional position information, it is possible to more accurately determine whether the visual marker has actually moved, reducing the possibility of misjudgment.
[0068] The reference 3D position information can be the average of multiple 3D position information of the visual marker within a preset measurement period. This average value can characterize the center point of the visual marker within the preset measurement period. This method of calculating the average value effectively balances the fluctuations during multiple measurements, extracts the stable position of the visual marker within this period, and can therefore be regarded as the center point coordinates of the visual marker. For the visual marker, the average of multiple 3D position information within the preset measurement period reflects the static center position of the visual marker within this measurement period.
[0069] It is understandable that, besides determining the static center position of the visual marker within this time period by calculating the average, other statistical methods can also be used to determine the static center position of the visual marker within this measurement time period, and this embodiment does not impose any limitations on this. In one example, the reference 3D position information is the median or weighted average of multiple 3D position information of the visual marker within the preset measurement time period. In another example, the smallest bounding box containing multiple 3D position information of the visual marker within the preset measurement time period can be found in 3D space, and the position information of the geometric center point of this bounding box can be used as the reference 3D position information.
[0070] For example, after obtaining the reference 3D position information, the electronic device can calculate the distance between the visual marker and the reference 3D position information based on each 3D position information of the visual marker within a preset measurement time period; then, based on the distances corresponding to the multiple 3D position information of the visual marker within the preset measurement time period, it can determine whether the visual marker has moved within the preset measurement time period. This embodiment, by calculating the distance between the 3D position information and the reference 3D position information at each moment, can dynamically track the movement trajectory of the visual marker, quickly identify the movement of the visual marker, and provide relatively accurate movement detection results.
[0071] In one possible implementation, the electronic device can calculate the standard deviation based on the distances corresponding to multiple three-dimensional positional information of the visual marker within a preset measurement time. If the standard deviation is greater than a preset threshold, it is determined that the visual marker has moved within the preset measurement time; otherwise, it is determined that the visual marker is stationary within the preset measurement time. In this embodiment, the standard deviation reflects the dispersion of the data. The larger the standard deviation, the greater the change in the three-dimensional positional information of the visual marker. Therefore, a standard deviation greater than the preset threshold can determine that the visual marker has moved. This method is particularly suitable for detecting continuous small-amplitude changes and can smoothly cope with short-term noise or sudden fluctuations.
[0072] In another possible implementation, if at least X distances among the distances corresponding to multiple 3D position information of the visual marker within a preset measurement time are all greater than a preset distance threshold, it is determined that the visual marker has moved within the preset measurement time; otherwise, it is determined that the visual marker is stationary within the preset measurement time. X can be specifically set according to the actual application scenario, and this embodiment does not impose any restrictions on it. In this embodiment, if the distance at multiple measurement moments exceeds the preset distance threshold during the measurement period, it indicates that the visual marker has undergone significant displacement. This method can filter out occasional small fluctuations, focus on larger motion changes, and reduce misjudgments.
[0073] As an example, if n visual markers are set on a rigid body, and n sets of 3D position information are collected at m time points within a preset measurement period (i.e., for each visual marker, there is a corresponding set of 3D position information), then for each visual marker, the mean of the m 3D position information collected at the m time points can be calculated as the reference 3D position information for that visual marker. Then, the distance between the m 3D position information of that visual marker and the reference 3D position information is calculated, and the standard deviation of the distance is calculated. If the standard deviation is greater than a preset threshold, it can be determined that the visual marker has moved significantly; otherwise, the visual marker is considered stationary. Alternatively, the distance between each of the m 3D position information of the visual marker and the reference 3D position information can be calculated, resulting in m distance values. If x of the m distance values are greater than a preset distance threshold, then the visual marker is considered to have moved significantly; otherwise, the visual marker is considered to be stationary. In this example, x can be half or more of m, but is not limited to this.
[0074] Of course, other methods can also be used for motion detection, such as calculating variance, calculating root mean square error, etc. This embodiment does not impose any restrictions on this.
[0075] In some embodiments, after obtaining the detection results of multiple visual markers for a preset measurement time, a comprehensive judgment can be made based on the detection results of the multiple visual markers for the preset measurement time to determine the detection result of the static state of the object to be detected. This embodiment, by integrating the judgments of multiple visual markers, results in a more stable overall detection result, effectively avoiding the erroneous influence of a single marker, thereby improving reliability.
[0076] In one possible implementation, if the number of visual markers indicating a stationary state is greater than a preset number, the electronic device can determine that the object to be detected is stationary; otherwise, the electronic device can determine that the object to be detected is non-stationary. The preset number can be specifically set according to the actual application scenario, and this embodiment does not impose any restrictions on it.
[0077] In another possible implementation, if the number of visual markers indicating that the detection results are stationary is greater than the number of visual markers indicating that the detection results are moving, the electronic device can determine that the object to be detected is stationary; otherwise, the electronic device can determine that the object to be detected is non-stationary.
[0078] Alternatively, the two implementation methods mentioned above can be combined for judgment, and this embodiment does not impose any restrictions on this.
[0079] In one exemplary embodiment, the object to be detected can be a camera in a virtual shooting scene, with a rigid body placed on the camera and multiple visual markers set on the rigid body. See also... Figure 3 The static state detection process consists of three steps:
[0080] The first process: matching and caching the three-dimensional position information of visual markers at various times.
[0081] For each rigid body, there is a custom local coordinate system, referred to as the B-frame, and the pose of the rigid body (called T-frame) is also defined. B W The B-frame is used to describe the position and orientation of a rigid body in the global coordinate system (W-frame), and can comprehensively reflect the state of the rigid body in three-dimensional space. For optical positioning systems, the pose of a rigid body is determined by the local three-dimensional position P of each visual marker on the rigid body in the B-frame. B and the global three-dimensional position P under the W system W It is derived from the correspondence between them, and the correspondence satisfies The pose of the rigid body can be considered to be consistent with the pose of the object to be detected in the previous text.
[0082] The initial matching results output by the optical positioning system may contain mismatches. If subsequent stationary state detection is performed based on these mismatched points, it may lead to erroneous detection results. The optical positioning system can perform optical positioning at each measurement moment, thereby outputting the following information: the pose information of the rigid body and the global 3D position of each visual marker. Therefore, for each visual marker, the global 3D position P at the same measurement moment can be determined based on the pose information of the rigid body at the same measurement moment. W Transferring to the B system, we obtain P. B1 and P B1 With pre-recorded P B The two are compared, and if the distance between them is greater than 1 cm, the global three-dimensional information of the visual mark at that measurement time is considered invalid.
[0083] When the distance between the two is less than or equal to 1 cm, the global 3D information of the visual marker at that measurement time is placed into the buffer. Each visual marker can correspond to one buffer. At the same time, the number of global 3D information cached in the buffer corresponding to the visual marker is judged. If it exceeds a certain number, it is removed to reduce memory usage. For example, if the buffer can cache 10 global 3D information, and the current buffer caches 10 global 3D information from 10 measurement times T1 to T10, and there is global 3D information at time T11 that also meets the above condition and can be stored in the buffer, then the global 3D information at time T1 is removed so that the global 3D information at time T11 can be cached.
[0084] Global 3D information measured at the same time points in each cache can be read based on a preset measurement duration for subsequent processing. For example, if the preset measurement duration is 10 time points, global 3D information measured at times T1 to T10 can be read from each cache. After the above processing, the global 3D information of some visual markers at certain times may be invalid. For example, for a certain visual marker, the global 3D information measured at time T1 is invalid. In this case, the global 3D information measured at times T2 to T10 can be read from the cache of that visual marker for subsequent processing.
[0085] The second process: preprocessing. The main function of preprocessing is to check whether the 3D position information of each visual marker within a preset time period (preset time window) meets the verification requirements and to perform smoothing operations on the 3D position information of the visual markers.
[0086] (1) Attribute verification: The attribute information of each visual marker at each measurement time can be obtained. If the attribute information of the visual marker at any measurement time does not meet the preset attribute conditions, the three-dimensional position information of the visual marker at that measurement time is discarded.
[0087] (2) Time Continuity Verification: Since the visual markers of a rigid body may be occluded when it moves within the field or interacts with its surroundings, to ensure the effectiveness of stationary detection, it is necessary to guarantee that the effective time percentage within the time window exceeds a certain value and that stationary detection is real-time. That is, it must satisfy B[-1].tt c <δt and Where B[-1].t represents the time of the last frame of the time window, t c Represents the current time, δt represents the tolerable time deviation, and C v C represents the number of valid 3D positional information representations of visual markers within a time window. a This represents the total number of time windows, and f represents the minimum ratio.
[0088] (3) Since visual markers are obtained through optical positioning, slight jitter in the two-dimensional position of the visual markers can cause fluctuations in the three-dimensional position, which may affect static verification. In order to reduce these fluctuations, low-pass filtering is applied to the multiple three-dimensional position information of each visual marker within the time window.
[0089] The third process: static detection process.
[0090] (1) Detection of each visual marker: The detection of visual markers mainly involves observing the changes in their positions within a time window, that is, the statistical results of their deviation from the three-dimensional center point. Let P be the three-dimensional position information of each visual marker within the time window. W(i) then the three-dimensional center point is the average value of the three-dimensional position information of the visual marker within the time window, that is, the three-dimensional center point. n represents the number of 3D positional information points of the visual marker within the time window, and i represents the i-th 3D positional information point of the visual marker. Then, the distance from each 3D positional information point of the visual marker to the center point is... Standard deviation is If the standard deviation is greater than the set threshold, the visual mark is considered to have moved (referred to as a moving visual mark); otherwise, the visual mark is considered to be stationary (referred to as a stationary visual mark).
[0091] (2) Since there are multiple visual markers on the rigid body, if the rigid body rotates or translates, the position of the visual markers will also change. Therefore, it is only necessary to count the static detection of the visual markers on the rigid body. For example, if the number of static visual markers is greater than the preset number and greater than the number of moving visual markers, the rigid body is considered to be in a static state, that is, the camera on which the rigid body is placed is also in a static state.
[0092] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.
[0093] In some embodiments, this specification also provides an electronic device, including: a processor; and a memory for storing processor-executable instructions; wherein the processor implements the method described in any one of the above embodiments by executing the executable instructions.
[0094] Figure 4 This is a schematic structural diagram of a device provided in an exemplary embodiment. Please refer to... Figure 4 At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, memory 408, and non-volatile memory 410, and may also include other hardware required for its functions. One or more embodiments of this specification can be implemented in software, for example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into memory 408 and then runs it. Of course, in addition to software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0095] Please refer to Figure 5 The static state detection device can be applied to, for example, Figure 4The device shown implements the technical solution described in this specification. The static state detection device may include:
[0096] The location information acquisition module 501 is used to acquire the three-dimensional location information of multiple visual markers associated with the object to be detected at various measurement times; the three-dimensional location information is obtained based on optical positioning of the multiple visual markers;
[0097] The visual marker detection module 502 is used to detect whether the visual marker has moved within the preset measurement time using multiple three-dimensional position information of the visual marker within the preset measurement time, and to obtain the detection result;
[0098] The object detection module 503 is used to determine the static state detection result of the object to be detected based on the detection results of the multiple visual markers at the preset measurement time.
[0099] In some embodiments, the visual marker detection module 502 is specifically used to perform statistical processing on multiple three-dimensional position information of the visual marker within a preset measurement time to obtain reference three-dimensional position information; and to determine whether the visual marker has moved within the preset measurement time based on the difference between each of the three-dimensional position information of the visual marker within the preset measurement time and the reference three-dimensional position information.
[0100] In some embodiments, the reference three-dimensional position information is the average value of multiple three-dimensional position information of the visual marker within a preset measurement time.
[0101] In some embodiments, the visual marker detection module 502 is specifically used to calculate the distance between the visual marker and the reference three-dimensional position information based on each of the three-dimensional position information of the visual marker within a preset measurement time; and to determine whether the visual marker has moved within the preset measurement time based on the distances corresponding to the multiple three-dimensional position information of the visual marker within the preset measurement time.
[0102] In some embodiments, the visual marker detection module 502 is specifically used to calculate the standard deviation based on the distances corresponding to the multiple three-dimensional position information of the visual marker within a preset measurement time; if the standard deviation is greater than a preset threshold, it is determined that the visual marker has moved within the preset measurement time; otherwise, it is determined that the visual marker is stationary within the preset measurement time.
[0103] In some embodiments, the object detection module 503 is specifically used to determine that the object to be detected is in a stationary state if the number of visual markers indicating that the detection result is in a stationary state is greater than a preset number, and / or the number of visual markers indicating that the detection result is in a stationary state is greater than the number of visual markers indicating that the detection result is in a moving state; otherwise, it determines that the object to be detected is in a non-stationary state.
[0104] In some embodiments, the three-dimensional position information includes the global three-dimensional position information of the visual marker in the global coordinate system;
[0105] The device further includes a preprocessing module for acquiring the pose information of the object to be detected at various measurement times, wherein the pose information is obtained based on optical positioning of the plurality of visual markers; for each visual marker, based on the pose information at the same measurement time, the global three-dimensional position information is converted into local three-dimensional position information in the local coordinate system where the plurality of visual markers are located; if the distance between the local three-dimensional position information and the pre-stored local three-dimensional position information is greater than a preset distance threshold, the global three-dimensional position information of the visual marker at that measurement time is discarded.
[0106] In some embodiments, the apparatus further includes a preprocessing module for obtaining attribute information of each visual marker at each measurement moment within a preset measurement duration based on the optical positioning process; if the attribute information at any measurement moment does not meet the preset attribute conditions, the three-dimensional position information of the visual marker at that measurement moment is discarded.
[0107] In some embodiments, the apparatus further includes a preprocessing module, configured to, for each of the visual markers, if the time difference between the measurement time of the last three-dimensional position information in the preset measurement duration and the current time is less than a preset time difference, execute a step of detecting whether the visual marker has moved within the preset measurement duration; and / or, for each of the visual markers, if the ratio between the number of multiple three-dimensional position information in the preset measurement duration and the expected number corresponding to the preset measurement duration is greater than a preset ratio, execute a step of detecting whether the visual marker has moved within the preset measurement duration.
[0108] In some embodiments, the apparatus further includes a preprocessing module for performing low-pass filtering on multiple three-dimensional position information of each of the visual markers within a preset measurement time.
[0109] In some embodiments, the plurality of visual markers are disposed on the object to be detected; or, the plurality of visual markers are disposed on a rigid body, the rigid body being placed on the object to be detected.
[0110] In some embodiments, the object to be detected includes at least one of a camera in a virtual shooting scene, an object being shot in a virtual shooting scene, and a person being shot in a virtual shooting scene.
[0111] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0112] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.
[0113] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0114] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.
[0115] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit the scope of one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this specification should be included within the protection scope of one or more embodiments of this specification.
Claims
1. A method for detecting a stationary state, the method comprising: Based on a preset measurement duration, read the three-dimensional position information of the same multiple measurement times in the buffer corresponding to multiple visual markers associated with the object to be detected; The three-dimensional position information is obtained based on optical positioning of the plurality of visual markers; the plurality of visual markers are disposed on the object to be detected; or, the plurality of visual markers are disposed on a rigid body, and the rigid body is placed on the object to be detected. For each visual marker, if the time difference between the measurement time of the last three-dimensional position information in the preset measurement time and the current time is less than the preset time difference, and the ratio between the number of multiple three-dimensional position information in the preset measurement time and the expected number corresponding to the preset measurement time is greater than the preset ratio, the visual marker is detected to have moved within the preset measurement time using the multiple three-dimensional position information of the visual marker within the preset measurement time, and a detection result is obtained. Based on the detection results of the multiple visual markers at the preset measurement time, the static state detection result of the object to be detected is determined.
2. The method according to claim 1, wherein detecting whether the visual marker has moved within the preset measurement time period using multiple three-dimensional position information of the visual marker within the preset measurement time period includes: Statistical processing is performed on multiple three-dimensional position information of the visual markers within a preset measurement time to obtain reference three-dimensional position information; Based on the differences between the three-dimensional position information of the visual marker and the reference three-dimensional position information within a preset measurement time, it is determined whether the visual marker has moved within the preset measurement time.
3. The method according to claim 2, wherein the reference three-dimensional position information is the average value of multiple three-dimensional position information of the visual marker within a preset measurement time.
4. The method according to claim 2, wherein determining whether the visual marker has moved within the preset measurement time period based on the difference between the three-dimensional position information of the visual marker and the reference three-dimensional position information within a preset measurement time period comprises: Based on the visual marker's three-dimensional position information and the reference three-dimensional position information within a preset measurement time, the distance between the two is calculated; Based on the distances corresponding to the multiple three-dimensional position information of the visual marker within a preset measurement time, it is determined whether the visual marker has moved within the preset measurement time.
5. The method according to claim 4, wherein determining whether the visual marker has moved within the preset measurement time period based on the distances corresponding to the multiple three-dimensional position information of the visual marker within the preset measurement time period comprises: The standard deviation is calculated based on the distances corresponding to the multiple three-dimensional position information of the visual markers within a preset measurement time. If the standard deviation is greater than a preset threshold, it is determined that the visual marker has moved within the preset measurement time. Otherwise, it is determined that the visual marker remains stationary during the preset measurement period.
6. The method according to claim 1, wherein determining the static state detection result of the object to be detected based on the detection results of the plurality of visual markers at the preset measurement time includes: If the number of visual markers indicating that the detection results are in a stationary state is greater than a preset number, and / or the number of visual markers indicating that the detection results are in a stationary state is greater than the number of visual markers indicating that the detection results are in a moving state, the object to be detected is determined to be in a stationary state. Otherwise, it is determined that the object to be detected is in a non-stationary state.
7. The method according to claim 1, wherein the three-dimensional position information includes the global three-dimensional position information of the visual marker in the global coordinate system; Before detecting whether a visual marker has moved within a preset measurement time period using multiple three-dimensional position information of the visual marker for each of the visual markers, the method further includes: The pose information of the object to be detected at each measurement time is obtained, and the pose information is obtained based on optical positioning of the multiple visual markers; For each of the aforementioned visual markers, based on the pose information at the same measurement time, the global three-dimensional position information is converted into local three-dimensional position information in the local coordinate system where the multiple visual markers are located; If the distance between the local 3D position information and the pre-stored local 3D position information is greater than a preset distance threshold, the global 3D position information of the visual marker at that measurement moment is discarded.
8. The method according to claim 1, further comprising, before detecting whether a visual marker has moved within a preset measurement time period using multiple three-dimensional position information of the visual marker within the preset measurement time period for each of the visual markers: For each of the aforementioned visual markers, attribute information of the visual markers at each measurement moment within the preset measurement duration is obtained based on the optical positioning process; If the attribute information at any measurement time does not meet the preset attribute conditions, the three-dimensional position information of the visual marker at that measurement time is discarded.
9. The method according to any one of claims 1 to 8, further comprising, before detecting whether the visual marker has moved within the preset measurement time using multiple three-dimensional position information of the visual marker within the preset measurement time for each of the visual markers: Low-pass filtering is applied to the multiple three-dimensional position information of each visual marker within a preset measurement time.
10. The method according to any one of claims 1 to 8, wherein the object to be detected includes at least one of a camera in a virtual shooting scene, an object being shot in a virtual shooting scene, and a person being shot in a virtual shooting scene.
11. An electronic device, comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as described in any one of claims 1-10 by executing the executable instructions.
12. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-10.
13. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-10.
Citation Information
Patent Citations
System for controlling an industrial tool by calculating the position thereof relative to an object moving on an assembly line
EP3187947A1
Apparatus and method for measuring runout
JP2011080962A
Position determination device, position determination system, position determination method and program
JP2014020986A
Image-based motion detection method
US20230215022A1