A measurement analysis method based on video measurement
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHERN JIANGSU PEOPLES HOSPITAL
- Filing Date
- 2026-05-13
- Publication Date
- 2026-06-09
AI Technical Summary
Existing technologies require the acquisition of a large number of complete static and dynamic images when measuring the structural state of an object, which makes the measurement of the object difficult. In particular, for some objects, it is difficult to acquire complete images and a large amount of object activity data is required.
By acquiring video of the object under test and the object's change scale, and using the normalized coordinates in the image coordinate system of the key points, an origin coordinate system is constructed to determine the physical coordinates of the marker points, and the activity characteristics of the marker points are output, thus realizing the measurement of the object's structural state.
It reduces the need for image data of the object under test, improves measurement efficiency, and enables quick and accurate determination of the object's structural state.
Smart Images

Figure CN122170838A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of photogrammetry technology, and more specifically to a measurement and analysis method based on video measurement. Background Technology
[0002] Existing technologies for measuring the structural state of an object (such as whether the object's structure is loose) often require acquiring numerous complete still images of the object under test. These still images are used as a calibration basis to calibrate the object. Then, moving images of the object are acquired, and a neural network model is used to analyze the object's structure and determine its structural state. However, for some objects, acquiring complete images is difficult, and using a neural network model requires a large amount of object movement data, making the measurement challenging. Summary of the Invention
[0003] This application provides a measurement and analysis method based on video measurement, which can obtain the structural state of the object under test through a small amount of video data and has high measurement efficiency.
[0004] This application provides a measurement and analysis method based on video measurement, which includes: Obtain a video of the object under test and an object change scale of the video. The video includes keypoints, which include reference points and marker points. The object change scale indicates the ratio between the physical length of the object and its pixel length. Obtain the normalized coordinates of the keypoints in the image coordinate system across multiple video frames. Construct an origin coordinate system using the physical coordinates of the reference points as the origin, and determine the physical coordinates of the marker points in the origin coordinate system. The physical coordinates of the keypoints in the image coordinate system are determined based on the normalized coordinates of the keypoints in the image coordinate system, the pixel count of the video, and the object change scale. Output the activity features of the marker points, which are determined based on the physical coordinates of the marker points in the origin coordinate system corresponding to the multiple video frames.
[0005] Compared to existing technologies that require a large amount of data for model training and require the test object to have a specific shape for calibration, the embodiments of this application can measure the structural state of the test object by collecting reference points and marker points of the object, and require less image data of the test object, without the need to collect a large number of moving images of the test object, which greatly improves the efficiency of the measurement of the test object. Attached Figure Description
[0006] The objectives, features, and advantages of the embodiments of this application will become readily understood by referring to the accompanying drawings and the detailed description of the embodiments. Wherein: Figure 1This is a flowchart illustrating the measurement and analysis method based on video measurement in an embodiment of this application; Figure 2 Another flowchart illustrating the video-based measurement and analysis method provided in this application embodiment; Figure 3 A schematic diagram of a video frame with key points marked; Figure 4 This is a schematic diagram of the temporal coordinates of key points in the image coordinate system; Figure 5 A schematic diagram of the normalized coordinates of key points in the image coordinate system of the measurement and analysis method based on video measurement provided in the embodiments of this application; Figure 6 Another flowchart illustrating the video-based measurement and analysis method provided in this application embodiment; Figure 7 This is a schematic diagram of the time-series coordinates of key points after the origin coordinate transformation; Figure 8 This is a schematic diagram of the time series fitting results of the key point x-coordinates after the transformation to the origin coordinate system. Figure 9 This is a schematic diagram of the time series fitting results of the key point x-coordinates after filtering out extreme points and transforming them into the origin coordinate system. Figure 10 for Figure 9 A schematic diagram of the line connecting the extreme points shown.
[0007] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0008] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects (e.g., a first marker and a second marker are represented as different markers, and so on), and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The division of modules in the embodiments of this application is merely a logical division; in actual applications, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be omitted or not performed. Additionally, the shown or discussed mutual coupling or direct coupling or communication connection may be through some interface, indirect coupling between modules, or electrical or other similar forms of communication connection, none of which are limited in the embodiments of this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed among multiple circuit modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiments of this application.
[0009] Reference Figure 1 , Figure 1 This is a flowchart illustrating a video-based measurement and analysis method provided in an embodiment of this application. The method can be executed by a measurement device and can be applied to scenarios involving the measurement of an object's structural state. This method can measure the object's structural state by acquiring reference points and marker points, and requires relatively little object image data, eliminating the need to acquire a large number of object motion images, thus significantly improving the efficiency of object structural state measurement. The method includes steps 101-104: Step 101: Obtain the video of the object to be tested and the object change scale of the video.
[0010] The video includes key points. Key points include reference points and marker points. The object scale indicates the ratio between the physical length of the object and its pixel length. It can be understood that pixel length represents the total length of pixels occupied by the corresponding object in the image, such as the number of pixels it occupies along its length.
[0011] The video of an object can be a video of the object's activity. When the object is active, the relative positions of the reference point and the marker point change, and the structural state of the object can be determined based on the measured changes.
[0012] There can be multiple keypoints, meaning each video frame contains multiple keypoints, which together form a keypoint set K. The normalized coordinates of the keypoints in the image coordinate system can be obtained through... This is represented by P, where P represents a keypoint and i represents the identifier of the keypoint. This represents the x-coordinate of the i-th keypoint in the z-th video frame. This represents the ordinate of the i-th keypoint in the z-th video frame. An image coordinate system can refer to a coordinate system with one vertex of the video frame as its origin. For example, with the bottom left corner of the video frame as the origin, the horizontal axis extends to the right from the origin within the video frame, and the vertical axis extends upwards from the origin within the video frame.
[0013] Taking the human body as an example, for the examination of scoliosis of the upper limb spine, the key point set K can be represented as: ,in The left shoulder, right shoulder, left elbow, right elbow, left hip, and right hip are represented in that order.
[0014] In some embodiments, the reference point is located on the main body of the object, and the marker point is located at the active position of the object and / or on the main body of the object; the number of marker points can be multiple. For example, taking a robotic arm as the object being tested, the reference point is located on the base of the robotic arm, the first marker point is located on the link of the robotic arm, and the second marker point is located on the base. The base is the main body of the robotic arm, which is relatively stable and has a small range of motion, while the link is located in the active area and has a larger range of motion.
[0015] The object scaling scale can be pre-obtained or received from an external device via a communication interface. Alternatively, it can be determined by acquiring the physical length of a portion of the object and then calculating the object scaling scale based on the pixel length of that portion in the video.
[0016] For example, refer to Figure 2 The object change scale is obtained through steps 201 to 203 below.
[0017] Step 201: Obtain the physical length L of the preset marker of the object to be tested.
[0018] Preset markers can be objects appearing in the same video as the object. These markers may or may not be part of the object. For example, in a scenario measuring the structural state of a robotic arm, a preset marker could be an end of the arm (or a link), or it could be a basketball next to the arm. The physical length of the preset marker can be obtained by measuring with a ruler or by reading from the product manual, etc.
[0019] Taking the human body as an example, when measuring scoliosis, the pre-defined marker can be the arm. By using the arm as the pre-defined marker instead of the body's height, the requirements for human movement can be reduced, making data acquisition easier.
[0020] Step 202: Obtain the pixel length of the preset marker in the video. .
[0021] The pixel length of the preset marker in the video can be determined by the number of pixels occupied by the preset marker and the pixel size of the video.
[0022] It is understandable that the pixel length of the preset marker in the video includes the pixel length of the preset marker in the horizontal direction and the pixel length in the vertical direction of the video.
[0023] Step 203: Determine the object change scale based on the physical length and pixel length of the preset marker. .
[0024] In monocular close-range photogrammetry, the relative positions of the object and the camera remain roughly constant; that is, the camera position remains roughly constant, and the object moves within a fixed, small area. In other words, the camera position is fixed when shooting video. When the angle between the extension direction of the preset marker and the camera's optical axis is less than a preset angle, and the preset marker occupies a larger proportion of pixels in the video than a preset proportion, and when the preset marker is nearly parallel to the camera's optical axis and occupies a large proportion of the frame, the ratio of pixel length to actual length can be considered a uniform constant throughout the entire video frame. The object change scale *r* can be the physical length of the preset marker. Divide by the pixel length of the preset marker For example, the preset included angle is 5 degrees and the preset ratio is 70%.
[0025] As another example, the object change scale can be the pixel length of the preset marker divided by the physical length of the preset marker.
[0026] In some embodiments, the pixel length of the preset marker in the video is recalculated every U video frames or after time Y. And update the object change scale. As shown in Formula 1.
[0027] Formula 1 Where t represents time.
[0028] Step 102: Obtain the normalized coordinates of key points in the image coordinate system in multiple video frames.
[0029] It is understood that a video is composed of video frames. In this embodiment of the application, the video frames can be all the video frames obtained in step 101, or they can be a portion of the video frames. For example, video frames can be obtained by sampling, for instance, sampling can be performed every 9 video frames, and then the normalized coordinates of the key points in the sampled video frames in the image coordinate system can be obtained.
[0030] For example, if the total number of video frames is N, and the frame counter z is equal to 0 to N, when z divided by s equals 0, that is, z mod s=0, the video frame is read, and the sampled video frame set F is represented by formula 2.
[0031] Formula 2 in, Indicates the first The image matrix of video frames, Indicates the sampling period. It can be in BGR, RGB, or other formats. If the video is in BGR or other formats, it can be converted to RGB format using color conversion functions. 's' can be equal to 10 or 20, meaning it captures data every 10 or 20 video frames.
[0032] For example, the image coordinate system used by the video frame obtained in step 102 is implemented as a two-dimensional coordinate system in pixels.
[0033] In some embodiments, step 102 can be implemented as step 1021.
[0034] Step 1021: Input multiple video frames into the pose estimation model to obtain the normalized coordinates of key points in the image coordinate system in multiple video frames.
[0035] For example, the pose estimation model can be implemented as a PVNet model, a DenseFusion model, etc. When the object to be measured is a human body, the model can be implemented as a human pose estimation model such as Blazepose or YOLOpose.
[0036] After inputting the video frames into the pose estimation model, the normalized coordinates of the key points in the video frames in the image coordinate system are obtained. For each video frame, the normalized coordinates of the keypoints in the image coordinate system can be obtained. The x-coordinates of the keypoints corresponding to multiple video frames in the image coordinate system can be represented as follows: The ordinate of key points corresponding to multiple video frames in the image coordinate system can be represented as: Where N is the number of video frames to be processed.
[0037] In some embodiments, key points can be points that have been processed from the points output by the attitude estimation model. For example, the points output by the attitude estimation model are coordinate points representing the left shoulder, right shoulder, left elbow, right elbow, left hip, and right hip, while the reference point is the midpoint of the line connecting the coordinates of the left hip and right hip, the first marker point is the midpoint of the line connecting the coordinates of the left shoulder and left elbow, the second marker point is the midpoint of the line connecting the coordinates of the right shoulder and right elbow, and the third marker point is the midpoint of the line connecting the coordinates of the left shoulder and right shoulder.
[0038] In some embodiments, circles, triangles, or other shapes can be drawn on the video frame to mark key points. For example, the location where the shape is drawn can be represented as... ,in Indicates image width Height. For example, see reference. Figure 3 , Figure 3 A schematic diagram of the marked video frame is shown. Figure 3 The locations of the left shoulder, right shoulder, left hip, and right hip are marked with green dots.
[0039] In some embodiments, low-quality detections are filtered out by setting a minimum detection confidence level. The minimum detection confidence level represents the lowest acceptable value for the model's accuracy in detecting keypoints. If the detection confidence level corresponding to the keypoint coordinates output by the model is less than the minimum detection confidence level, the keypoint is discarded. This ensures that the detected keypoints have high reliability.
[0040] For example, the minimum detection confidence level is set to 0.5.
[0041] In some embodiments, step 1022 may be included after step 102.
[0042] Step 1022: Construct a keypoint temporal sequence by normalizing the coordinates of each keypoint across multiple video frames.
[0043] Referring to Formula 3, the raw sampled data is transformed into a structured time series. Formula 3 is used to extract the time axis sequence T. This converts discrete frame numbers into actual time scales, facilitating subsequent analysis of the dynamic changes of signals over time.
[0044] Formula 3 Where T represents the time series, This represents the nth time in the time series, where n is an integer variable used to iterate through the time series. Indicates the frame sampling period. N represents the video frame rate, and N represents the total time in the time series.
[0045] For example, t is in the unit "seconds", each (e.g., s=10) frames are sampled once. The frame rate is 30Hz, and the time interval is approximately 0.333s. Figure 4 This diagram illustrates the normalized coordinates of key points in the image coordinate system, with the horizontal axis representing time and the vertical axis representing the x-coordinate of the key points. The x-coordinate of the reference point is represented by the red curve, and the x-coordinate of the marker point is represented by the blue curve.
[0046] In some embodiments, step 102 may be followed by step 1023.
[0047] Step 1023: Output the normalized coordinates of the key points in the image coordinate system.
[0048] For example, the video frame number and the normalized coordinates of key points in the image coordinate system are serialized into strings and written to an output file to achieve output. Referring to Formula 4, each line of data in the output... Including time series The normalized coordinates of the key points in the image coordinate system.
[0049] Formula 4 in, Represents the set of key points. This represents the nth time in the time series.
[0050] For example, refer to Figure 5 , Figure 5 This is a schematic diagram showing the normalized coordinates of the keypoints in the image coordinate system. Each line of the output includes the line number and the normalized coordinates of the keypoints in the image coordinate system; the line number can indicate the video frame number.
[0051] In some embodiments, step 1024 is included before step 103 is performed.
[0052] Step 1024: Filter the key point coordinate time series to obtain the filtered key point coordinate time series.
[0053] Subsequent processing of the keypoint coordinate time series can be implemented as processing of the filtered keypoint coordinate time series.
[0054] For example, referring to Formula 5, a smoothing filter is used to filter the time series of key point coordinates to reduce noise.
[0055] Formula 5 in, This represents the coordinates of the i-th keypoint in the image coordinate system at the n-th video frame, where k represents a window size parameter. For example, if k=2, the smoothed coordinates are... =The coordinates of the five points -2, -1, 0, 1, 2, where j represents an integer variable and n represents the sequence number in the video within the time sequence. This can be the x-coordinate of the keypoint. Similarly, the y-coordinate of the keypoint can be obtained. .
[0056] Therefore, referring to Formula 6, the filtered keypoint coordinate time series is obtained. .
[0057] Formula 6 Where m represents the m-th point in the sequence, and M represents the number of coordinates in the time series of the filtered keypoint coordinates. This represents the timestamp corresponding to the m-th video frame in the filtered sequence. This represents the coordinates of the i-th keypoint corresponding to the m-th video frame in the filtered sequence. The key point coordinate time sequence obtained from the above steps can be used for subsequent analysis of the object's structural state.
[0058] Step 103: Construct an origin coordinate system using the physical coordinates of the reference point in the image coordinate system as the origin, and determine the physical coordinates of the marker point in the origin coordinate system. In other words, determine the physical coordinates of the marker point after the origin coordinate transformation.
[0059] The physical coordinates of key points in the image coordinate system are determined based on the normalized coordinates of the key points, the pixel count of the video, and the object scaling factor.
[0060] Reference Figure 6 In some embodiments, the physical coordinates of key points can be determined through steps 301-304.
[0061] Step 301: Normalize the x-coordinates of key points in the image coordinate system. Total number of horizontal pixels in the video Multiply to obtain the horizontal image pixel coordinates of the key points. As shown in Formula 7.
[0062] Formula 7 Step 302: Normalize the ordinates of the key points in the image coordinate system. Total number of vertical pixels in the video Multiply to obtain the vertical image pixel coordinates of the key points. For example, as in formula 8.
[0063] Formula 8 Step 303: Set the horizontal image pixel coordinates of the key points. Scale of changes to the object Multiplying them yields the lateral physical coordinates of the key points in the image coordinate system. As shown in Formula 9.
[0064] Formula 9 Step 304: Convert the vertical image pixel coordinates of the key points. Scale of changes to the object Multiplying them yields the vertical physical coordinates of the key points in the image coordinate system. For example, as in Formula 10.
[0065] Formula 10 Thus, each key point can be obtained. Keypoint temporal sequence formed by physical coordinates in the image coordinate system across multiple video frames .
[0066] Reference Figure 7 , Figure 7 This diagram illustrates the physical coordinates of key points after the origin coordinate transformation. The horizontal axis represents time, and the vertical axis represents the horizontal physical coordinates of the key points. The horizontal physical coordinates of the reference point are represented by a red curve, while the horizontal physical coordinates of the marker point are represented by a green curve.
[0067] Taking the human body as an example, in the examination of scoliosis of the upper limb spine, the midpoint of the line connecting the left and right hips is selected in each sampled video frame. (An example of a reference point) The floating origin of the dynamic torso reference coordinate system (an example of an origin coordinate system) is referred to Formula 11.
[0068] Formula 11 in, This represents the m-th video frame. Represents the lateral physical coordinates of the left hip. This represents the lateral physical coordinates of the right hip. This represents the longitudinal physical coordinates of the left hip. This represents the longitudinal physical coordinates of the left hip.
[0069] The midpoint of the line connecting the left and right hips is least affected by breathing and upper limb swinging, and has high biomechanical stability. Therefore, it was selected as the reference point to represent the overall translation and rotation of the trunk.
[0070] Refer to Formula 12 to transform the physical coordinates in the image coordinate system of the marked points into the physical coordinates in the origin coordinate system.
[0071] Formula 12 in, This represents the m-th video frame. This represents the physical coordinates of the marked point in the origin coordinate system. This represents the physical coordinates of the marked point in the image coordinate system. This represents the physical coordinates of the reference point in the image coordinate system.
[0072] Physical coordinates of the marked point in the origin coordinate system ( The coordinates can be obtained from formulas 13 and 14, where formula 13 provides the horizontal physical coordinates of the marked point in the origin coordinate system, and formula 14 provides the vertical physical coordinates of the marked point in the origin coordinate system.
[0073] Formula 13 Formula 14 in, This represents the m-th video frame. This represents the horizontal physical coordinates of the marked point after transformation to the origin coordinate system. Indicates the scale of changes in the object. This represents the horizontal image pixel coordinates of the marker point in the global camera pixel coordinate system. This represents the horizontal image pixel coordinates of the reference point in the global camera pixel coordinate system. This represents the longitudinal physical coordinates of the marked point after transformation to the origin coordinate system. This represents the vertical image pixel coordinates of the marker point in the global camera pixel coordinate system. This represents the vertical image pixel coordinates of the reference point in the global camera pixel coordinate system.
[0074] Taking the human body as the test subject as an example, the reference point is the midpoint of the hip, and the marker point is the midpoint of the shoulder. In the scenario of measuring scoliosis, the reference point is... Figure 7 The transformed trajectory of the marker points It eliminates compensatory components such as trunk translation, rotation, lateral bending, and respiratory fluctuations, thus meeting the measurement needs of "pure upper limb movement" of the human body.
[0075] Step 104: Output the activity characteristics of the marked points.
[0076] The activity characteristics of the marker points are determined based on the physical coordinates of the marker points in the origin coordinate system corresponding to multiple video frames.
[0077] For example, the activity characteristics of a marker point include activity period, activity amplitude, and activity speed.
[0078] In some embodiments, steps 401 to 405 are further included after step 103.
[0079] Step 401: Construct a keypoint temporal sequence by combining the physical coordinates of each keypoint across multiple video frames.
[0080] Step 402: Use a higher-order polynomial to perform least-squares fitting on the time series of physical coordinates of the marked points to obtain the fitting result.
[0081] Taking a polynomial fitting with a highest order of 25 as an example, refer to Formula 15. It can be understood that the highest order can also be 21, 23, etc.
[0082] Formula 15 in, For coefficients, It can be done Solving for the minimum sum of squared residuals involves adjusting the polynomial coefficients. This minimizes the sum of squares.
[0083] Step 403: Extract the peaks, troughs, time difference sequences between adjacent peaks, and time difference sequences between adjacent troughs from the fitting results.
[0084] For example, the critical point can be obtained by calculating the first derivative of the fitting result, referring to Formula 16.
[0085] Formula 16 in, The time representing the critical point (i.e., peak and trough). Expressing the request First derivative expression The root when it equals 0.
[0086] Then calculate the second derivative of the fitting result. If the second derivative is less than zero, the critical point is a peak; if the second derivative is greater than zero, the critical point is a trough.
[0087] Time difference sequence of adjacent peaks and troughs like Figure 8 As shown. Figure 8This is a schematic diagram of the time-series fitting results of the physical coordinates of key points after transformation to the origin coordinate system. The horizontal axis represents time, and the vertical axis represents the lateral physical coordinates of the key points. , Figure 8 The physical coordinates of the marked points are represented by the blue curve, and the physical coordinates of the marked points after polynomial fitting are represented by the purple curve. Red dots represent peak values, and blue dots represent valley values.
[0088] Step 404: Calculate the amplitude difference between adjacent extreme points and use the average amplitude difference as a threshold. If the amplitude difference between an extreme point and the extreme points before and after it is greater than the threshold, then the extreme point is considered a valid extreme point.
[0089] Refer to Formula 17 to obtain the amplitude difference between adjacent extreme points.
[0090] Formula 17 This represents the "displacement value" of the i-th extreme point, i.e. Figure 8 The fitted curve shown represents the horizontal physical coordinates of the marked points at the extreme time ti. This represents the "displacement value" of the (i+1)th extreme point, i.e. Figure 8 The fitted curve shown represents the horizontal physical coordinates of the marker point at the extreme time ti+1. It represents the absolute value of the displacement change between two adjacent extreme values.
[0091] Taking the human body as the test subject as an example, This represents the lateral displacement of the shoulder relative to the torso at time i. This represents the lateral displacement of the shoulder relative to the torso at time i+1. It indicates the range of motion of the shoulder.
[0092] In other embodiments, the threshold may also be determined in other ways, such as using the mode of the amplitude difference as the threshold. This application does not limit the method of determining the threshold.
[0093] Reference Figure 9 , Figure 9 The red and blue dots in the image represent the valid extreme points retained after filtering out the extreme points. Figure 8 Compared to the previous version, it has one less peak and one less trough.
[0094] Step 405: Determine the activity characteristics of the marker points through effective extreme points.
[0095] Once the effective extreme points are determined, the activity characteristics of the marker points can be determined based on these points. These characteristics include the activity period, activity amplitude, and activity speed. Referring to Formula 18, the activity period... This can be the time between two adjacent peaks or the time between two adjacent troughs. Refer to Formula 17 for the activity amplitude. This can be the distance between adjacent peaks and troughs. Refer to Formula 19 for the activity speed. It can be the ratio of the activity's magnitude to its duration.
[0096] Formula 18 Formula 19 Reference Figure 10 , Figure 10 for Figure 9 The diagram showing the connection between the extreme points is obtained by... Figure 10 It can more intuitively show the time difference between extreme points, thereby determining whether the speed of the object under test has decayed, and evaluating the performance of the object under test based on its actual performance.
[0097] In this embodiment, by setting the reference point as the origin of the physical coordinate system where the marker point is located, an origin coordinate system is constructed. This allows the determined physical coordinates of the marker point in the video to better reflect the movement of the object's active parts relative to the main body of the object, reducing interference from the overall movement of the object. By converting the pixel coordinates of the marker point into physical coordinates—that is, by measuring the physical coordinates of the marker point using the length of the object's world—the physical coordinates of the marker point can be compared with the length of the physical world. This allows for a determination of whether the activity characteristics of the marker point match the activity characteristics of the object when its structure is intact, thus obtaining the object's structural state. This embodiment can determine structural loosening by observing the relative positions of the various structures within the structural object during its operation, facilitating the acquisition of the structural state by personnel. Through this embodiment, it is not necessary to train the object state estimation model using a large number of complete object images; the estimation results of the object state can be obtained quickly with less data.
[0098] By calculating the activity cycle of the marker points, or the time between adjacent extreme points, we can determine whether the movement rhythm of the object under test can be stable and smooth. By calculating the actual physical distance between adjacent extreme points, we can determine the actual activity range of the marker points, thereby determining whether the movement ability of the object under test meets the standard.
[0099] In some embodiments of this application, the object to be tested includes a first marker and a second marker, and the difference in physical length between the first marker and the second marker is less than a preset length. The preset marker is determined through steps 501-503.
[0100] For example, the preset length is 3%, 5%, etc., of the length of the first marker. Taking the human body as an example, the first marker can be the left upper arm, and the second marker can be the right upper arm.
[0101] Step 501: Obtain the pixel length of the first marker and the pixel length of the second marker in multiple video frames.
[0102] For example, a keypoint time sequence is obtained from the first y minutes or the first u video frames of the video. The average distance between keypoint 1 and keypoint 2 in each video frame is calculated, as well as the average distance between keypoint 3 and keypoint 4 in each video frame. The average distance between keypoint 1 and keypoint 2 is used as the pixel length of the first marker, and the average distance between keypoint 3 and keypoint 4 is used as the pixel length of the second marker.
[0103] Determine the pixel length of the first marker using Formula 20. .
[0104] Formula 20 in, This represents the pixel coordinates of keypoint 1. This represents the pixel coordinates of keypoint 2.
[0105] Determine the pixel length of the second marker by referring to Formula 21. .
[0106] Formula 21 in, This represents the pixel coordinates of keypoint 3. This represents the pixel coordinates of keypoint 4.
[0107] Taking the human body as the test subject, key point 1 is the left elbow, key point 2 is the left shoulder, key point 3 is the right elbow, and key point 4 is the right shoulder.
[0108] Referring to formulas 22 and 23, determine the average distance between keypoint 1 and keypoint 2 in the first y minutes or the first u video frames. The average distance between keypoint 3 and keypoint 4 .
[0109] Formula 22 Formula 23 in, Indicates the number of video frames.
[0110] Step 502: If the pixel length of the first marker is greater than the pixel length of the second marker, then the first marker is used as the preset marker.
[0111] Step 503: If the pixel length of the first marker is less than the pixel length of the second marker, then the second marker is used as the preset marker.
[0112] Since the lengths of the first and second markers are relatively close, the marker with the larger pixel length is closer to the lens, resulting in minimal perspective scaling and the highest calibration accuracy. Therefore, a preset marker is set based on its pixel length.
[0113] By taking video frames from the first y minutes or the first u video frames to determine preset markers and calculating the change scale, the time consumed when measuring the structural state of an object can be reduced.
[0114] In some embodiments of this application, the marker points include a first marker point and a second marker point, which are symmetrically arranged on the object. The fitting result includes the fitting result of the first marker point and the fitting result of the second marker point, and the method further includes steps 406 and 407.
[0115] For example, when the object is a robot, the first marker point of the object can be set at a first distance from the left fingertip of the robot's left arm, and the second marker point of the object can be set at a first distance from the right fingertip of the robot's right arm.
[0116] Setting symmetry on an object can be understood as the first and second marker points of the object being symmetrical in a natural state according to the object's structure, rather than the first and second marker points being able to achieve a symmetrical state or always being symmetrical when the object is active.
[0117] Step 406: Calculate the difference in activity speed between the first marker point and the second marker point.
[0118] Step 407: Calculate the activity frequencies of the first and second marker points respectively.
[0119] For example, referring to Formula 19, the activity velocities of the first and second marker points are calculated respectively. For instance, the activity velocity of the first marker point is... The activity speed of the second marker point is The difference in activity speed S can be determined by referring to Formula 24. As another example, the difference in activity speed S can be determined by referring to Formula 25.
[0120] = - Formula 24 Formula 25 in, This represents the average activity velocity of multiple first marker points. This represents the average activity velocity of multiple second marker points.
[0121] By performing a Fourier transform on the physical coordinates of the first marker point, the main peak frequency of the first marker point can be obtained. The activity frequency of the first marker point can be obtained by performing a Fourier transform on the physical coordinates of the second marker point, which yields the main peak frequency of the second marker point. That is, the activity frequency of the second marker point. If and If the value falls within the preset range, the structural state of the test object is considered normal. The preset range can be set according to the type of test object. Taking the human body as an example, for the examination of scoliosis of the upper limb spine, the preset range is 0.5 to 2 Hz.
[0122] By calculating the difference in movement velocities and the movement frequency of the first and second marker points, the structural symmetry and stability of the object under test can be determined. Determining the movement velocities of the marker points determines whether the motion of the object is smooth and whether the movements of the first and second marker points are synchronized and symmetrical. For example, if the difference in velocity between the first and second marker points is less than 5% of the first marker point's movement velocity, then the movements of the first and second marker points are synchronized; otherwise, they are asynchronous. Determining the movement frequencies of the marker points determines whether the motion of the first and second marker points is stable. For example, if the difference in frequency between the first and second marker points is less than 5% of the first marker point's movement frequency, then the motion of the first and second marker points is stable; otherwise, it is unstable.
[0123] In some embodiments of this application, steps 408 to 411 are included after step 405.
[0124] Step 408: Calculate the instantaneous phase of the reference point transformation based on the reference point time series.
[0125] Step 409: Calculate the instantaneous phase of the transformation of the reference point based on the time sequence of the marked points.
[0126] Step 410: Obtain the phase difference based on the instantaneous phase of the transformation of the reference point and the instantaneous phase of the transformation of the marker point.
[0127] Step 411: If the phase difference is greater than the preset difference value, output a prompt message indicating that the correlation between the marker point and the reference point is too large.
[0128] For example, the instantaneous phases of the reference point and the marker point are calculated using Hilbert transform. In some embodiments, the reference point time series and the marker point time series are sequences in the same coordinate system. For example, both are time series composed of coordinates that have not undergone origin coordinate transformation, or both are in the camera absolute coordinate system or the world coordinate system.
[0129] Understandably, the above fitting result can be achieved.
[0130] For example, the preset difference is 45 degrees.
[0131] Excessive correlation between the marker point and the reference point indicates that the movement of the marker point is significantly influenced by the reference point, suggesting insufficient stability of the reference point or insufficient flexibility of the marker point. In the field of robotic arms, this may indicate inappropriate connections between parts, such as excessive coupling. In the field of scoliosis measurement, it may indicate spinal compensation.
[0132] In some embodiments of this application, step 4011 is included before step 401. Step 401 can be implemented as step 4012.
[0133] Step 4011: Perform polynomial least squares fitting on the key point time series within a local sliding window to obtain the processed key point time series.
[0134] It is understandable that the key point time series can be implemented as a reference point time series and a marker point time series.
[0135] For example, in order to reduce outliers in the key point time series and reduce the impact of high-frequency jitter on the calculation results, (Savitzky-Golay, SG) filters are applied to the horizontal and vertical axes of the key point time series, respectively. While suppressing high-frequency jitter, the true motion extreme points and periodic characteristics are preserved, as shown in Formula 26.
[0136] Formula 26 in, This represents the filtered physical coordinates of the key points. The least squares convolution coefficients of the SG filter are represented, and k represents an integer variable used to iterate through the current time step. The data window is centered, and E represents the width of the filter window on one side. For example, E = 10, 12, or 15, etc.
[0137] Step 4012: Use a higher-order polynomial to perform least-squares fitting on the processed marker point time series.
[0138] By performing least-squares fitting on the key point time series, filtering of the key point time series can be achieved, such as removing high-frequency noise.
[0139] The above embodiments describe first converting the coordinates in the image coordinate system from pixel coordinates to physical coordinates, and then converting the image coordinate system to the origin coordinate system. In some other embodiments of this application, the coordinate system of the video frame can also be transformed from the image coordinate system to the origin coordinate system first, and then the coordinates of the key points can be transformed from pixel coordinates to physical coordinates.
[0140] The technical solutions provided in the embodiments of this application have been described in detail above. Specific examples have been used in the embodiments of this application to illustrate the principles and implementation methods of the embodiments of this application. The description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of the embodiments of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments of this application. Therefore, the content of this specification should not be construed as a limitation on the embodiments of this application.
Claims
1. A measurement and analysis method based on video measurement, characterized in that, The method includes: Acquire a video of the object to be tested and an object change scale of the video; the video includes key points, the key points include reference points and marker points, and the object change scale indicates the ratio between the physical length of the object and the pixel length of the object; Obtain the normalized coordinates of the key points in the image coordinate system in multiple video frames of the video; The physical coordinates of the reference point in the image coordinate system are used as the origin to construct the origin coordinate system, and the physical coordinates of the marker point in the origin coordinate system are determined. The physical coordinates of the key point in the image coordinate system are determined based on the normalized coordinates of the key point in the image coordinate system, the pixels of the video, and the object change scale. The activity characteristics of the marker points are output, which are determined based on the physical coordinates of the marker points in the origin coordinate system corresponding to the multiple video frames.
2. The method according to claim 1, characterized in that, The object change scale was obtained through the following method: Obtain the physical length of the preset marker of the object to be tested; Obtain the pixel length of the preset marker in the video; The object change scale is determined based on the physical length and pixel length of the preset marker.
3. The method according to claim 2, characterized in that, The object to be tested includes a first marker and a second marker, wherein the difference in physical length between the first marker and the second marker is less than a preset length; The preset marker is determined by the following method: Obtain the pixel length of the first marker and the pixel length of the second marker in the plurality of video frames; If the pixel length of the first marker is greater than the pixel length of the second marker, then the first marker is used as the preset marker; If the pixel length of the first marker is less than the pixel length of the second marker, then the second marker is used as the preset marker.
4. The method according to claim 1, characterized in that, The step of obtaining the coordinates of the key points in the image coordinate system in multiple video frames of the video includes: Input the multiple video frames into the pose estimation model to obtain the normalized coordinates of the key points in the multiple video frames in the image coordinate system; The physical coordinates of the key points in the image coordinate system are determined in the following way: Multiply the normalized horizontal coordinate of the key point in the image coordinate system by the horizontal pixel coordinate of the video to obtain the horizontal image pixel coordinate of the key point in the image coordinate system. Multiply the normalized ordinate of the key point in the image coordinate system by the vertical pixel coordinate of the video to obtain the vertical image pixel coordinate of the key point; Multiply the horizontal image pixel coordinates of the key point by the object's change scale to obtain the horizontal physical coordinates of the key point; The vertical image pixel coordinates of the key point are multiplied by the object's change scale to obtain the vertical physical coordinates of the key point.
5. The method according to claim 1, characterized in that, After determining the physical coordinates of the marked point in the origin coordinate system, the method further includes: The physical coordinates of each key point in the multiple video frames are used to form a key point temporal sequence. The physical coordinate time series of the marked points is fitted using a higher-order polynomial to obtain the fitting result. Extract the peaks, troughs, time difference sequences between adjacent peaks, and time difference sequences between adjacent troughs from the fitting results; Calculate the amplitude difference between adjacent extreme points, and use the mean of the amplitude difference as a threshold. If the amplitude difference between an extreme point and its adjacent extreme points is greater than the threshold, then the extreme point is considered a valid extreme point. The activity characteristics of the marker point are determined by the effective extreme points, including the activity period, activity amplitude, and activity speed.
6. The method according to claim 5, characterized in that, The marker points include a first marker point and a second marker point, which are symmetrically arranged on the object. Before outputting the activity features of the marked points, the method further includes: Calculate the difference in activity speed between the first marker point and the second marker point respectively; Calculate the activity frequency of the first marker point and the second marker point respectively.
7. The method according to claim 6, characterized in that, Before outputting the activity features of the marked points, the method further includes: Calculate the instantaneous phase of the reference point based on the time sequence of the reference point; Calculate the instantaneous phase of the transformed marker points based on the time sequence of the marker points; The phase difference is obtained based on the instantaneous phase of the transformation of the reference point and the instantaneous phase of the transformation of the marker point; The step of outputting the activity characteristics of the marker point further includes: if the phase difference is greater than a preset difference value, outputting a prompt message, the prompt message indicating that the correlation between the marker point and the reference point is too large.
8. The method according to any one of claims 5-7, characterized in that, After constructing a keypoint temporal sequence from the physical coordinates of each keypoint in the multiple video frames, the method further includes: The key point time series is fitted with polynomial least squares within a local sliding window to obtain the processed key point time series. The step of performing least-squares fitting on the time series of the marked points using higher-order polynomials includes: performing least-squares fitting on the processed time series of the marked points using higher-order polynomials.
9. The method according to claim 1, characterized in that, Before constructing an origin coordinate system by using the physical coordinates of the reference point in the image coordinate system as the origin, and determining the physical coordinates of the marker point in the origin coordinate system, the method further includes: filtering the key point coordinate time sequence to obtain the filtered key point coordinate time sequence.
10. The method according to claim 2 or 3, characterized in that, The method is applied to monocular close-range photogrammetry. When the video is captured, the position of the camera is fixed, the angle between the extension direction of the preset marker and the optical axis of the camera is less than a preset angle, and the pixel ratio of the preset marker in the video is greater than a preset ratio.