Unattended vision detection method and system based on image recognition

By collecting and integrating image data in an unattended vision detection system, establishing basic geometric delineation elements and multi-view fusion scene state point sets, the spatial positioning error and multi-view data fusion problems in the prior art are solved, and the stability and accuracy of the detection results are improved.

CN120014054AActive Publication Date: 2025-05-16SHANGHAI CANGAO TRADE CO LTD

Patent Information

Application Number
CN202510457628.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-16
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing unattended vision detection method based on image recognition is prone to spatial positioning errors when complex spatial structures and user relative positions change, resulting in poor stability of vision assessment results and lack of fusion processing of multi-view data, which affects the reliability and accuracy of detection results.

Method used

By collecting image data from the vision detection site, a basic geometric delineation element is established, and the initial scene geometric structure is constructed by collecting image data from the vision detection site, and combining the spatial integration of the user's vision area and the boundary of the vision detection screen to build the initial scene geometric structure. At the same time, track the camera pose changes, calculate the camera's relative pose sequence, build a three-dimensional spatial coordinate system, and generate a dynamic spatial vector of the user's screen. Obtain multi-view image data, identify the three-dimensional position information of user behavior feature points and key elements of the scene, perform cross-view data mapping and fusion, establish a multi-view fusion scene state point set, match user behavior feature points and scene elements under different perspective angles, calculate spatial position difference values, and judge spatial state consistency.

Benefits of technology

It improves the stability and accuracy of vision detection results, enhances the consistency and positioning accuracy of multi-view data, improves the reliability and automation of spatial state consistency judgment results, and ensures the effectiveness of the unattended vision detection process and the authenticity and accuracy of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014054A_ABST
    Figure CN120014054A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to an unattended vision detection method and system based on image recognition, and the method comprises the following steps: collecting vision detection field image data, carrying out scene structure line detection and plane detection to obtain a line segment set and plane parameters, and establishing a basic geometric description element; and based on the basic geometric description element, combining the detected user sight line area coordinate and the vision detection screen boundary coordinate to carry out spatial integration, and obtaining the geometric structure information of the initial scene. According to the method, scene structure line detection and plane detection are carried out by collecting image data of a vision detection site, a line segment set and plane parameters are obtained, a basic geometric description element is established, and construction of an initial scene geometric structure is realized by combining spatial integration of a user sight area and a vision detection screen boundary; and the stability of on-site space structure analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to an unattended vision detection method and system based on image recognition. Background Art

[0002] The unattended vision detection method based on image recognition is a method that automatically analyzes and evaluates the user's vision status through computer vision. Its main purpose is to use image recognition technology to automatically perform vision detection and evaluation without human supervision.

[0003] Existing technologies only rely on basic computer vision methods to automatically analyze and evaluate the user's vision status. In the case of complex spatial structures or frequent changes in the user's relative position, it is easy to cause large errors in spatial positioning, and it is difficult to effectively deal with the real-time spatial instability caused by changes in line of sight and camera angle of view, resulting in poor stability of vision assessment results. At the same time, there is a lack of fusion processing of multi-view data, making it difficult to identify and correct the spatial position differences between user posture and line of sight direction, resulting in a lack of spatial consistency verification of scene analysis results, and the evaluation conclusions are easily affected by external interference factors, which reduces the reliability and accuracy of the detection results. Therefore, improvements are needed. Summary of the invention

[0004] The purpose of the present invention is to solve the shortcomings existing in the prior art and to propose an unattended vision detection method and system based on image recognition.

[0005] In order to achieve the above object, the present invention adopts the following technical solution, an unattended vision detection method based on image recognition, comprising the following steps: Collecting on-site image data of vision detection, performing scene structure line detection and plane detection to obtain line segment sets and plane parameters, and establishing basic geometric depiction elements; based on the basic geometric depiction elements, spatial integration is performed in combination with the user sight area coordinates obtained by the detection and the vision detection screen boundary coordinates to obtain initial scene geometric structure information; Based on the initial scene geometric structure information, tracking the camera posture changes in continuous image frames, calculating and obtaining the displacement and rotation parameters of the camera relative to the initial geometric structure, obtaining a camera relative posture sequence, and based on the camera relative posture sequence, constructing a three-dimensional space coordinate system and solving the coordinate transformation from the user's sight area to the vision detection screen to generate a user screen dynamic space vector; Acquire multi-view image data of the vision test site, identify the user behavior feature points and the three-dimensional position information of the key elements of the scene at each view, compile them into a single-view element positioning table, perform cross-view data mapping and fusion based on the unified coordinate system defined by the single-view element positioning table and the user screen dynamic space vector, and establish a multi-view fusion scene state point set; Based on the multi-perspective fusion scene state point set, the corresponding user behavior feature points and scene elements under different perspectives are matched, the spatial position difference value in the unified coordinate system is calculated, the spatial position difference value is compared with the preset spatial tolerance threshold, and the spatial state consistency judgment result is obtained.

[0006] Preferably, the steps of obtaining the basic geometric depiction element are: Collecting image data of the vision detection scene, performing image edge gradient analysis based on the collected image data of the vision detection scene to calculate the pixel gradient change of the edge point, extracting the edge contour at the boundary of each object in the scene, combining the edge points to form structural line segments, and generating a vision detection scene structural line segment set; Based on the vision detection scene structure line segment set, extract the spatial coordinate data of all line segment endpoints in the line segment set, determine the coplanarity of the line segment endpoint spatial coordinates, select the coplanar point group according to the line segment endpoint coordinate coplanarity, fit and calculate the plane equation parameters corresponding to the spatial position of the coplanar point group, and generate the plane parameters corresponding to the structure line segment; Based on the plane parameters corresponding to the structural line segments, the line segment endpoint coordinates in the visual inspection site structural line segment set are called, and the line segment endpoint spatial coordinates and the corresponding plane parameters are spatially projected one by one to determine the spatial position relationship between the line segment and the corresponding plane, and establish a spatial mapping relationship between the line segment endpoint coordinates and the plane parameters to generate basic geometric depiction elements.

[0007] Preferably, the steps of acquiring the initial scene geometric structure information are: Based on the basic geometric depiction element, the line segment endpoint coordinates and plane parameters in the basic geometric depiction element are called, combined with the vision detection screen boundary coordinates, spatial coordinate transformation and projection calculation are performed to determine the spatial position relationship of the vision detection screen in the basic geometric depiction element, and generate the screen boundary space mapping coordinates; Based on the screen boundary space mapping coordinates, the coordinates of the user's sight area are obtained, coordinate system alignment and space coordinate transformation are performed, the spatial position relationship between the user's sight area and the screen boundary is integrated, a mapping correspondence relationship between the sight area coordinates and the screen boundary coordinates is established, and the sight screen integrated space coordinates are generated; Based on the line-of-sight screen integrated space coordinates, a spatial matching relationship between the line-of-sight screen integrated space coordinates and basic geometric depiction elements is determined to obtain initial scene geometric structure information.

[0008] Preferably, the steps of acquiring the relative position sequence of the camera are: Based on the initial scene geometric structure information, the coordinates of the three-dimensional boundary points of the user's sight area and the boundary corner points of the vision detection screen are obtained, and the matching errors and reprojection offsets of the corresponding boundary points between the image frames are calculated by combining the image acquisition timestamps and camera internal parameters recorded in the continuous image frames to generate a continuous image frame matching error sequence; Calculating the camera posture change angle value according to the continuous image frame matching error sequence; According to the camera posture change angle value, combined with the three-dimensional offset of the boundary points of each frame image in the continuous image frame matching error sequence, the three-dimensional rotation axis offset and direction vector change between adjacent frames are obtained to obtain the camera relative posture sequence.

[0009] Preferably, the step of acquiring the user screen dynamic space vector is: Based on the relative position sequence of the camera, the rotation matrix and translation vector of each frame are obtained, and the focal length, image center coordinates, pixel size and image size parameters in the camera intrinsic parameters are combined to sequentially reconstruct the coordinates and space vectors of the starting point coordinates of the user's sight area and the screen corner coordinates to generate an initial three-dimensional vector group from the user's sight to the screen; Calculate the spatial transformation value from the user's sight line vector to the target position on the screen according to the initial three-dimensional vector group from the user's sight line to the screen; Based on the spatial transformation value from the user's sight line vector to the screen target position, vector direction normalization and projection difference correction are performed to obtain the user screen dynamic space vector.

[0010] Preferably, the steps of obtaining the single-view element positioning table are: Acquire multi-view image data of the vision testing site, extract pixel information of the user behavior feature area and the scene key element area in the multi-view image data of the vision testing site, identify the three-dimensional position information of the user behavior feature points through spatial mapping, and generate a set of spatial coordinates of the user behavior feature points; Based on the spatial coordinate set of the user behavior feature points, the contour edges of the key elements of the scene in the multi-view image data of the vision detection scene are identified, and spatial coordinate matching and three-dimensional space fitting are performed to generate a three-dimensional space position coordinate set of the key elements of the scene; Based on the three-dimensional space position coordinate set of the scene key elements, the three-dimensional space position coordinates of the user behavior feature points and the scene key elements under the corresponding viewing angle are summarized to establish a single-view element positioning table.

[0011] Preferably, the steps of acquiring the multi-view fusion scene state point set are: Based on the single-view element positioning table, extract the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the image frame number, the camera internal parameter matrix and the user screen dynamic space vector coordinate endpoints under each view, and uniformly convert them to the unified coordinate system defined by the user screen dynamic space vector to generate a multi-view three-dimensional positioning matrix set in the unified coordinate system; Calculating a fusion stabilization offset of each perspective according to the multi-perspective three-dimensional positioning matrix set in the unified coordinate system; Based on the fused stable offset, traverse the positioning point pairs of each perspective, prioritize the spatial position fusion in the order of the fused stable offset from small to large, select the point pairs that meet the spatial overlap conditions and merge them into a unified spatial point position, and establish a multi-perspective fusion scene state point set.

[0012] Preferably, the steps of obtaining the spatial state consistency determination result are: Based on the multi-view fusion scene state point set, the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the camera optical center position, the camera visual axis vector, the user screen dynamic space vector direction and the user behavior feature point orientation vector corresponding to each view are extracted to generate a multi-view point direction comparison matrix; Calculate the spatial position difference value in the unified coordinate system according to the multi-view point direction comparison matrix; Based on the spatial position difference value, a comparison is performed one by one with the set spatial tolerance threshold, all matching point pairs that meet the spatial tolerance threshold are screened, and a unified judgment result is output based on the consistent matching of the point pairs in all viewing angles to obtain the spatial state consistency judgment result.

[0013] The present invention provides an unattended vision detection system, comprising: The data acquisition module performs image acquisition based on the image data of the vision test site, obtains multi-angle and multi-view image data of the vision test site, performs scene structure line and plane detection, and constructs a scene geometric feature set by extracting line segment sets and plane parameters; The scene modeling module performs spatial integration based on the scene geometric feature set, and obtains the initial scene geometric structure by combining the coordinates of the user's sight area obtained by detection with the coordinates of the boundary of the vision detection screen. According to the initial scene geometric structure, the camera posture changes in continuous image frames are tracked, and the relative posture of the camera is calculated and obtained, and then a three-dimensional space coordinate system is constructed, and a camera relative posture sequence is generated. The coordinate transformation from the user's sight area to the vision detection screen is solved to obtain the user screen dynamic space vector; The gaze tracking module tracks the posture changes of continuous image frames based on the camera relative posture sequence and the user screen dynamic space vector, calculates the dynamic changes of gaze direction and user behavior, and generates a user gaze tracking sequence; The multi-view fusion module, based on the multi-view image data of the vision detection site, identifies the user behavior feature points and key elements of the scene at each view, obtains the three-dimensional space position data, and compiles it into a single-view element positioning table; combines the user's line of sight tracking sequence with the single-view element positioning table, performs cross-view data mapping and fusion, and generates a multi-view fusion scene state point set; The spatial consistency module matches the user behavior feature points and scene elements under different perspectives based on the multi-perspective fusion scene state point set, calculates the spatial position difference value, compares the spatial position difference value with the preset spatial tolerance threshold, judges the consistency, and obtains the spatial consistency judgment result.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are: The present invention collects image data of the vision detection site to perform scene structure line detection and plane detection, obtains line segment sets and plane parameters, establishes basic geometric descriptors, and realizes the construction of the initial scene geometric structure by combining the spatial integration of the user's sight area and the boundary of the vision detection screen, thereby improving the stability of the on-site spatial structure analysis; further tracks the camera posture changes in continuous image frames, calculates the displacement and rotation parameters in real time, generates a camera relative posture sequence, realizes the accurate dynamic space vector construction between the user's sight area and the screen, and reduces the detection error caused by posture changes in dynamic scenes; at the same time, obtains multi-view image data to identify the spatial positions of user behavior feature points and scene key elements under each perspective, constructs a single-view element positioning table and integrates them across perspectives, thereby improving the consistency and positioning accuracy of multi-view data; finally, the spatial position difference value under a unified coordinate system is accurately calculated and the spatial tolerance threshold is judged, thereby improving the reliability and automation of the spatial state consistency judgment result, and ensuring the accuracy and intelligence of the overall vision detection process, and the spatial state consistency judgment result is used to confirm whether the user's sight area and the vision detection screen display area are accurately coincident in the multi-view image, and to determine whether the user maintains the correct posture and sight direction during the detection process. By calculating the position difference between the user behavior feature points and the vision detection screen elements under different viewing angles, it is determined whether they are within the predetermined tolerance range. If the position difference exceeds the tolerance range, it means that the user's line of sight has deviated from the detection screen or the posture is abnormal. At this time, the current invalid detection data can be excluded, thereby ensuring the effectiveness of the unattended vision detection process and the authenticity and accuracy of the evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0017] See also Figure 1 The present invention provides a technical solution, an unattended vision detection method based on image recognition, comprising the following steps: Collect image data from the vision test site, perform scene structure line detection and plane detection to obtain line segment sets and plane parameters, and establish basic geometric descriptors; based on the basic geometric descriptors, combine the user's sight area coordinates obtained from the detection with the vision test screen boundary coordinates to perform spatial integration to obtain the initial scene geometric structure information; Based on the initial scene geometry information, the camera posture changes in continuous image frames are tracked, the displacement and rotation parameters of the camera relative to the initial geometry are calculated, and the camera relative posture sequence is obtained. Based on the camera relative posture sequence, a three-dimensional space coordinate system is constructed and the coordinate transformation from the user's line of sight to the vision detection screen is solved to generate the user's screen dynamic space vector. Acquire multi-view image data from the vision test site, identify the user behavior feature points and the three-dimensional position information of the key elements of the scene at each view, compile them into a single-view element positioning table, and perform cross-view data mapping and fusion based on the unified coordinate system defined by the single-view element positioning table and the user screen dynamic space vector to establish a multi-view fusion scene state point set; Based on the multi-perspective fusion scene state point set, the corresponding user behavior feature points and scene elements under different perspectives are matched, the spatial position difference value in the unified coordinate system is calculated, and the spatial position difference value is compared with the preset spatial tolerance threshold to obtain the spatial state consistency judgment result.

[0018] The steps to obtain the basic geometric descriptor are: Collecting image data of the vision detection scene, performing image edge gradient analysis based on the collected image data of the vision detection scene to calculate the pixel gradient change of the edge point, extracting the edge contour at the boundary of each object in the scene, combining the edge points to form structural line segments, and generating a vision detection scene structural line segment set; Based on the visual inspection site structure line segment set, the spatial coordinate data of all line segment endpoints in the line segment set are extracted, the coplanarity of the line segment endpoints' spatial coordinates is determined, the coplanar point group is selected according to the line segment endpoints' coordinate coplanarity, the plane equation parameters corresponding to the coplanar point group's spatial position are calculated by fitting, and the plane parameters corresponding to the structure line segment are generated; Based on the plane parameters corresponding to the structural line segments, the line segment endpoint coordinates in the visual inspection site structural line segment set are called, and the line segment endpoint spatial coordinates and the corresponding plane parameters are spatially projected one by one to determine the spatial position relationship between the line segment and the corresponding plane. The spatial mapping relationship between the line segment endpoint coordinates and the plane parameters is established to generate basic geometric depiction elements.

[0019] Specifically, pixel-level edge gradient calculation is performed based on the acquired image data. First, an image with a resolution of about 1920×1080 is collected in the visible light range and converted into grayscale form frame by frame. Then, the horizontal and vertical differential operators are selected to calculate the gradient value of each pixel and generate a gradient matrix. If the gradient amplitude of a pixel point is If the number of pixels exceeds 25, it is determined as a potential edge point. The threshold of 25 is determined by using different thresholds from 15 to 30 in the test samples and comparing the edge integrity. Then, non-maximum suppression is performed to eliminate pixels with weak gradient response and overlapping with the main edge. Then, the continuity is adjusted by setting morphological operations such as closing and opening operations to merge adjacent edges. A minimum pixel number threshold of 5 is introduced when judging the edge chain length. This value is calculated by calculating the edge chain length distribution for 20 test images and combined with visual judgment settings. Edge chains with a length of less than 5 are not recorded. After completion, the remaining continuous edge chains are regarded as preliminary line segments and their starting and ending coordinates are identified. By traversing these starting and ending coordinates, an index list that can distinguish adjacent or overlapping line segments is formed. When adjacent line segments intersect and form an approximate straight line, the coordinate difference method is used to determine whether to merge. If the length of the merged line segment is greater than the aforementioned minimum pixel number threshold, the merged new line segment is retained. Finally, a set of available line segments is obtained and uniformly recorded as a set of structural line segments on the vision detection site.

[0020] Based on the obtained visual inspection site structure line segment set, the starting and ending point three-dimensional coordinate data of the line segment are read one by one, and the endpoint coordinates of each line segment are firstly calibrated for preliminary accuracy. By comparing with the set spatial error range such as 0.00 to 0.05 and excluding erroneous data beyond this range, all endpoint coordinates are collected into the same list after calibration and the coplanarity judgment rule is used to detect the endpoint distribution. The coplanarity judgment introduces a distance judgment threshold It is used to measure the maximum allowable deviation of the distance from a point to a plane. The final determination is made by increasing the value from 0.001 to 0.02 in a trial manner and observing the stability of the endpoint classification results. The set within is regarded as a coplanar point group, and the coplanar point group is fitted with a plane using the least squares method to obtain the plane equation parameters , x, y, and z represent the coordinate values ​​of the endpoints of the line segment along the x, y, and z axes in three-dimensional space, respectively. The plane parameters are determined by performing eigenvalue decomposition on the point coordinate matrices respectively. To avoid ambiguous references, the endpoints of the line segments that have been determined must be matched one by one. If the two endpoints of the line segment belong to the same coplanar point group, the plane equation parameters are recorded in the corresponding line segment entry. After completing the processing of all line segments, the plane parameters corresponding to the structural line segments can be obtained.

[0021] Based on the plane parameters corresponding to the structural line segments, the endpoint coordinates of the visual inspection site structural line segment set are called to perform spatial projection calculation. For each line segment, first calculate the plane equation of the line segment. Extract As the normal vector, the endpoint coordinates are used in the projection calculation Make a vertical projection in the direction of the normal vector and get the projection coordinates For the two endpoints of the same line segment, the same plane normal vector is used to project the line segment's mapping position in the plane. The distance difference threshold is used to compare the difference between the endpoint projection result and the original three-dimensional coordinates of the line segment. To determine whether there is a significant deviation, After testing step by step in the range of 0.005 to 0.03, if the projection deviation is less than A corresponding record is then established for the line segment mapping and the plane parameters. If multiple plane parameters correspond to the same line segment endpoint, confirmation is made based on the solution with the smallest line segment projection distance. Finally, after all endpoints are projected and the mapping relationship is established, the corresponding information between the line segment and the plane is summarized into a spatial mapping index table. The index table contains the coordinates of each line segment endpoint and the projection data of the corresponding plane equation parameters. After summarizing this information, the basic geometric depiction element is obtained.

[0022] The steps to obtain the initial scene geometry information are: Based on the basic geometric depiction element, the line segment endpoint coordinates and plane parameters in the basic geometric depiction element are called, combined with the vision detection screen boundary coordinates, spatial coordinate transformation and projection calculation are performed to determine the spatial position relationship of the vision detection screen in the basic geometric depiction element, and generate the screen boundary space mapping coordinates; Based on the screen boundary space mapping coordinates, the coordinates of the user's sight area are obtained, the coordinate system is aligned and the space coordinates are transformed, the spatial position relationship between the user's sight area and the screen boundary is integrated, the mapping correspondence between the sight area coordinates and the screen boundary coordinates is established, and the sight screen integrated space coordinates are generated; Based on the line-of-sight screen integrated space coordinates, the spatial matching relationship between the line-of-sight screen integrated space coordinates and the basic geometric depiction elements is determined to obtain the initial scene geometric structure information.

[0023] Specifically, based on the basic geometric descriptors obtained above, the three-dimensional coordinates corresponding to the line segment endpoints and their plane equation parameters are first retrieved from them to define the transformation matrix and prepare to transform and project the coordinates of the vision detection screen boundary. When executing specifically, the coordinate set of the line segment endpoints is first read and compared with the plane parameters. Corresponding matching, by Normalize and construct a rotation matrix to perform affine transformation on the subsequent screen boundary coordinates. If any coordinate component exceeds the pre-established valid range during the transformation, for example to The data is discarded in the interval to avoid positioning deviation. The interval is determined by collecting the coordinate distribution in the sample scene and observing the maximum and minimum values ​​and leaving a certain margin. Then the boundary coordinates of the four corners of the visual detection screen are Substitute the transformation matrix into the new coordinates after projection. If the absolute value of the coordinates of any vertex after projection is greater than 50, it means that the mapping relationship between the vision detection screen and the basic geometric depiction element is unstable. At this time, it is necessary to confirm again whether the input plane parameters and the line segment endpoint coordinates meet the projection requirements and recalculate the rotation matrix and displacement vector based on the results. Finally, after completing this part of the transformation, the screen boundary space mapping coordinates are formed.

[0024] Based on the obtained screen boundary space mapping coordinates, read the user's line of sight area coordinates and perform coordinate system alignment and space coordinate transformation. When executing, first clarify the collection method of the user's line of sight area coordinates. Here, a binocular camera combined with an eye tracking model is used to identify the user's pupil center and line of sight direction. The model is trained using a convolutional neural network and contains five convolutional layers and two fully connected layers. The input data is the normalized binocular image frame data. , the size of each frame image is fixed to The output is the estimated value of pupil coordinates and gaze direction vector. The training is carried out on a public dataset containing 20,000 labeled pupil positions. The stochastic gradient descent method is used for 100 cycles and the learning rate is set to 0.001. The Euclidean distance loss between the pupil position and the labeled true value is continuously minimized. To update the network weights, Indicates the index number of the image frame. Indicates The horizontal coordinate value of the pupil center point predicted in the frame image, Indicates The vertical coordinate value of the pupil center point predicted in the frame image, m is the total number of image frames involved in the training, For the The horizontal coordinate value of the pupil center point actually marked in the frame image, For the The vertical coordinate value of the pupil center point actually marked in the frame image. In the inference stage, the real-time acquired binocular image frame is input into the trained convolutional neural network and the pupil center position is output. and the sight direction vector Then, the above coordinates are triangulated according to the binocular baseline distance calibrated by the system to obtain the starting point of the three-dimensional line of sight. In the alignment process, the three-dimensional coordinates are first mapped to the same coordinate system as the screen boundary space mapping coordinates, and compared with the pre-established line of sight direction offset threshold of 0.02. This threshold is selected by counting the line of sight samples of five groups of users in the experiment. If the offset exceeds 0.02, the binocular image is recaptured and inferred again. Finally, the aligned user line of sight area coordinates are matched one by one with the screen boundary coordinates, and the mapping correspondence between the line of sight area coordinates and the screen boundary coordinates is merged to generate the line of sight screen integrated space coordinates.

[0025] Based on the above-mentioned line of sight screen integrated space coordinates, compare the coordinate mapping records in the basic geometric depiction element and confirm the corresponding index between the two, and set an allowable offset threshold when matching the line of sight screen integrated space coordinate segments with the corresponding plane coordinates. The threshold is based on the distribution range of most spatial position errors in previous scene tests, which is approximately 0.01 to 0.04. Take 0.05 to ensure coverage of large fluctuations. If the distance between a coordinate and the theoretical position calculated by the plane equation exceeds Then skip the coordinate pair and review the source of the abnormal item, then compare the remaining valid coordinates one by one, complete the spatial mapping verification, count all the successfully matched coordinates, summarize and output the initial scene geometry information.

[0026] The steps to obtain the camera relative pose sequence are: Based on the initial scene geometry information, the coordinates of the three-dimensional boundary points of the user's sight area and the boundary corner points of the vision detection screen are obtained. Combined with the image acquisition timestamps and camera internal parameters recorded in the continuous image frames, the matching errors and reprojection offsets of the corresponding boundary points between the image frames are calculated to generate a continuous image frame matching error sequence. According to the matching error sequence of continuous image frames, the camera posture change angle value is calculated. The calculation formula is: ; in, For the The camera posture change angle of the frame image, , For the The center pixel coordinates of the frame image, , is the center pixel coordinate of the image of the initial frame, For the The vertical resolution of the frame image, For the The horizontal resolution of the frame image, For the The horizontal viewing angle of the frame image, For the The vertical viewing angle of the frame image, For the The camera focal length of the frame image; According to the camera posture change angle value and the three-dimensional offset of the boundary points of each frame image in the continuous image frame matching error sequence, the three-dimensional rotation axis offset and direction vector change between adjacent frames are obtained to obtain the camera relative posture sequence.

[0027] Specifically, based on the initial scene geometry information, first summarize the coordinates of the corner points of the vision detection screen boundary and the three-dimensional boundary points of the user's line of sight area, and collect multiple frames of images in a monitoring environment equipped with a high-resolution imaging device as the data source, read and record the acquisition timestamp of each frame of the image together with the camera's intrinsic parameters, and obtain the focal length, principal point coordinates, distortion coefficient and other data in the camera's intrinsic parameters through a calibration process in advance. These calibration data refer to the calibration procedure given in the imaging system's instructions for use. For example, collect at least 50 images within a range of 1 to 3 meters from the calibration plane and then solve the camera parameters to obtain a relatively stable intrinsic parameter matrix. Then, compare the coordinates of the corresponding boundary points between adjacent frames in the same target scene multiple times to obtain the difference at the pixel level. If the pixel deviation of any corner point is greater than the event First, a critical threshold of 5 pixels is set (this threshold is selected after testing in the range of 10 to 20 pixels and checking the matching error rate). This part of the matching pairs is classified as abnormal matches for additional verification. In addition, when calculating the reprojection, the camera intrinsic parameter projection model needs to be applied to the coordinates of each boundary point, and the three-dimensional coordinates are projected onto the image plane according to the pinhole imaging model to generate the projection coordinates. The difference between the projection and the actual pixel coordinates is the reprojection offset. If the reprojection offset exceeds the established calibration range of 0 to 2 pixels, the matching point is marked as an over-offset point for statistics, and the remaining matching points in the same frame image are selected to calculate the mean square value of the overall offset, so as to construct a continuous image frame matching error sequence. Finally, the matching error information of all frames is summarized in chronological order to obtain a continuous image frame matching error sequence.

[0028] formula: The benefit of the formula is that it comprehensively considers factors such as the camera center pixel offset, actual resolution, camera viewing angle and focal length, and can present the pose changes between image frames with more intuitive angle values, thereby more clearly expressing the dynamic changes of the camera pose in subsequent steps.

[0029] The steps to obtain the parameters are as follows: The pixel coordinate system of the frame image reads the horizontal coordinate value of the actual center pixel position. The horizontal coordinate value is taken from the image with a pre-calibrated resolution range of 0 to 4000. The specific method is to record the actual pixel width of each frame image in an environment with a higher shooting resolution, thereby obtaining a pixel width sequence , will The actual width of the frame image is recorded as , and then find the center pixel column index in this frame image , the actual value of the column index is set to For example, in a test scene, the width of the fifth frame image is 1920 pixels. , if we observe that there are 5-pixel black borders on the left and right sides of the image, the actual effective screen width is 1910 pixels, so That is The specific value of .

[0030] The steps to obtain the parameters are as follows: The vertical pixel coordinates of the center of the frame image are located in combination with the shooting resolution. In an image with a height of more than 3000 pixels, the vertical pixel midline is marked as , and then observe whether there are upper and lower black edges or embedded graphic areas. If so, subtract and adjust according to the actual number of pixel rows of the upper or lower edge detected, so that the final value can accurately correspond to the center line. In addition, in the panoramic shooting environment, some images have custom text or border information, and their pixel rows need to be counted and subtracted. Finally, record the obtained For example, if the image height is 1080 pixels in a certain shot and there are 7 pixels of blank space above and 3 pixels of blank space below, then the effective height is , center ordinate That is regarded as .

[0031] The parameter acquisition step is to read the horizontal coordinates of the central pixel of the initial frame image. The initial frame is fixed as frame number 1 at the beginning of shooting. For the initial frame, it is also necessary to first count its actual resolution width and possible black edges or nested areas. To ensure accuracy, a series of feature points should be extracted from the initial frame and the actual effective width of the left and right boundaries of the image should be solved, and then divided by 2 to obtain the central line index, which is recorded as , if the width of the initial frame in an actual shooting is 2560 pixels, and there is no black border correction on the top, bottom or left and right, then .

[0032] The parameter acquisition step is to divide the initial frame image height by 2 to obtain the longitudinal midline position corresponding to the central pixel longitudinal coordinate of the initial frame image. If it is detected that there are dozens of pixels above or below the screen at the time of shooting, these occupied areas need to be removed before the midline positioning and then the midline solution is performed. For example, in a certain detection, the height is 1440 pixels and only 5 pixels of blank area are left above, and there is no blank area below, then the effective image height is ,correspond .

[0033] The steps to obtain the parameters are: read The vertical resolution of the frame image is combined with the system calibration information. The vertical resolution can usually be automatically identified as a value between 480 and 3000 when the imaging device is shooting. The specific record is as follows .

[0034] The steps to obtain the parameters are: read The horizontal resolution of the frame image.

[0035] The steps to obtain the parameters are as follows: The horizontal viewing angle information of the frame image is usually determined by the physical camera lens structure and focal length. The image is taken point by point on a plane fixed at 1 to 3 meters away from the lens, and the horizontal range covered in the captured image is recorded as arc or angle values. These angles are averaged and compared with the accuracy correction value to obtain the final value. For example, when the focal length of a lens is 10 mm, the actual horizontal viewing angle is measured to be 85 degrees. 85 degrees is converted into 1.48353 radians for recording.

[0036] The steps to obtain the parameters are as follows: Similar to the original, but changed to vertical viewing angle, through the same field shooting and mapping method to confirm that the vertical viewing angle is about 60 degrees when the lens focal length is 10 mm, that is, 1.0472 arc, thus obtained And record.

[0037] The steps to obtain the parameters are as follows: The camera focal length of the frame image, obtained from lens specifications or software and hardware queries.

[0038] Calculation process: In a demonstration example, let The specific parameter values ​​corresponding to the frame image are , , , , , , , , Millimeters, calculate the numerator first : ; ; Take the absolute value , then calculate the denominator : ; ; Ratio the numerator to the denominator , and then find the inverse tangent radians, convert them to degrees , get the angle value of the camera posture change in this frame .

[0039] The result shows that there is a significant field of view deviation between the camera in this frame and the initial frame. When subsequently processing the three-dimensional rotation axis offset and direction vector change of adjacent frames, this angle value can be used as one of the inputs and combined with the three-dimensional offset of the boundary points in the continuous image frame matching error sequence to finally obtain the camera relative pose sequence.

[0040] The steps to obtain the user screen dynamic space vector are: Based on the relative position sequence of the camera, the rotation matrix and translation vector of each frame are obtained. Combined with the focal length, image center coordinates, pixel size and image size parameters in the camera intrinsic parameters, the coordinates and spatial vectors of the starting point coordinates of the user's line of sight area and the screen corner coordinates are reconstructed in turn to generate the initial three-dimensional vector group from the user's line of sight to the screen; According to the initial three-dimensional vector group from the user's line of sight to the screen, the spatial transformation value from the user's line of sight vector to the target position on the screen is calculated. The calculation formula is: ; in, is the spatial transformation value from the user's sight vector to the target position on the screen, is the starting point coordinate of the user's sight vector, is the target screen corner coordinate, is the camera focal length, is the coordinate of the center point of the image, is the image width, is the image height; Based on the spatial transformation value from the user's line of sight vector to the screen target position, vector direction normalization and projection difference correction are performed to obtain the user's screen dynamic space vector.

[0041] Specifically, the rotation matrix and translation vector corresponding to each frame are read based on the relative position sequence of the camera. These rotation matrices and translation vectors are first uniformly numbered and matched in chronological order, and the coordinate transformation parameters recorded therein are retrieved. The number can be assigned to each frame of the image in sequence in the shooting environment for subsequent comparison. Each rotation matrix is ​​converted from the reference coordinate system and the camera lens coordinate system obtained by on-site calibration. It is necessary to collect multiple sets of feature point positions during calibration and solve them frame by frame to obtain the matrix. Then the translation vector is calculated based on the relative offset between the actual position mark and the camera optical center. These information with time series labels are sorted and corresponded one by one to the starting point coordinates of the user's line of sight area and the screen corner coordinates. Then, according to the recorded pixel size and image size parameters, coordinate system mapping and three-dimensional vector stitching operations are performed on these coordinate values, and the camera is converted using Euler angle conversion or quaternion transformation. The coordinate system is the same as the global coordinate system of the scene. Then, the coordinate landing point is verified in the shooting space to be within a preset valid range, for example, within the horizontal coordinate range of -500.0 to 500.0. Points beyond this range are additionally compared or discarded. After completing the range verification, the starting coordinates of the user's line of sight area and the screen corner coordinates are respectively processed by three-dimensional vectorization. Each vector carries a corresponding time index and a rotation and translation correction value. If it is detected in a certain operation that the cumulative deviation of the vector direction and the coordinate is greater than the preset 5.0-degree direction deviation threshold, the rotation matrix of the most recent frames needs to be read again for verification. The threshold of 5.0 degrees is obtained by counting the line of sight direction offset values ​​and taking the median plus a 1-degree margin when testing multiple scenes. Finally, after accumulating these steps, the starting coordinates of the user's line of sight area and the screen corner coordinates can be converted from a simple three-dimensional point set to a continuous initial three-dimensional vector group.

[0042] formula The benefit of the formula is that it analyzes the coordinate difference between the user's sight vector and the screen target position in segments, and combines the focal length and image width and height parameters at the same time. Through the joint action of the logarithmic function and the inverse tangent function, the horizontal and vertical components are integrated into a unified spatial transformation value.

[0043] The steps to obtain the parameters are to extract the horizontal component from the starting coordinates of the user's line of sight vector. It is necessary to first deploy an eye tracking device at the shooting site and record the user's eye position. The pupil center coordinates are read using a calibration method based on the binocular imaging principle and marked as The lateral component of The specific calibration method is to take binocular photos within a certain range from the user and track the pupil center frame by frame, and form a set of horizontal coordinate data through multiple measurements. The coordinate system is compared with the pre-calibrated three-dimensional coordinate system. If there is a discrepancy, the numerical correction is performed based on the field survey data. For example, in one calibration, the user's eye is measured to be about 50.3 cm away from the reference origin in the X direction. .

[0044] The steps for obtaining the parameters are as follows: the lateral component of the image center coordinates needs to determine the camera principal point position when the imaging device is initialized and converted into the reference value under the same three-dimensional scene. First, the resolution midpoint of the lens center in the image plane coordinates is recorded as , and then adapt it to the real object coordinates collected in the scene calibration. If the lens is facing forward, it corresponds to a world coordinate origin offset value, which is corrected as For example, if the lens is confirmed to be offset by 2.0 cm in the X direction during calibration, this information is incorporated into the coordinate value of the image center, so that .

[0045] The steps for obtaining the parameters are as follows: the focal length value of the camera is obtained after reading the internal parameters of the camera or related documents.

[0046] The parameter acquisition steps are as follows: the pixel size of the image width is directly read through the resolution information of the imaging device. In one verification, an image with a resolution of 1920×1080 is captured, so .

[0047] The steps to obtain the parameters are as follows: Similarly, the pixel size used to represent the image height, for example, the image height is nominally 1080 pixels when shooting on the spot. .

[0048] The steps of obtaining the parameters are as follows: extract the lateral component from the coordinates of the corner points of the target screen. In order to make the coordinates of the corner points consistent with the shooting coordinate system, the screen position needs to be measured in three dimensions. Obvious marks can be attached to the front of the screen and the X-direction coordinates can be obtained by using a laser rangefinder and multi-angle shooting. The X-direction coordinates are then converted into specific values ​​in the scene coordinate system. If the measured coordinates deviate from the pre-designed drawings by no more than 2 cm, they are recorded as For example, if the X coordinate of the upper left corner of a display is 120.4 cm measured from the scene origin, then .

[0049] Calculation process: In one example, the collected , , , , , ; Calculate the numerator first , and then multiply it by get ; because , so the molecular part is ; Then , then Performing absolute value ,calculate ; Again Do the evaluation, , Radians, multiply the two , then square , add the two parts together , take its square root and we get .

[0050] The result shows that in this example, the spatial transformation value between the starting point of the user's line of sight vector and the target position on the screen is approximately 2.175. When the value is too large, it means that the difference between the starting point and the target point in the X direction or the aspect ratio is obvious. If the value is small, it means that the distance or difference between the two is relatively low. The result can be used to generate the final user screen dynamic space vector by subsequent normalization of the vector direction and projection difference correction.

[0051] Based on the spatial transformation value from the user's line of sight vector to the target position on the screen, the value is first segmented and detected and normalized in combination with the direction vector of the user's line of sight vector. The direction vector used comes from the previously recorded three-dimensional coordinate difference. The three-dimensional difference needs to be squared and then squared in the X, Y, and Z directions to obtain the vector modulus. Then each component is divided by this modulus to normalize the overall length of the vector to 1. If a component is detected to exceed the established absolute threshold of 500.0 when calculating the modulus, it is necessary to recheck the previous observation data for calibration errors or system abnormalities. This threshold of 500.0 refers to the scene scale. Select the middle segment in the range of -1000.0 to 1000.0 and perform After multiple dynamic tests, half of the value is taken. The vector components that exceed the threshold are recorded and then manually reviewed and deleted or corrected as appropriate. After normalization, the projection differences are compared, and the position difference between the screen plane and the user's line of sight vector in the two-dimensional projection is calculated in the same coordinate system. If it is found that the projection distance between the two exceeds the previously statistical average offset of 3.0 pixels, the user's head posture tracking information at the same time is retrieved to confirm whether the head is quickly turned or blocked. Once the cause is locked, the coordinates can be updated in combination with the actual scene record and the projection difference can be calculated again. The final qualified data will be summarized and marked as valid vector information, thereby obtaining the user's screen dynamic space vector.

[0052] The steps to obtain the single-view feature positioning table are: Acquire multi-view image data of the vision testing site, extract pixel information of the user behavior feature area and the scene key element area in the multi-view image data of the vision testing site, identify the three-dimensional position information of the user behavior feature points through spatial mapping, and generate a set of spatial coordinates of the user behavior feature points; Based on the spatial coordinate set of user behavior feature points, identify the contour edges of key elements of the scene in the multi-view image data of the vision detection site, and perform spatial coordinate matching and three-dimensional space fitting to generate a three-dimensional spatial position coordinate set of key elements of the scene; Based on the three-dimensional spatial position coordinate set of the scene's key elements, the three-dimensional spatial position coordinates of the user behavior feature points and the scene's key elements under the corresponding perspective are summarized to establish a single-perspective element positioning table.

[0053] Specifically, obtain multi-view image data at the vision test site, first initialize and check the shooting position and optical parameters for each view, measure the distance between the camera and each calibration point of the scene to ensure that all views are in the same reference coordinate system, and then extract pixel information involving user behavior feature areas and scene key element areas from each frame of the image. When performing the extraction, determine whether the pixel falls within the specified user activity color interval or key element contour interval based on the brightness range of 0 to 255 and the color value distribution range obtained in advance. If the pixel feature value does not match the above range, it will be discarded to avoid mixing with interfering data. At the same time, to ensure the accuracy of the extraction process, a pixel connectivity threshold of 10 needs to be set in each view. This 10 is obtained by trying values ​​from 5 to 20 in multiple sample images and observing the user's body or The connected area on the contour of the key element is selected after the effect. Only when the number of adjacent pixels reaches 10, they are merged into the same connected block and marked as suspicious areas. After completing the marking of suspicious areas in multi-view data, the image frame number and its pixel position corresponding to each view are converted one by one. The conversion method is to map the pixel coordinates to the three-dimensional coordinate system according to the camera intrinsic parameter matching method. The coordinates whose values ​​exceed the set valid range of -1000.0 to 1000.0 during the mapping process are considered invalid and excluded. The remaining three-dimensional coordinate points are grouped in a spatial clustering manner. The same feature area of ​​different viewpoints is considered to be part of the same object if the difference in X, Y, and Z coordinates is less than 5.0. Finally, the three-dimensional coordinates corresponding to all confirmed user activity pixel blocks are recorded to form a three-dimensional coordinate set of user behavior feature points.

[0054] Based on the spatial coordinate set of user behavior feature points, when obtaining the contour edges of key elements of the scene appearing in the multi-view image data of the vision test site, we must first select the color or shape feature values ​​that can represent obvious edges based on the characteristics of the key elements of the scene, such as the edge of the screen, the desktop or the outline of the device shell, and confirm them frame by frame in the multi-view image. Use the same method to filter out pixel blocks that meet the contours of the key elements at the pixel level and convert them to a three-dimensional coordinate system for grouping. If there is a local continuous blank greater than 10 cm on the edge of the key element of the scene in the X, Y or Z direction, the contour segment needs to be segmented. This 10 cm comes from the actual measurement. The upper limit of the spacing of the surface continuity of the key elements is observed. After the segmentation is completed, the 3D points of each contour segment need to be matched with the spatial coordinate set of the user behavior feature points. The specific matching method is to calculate the Euclidean distance in the X, Y, and Z directions and compare whether they are in the range of 0.00 to 2.00 cm. If they exceed 2.00 cm, they are not considered to be the same object. All successfully matched contour edge points are then fitted by least squares to generate fitted 3D edge data to express the complete spatial position of the scene key elements. These fitting results are then merged according to the physical area adjacent criterion to obtain the 3D spatial position coordinate set of the scene key elements.

[0055] Based on the three-dimensional spatial position coordinate set of the scene's key elements, the user behavior feature points and the three-dimensional coordinates of the scene's key elements at the same viewing angle are integrated. Each viewing angle is first numbered separately and the feature points and key element coordinates at that viewing angle are uniformly converted to the scene's global coordinate system. If it is detected that the coordinates of some of the data exceed the boundary range of -2000.0 to 2000.0 in the X or Y direction, it is necessary to compare the original image and calibration data for abnormal interference or calibration inaccuracy, because the actual test stipulates that the effective interaction area in the scene is mainly concentrated between -2000.0 and 2000.0. If an error is confirmed, these abnormal data points can be directly deleted, and then the three-dimensional coordinates of all feature points and key elements at the same viewing angle are recorded in turn, and these coordinates and the viewing angle numbers associated with them are indexed together. In this way, a single-view feature positioning table can be established.

[0056] The steps for obtaining the multi-view fusion scene state point set are: Based on the single-view feature positioning table, extract the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the image frame number, the camera internal parameter matrix and the user screen dynamic space vector coordinate endpoints under each view, and uniformly convert them to the unified coordinate system defined by the user screen dynamic space vector to generate a multi-view three-dimensional positioning matrix set in the unified coordinate system; According to the multi-view 3D positioning matrix set in the unified coordinate system, the fusion stability offset of each view is calculated. The calculation formula is: ; in, For the Perspective fusion stabilization offset, is the coordinate of the user behavior feature point, are the coordinates of the key elements of the scene, is the screen vector end point coordinate, is the camera optical center coordinate, is the perspective number, is the image frame index; Based on the fusion stable offset, the point pairs of each perspective are traversed, and the spatial position fusion is prioritized in the order of the fusion stable offset from small to large. The point pairs that meet the spatial overlap conditions are selected and merged into a unified spatial point position to establish a multi-perspective fusion scene state point set.

[0057] Specifically, when extracting the three-dimensional coordinates of the user behavior feature points of each perspective, the three-dimensional coordinates of the scene key elements, the image frame number and other numerical contents based on the single-perspective element positioning table and converting them into a unified coordinate system defined by the user screen dynamic space vector, first compare them item by item according to the recorded camera intrinsic parameter matrix and the resolution of each frame of the image during the shooting process and the shooting visual axis angle. If the intrinsic parameter matrix of a certain perspective does not match the actual resolution information, it is necessary to compare the original calibration data again and check the resolution correction process. After confirmation, the coordinates of the user behavior feature points and the coordinates of the scene key elements are read out of the X, Y, and Z components respectively, and combined with the corresponding image frame number and the endpoints of the user screen dynamic space vector coordinates obtained in the early stage to form a basic data set within a perspective. Here, when putting these data sets and coordinates into the same coordinate system according to the perspective number, it is necessary to first take the camera origin or optical center position into consideration. If the light of a certain perspective is detected If the center of the image is offset from the origin of the unified coordinate system by a large margin exceeding 2 meters, it is necessary to record the offset value and perform vector compensation. The 2-meter range is the detection standard selected for general vision detection scenes through indoor tests. If it exceeds this range, projection offset is very likely to occur. After the correction is completed, the projection matrix or the rotation and translation matrix must be applied to the three-dimensional coordinates of the above-mentioned feature points, and each point must be checked to see if the transformed values ​​fall within the global range of -500 to 500. If the number of points exceeds this range, it is necessary to check the camera posture or the distribution area of ​​the detection scene to avoid mismatches or abnormal coordinates that may cause subsequent data interference. Finally, the coordinates of the user behavior feature points of each perspective, the coordinates of the key elements of the scene, the image frame number, and the endpoints of the user screen dynamic space vector coordinates are mapped to the unified coordinate system and recombined into a multi-perspective three-dimensional positioning matrix. If other data comparisons are needed later, the matrix structure can be expanded to generate a multi-perspective three-dimensional positioning matrix set in the unified coordinate system.

[0058] formula: ,The benefit of the formula is that it integrates multiple spatial relationships such as ,user behavior feature points and key elements of the scene as well as the ,camera optical center and the screen vector endpoint, and calculates the ,stable offset degree of positioning data at different ,viewing angles by combining the cross product and the Euclidean distance.

[0059] The steps for obtaining the parameters are as follows: the coordinates of the user behavior feature points are in the The three-dimensional position under each perspective is first captured by multiple cameras arranged on site and the corresponding image frame numbers are aligned with the feature point recognition results. Each user behavior area is marked on the pixel plane and mapped to the global coordinate system. During mapping, the calibrated camera focal length and camera posture need to be collected, and the pixel coordinates are converted into spatial coordinates. If the coordinate value exceeds the range of -1000 to 1000, the point needs to be re-aligned to confirm that the final three-dimensional coordinates fall within the specified interval and are stored as For example, in a scenario test The facial feature area captured by the perspective is located at the pixel coordinates (350, 250), and the three-dimensional coordinates are (120.5, 48.2, 80.0) after projection. At this time, it is detected that all these values ​​are within [-1000, 1000], which can be recorded as .

[0060] The steps for obtaining parameters are as follows: the coordinates of the key elements of the scene are in the To determine the three-dimensional position under a certain viewing angle, it is necessary to extract the contours of known key elements (such as the screen edge and the device shell) and combine them with the calibration information to obtain their coordinates in a unified coordinate system. For example, the coordinates of a vertex on the side of a display screen detected at a certain viewing angle are (220.0, 55.0, 5.0), which is called .

[0061] The steps to obtain the parameters are as follows: the screen vector end point coordinates are in the The terminal position of the user screen dynamic space vector captured by the camera under each viewing angle is reflected. It is necessary to combine the user screen dynamic space vector calculated previously and transform it under this viewing angle. The specific method is: starting from the starting coordinate of the user screen dynamic space vector, add the vector length along the vector direction to obtain the terminal coordinate. If it is found in actual measurement or tracking that the terminal coordinate is offset by more than 10 cm when aligned with the scene, it is necessary to check the angle difference between the vector direction and the actual inclination of the screen. The 10 cm here is due to the observation in a series of tests that less than 10 cm is often considered to be more accurately aligned. If it is indeed a viewing angle problem, re-read the camera posture correction and then perform the vector terminal calculation. After completion, it can be obtained For example, in a calculation, the direction vector of the screen vector is (1.5, 2.0, 0.8), the length is set to 3.5 meters, the starting point is (0, 0, 0), then the end is (1.5×3.5, 2.0×3.5, 0.8×3.5)=(5.25, 7.0, 2.8), and it is translated to a reference point in the scene according to the coordinates. If the reference point is (100, 100, 0), then .

[0062] The steps for obtaining the parameters are as follows: the camera optical center coordinates are in the The three-dimensional position under the perspective, such as When , the optical center position is measured (0, 150.0, 100.0).

[0063] The steps to obtain the parameters are as follows: number the viewing angles. Each viewing angle is determined by shooting with one camera or the same camera at different positions / angles. When setting up on site, you can mark the order at the camera bracket or gimbal position. If a total of 8 cameras are set, Numbered from 1 to 8.

[0064] The steps for obtaining the parameters are as follows: the image frame index is in the A set of serial numbers under each viewing angle. Each viewing angle may capture multiple frames of images, and each frame needs to be processed separately. The frame index can be converted from the shooting timestamp. For example, the cumulative number of frames after the start of shooting is recorded as 1, 2, 3..., or the internal serial number of the camera is directly taken. If the shooting rate of a camera is 30 frames / second, the data captured at the 2nd second is marked as frame 60, and the corresponding Perspective .

[0065] Calculation process: In an example, let , , , , , , first calculate the numerator : ; ; Take the cross product of these two vectors: ; The crossover operation is expanded into , and then we get the result after we put in the value (the calculation example omits the intermediate steps): ; It can be calculated step by step Approximately , and then calculate the norm of the vector , and then divided by ,in , , add the two together , first calculate the norm of the previous cross product result , the molecule is , the denominator , the ratio is about , the square of this part is about , plus , take this part and do simple addition , and then take the square root to get .

[0066] The result shows that the fusion stability offset of view number 3 is about 38.94. When the value is small, it means that the user behavior feature points are relatively consistent with the key elements of the scene and the screen vector endpoints in space. If the value is large, it suggests that there may be measurement or matching errors. Subsequently, this value will be compared with other viewpoints to judge the fusion stability and determine the point pair merging strategy.

[0067] Based on the above fused stable offset, when traversing each perspective positioning point pair, first Sort in ascending order and create a mapping record for each view number and its corresponding offset. If some If it is greater than 40.0, you should check whether there is a significant difference in the calibration parameters or shooting timing used for this perspective. This 40.0 is an obvious mismatch limit summarized by multiple measurements in actual scenes. When it is detected that the value is greater than 40.0, the perspective point pair will be separated and left for further manual confirmation. For the remaining point pairs with values ​​less than or equal to 40.0, position fusion is performed in ascending order, and similar coordinates are averaged or weighted merged. If the distance between a certain coordinate and an existing coordinate is less than 2.0 cm, it is considered to be mergeable, otherwise it is retained as an independent point. In this way, the repeated points in multiple shooting perspectives can be merged into a single coordinate one by one, realizing the unified positioning processing of key elements and user behavior feature points under different camera perspectives. After completing the processing of all point pairs, the multi-perspective fusion scene state point set can be obtained.

[0068] The steps for obtaining the spatial state consistency determination result are as follows: Based on the multi-view fusion scene state point set, the corresponding three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the camera optical center position, the camera visual axis vector, the user screen dynamic space vector direction and the user behavior feature point orientation vector are extracted at each view to generate a multi-view point direction comparison matrix; According to the multi-view point direction comparison matrix, the spatial position difference value in the unified coordinate system is calculated. The calculation formula is: ; in, is the spatial position difference value, and are the coordinates of the user behavior feature points under two perspectives, is the coordinates of the key elements of the scene at the current perspective, is the camera optical center coordinate, is the camera viewing axis vector, It is the direction vector of the user screen dynamic space vector; Based on the spatial position difference value, a comparison is performed one by one with the set spatial tolerance threshold, all matching point pairs that meet the spatial tolerance threshold are screened, and a unified judgment result is output based on the consistent matching of the point pairs in all viewing angles to obtain the spatial state consistency judgment result.

[0069] Specifically, when extracting the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the camera optical center position, the camera visual axis vector, the user screen dynamic space vector direction and the user behavior feature point orientation vector at each viewing angle based on the multi-view fusion scene state point set, it is necessary to first read the previously recorded calibration data in each shooting viewing angle and calibrate the optical center and visual axis position of the camera at that viewing angle. To ensure the accuracy of the data, several reference objects with known coordinates can be placed in the shooting environment and the deviation between the coordinates measured by the camera and the reference value can be checked after shooting. When the deviation is greater than the 3 cm range determined in advance based on experimental statistics, the camera external parameters should be recalibrated and updated. Then, the coordinates of the user behavior feature points and the coordinates of the scene key elements are compared in the X, Y, and Z directions respectively and associated with the camera optical center position and the camera visual axis vector. When the distance value between a point and the optical center in the X direction exceeds the set maximum distance threshold of 600.0, it is necessary to check whether there is a long-range lens or scene occlusion in the shooting script. The 600.0 is obtained by measuring the size of the indoor vision test site with a certain margin. If it is verified to be a long-range shot, it is necessary to read additional shooting position information for position conversion to avoid error accumulation when the coordinates or direction vectors are subsequently merged. Then, the orientation vector of the feature point is compared with the screen vector according to the recorded user screen dynamic space vector direction option. A separate label is given to the situation where the deviation exceeds 5 degrees. The 5 degrees is obtained by selecting the median plus 2 degrees from the statistical sequence of relative error angles of each camera position in multi-view shooting. Subsequently, all multi-view points that meet the data quality requirements are uniformly recorded in an index table, and the correspondence between the camera's visual axis vector and the camera's optical center position is marked one by one in the index table. The mismatched or abnormal records are reviewed. After these steps are completed, a multi-view point direction comparison matrix covering the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the camera optical center position, the camera's visual axis vector, the user screen dynamic space vector direction, and the user behavior feature point orientation vector can be obtained.

[0070] formula The benefit of this method is that it incorporates multiple factors such as user behavior feature points of two perspectives, camera visual axis, scene key elements, and user screen dynamic space vector direction into the same evaluation index under a unified coordinate system. With the help of a combination of vector cross product, vector dot product, and Euclidean distance, the difference amount includes both directional and positional measurements, which is very intuitive for identifying the consistency of multi-perspective data.

[0071] The steps to obtain the parameter are as follows: The acquisition process of the three-dimensional coordinates of the user behavior feature points under each perspective relies on multiple cameras on site to simultaneously capture and record the user activity scene. For example, in a certain measurement phase, after multi-camera fusion, it is known that the user's location coordinates (120.0, 55.5, 98.0) are all within the range of [-500, 500], which can be registered as or Equivalent number.

[0072] The steps to obtain the parameter are as follows: Corresponding but from another perspective.

[0073] The steps to obtain the parameters are as follows: The camera's viewing axis vector of a viewing angle is generated by positioning the lens direction during camera calibration. This is done by measuring the angles between the camera and several points in the scene with marked coordinates and combining the three-dimensional coordinate difference to obtain the X, Y, and Z components of the viewing axis direction. For example, (0.866, 0.0, 0.5) represents a direction that is 30 degrees tilted in the horizontal plane and tilted upward.

[0074] The steps to obtain the parameters are: current viewing angle The coordinates of the key elements of the scene under the scene are mainly for elements with stable and unchanging positions, such as the screen surface and the fixed structure of the device. These elements need to be 3D scanned or laser measured when arranging the scene, and their coordinate values ​​​​are registered in the scene database. After that, this coordinate is combined with the camera intrinsic parameter to form the world coordinate mapping under this perspective.

[0075] The steps to obtain the parameters are as follows: The optical center coordinates of the camera at each viewing angle are usually measured repeatedly on site using a laser rangefinder or a special calibration fixture to repeatedly measure the offset of the camera at the origin of the 3D scene. After the measurement, it is recorded as (ox, oy, oz) corresponding to , for example, in one arrangement, the displacement is (10.0, 0.0, 50.0) mm, then =(0.01, 0.0, 0.05) m.

[0076] The steps for obtaining the parameters are as follows: the direction vector of the user's screen dynamic space vector corresponds to the screen vector direction calculated and recorded in the previous step. When the direction component needs to be determined, the orientation of the edge of the display device can be measured in the static calibration stage, and then the relative orientation between the user and the display device is obtained based on real-time monitoring. On this basis, the unit vector components in the X, Y, and Z directions are calculated.

[0077] Calculation process: In an example calculation, select , , , , , ; Calculate first ,in ,and Do the cross multiplication , norm ; Then calculate the vector and The dot product of , norm , after normalization , , norm , normalized to approximately , the dot product of the two , absolute value ; Then calculate , , , its norm , and finally add the three parts together ,Right now .

[0078] The result shows that there is a certain degree of spatial difference between the user feature points of the two perspectives in this example scenario. The larger the value, the more obvious the deviation in position or direction. This result can be compared with the pre-defined spatial tolerance threshold to determine whether it meets the matching requirements.

[0079] Based on the spatial position difference value, each pair of user behavior feature points and scene key elements are checked against the camera optical center. A spatial tolerance threshold is required during the on-site survey phase to measure whether the coordinate difference is acceptable. The threshold can be measured and recorded based on the size of the site and the user's activity range. For example, in an indoor site, the position coordinates are limited to a range of -500 to 500 meters. Combined with the previous data statistics, a threshold interval of 40.0 to 60.0 is set and fixed to 50.0 after multiple tests. When a pair of difference values ​​is detected to be greater than 50.0, a check is made for the pair of feature points in the comparison list. Mark and wait for subsequent verification. If it is less than or equal to 50.0, it is temporarily classified as within the matching range. Then, the points within the matching range are checked for consistency from each perspective. If the same feature point does not exceed the above threshold in multiple perspectives, it will be merged into the same global coordinate point. If there is a problem of large differences between some perspectives, it is necessary to conduct in-depth comparison in camera calibration or user movement trajectory records to avoid subsequent integration deviations due to inconsistent data. All qualified points are recorded one by one and summarized in the final comparison result index table. Finally, the spatial state consistency judgment result can be obtained.

[0080] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.

Claims

1. An unattended vision detection method based on image recognition, characterized in that: The following steps are involved: Collect image data from the vision test site, perform scene structure line detection and plane detection to obtain line segment sets and plane parameters, and establish basic geometric depiction elements; Based on the basic geometric depiction element, spatial integration is performed by combining the coordinates of the user's sight area obtained by detection with the coordinates of the vision detection screen boundary to obtain the initial scene geometric structure information; Based on the initial scene geometric structure information, tracking the camera posture changes in continuous image frames, calculating and obtaining the displacement and rotation parameters of the camera relative to the initial geometric structure, obtaining a camera relative posture sequence, and based on the camera relative posture sequence, constructing a three-dimensional space coordinate system and solving the coordinate transformation from the user's sight area to the vision detection screen to generate a user screen dynamic space vector; Acquire multi-view image data of the vision test site, identify the user behavior feature points and the three-dimensional position information of the key elements of the scene at each view, compile them into a single-view element positioning table, perform cross-view data mapping and fusion based on the unified coordinate system defined by the single-view element positioning table and the user screen dynamic space vector, and establish a multi-view fusion scene state point set; Based on the multi-perspective fusion scene state point set, the corresponding user behavior feature points and scene elements under different perspectives are matched, the spatial position difference value in the unified coordinate system is calculated, the spatial position difference value is compared with the preset spatial tolerance threshold, and the spatial state consistency judgment result is obtained.

2. The unattended vision detection method based on image recognition according to claim 1, characterized in that: The steps for obtaining the basic geometric depiction element are: Collecting image data of the vision detection scene, performing image edge gradient analysis based on the collected image data of the vision detection scene to calculate the pixel gradient change of the edge point, extracting the edge contour at the boundary of each object in the scene, combining the edge points to form structural line segments, and generating a vision detection scene structural line segment set; Based on the vision detection scene structure line segment set, extract the spatial coordinate data of all line segment endpoints in the line segment set, determine the coplanarity of the line segment endpoint spatial coordinates, select the coplanar point group according to the line segment endpoint coordinate coplanarity, fit and calculate the plane equation parameters corresponding to the spatial position of the coplanar point group, and generate the plane parameters corresponding to the structure line segment; Based on the plane parameters corresponding to the structural line segments, the line segment endpoint coordinates in the visual inspection site structural line segment set are called, and the line segment endpoint spatial coordinates and the corresponding plane parameters are spatially projected one by one to determine the spatial position relationship between the line segment and the corresponding plane, and establish a spatial mapping relationship between the line segment endpoint coordinates and the plane parameters to generate basic geometric depiction elements.

3. The unattended vision detection method based on image recognition according to claim 1, characterized in that: The steps for obtaining the initial scene geometric structure information are as follows: Based on the basic geometric depiction element, the line segment endpoint coordinates and plane parameters in the basic geometric depiction element are called, combined with the vision detection screen boundary coordinates, spatial coordinate transformation and projection calculation are performed to determine the spatial position relationship of the vision detection screen in the basic geometric depiction element, and generate the screen boundary space mapping coordinates; Based on the screen boundary space mapping coordinates, the coordinates of the user's sight area are obtained, coordinate system alignment and space coordinate transformation are performed, the spatial position relationship between the user's sight area and the screen boundary is integrated, a mapping correspondence relationship between the sight area coordinates and the screen boundary coordinates is established, and the sight screen integrated space coordinates are generated; Based on the line-of-sight screen integrated space coordinates, a spatial matching relationship between the line-of-sight screen integrated space coordinates and basic geometric depiction elements is determined to obtain initial scene geometric structure information.

4. The unattended vision detection method based on image recognition according to claim 1, characterized in that: The steps for obtaining the relative position sequence of the camera are: Based on the initial scene geometric structure information, the coordinates of the three-dimensional boundary points of the user's sight area and the boundary corner points of the vision detection screen are obtained, and the matching errors and reprojection offsets of the corresponding boundary points between the image frames are calculated by combining the image acquisition timestamps and camera internal parameters recorded in the continuous image frames to generate a continuous image frame matching error sequence; Calculating the camera posture change angle value according to the continuous image frame matching error sequence; According to the camera posture change angle value, combined with the three-dimensional offset of the boundary points of each frame image in the continuous image frame matching error sequence, the three-dimensional rotation axis offset and direction vector change between adjacent frames are obtained to obtain the camera relative posture sequence.

5. The unattended vision detection method based on image recognition according to claim 1, characterized in that: The steps for obtaining the user screen dynamic space vector are: Based on the relative position sequence of the camera, the rotation matrix and translation vector of each frame are obtained, and the focal length, image center coordinates, pixel size and image size parameters in the camera intrinsic parameters are combined to sequentially reconstruct the coordinates and space vectors of the starting point coordinates of the user's sight area and the screen corner coordinates to generate an initial three-dimensional vector group from the user's sight to the screen; Calculate the spatial transformation value from the user's sight line vector to the target position on the screen according to the initial three-dimensional vector group from the user's sight line to the screen; Based on the spatial transformation value from the user's sight line vector to the screen target position, vector direction normalization and projection difference correction are performed to obtain the user screen dynamic space vector.

6. The unattended vision detection method based on image recognition according to claim 1, characterized in that: The steps for obtaining the single-view element positioning table are as follows: Acquire multi-view image data of the vision testing site, extract pixel information of the user behavior feature area and the scene key element area in the multi-view image data of the vision testing site, identify the three-dimensional position information of the user behavior feature points through spatial mapping, and generate a set of spatial coordinates of the user behavior feature points; Based on the spatial coordinate set of the user behavior feature points, the contour edges of the key elements of the scene in the multi-view image data of the vision detection scene are identified, and spatial coordinate matching and three-dimensional space fitting are performed to generate a three-dimensional space position coordinate set of the key elements of the scene; Based on the three-dimensional space position coordinate set of the scene key elements, the three-dimensional space position coordinates of the user behavior feature points and the scene key elements under the corresponding viewing angle are summarized to establish a single-view element positioning table.

7. The unattended vision detection method based on image recognition according to claim 1, characterized in that: The steps for obtaining the multi-view fusion scene state point set are: Based on the single-view element positioning table, extract the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the image frame number, the camera internal parameter matrix and the user screen dynamic space vector coordinate endpoints under each view, and uniformly convert them to the unified coordinate system defined by the user screen dynamic space vector to generate a multi-view three-dimensional positioning matrix set in the unified coordinate system; Calculating a fusion stabilization offset of each perspective according to the multi-perspective three-dimensional positioning matrix set in the unified coordinate system; Based on the fused stable offset, traverse the positioning point pairs of each perspective, prioritize the spatial position fusion in the order of the fused stable offset from small to large, select the point pairs that meet the spatial overlap conditions and merge them into a unified spatial point position, and establish a multi-perspective fusion scene state point set.

8. The unattended vision detection method based on image recognition according to claim 1, characterized in that: The steps for obtaining the spatial state consistency determination result are: Based on the multi-view fusion scene state point set, the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the camera optical center position, the camera visual axis vector, the user screen dynamic space vector direction and the user behavior feature point orientation vector corresponding to each view are extracted to generate a multi-view point direction comparison matrix; Calculate the spatial position difference value in the unified coordinate system according to the multi-view point direction comparison matrix; Based on the spatial position difference value, a comparison is performed one by one with the set spatial tolerance threshold, all matching point pairs that meet the spatial tolerance threshold are screened, and a unified judgment result is output based on the consistent matching of the point pairs in all viewing angles to obtain the spatial state consistency judgment result.

9. The unattended vision detection system according to any one of claims 1 to 8, characterized in that: include: The data acquisition module performs image acquisition based on the image data of the vision test site, obtains multi-angle and multi-view image data of the vision test site, performs scene structure line and plane detection, and constructs a scene geometric feature set by extracting line segment sets and plane parameters; The scene modeling module performs spatial integration based on the scene geometric feature set, and obtains the initial scene geometric structure by combining the user's visual area coordinates obtained by detection with the vision detection screen boundary coordinates; According to the geometric structure of the initial scene, the camera posture changes in continuous image frames are tracked, the relative posture of the camera is calculated, and then a three-dimensional space coordinate system is constructed to generate a camera relative posture sequence, and the coordinate transformation from the user's line of sight to the vision detection screen is solved to obtain the user's screen dynamic space vector; The gaze tracking module tracks the posture changes of continuous image frames based on the camera relative posture sequence and the user screen dynamic space vector, calculates the dynamic changes of gaze direction and user behavior, and generates a user gaze tracking sequence; The multi-view fusion module, based on the multi-view image data of the vision detection site, identifies the user behavior feature points and key elements of the scene at each view, obtains the three-dimensional space position data, and compiles it into a single-view element positioning table; combines the user's line of sight tracking sequence with the single-view element positioning table, performs cross-view data mapping and fusion, and generates a multi-view fusion scene state point set; The spatial consistency module matches the user behavior feature points and scene elements under different perspectives based on the multi-perspective fusion scene state point set, calculates the spatial position difference value, compares the spatial position difference value with the preset spatial tolerance threshold, judges the consistency, and obtains the spatial consistency judgment result.

Citation Information

Patent Citations

  • Multi-angle consistent plane detection and analysis method for monocular video scene three dimensional structure

    CN106570507A

  • Unattended group vision screening device

    CN111839453A

  • Multi-view human behavior recognition method and system under edge computing architecture

    CN113743221A

  • Multi-view target detection and model training method, system, equipment and medium

    CN118314497A

  • Method and device for constructing visual point cloud map

    WO2022002150A1

Cited By

  • Automatic detection method and device for mold penetration and coplanar defects of visual system

    CN120179566A

  • Gamma camera resolution imaging data processing method and system

    CN120563327A

  • Dynamic interference-oriented camera array online collaborative calibration method and system

    CN121120802A

  • A camera array online cooperative calibration method and system for dynamic interference

    CN121120802B

  • Bonding full-automatic detection system and method based on image recognition

    CN121639637A