An unattended vision detection method and system based on image recognition

By processing image data and fusion of multi-view angles at the vision detection site, the problems of spatial positioning error and line of sight changes in unattended vision detection are solved, and more stable and accurate vision detection results are achieved, ensuring the effectiveness and automation of the detection process.

CN120014054BActive Publication Date: 2025-07-08SHANGHAI CANGAO TRADE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510457628.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-08
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing unattended vision detection method based on image recognition has real-time spatial instability caused by large spatial positioning errors, changes in line of sight and changes in camera perspectives under complex spatial recognition. It is difficult to achieve effective fusion and correction of multi-view data, affecting the reliability and accuracy of detection results.

Method used

By collecting on-site image data for vision detection, a basic geometric delineation element is established, combining the spatial integration of the user's line of sight area and the boundary of the vision detection screen, tracking camera posture changes, obtaining relative pose sequences, building a three-dimensional spatial coordinate system, identifying user behavior feature points and key scene elements in multiple perspectives, mapping and fusion of cross-view data, calculating spatial position difference values and comparing them with tolerance thresholds, and judging the consistency of spatial state.

Benefits of technology

It improves the stability and accuracy of the vision detection process, ensures the effectiveness of unattended vision detection and the authenticity of evaluation results, can eliminate invalid detection data, and improves the intelligence and automation of the detection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014054B_ABST
    Figure CN120014054B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image recognition technology, and specifically to an unattended vision detection method and system based on image recognition, including the following steps: collecting image data of the vision detection site, performing scene structure line detection and plane detection to obtain a line segment set and plane parameters, and establishing basic geometric description elements; based on the basic geometric description elements, combining the coordinates of the user's line of sight area obtained by detection and the coordinates of the boundary of the vision detection screen for spatial integration to obtain initial scene geometric structure information. The present invention improves the stability of on-site spatial structure analysis by collecting image data of the vision detection site, performing scene structure line detection and plane detection, obtaining a line segment set and plane parameters, establishing basic geometric description elements, and realizing the construction of the initial scene geometric structure through the spatial integration of the user's line of sight area and the boundary of the vision detection screen.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to an unattended vision detection method and system based on image recognition. Background Art

[0002] The unattended vision detection method based on image recognition is a method that automatically analyzes and evaluates the vision state of users through computer vision. Its main purpose is to automatically perform vision detection and evaluation using image recognition technology without human supervision.

[0003] The existing technology only relies on basic computer vision methods to automatically analyze and evaluate the vision state of users. In the case of complex spatial structures or frequent changes in the relative positions of users, it is easy to cause large errors in spatial positioning, and it is difficult to effectively cope with the real-time spatial instability caused by changes in the line of sight and the camera perspective, resulting in poor stability of the vision evaluation results; at the same time, the lack of fusion processing of multi-view data makes it difficult to identify and correct the spatial position differences between the user's posture and the line of sight direction, resulting in the lack of spatial consistency verification of the scene analysis results, and the evaluation conclusions are easily affected by external interference factors, reducing the reliability and accuracy of the detection results. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of the present invention is to solve the shortcomings existing in the prior art, and to propose an unattended vision detection method and system based on image recognition.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions. An unattended vision detection method based on image recognition includes the following steps:

[0006] Collect image data of the vision detection site, perform scene structure line detection and plane detection to obtain a line segment set and plane parameters, and establish basic geometric descriptors; based on the basic geometric descriptors, combine the detected user line-of-sight area coordinates and the vision detection screen boundary coordinates for spatial integration to obtain initial scene geometric structure information;

[0007] Based on the initial scene geometric structure information, track the camera pose changes in consecutive image frames, calculate and obtain the displacement and rotation parameters of the camera relative to the initial geometry, obtain the camera relative pose sequence, and based on the camera relative pose sequence, construct a three-dimensional space coordinate system and solve the coordinate transformation from the user line-of-sight area to the vision detection screen to generate a user screen dynamic space vector;

[0008] Obtain multi-view image data of the vision detection site, respectively identify the three-dimensional position information of user behavior feature points and scene key elements under each view, collect them into a single-view element positioning table, and based on the unified coordinate system defined by the single-view element positioning table and the dynamic space vector of the user screen, perform cross-view data mapping and fusion to establish a multi-view fusion scene state point set;

[0009] Based on the multi-view fusion scene state point set, match the corresponding user behavior feature points and scene elements under different views, calculate the spatial position difference value in the unified coordinate system, compare the spatial position difference value with the preset spatial tolerance threshold for judgment, and obtain the spatial state consistency judgment result.

[0010] Preferably, the obtaining step of the basic geometric description element is as follows:

[0011] Collect image data of the vision detection site, based on the collected image data of the vision detection site, perform image edge gradient analysis to calculate the edge point pixel gradient change, extract the edge contours at the boundaries of each object in the scene, combine the edge points to form structural line segments, and generate a set of structural line segments of the vision detection site;

[0012] Based on the set of structural line segments of the vision detection site, extract the spatial coordinate data of all line segment endpoints in the line segment set, judge the coplanarity of the line segment endpoint spatial coordinates, screen the coplanar point groups according to the coplanarity of the line segment endpoint coordinates, fit and calculate the plane equation parameters corresponding to the spatial positions of the coplanar point groups, and generate the plane parameters corresponding to the structural line segments;

[0013] Based on the plane parameters corresponding to the structural line segments, call the line segment endpoint coordinates in the set of structural line segments of the vision detection site, perform spatial projection calculations on the line segment endpoint spatial coordinates and the corresponding plane parameters one by one, determine the spatial position relationship between the line segment and the corresponding plane, establish the spatial mapping relationship between the line segment endpoint coordinates and the plane parameters, and generate the basic geometric description element.

[0014] Preferably, the obtaining step of the initial scene geometric structure information is as follows:

[0015] Based on the basic geometric description element, call the line segment endpoint coordinates and plane parameters in the basic geometric description element, combine the vision detection screen boundary coordinates, perform spatial coordinate transformation and projection calculations, determine the spatial position relationship of the vision detection screen in the basic geometric description element, and generate the screen boundary spatial mapping coordinates;

[0016] Based on the screen boundary spatial mapping coordinates, obtain the user's line of sight area coordinates, perform coordinate system alignment and spatial coordinate transformation, integrate the spatial position relationship between the user's line of sight area and the screen boundary, establish the mapping correspondence between the line of sight area coordinates and the screen boundary coordinates, and generate the line of sight-screen integrated spatial coordinates;

[0017] Integrate spatial coordinates based on the line-of-sight screen to determine the spatial matching relationship between the integrated spatial coordinates of the line-of-sight screen and the basic geometric description elements, and obtain the initial scene geometric structure information.

[0018] Preferably, the step of obtaining the relative pose sequence of the camera is as follows:

[0019] Based on the initial scene geometric structure information, obtain the three-dimensional boundary points of the user's line-of-sight area and the boundary corner point coordinates of the vision detection screen. Combine the image acquisition timestamps recorded in consecutive image frames and the camera internal parameters, calculate the matching error and reprojection offset of the corresponding boundary points between image frames, and generate a sequence of matching errors for consecutive image frames;

[0020] Calculate the camera attitude change angle value according to the sequence of matching errors of consecutive image frames;

[0021] According to the camera attitude change angle value, combine the three-dimensional offset of the boundary points of each frame in the sequence of matching errors of consecutive image frames to obtain the three-dimensional rotation axis offset and direction vector change between adjacent frames, and obtain the relative pose sequence of the camera.

[0022] Preferably, the step of obtaining the dynamic spatial vector of the user screen is as follows:

[0023] Based on the relative pose sequence of the camera, obtain the rotation matrix and translation vector of each frame. Combine the focal length, image center coordinates, pixel size, and image size parameters in the camera internal parameters, and reconstruct the coordinates and spatial vectors of the starting point coordinates of the user's line-of-sight area and the screen corner point coordinates in turn to generate an initial three-dimensional vector group from the user's line of sight to the screen;

[0024] Calculate the spatial transformation value from the user's line-of-sight vector to the target position on the screen according to the initial three-dimensional vector group from the user's line of sight to the screen;

[0025] Based on the spatial transformation value from the user's line-of-sight vector to the target position on the screen, perform vector direction normalization and projection difference correction to obtain the dynamic spatial vector of the user screen.

[0026] Preferably, the step of obtaining the single-viewpoint element positioning table is as follows:

[0027] Obtain multi-view image data of the vision detection site, extract the pixel information of the user behavior feature area and the scene key element area in the multi-view image data of the vision detection site, and identify the three-dimensional position information of the user behavior feature points through spatial mapping to generate a set of spatial coordinates of the user behavior feature points;

[0028] Based on the set of spatial coordinates of the user behavior feature points, identify the contour edges of the scene key elements in the multi-view image data at the vision detection site, perform spatial coordinate matching and three-dimensional space fitting, and generate a set of three-dimensional spatial position coordinates of the scene key elements;

[0029] Based on the set of three-dimensional spatial position coordinates of the scene key elements, summarize the three-dimensional spatial position coordinates of the user behavior feature points and the scene key elements under the corresponding views, and establish a single-view element positioning table.

[0030] Preferably, the steps for obtaining the multi-view fusion scene state point set are as follows:

[0031] Based on the single-view element positioning table, extract the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the image frame numbers, the camera internal parameter matrices, and the end points of the user screen dynamic space vector coordinates under each view, and uniformly transform them to the unified coordinate system defined by the user screen dynamic space vector to generate a multi-view three-dimensional positioning matrix set within the unified coordinate system;

[0032] According to the multi-view three-dimensional positioning matrix set within the unified coordinate system, calculate the fusion stable offset of each view;

[0033] Based on the fusion stable offset, traverse the positioning point pairs of each view, perform a priority sorting of spatial position fusion in ascending order of the fusion stable offset, screen out the point pairs that meet the spatial coincidence degree condition and merge them into unified spatial point positions, and establish a multi-view fusion scene state point set.

[0034] Preferably, the steps for obtaining the determination result of spatial state consistency are as follows:

[0035] Based on the multi-view fusion scene state point set, extract the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the camera optical center positions, the camera optical axis vectors, the user screen dynamic space vector directions, and the user behavior feature point orientation vectors corresponding to each view, and generate a multi-view point position direction comparison matrix;

[0036] According to the multi-view point position direction comparison matrix, calculate the spatial position difference value in the unified coordinate system;

[0037] Based on the spatial position difference value, perform a pairwise comparison with the preset spatial tolerance threshold, screen out all matching point pairs within the spatial tolerance threshold range, and output a unified judgment result according to the consistent matching situation of the point pairs in all views to obtain the determination result of spatial state consistency.

[0038] The present invention provides an unattended vision detection system, including:

[0039] The data acquisition module performs image acquisition based on the on-site image data of vision detection, obtains multi-angle and multi-view image data of the vision detection site, conducts scene structure line and plane detection, constructs a scene geometric feature set by extracting line segment sets and plane parameters;

[0040] The scene modeling module performs spatial integration based on the scene geometric feature set. By combining the coordinates of the user's line-of-sight area obtained from detection with the coordinates of the boundary of the vision detection screen, an initial scene geometric structure is obtained; according to the initial scene geometric structure, the camera pose changes in consecutive image frames are tracked, the relative pose of the camera is calculated and obtained, and then a three-dimensional space coordinate system is constructed to generate a sequence of the relative poses of the camera, and the coordinate transformation from the user's line-of-sight area to the vision detection screen is solved to obtain the dynamic spatial vector of the user's screen;

[0041] The line-of-sight tracking module tracks the pose changes of consecutive image frames based on the sequence of the relative poses of the camera and the dynamic spatial vector of the user's screen, calculates the dynamic changes in the line-of-sight direction and the user's behavior, and generates a user line-of-sight tracking sequence;

[0042] The multi-view fusion module identifies the user behavior feature points and the key elements of the scene from multiple perspectives based on the multi-view image data of the vision detection site, obtains the three-dimensional space position data, and aggregates them into a single-view element positioning table; combines the user line-of-sight tracking sequence with the single-view element positioning table to perform cross-view data mapping and fusion, and generates a set of multi-view fusion scene state points;

[0043] The spatial consistency module matches the user behavior feature points and the scene elements from different perspectives based on the set of multi-view fusion scene state points, calculates the spatial position difference value, compares the spatial position difference value with the preset spatial tolerance threshold to judge the consistency, and obtains the spatial consistency determination result.

[0044] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0045] The present invention detects scene structure lines and planes by collecting image data at the vision detection site, obtains a line segment set and plane parameters, establishes basic geometric description elements, combines the spatial integration of the user's line of sight area and the boundary coordinates of the vision detection screen to construct the initial scene geometric structure, and improves the stability of on-site spatial structure analysis; further tracks the changes in the camera pose in consecutive image frames, calculates the displacement and rotation parameters in real time, generates a sequence of relative camera poses, and realizes the accurate dynamic spatial vector construction between the user's line of sight area and the screen, reducing the detection error caused by pose changes in the dynamic scene; at the same time, obtains multi-view image data to identify the spatial positions of user behavior feature points and scene key elements in each view, constructs a single-view element positioning table and fuses it across views to improve the consistency and positioning accuracy of multi-view data; finally, accurately calculates the spatial position difference value in the unified coordinate system and judges the spatial tolerance threshold, improving the reliability and automation of the spatial state consistency judgment result, ensuring the accuracy and intelligence of the overall vision detection process, and the spatial state consistency judgment result is used to confirm whether the positions of the user's line of sight area and the vision detection screen display area are exactly coincident in multi-view images, and to judge whether the user maintains the correct pose and line of sight direction during the detection process. By calculating the position difference between the user behavior feature points and the vision detection screen elements in different views and judging whether it is within the predetermined tolerance range, if the position difference exceeds the tolerance range, it means that the user's line of sight deviates from the detection screen or the pose is abnormal. At this time, the current invalid detection data can be excluded, thereby ensuring the effectiveness of the unattended vision detection process and the authenticity and accuracy of the evaluation result. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0048] Please refer to Figure 1 , the present invention provides a technical solution, an unattended vision detection method based on image recognition, including the following steps:

[0049] Collect image data at the vision detection site, perform scene structure line detection and plane detection to obtain a line segment set and plane parameters, and establish basic geometric description elements; based on the basic geometric description elements, combine the coordinates of the user's line of sight area obtained by detection and the boundary coordinates of the vision detection screen for spatial integration to obtain the initial scene geometric structure information;

[0050] Based on the initial scene geometric structure information, track the camera pose changes in consecutive image frames, calculate and obtain the displacement and rotation parameters of the camera relative to the initial geometry, obtain the camera relative pose sequence, and based on the camera relative pose sequence, construct a three-dimensional space coordinate system and solve the coordinate transformation from the user's line of sight area to the vision detection screen to generate the user screen dynamic space vector;

[0051] Obtain multi-view image data of the vision detection site, respectively identify the three-dimensional position information of the user behavior feature points and scene key elements at each view, gather them into a single-view element positioning table, and based on the unified coordinate system defined by the single-view element positioning table and the user screen dynamic space vector, perform cross-view data mapping and fusion to establish a multi-view fusion scene state point set;

[0052] Based on the multi-view fusion scene state point set, match the corresponding user behavior feature points and scene elements under different views, calculate the spatial position difference value in the unified coordinate system, compare the spatial position difference value with the preset spatial tolerance threshold for judgment, and obtain the spatial state consistency determination result.

[0053] The steps for obtaining the basic geometric descriptors are as follows:

[0054] Collect the image data of the vision detection site. Based on the collected image data of the vision detection site, perform image edge gradient analysis to calculate the pixel gradient changes of the edge points, extract the edge contours at the boundaries of each object in the scene, combine the edge points to form structural line segments, and generate a set of structural line segments for the vision detection site;

[0055] Based on the set of structural line segments for the vision detection site, extract the spatial coordinate data of all the endpoints of the line segments in the set, judge the coplanarity of the spatial coordinates of the line segment endpoints, screen the coplanar point groups according to the coplanarity of the line segment endpoint coordinates, fit and calculate the plane equation parameters corresponding to the spatial positions of the coplanar point groups, and generate the plane parameters corresponding to the structural line segments;

[0056] Based on the plane parameters corresponding to the structural line segments, call the endpoint coordinates of the line segments in the set of structural line segments for the vision detection site, perform spatial projection calculations on the spatial coordinates of the line segment endpoints and the corresponding plane parameters one by one, determine the spatial position relationship between the line segments and the corresponding planes, establish the spatial mapping relationship between the line segment endpoint coordinates and the plane parameters, and generate the basic geometric descriptors.

[0057] Specifically, perform pixel-level edge gradient calculation based on the acquired image data. First, collect images with a resolution of approximately 1920×1080 in the visible light range and convert them to grayscale form frame by frame. Then, select the differential operators in the horizontal and vertical directions to calculate the gradient value of each pixel and generate a gradient matrix. If the gradient amplitude of a certain pixel point If it exceeds 25, it is determined as a potential edge point. This threshold of 25 is determined by successively using different thresholds from 15 to 30 in the test samples and comparing the edge integrity. Then, non-maximum suppression processing is performed to eliminate the pixel points with weak gradient response and overlapping with the main edge. Next, morphological operations of closing and opening are set to adjust the continuity to merge adjacent edges. When judging the edge chain length, a minimum pixel number threshold of 5 is introduced. This value is set by calculating the edge chain length distribution for 20 test images and combining visual judgment. The edge chains with a length less than 5 are not recorded. After completion, the remaining continuous edge chains are regarded as preliminary line segments and their starting and ending coordinates are identified. By traversing these starting and ending coordinates, an index list that can distinguish adjacent or overlapping line segments is formed. When adjacent line segments intersect and form an approximate straight line, the coordinate difference method is used to judge whether to merge. If the length of the merged line segment is greater than the aforementioned minimum pixel number threshold, the new merged line segment is retained. Finally, a set of available line segments is obtained and uniformly recorded as the line segment set of the visual acuity detection site structure.

[0058] Based on the obtained line segment set of the visual acuity detection site structure, the three-dimensional coordinate data of the starting and ending points of each line segment is read one by one. First, a preliminary accuracy correction is performed on the endpoint coordinates of each line segment. By comparing the set spatial error range such as 0.00 to 0.05 and excluding the error data outside this range. After correction, all endpoint coordinates are collected into the same list and the coplanarity judgment rule is used to detect the distribution of the endpoints. The coplanarity judgment introduces a distance judgment threshold used to measure the maximum allowable deviation of the distance from a point to a plane. This is determined by successively increasing from 0.001 to 0.02 by the experimental method and observing the stability of the endpoint classification results. For a set containing at least three non-coincident points and the distance from all points to a certain plane is within it is regarded as a coplanar point group. The coplanar point group is fitted with a plane by the least square method and the plane equation parameters are obtained. x, y, and z respectively represent the coordinate values of the line segment endpoints in the three-dimensional space along the x, y, and z axis directions, where are respectively determined by performing eigenvalue decomposition on the point coordinate matrix. To avoid ambiguous references, it is also necessary to correspond one by one to the line segment endpoints that have been judged. If the two endpoints of a line segment both belong to the same coplanar point group, the plane equation parameters are recorded in the corresponding line segment entry. After processing all the line segments, the plane parameters corresponding to the structural line segments can be obtained.

[0059] Based on the plane parameters corresponding to the structural line segments, the spatial projection calculation is performed by calling the endpoint coordinates in the line segment set of the visual acuity detection site structure. For each line segment, first extract from its plane equation as the normal vector. In the projection calculation, the endpoint coordinates are projected perpendicularly in the direction of the normal vector to obtain the projection coordinates For the two endpoints of the same line segment, after projection with the same plane normal vector, the mapped position of the line segment in the plane is obtained. When comparing the difference between the endpoint projection results and the original three-dimensional coordinates of the line segment, a distance difference threshold is used. Determine whether there is a significant deviation. This is selected after testing step by step in the range of 0.005 to 0.03. If the projection deviation is less than then a corresponding record of the mapping of this line segment and the plane parameters is established. Subsequently, if multiple plane parameters correspond to the endpoints of the same line segment, it is confirmed based on the scheme with the smallest projection distance of the line segment. Finally, after all endpoints are projected and the mapping relationship is established, the correspondence information between the line segment and the plane is summarized into a spatial mapping index table. This index table contains the projection data of the endpoint coordinates of each line segment and the corresponding plane equation parameters. After summarizing this information, the basic geometric description elements are obtained.

[0060] The steps for obtaining the initial scene geometric structure information are as follows:

[0061] Based on the basic geometric description elements, the endpoint coordinates of the line segments and the plane parameters in the basic geometric description elements are called, and combined with the boundary coordinates of the vision detection screen, spatial coordinate transformation and projection calculation are performed to determine the spatial position relationship of the vision detection screen in the basic geometric description elements, and the spatial mapping coordinates of the screen boundary are generated.

[0062] Based on the spatial mapping coordinates of the screen boundary, the coordinates of the user's line of sight area are obtained, coordinate system alignment and spatial coordinate transformation are performed, the spatial position relationship between the user's line of sight area and the screen boundary is integrated, and a mapping correspondence relationship between the coordinates of the line of sight area and the screen boundary coordinates is established, and the integrated spatial coordinates of the line of sight and the screen are generated.

[0063] Based on the integrated spatial coordinates of the line of sight and the screen, the spatial matching relationship between the integrated spatial coordinates of the line of sight and the screen and the basic geometric description elements is determined, and the initial scene geometric structure information is obtained.

[0064] Specifically, based on the basic geometric description elements obtained previously, first retrieve the three-dimensional coordinates corresponding to the line segment endpoints and their plane equation parameters from them to define the transformation matrix and prepare to transform and project the boundary coordinates of the vision detection screen. When specifically executing, first read the coordinate set of the line segment endpoints and match them with the plane parameters correspondingly. By normalizing the plane normal vector and constructing a rotation matrix to perform an affine transformation on the subsequent screen boundary coordinates. If there are coordinate components exceeding the pre-established effective range during the conversion process, for example to interval, then discard this data to avoid positioning deviation. This interval is determined by collecting the coordinate distribution in the sample scene, observing the maximum and minimum values, and then leaving a certain margin. Subsequently, the boundary coordinates of the four corners of the vision detection screen Successively substitute into the transformation matrix and obtain the new coordinates after projection , if the absolute value of the coordinates of any vertex after projection is greater than 50, it indicates that the mapping relationship between the vision detection screen and the basic geometric description element is unstable. At this time, it is necessary to confirm again whether the input plane parameters and the line segment endpoint coordinates meet the projection requirements and recalculate the rotation matrix and displacement vector according to the results. Finally, after completing this part of the transformation, the screen boundary space mapping coordinates are formed.

[0065] Based on the obtained screen boundary space mapping coordinates, read the user's line-of-sight area coordinates, perform coordinate system alignment and space coordinate transformation. When executing, first clarify the acquisition method of the user's line-of-sight area coordinates. Here, a binocular camera combined with an eye movement tracking model is used to identify the pupil center and the line-of-sight direction of the user. This model is trained using a convolutional neural network and contains five convolutional layers and two fully connected layers. The input data is the normalized binocular image frame data , the size of each frame of image is fixed at pixels, and the output is the estimated values of the pupil coordinates and the line-of-sight direction vector. During training, it is carried out on a public data set containing 20,000 labeled pupil positions, using the stochastic gradient descent method to iterate 100 epochs and setting the learning rate to 0.001. By continuously minimizing the Euclidean distance loss between the pupil position and the labeled true value to update the network weights, represents the index number of the image frame, represents the coordinate value of the predicted pupil center point in the horizontal direction in the th frame of image, represents the coordinate value of the predicted pupil center point in the vertical direction in the th frame of image, m is the total number of image frames participating in training, is the horizontal direction coordinate value of the true labeled pupil center point in the th frame of image, is the vertical direction coordinate value of the true labeled pupil center point in the th frame of image. In the inference stage, the real-time collected binocular image frame is input into the trained convolutional neural network and the pupil center position and the line-of-sight direction vector , during the alignment process, first map this three-dimensional coordinate to the same coordinate system as the screen boundary space mapping coordinate, and compare it according to the preset line-of-sight direction offset threshold of 0.02. This threshold is selected after statistically analyzing five groups of user line-of-sight samples in the experiment. If the offset exceeds 0.02, collect the binocular images again and perform inference again. Finally, the aligned user line-of-sight area coordinates correspond one by one to the screen boundary coordinates, and the mapping correspondence between the line-of-sight area coordinates and the screen boundary coordinates is merged to generate the line-of-sight screen integration space coordinates.

[0066] Based on the above line-of-sight screen integration space coordinates, compare each coordinate mapping record in the basic geometric description element and confirm the corresponding index between the two. When matching the line-of-sight screen integration space coordinates segment by segment with the corresponding plane coordinates, set an allowable offset threshold , and this threshold refers to the range of most spatial position errors in previous scene tests, which is approximately 0.01 to 0.04. Finally, take it as 0.05 to ensure covering a large fluctuation. If the distance between a certain coordinate and the theoretical position calculated by the plane equation exceeds , then skip this coordinate pair and review the source of the abnormal item, and then compare the remaining valid coordinates one by one. After completing the spatial mapping verification, count all the coordinates that match successfully, summarize and output the initial scene geometric structure information.

[0067] The steps to obtain the relative pose sequence of the camera are as follows:

[0068] Based on the initial scene geometric structure information, obtain the three-dimensional boundary points of the user's line-of-sight area and the corner point coordinates of the vision detection screen boundary. Combine the image acquisition timestamps and camera internal parameters recorded in the consecutive image frames, calculate the matching error and reprojection offset of the corresponding boundary points between the image frames, and generate a consecutive image frame matching error sequence;

[0069] According to the consecutive image frame matching error sequence, calculate the camera pose change angle value. The calculation formula is:

[0070] ;

[0071] Among them, is the camera pose change angle of the th frame image, , are the central pixel coordinates of the th frame image, , are the central pixel coordinates of the initial frame image, is the vertical resolution of the th frame image, is the horizontal resolution of the th frame image, is the horizontal viewing angle of the th frame image, is the vertical viewing angle of the th frame image, is the camera focal length of the th frame image;

[0072] According to the camera pose change angle value, combined with the three-dimensional offset of the boundary points of each frame image in the continuous image frame matching error sequence, the three-dimensional rotation axis offset and direction vector change between adjacent frames are obtained, and the relative camera pose sequence is obtained.

[0073] Specifically, based on the initial scene geometric structure information, first summarize the boundary corner point coordinates of the vision detection screen and the three-dimensional boundary points of the user's line of sight area. By collecting multiple frames of images as data sources in a monitoring environment equipped with a high-resolution imaging device, the acquisition timestamp of each frame of image and the camera internal parameters are read and recorded together. For the data such as the focal length, principal point coordinates, and distortion coefficients in the camera internal parameters, they are obtained through a calibration process in advance. These calibration data refer to the calibration procedure given in the imaging system user manual. For example, at least 50 images are collected within the range of 1 meter to 3 meters from the calibration plane and then the camera parameters are solved to obtain a relatively stable internal parameter matrix. Then, the corresponding boundary point coordinates between adjacent frames in the same target scene are compared multiple times to obtain the difference at the pixel level. If the pixel deviation of any corner point is greater than the pre-set critical threshold of 5 pixels (this threshold is selected by testing in the range of 10 to 20 pixels and checking the matching error rate), then this part of the matching pair is classified as an abnormal match for additional verification. In addition, when calculating the reprojection, the camera internal parameter projection model needs to be applied to each boundary point coordinate, and the three-dimensional coordinates are projected onto the image plane according to the pinhole imaging model to generate the projection coordinates. The difference between the projection and the actual pixel coordinates is the reprojection offset. If the reprojection offset exceeds the set calibration range of 0 to 2 pixels, then the matching point is marked as an over-offset point for statistics, and the mean square value of the overall offset is calculated by selecting the remaining matching points in the same frame of image, so as to construct the continuous image frame matching error sequence. Finally, the matching error information of all frames is summarized in chronological order to obtain the continuous image frame matching error sequence.

[0074] Formula: , the advantage of the formula is that it comprehensively considers factors such as the camera center pixel offset, actual resolution, camera viewing angle, and focal length, and can present the pose change amount between image frames with a more intuitive angle value, so as to more clearly express the dynamic change law of the camera pose in the subsequent steps.

[0075] The steps for obtaining the parameter are based on the For the pixel coordinate system of the frame image, read the horizontal coordinate value of the actual center pixel position. This horizontal coordinate value is obtained from an image with a pre-calibrated resolution range of 0 to 4000. Specifically, in an environment with a relatively high shooting resolution, record the actual pixel width of each frame image to obtain a pixel width sequence. , denote the actual width of the -th frame image as , then find the index of the most central pixel column in this frame image, and set the actual value of this column index as . For example, in a test scenario, the width of the 5-th frame image taken is 1920 pixels, then . At this time, if there are 5 pixels of black edges on both the left and right sides of the image, the actual effective picture width is 1910 pixels. So is the specific value of .

[0076] The steps for obtaining the parameter are as follows: For the vertical pixel coordinate at the center of the -th frame image, combine the shooting resolution for positioning. In an image with a height that may reach more than 3000 pixels, mark the vertical pixel midline as . Then observe whether there are upper and lower black edges or embedded graphic areas. If so, make subtraction adjustments according to the number of pixel rows of the actually detected upper or lower edges to make the final value accurately correspond to the midline. In addition, in a panoramic shooting environment, some images have custom text or border information, and the number of their pixel rows needs to be counted and deducted. Finally, record and obtain . For example, in a certain shooting, the height of the image is 1080 pixels, there are 7 pixels in the upper blank area and 3 pixels in the lower blank area, then the effective height , and the central vertical coordinate is regarded as .

[0077] The steps for obtaining the parameter are as follows: Read the horizontal coordinate of the center pixel of the initial frame image. This initial frame is fixed as the frame with serial number 1 at the start of shooting. For the initial frame, it is also necessary to first count its actual resolution width and information such as possible black edges or nested areas. To ensure accuracy, a series of feature points need to be extracted from the initial frame and the actual effective width of the left and right boundaries of the image is solved, and then it is divided by 2 to obtain the central row index, and record this index value as . For example, in an actual shooting, the width of the initial frame is statistically 2560 pixels, and there is no black edge correction on the upper, lower, left, or right, then .

[0078] The steps to obtain the parameters are as follows. For the vertical coordinate of the central pixel of the initial frame image, divide the height of the initial frame image by 2 to obtain the vertical midline position. If it is detected that there are dozens of pixel occupancies above or below the picture at the shooting moment, these occupancy areas need to be removed before midline positioning and then the midline is solved. For example, in a certain detection, the height is 1440 pixels and there is only a 5-pixel blank area above, and no blank area below, then the effective image height is , corresponding to .

[0079] The steps to obtain the parameters are as follows. Read the vertical resolution of the frame image and combine it with the system calibration information. The vertical resolution can usually be automatically recognized within a numerical range of 480 to 3000 when the imaging device takes pictures. The specific record is as .

[0080] The steps to obtain the parameters are as follows. Read the horizontal resolution of the frame image.

[0081] The steps to obtain the parameters are as follows. The horizontal viewing angle information of the frame image is usually determined by the physical camera lens structure and the focal length. Shoot point by point on a plane 1 meter to 3 meters away from the fixed lens, and record the horizontal range covered in the shooting picture as a radian or angle value. After averaging these angles and comparing them with the accuracy correction value, finally obtain , for example, when the focal length of a certain lens is 10 mm, the measured actual horizontal viewing angle is 85 degrees, and 85 degrees is converted to 1.48353 radians for recording.

[0082] The steps to obtain the parameters are as follows. Similar to , but change to the vertical viewing angle. Through the same on-site shooting and surveying method, it is confirmed that when the focal length of the lens is 10 mm, the vertical viewing angle is about 60 degrees, that is, 1.0472 radians. Thus, obtain and record.

[0083] The steps to obtain the parameters are as follows. The camera focal length of the frame image is obtained by querying the lens specifications or software and hardware.

[0084] Calculation process: In a demonstration example, let the specific values of the parameters corresponding to the frame image be , , , , , , , , millimeters, first calculate the numerator :

[0085] ;

[0086] ;

[0087] Take the absolute value as , then calculate the denominator :

[0088] ;

[0089] ;

[0090] Divide the numerator by the denominator , and then find the arctangent radians, convert it to degrees , to obtain the camera pose change angle value of this frame .

[0091] This result indicates that there is a relatively obvious field of view deviation between the camera in this frame and the initial frame. When processing the three-dimensional rotation axis offset and direction vector change of adjacent frames in the follow-up, this angle value can be used as one of the inputs, combined with the three-dimensional offset of the boundary points in the continuous image frame matching error sequence, and then finally obtain the relative pose sequence of the camera.

[0092] The steps to obtain the dynamic space vector of the user screen are as follows:

[0093] Based on the relative pose sequence of the camera, obtain the rotation matrix and translation vector of each frame, and combine the focal length, image center coordinates, pixel size, and image size parameters in the camera internal parameters to reconstruct the coordinates and space vectors of the starting coordinates of the user's line of sight area and the screen corner coordinates in turn, and generate the initial three-dimensional vector group from the user's line of sight to the screen;

[0094] According to the initial three-dimensional vector group from the user's line of sight to the screen, calculate the space transformation value from the user's line of sight vector to the target position on the screen. The calculation formula is:

[0095] ;

[0096] Among them, is the space transformation value from the user's line of sight vector to the target position on the screen, is the starting coordinate of the user's line of sight vector, is the target screen corner coordinate, is the camera focal length, is the image center point coordinate, is the image width, is the image height;

[0097] Based on the spatial transformation value from the user's line-of-sight vector to the target position on the screen, perform vector direction normalization and projection difference correction to obtain the user's screen dynamic spatial vector.

[0098] Specifically, based on the relative pose sequence of the camera, read the corresponding rotation matrix and translation vector for each frame. First, uniformly number these rotation matrices and translation vectors and match them in chronological order. At the same time, retrieve the recorded coordinate transformation parameters. This number can be used to assign values to each frame of the image in sequence in the shooting environment for subsequent comparison. Each rotation matrix is obtained by the mutual transformation between the reference coordinate system obtained by on-site calibration and the camera lens coordinate system. It is necessary to collect multiple sets of feature point positions during calibration and solve this matrix frame by frame. Then, the translation vector is calculated based on the relative offset between the on-site position identifier and the camera optical center. After organizing this information with time sequence tags and corresponding them one by one to the starting coordinates of the user's line-of-sight area and the corner coordinates of the screen, then according to the recorded pixel size and image size parameters, through performing coordinate system mapping and three-dimensional vector splicing operations on these coordinate values, use the Euler angle transformation or quaternion transformation method to unify the camera coordinate system and the scene global coordinate system. Then, check whether the coordinate landing point is within the pre-set valid range in the shooting space. For example, in the horizontal coordinate interval from -500.0 to 500.0, perform additional comparison or discard the points outside this interval. After completing this range check, perform three-dimensional vectorization processing on the starting coordinates of the user's line-of-sight area and the corner coordinates of the screen respectively. Each vector has a corresponding time index and rotation and translation correction values. If it is detected in a certain operation that the cumulative vector direction and coordinate deviation are greater than the pre-set 5.0-degree direction deviation threshold, then it is necessary to read the rotation matrices of the recent few frames again for verification. This threshold of 5.0 degrees is obtained by statistically calculating the line-of-sight direction offset values in multiple scenarios and taking the median plus a 1-degree margin. Finally, after accumulating these steps, the starting coordinates of the user's line-of-sight area and the corner coordinates of the screen can be transformed from a simple three-dimensional point set into a continuous initial three-dimensional vector group.

[0099] Formula , the benefit of the formula is to perform segmented analysis on the coordinate difference between the user's line-of-sight vector and the target position on the screen. At the same time, combined with parameters such as focal length and image width and height, through the combined action of the logarithmic function and the arctangent function, the horizontal and vertical components are integrated into a unified spatial transformation value.

[0100] The steps to obtain the parameter are as follows: Extract the horizontal component from the starting coordinates of the user's line-of-sight vector. It is necessary to first deploy an eye movement tracking device at the shooting site and record the position of the user's eyes. Use the calibration method based on the binocular imaging principle to read the pupil center coordinates, and record this coordinate as the horizontal component in Specific calibration method is to perform binocular shooting within a certain range from the user and track the pupil center frame by frame, and form a set of horizontal coordinate data through multiple measurements and compare it with the pre-calibrated three-dimensional coordinate system. If there is a difference, numerical correction is performed according to the field survey data. For example, in a calibration, it is measured that the user's eye is about 50.3 cm from the reference origin in the X direction. At this time, .

[0101] The acquisition steps of the parameter are as follows. The horizontal component of the image center point coordinates needs to determine the position of the camera principal point when the imaging device is initialized and convert it into a reference value in the same three-dimensional scene. First, record the resolution midpoint of the lens center in the image plane coordinates as , and then match it with the physical coordinates collected in the scene calibration. If there is a world coordinate origin offset value when the lens is facing directly forward, let it be used as the basic value after correction. For example, if the lens is confirmed to be offset 2.0 cm in the X direction during calibration, then this information is merged into the coordinate value of the image center, so that .

[0102] The acquisition steps of the parameter are as follows. The camera focal length value is obtained after reading the internal parameters of the camera or relevant documents.

[0103] The acquisition steps of the parameter are as follows. The pixel size of the image width is directly read through the resolution information of the imaging device. In a verification, an image with a resolution of 1920×1080 is captured. Then .

[0104] The acquisition steps of the parameter are as follows. Similar to , it is used to represent the pixel size of the image height. For example, when shooting on site, the nominal image height is 1080 pixels, .

[0105] The acquisition steps of the parameter are as follows. Extract the horizontal component from the target screen corner point coordinates. To make the corner point coordinates consistent with the shooting coordinate system, three-dimensional positioning measurement of the screen position is required. Obvious marks can be attached to the front of the screen and the X-direction coordinate can be obtained using a laser rangefinder and multi-view shooting, and then converted into a specific value in the scene coordinate system. If the measured coordinate deviation from the pre-designed drawing does not exceed 2 cm, it is recorded as , for example, the X coordinate of the upper left corner of a certain display screen is measured to be 120.4 cm in the direction of the scene origin. Then .

[0106] Calculation process: In an example, , , , , , ;

[0107] Calculate the numerator first , and then multiply it by get ;

[0108] because , so the molecular part is ;

[0109] Then , then Performing absolute value ,calculate ;

[0110] Again Do the evaluation, , Radians, multiply the two , then square , add the two parts together , take the square root of .

[0111] The result shows that in this example, the spatial transformation value between the starting point of the user's line of sight vector and the target position on the screen is approximately 2.175. When the value is too large, it means that the difference between the starting point and the target point in the X direction or the aspect ratio is obvious. If the value is small, it means that the distance or difference between the two is relatively low. The result can be used to generate the final user screen dynamic space vector by subsequent normalization of the vector direction and projection difference correction.

[0112] Based on the spatial transformation value from the user's line-of-sight vector to the target position on the screen, first perform piecewise detection on this value and execute a normalization operation in combination with the direction vector of the user's line-of-sight vector. The used direction vector comes from the previously recorded three-dimensional coordinate differences. It is necessary to first calculate the square sum of the three-dimensional differences in the X, Y, and Z directions respectively and then take the square root to obtain the modulus length of the vector. Then divide each component by this modulus length to normalize the overall length of the vector to 1. If it is detected that a certain component exceeds the set absolute threshold of 500.0 when calculating the modulus length, it is necessary to recheck whether there are calibration errors or system abnormalities in the previous observation data. This threshold of 500.0 is selected from the middle section within the range of -1000.0 to 1000.0 with reference to the scene scale and the half value is taken after multiple dynamic tests. After recording the vector components that exceed this threshold, manual intervention is required for rechecking and deleting or correcting as appropriate. After normalization, the projection differences are compared. Calculate the position difference between the screen plane and the user's line-of-sight vector in the two-dimensional projection within the same coordinate system. If it is found that the projection distance between the two exceeds the previously statistically averaged offset of 3.0 pixels, retrieve the user's head pose tracking information at the same moment to confirm whether there is a rapid head rotation or occlusion. Once the cause is locked, the coordinate update can be performed in combination with the actual scene record and the projection difference is calculated again. For the finally qualified data, it is summarized and marked as valid vector information, thus obtaining the user's screen dynamic space vector.

[0113] The steps for obtaining the single-viewpoint element positioning table are as follows:

[0114] Obtain multi-viewpoint image data of the visual acuity detection site, extract the pixel information of the user behavior feature area and the scene key element area in the multi-viewpoint image data of the visual acuity detection site, and identify the three-dimensional position information of the user behavior feature points through spatial mapping to generate a set of three-dimensional coordinates of the user behavior feature points;

[0115] Based on the set of three-dimensional coordinates of the user behavior feature points, identify the contour edges of the scene key elements in the multi-viewpoint image data of the visual acuity detection site, and perform spatial coordinate matching and three-dimensional space fitting to generate a set of three-dimensional position coordinates of the scene key elements;

[0116] Based on the set of three-dimensional position coordinates of the scene key elements, summarize the three-dimensional position coordinates of the user behavior feature points and the scene key elements under the corresponding viewpoints, and establish a single-viewpoint element positioning table.

[0117] Specifically, to obtain multi-perspective image data at the vision detection site, first perform initialization checks on the shooting positions and optical parameters for each perspective. By measuring the distances between the camera and each calibration point in the scene, ensure that all perspectives are in the same reference coordinate system. Then, extract the pixel information of the areas related to user behavior features and key elements in the scene from each frame of the image. When performing this extraction, based on the previously statistically obtained brightness range from 0 to 255 and the color value distribution range, determine whether the pixel falls within the specified user activity color interval or the key element contour interval. If the pixel feature value does not match the above ranges, exclude it to avoid mixing in interfering data. At the same time, to ensure the accuracy of the extraction process, a pixel connectivity threshold of 10 needs to be set within each perspective. This 10 is selected by trying values from 5 to 20 in multiple sample images and observing the connectivity region effect on the user's body shape or the key element contour. Only when the number of adjacent pixels reaches 10 will they be merged into the same connected block and marked as a suspicious area. After marking the suspicious areas in the multi-perspective data, perform a one-by-one coordinate transformation on the image frame numbers and their pixel positions corresponding to each perspective. The transformation method is to map the pixel coordinates to the three-dimensional coordinate system according to the camera internal parameter matching method. Coordinates with values exceeding the set valid interval of -1000.0 to 1000.0 during the mapping process are regarded as invalid and excluded. The remaining three-dimensional coordinate points are then grouped by spatial clustering. If the differences in the X, Y, and Z coordinates of the same feature area in different perspectives are all less than 5.0, they are regarded as part of the same object. Finally, record the three-dimensional coordinates corresponding to all the confirmed user activity pixel blocks to form a three-dimensional coordinate set of user behavior feature points.

[0118] When obtaining the contour edges of the scene key elements in the multi-view image data at the vision detection site based on the set of spatial coordinates of the user behavior feature points, first, according to the characteristics of the scene key elements, such as the screen edge, the desktop or the device shell contour, etc., select the color or shape feature values that can represent the obvious edges and confirm them frame by frame in the multi-view images. Use the same method to filter out the pixel blocks that conform to the key element contours at the pixel level and convert them to the three-dimensional coordinate system for grouping. If there is a local continuous blank greater than 10 cm in the X, Y, or Z direction of the scene key element edge, then this section of the contour needs to be split. This 10 cm comes from the upper limit of the spacing for observing the surface continuity of the key element during actual measurement. After the splitting is completed, the three-dimensional points of each section of the contour need to be matched with the set of spatial coordinates of the user behavior feature points. The specific matching method is to calculate the Euclidean distance in the X, Y, and Z directions and compare whether it is within the range of 0.00 to 2.00 cm. If it exceeds 2.00 cm, it is not regarded as the same object. Then, all the successfully matched contour edge points are subjected to the least squares fitting to generate the three-dimensional edge data after fitting, which is used to express the complete spatial position of the scene key element. Then, according to the adjacent physical area criterion, these fitting results are combined to obtain the set of three-dimensional spatial position coordinates of the scene key element.

[0119] Based on the set of three-dimensional spatial position coordinates of the scene key element, integrate the three-dimensional coordinates of the user behavior feature points and the scene key element under the same view. First, number each view separately and uniformly transform the feature points and key element coordinates under that view into the scene global coordinate system. If it is detected that the coordinates of some of this data exceed the boundary range of -2000.0 to 2000.0 in the X or Y direction, then it is necessary to compare whether there are abnormal interferences or calibration inaccuracies in the original image and the calibration data. Because in the actual test, it is stipulated that the effective interaction area within the scene is mainly concentrated between -2000.0 and 2000.0. If it is confirmed that there are errors, these abnormal data points can be directly deleted. Subsequently, record the three-dimensional coordinates of all the feature points and key elements in the same view in sequence, and compile the indexes of these coordinates and their associated view numbers together. In this way, a single-view element positioning table can be established.

[0120] The steps for obtaining the multi-view fusion scene state point set are as follows:

[0121] Based on the single-view element positioning table, extract the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the image frame numbers, the camera internal parameter matrix, and the end points of the user screen dynamic space vector coordinates under each view, and uniformly transform them to the unified coordinate system defined by the user screen dynamic space vector to generate the multi-view three-dimensional positioning matrix set within the unified coordinate system;

[0122] According to the multi-view three-dimensional positioning matrix set within the unified coordinate system, calculate the fusion stable offset of each view. The calculation formula is:

[0123] ;

[0124] wherein, is the th perspective fusion stable offset, are the coordinates of the user behavior feature points, are the coordinates of the key elements of the scene, are the coordinates of the end point of the screen vector, are the coordinates of the camera optical center, is the perspective number, is the image frame index;

[0125] Based on the fusion stable offset, traverse the perspective positioning point pairs, perform spatial position fusion priority sorting in ascending order of the fusion stable offset, screen the point pairs that meet the spatial coincidence degree condition and merge them into a unified spatial point position, and establish a multi-perspective fusion scene state point set.

[0126] Specifically, when extracting the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the key elements of the scene, the image frame numbers, etc. of each perspective based on the single-perspective element positioning table and uniformly converting them to the unified coordinate system defined by the user screen dynamic space vector, first check item by item according to the recorded camera internal parameter matrix and the resolution and shooting axis angle of each frame of image during the shooting process. If the internal parameter matrix of a certain perspective does not match the actual resolution information, it is necessary to compare the original calibration data again and check the resolution correction process. After confirmation, read out the X, Y, and Z components of the user behavior feature point coordinates and the key element coordinates of the scene respectively, and combine the corresponding image frame numbers with the end points of the user screen dynamic space vector obtained in the early stage to form a basic data set within one perspective. When putting these data sets and coordinates into the same coordinate system according to the perspective number, it is necessary to first consider the position of the camera origin or optical center. If it is detected that the optical center of a certain perspective has a large offset from the unified coordinate origin exceeding the range of 2 meters, the offset value needs to be recorded and vector compensation is performed. This 2-meter range is the detection standard selected for the general vision detection scene through indoor testing. Once exceeded, projection offset is very likely to occur. After calibration, the projection matrix or rotation translation matrix also needs to be applied to the three-dimensional coordinates of the above-mentioned feature points, and check whether each point falls within the global range of -500 to 500 after transformation. If the number of points exceeds this range, it is necessary to check the camera pose or the distribution area of the detection scene again to avoid subsequent data interference caused by mismatch or abnormal coordinates. Finally, map the user behavior feature point coordinates, the key element coordinates of the scene, the image frame numbers, and the end points of the user screen dynamic space vector of each perspective to the unified coordinate system and recombine them into a multi-perspective three-dimensional positioning matrix. If other data comparison is needed later, the matrix structure can be continuously expanded, thereby generating a multi-perspective three-dimensional positioning matrix set within the unified coordinate system.

[0127] Formula: , The advantage of the formula is that it synthesizes multiple spatial relationships such as user behavior feature points, scene key elements, camera optical center, and the end point of the screen vector, and calculates the stable offset degree of different perspective positioning data through the combination of cross product and Euclidean distance.

[0128] The steps to obtain the parameters are as follows: the three-dimensional position of the user behavior feature point coordinates in the th perspective. First, multiple cameras arranged on-site are used to take pictures and align the corresponding image frame numbers with the feature point recognition results. After marking each user behavior area on the pixel plane, it is mapped to the global coordinate system. When mapping, the calibrated camera focal length and camera pose need to be collected to convert the pixel coordinates into spatial coordinates. If the coordinate values exceed the range of -1000 to 1000, the point needs to be re-compared. After confirming that the final three-dimensional coordinates fall within the specified range, they are stored as , for example, in a scene test, the face feature area captured in the th perspective is located at pixel coordinates (350, 250). After projection calculation, the three-dimensional coordinates are (120.5, 48.2, 80.0). At this time, it is detected that all these values are within [-1000, 1000], and it can be recorded as .

[0129] The steps to obtain the parameters are as follows: the three-dimensional position of the scene key element coordinates in the th perspective. It is necessary to extract the contours of the known key elements (such as the screen edge, device housing) and combine the calibration information to obtain their coordinates in the unified coordinate system. For example, the coordinates of a vertex on the side of a display screen detected in a certain perspective are (220.0, 55.0, 5.0), which is called .

[0130] The steps to obtain the parameters are as follows: the end point coordinates of the screen vector in the th perspective reflect the terminal position of the user screen dynamic space vector captured by the camera. It is necessary to combine the previously calculated user screen dynamic space vector and transform it in this perspective. The specific method is as follows: starting from the starting coordinates of the user screen dynamic space vector, add the vector length along the vector direction to obtain the end coordinates. If it is found in actual measurement or tracking that the offset of this end coordinate during alignment with the scene is greater than 10 cm, it is necessary to check the angle difference between the vector direction and the actual inclination of the screen. Here, 10 cm is due to the fact that in a series of tests, an offset less than 10 cm is often considered to be accurately aligned. If it is indeed a perspective problem, the camera pose is re-read and corrected, and then the vector end calculation is performed again. After completion, , for example, in a calculation, the direction vector of the screen vector is (1.5, 2.0, 0.8), the length is set to 3.5 meters, and the starting point is (0, 0, 0). Then the end point is (1.5 × 3.5, 2.0 × 3.5, 0.8 × 3.5) = (5.25, 7.0, 2.8). Translate it to a certain reference point in the scene according to the coordinates. If the reference point is (100, 100, 0), then .

[0131] The steps for obtaining the parameters are as follows: the three-dimensional position of the camera optical center coordinate in the th view angle. For example, in the view angle , the measured optical center position is (0, 150.0, 100.0).

[0132] The steps for obtaining the parameters are as follows: is the view angle number. Each view angle is determined by a camera or the same camera taking pictures at different positions / angles. During on-site setting, sequential marks can be made at the camera bracket or pan-tilt position. If a total of 8 cameras are set, then number them sequentially from 1 to 8.

[0133] The steps for obtaining the parameters are as follows: a set of serial numbers of the image frame index in the th view angle. Each view angle may capture multiple frames of images, and each frame needs to be processed separately. The frame index can be converted from the shooting timestamp. For example, the cumulative number of frames after the start of shooting is recorded as 1, 2, 3…, or directly take the internal serial number of the camera. If the shooting rate of a certain camera is 30 frames per second, the data captured at the 2nd second is marked as frame 60. At this time, the corresponding view angle of .

[0134] Calculation process: In an example, let , , , , , . First, calculate the numerator :

[0135] ;

[0136] ;

[0137] Take the cross product of these two vectors:

[0138] ;

[0139] The cross operation is expanded successively as , and after substituting the values, the result is obtained (an example of omitting the intermediate steps in the calculation):

[0140] ;

[0141] It can be calculated step by step to get Approximately , and then calculate the norm of this vector , and then divide by , where , , add the two , first calculate the norm of the previous cross product result , the numerator is , the denominator , the ratio is approximately , the square of this part is approximately , and then add , take this part to do a simple addition , and then take the square root to get .

[0142] This result shows that the fusion stable offset of view number 3 is about 38.94. When this value is small, it means that the user behavior feature points, the key elements of the scene and the endpoints of the screen vector are relatively consistent in space. If the value is very large, it may imply measurement or matching errors. Subsequently, this value will be compared with other views to judge the fusion stability and determine the point pair merging strategy.

[0143] Based on the above fusion stable offset, when traversing the view positioning point pairs, first sort all in ascending order, and establish a mapping record for each view number and its corresponding offset. If it is found in the sorting result that some is greater than 40.0, it is necessary to check whether there are significant differences in the calibration parameters or shooting times used for this view. This 40.0 is an obvious mismatch limit summarized from multiple measurements in the actual scene. In the case of detecting a value greater than 40.0, the view point pair will be separated and left for further manual confirmation. For the remaining point pairs with values lower than or equal to 40.0, position fusion will be performed in ascending order, averaging or weighted merging the similar coordinates. If the distance between a certain coordinate and the existing coordinates is less than 2.0 cm, it is considered mergable, otherwise it is retained as an independent point. In this way, the repeated points in multiple shooting views can be gradually merged into a single coordinate, realizing the unified positioning processing of key elements and user behavior feature points under different camera views. After all point pairs are processed, the multi-view fusion scene state point set can be obtained.

[0144] The steps to obtain the spatial state consistency determination result are as follows:

[0145] Based on the scene state point set with multi-view fusion, extract the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the key scene elements, the camera optical center position, the camera optical axis vector, the user screen dynamic space vector direction, and the user behavior feature point orientation vector corresponding to each view, and generate a multi-view point position direction comparison matrix;

[0146] According to the multi-view point position direction comparison matrix, calculate the spatial position difference value in the unified coordinate system. The calculation formula is:

[0147] ;

[0148] where, is the spatial position difference value, and are the coordinates of the user behavior feature points in two views respectively, is the coordinate of the key scene element in the current view, is the camera optical center coordinate, is the camera optical axis vector, is the user screen dynamic space vector direction vector;

[0149] Based on the spatial position difference value, compare it with the preset spatial tolerance threshold one by one, screen all the matching point pairs that meet the spatial tolerance threshold range, and output the unified judgment result according to the consistent matching situation of the point pairs in all views to obtain the spatial state consistency judgment result.

[0150] Specifically, when extracting the three-dimensional coordinates of user behavior feature points, the three-dimensional coordinates of key scene elements, the camera optical center position, the camera optical axis vector, and the user screen dynamic space vector direction and the user behavior feature point orientation vector based on the multi-view fusion scene state point set, it is necessary to first read the previously recorded calibration data in each shooting view and calibrate the optical center and optical axis position of the camera in this view. To ensure data accuracy, several reference objects with known coordinates can be placed in the shooting environment and the deviation between the coordinates measured by the camera and the reference value can be checked after shooting. When the deviation is greater than the 3-cm range determined in advance based on experimental statistics, recalibration should be performed and the external camera parameters should be updated. Subsequently, the coordinates of user behavior feature points and the coordinates of key scene elements are expanded and compared in the X, Y, and Z directions respectively, and are associated with the camera optical center position and the camera optical axis vector. When the distance value between a certain point and the optical center in the X direction exceeds the set maximum distance threshold of 600.0, it is necessary to check whether there are long-shot scenes or scene occlusions in the shooting script. This 600.0 is obtained by leaving a certain margin after measuring the size of the indoor vision detection site. If it is verified that it is a long-shot scene, additional shooting position information needs to be read for position conversion to avoid error accumulation when combining coordinates or direction vectors subsequently. Then, according to the recorded user screen dynamic space vector direction options, the orientation vector of the feature points is compared with the screen vector, and a separate label is assigned to the situation where the deviation exceeds 5 degrees. This 5 degrees is obtained by selecting the median plus 2 degrees from the statistical sequence of relative error angles of each camera position in multi-view shooting. Subsequently, all multi-view points that meet the data quality requirements are uniformly recorded in an index table, and the corresponding relationship between the camera optical axis vector and the camera optical center position is marked item by item in this index table, and the mismatched or abnormal records are rechecked. After these steps are completed, a multi-view point direction comparison matrix covering the three-dimensional coordinates of user behavior feature points, the three-dimensional coordinates of key scene elements, the camera optical center position, the camera optical axis vector, the user screen dynamic space vector direction, and the user behavior feature point orientation vector can be obtained.

[0151] Formula The advantage of is that it incorporates multiple factors such as user behavior feature points, camera optical axes, key scene elements, and user screen dynamic space vector directions from two views into the same evaluation index in a unified coordinate system. By means of the combination of vector cross product, vector dot product, and Euclidean distance, the difference quantity includes both directional measurement and positional measurement, which is very intuitive for identifying the consistency of multi-view data.

[0152] The steps for obtaining the parameter are as follows. This parameter represents the The three-dimensional coordinates of the user behavior feature points from a certain perspective. The acquisition process relies on multiple cameras on-site to simultaneously capture and record the user's activity scene. For example, in a certain measurement session, through multi-camera fusion, it is known that the user's location coordinates (120.0, 55.5, 98.0) are all within the interval [-500, 500], and they can be registered as or and other corresponding numbers.

[0153] The steps for obtaining the parameter are that this parameter corresponds to but comes from another perspective.

[0154] The steps for obtaining the parameter are as follows. The camera optical axis vector of the th perspective is generated by positioning the lens orientation during camera calibration. The method is to measure the angles between the camera and several points with known coordinates in the scene and combine the three-dimensional coordinate differences to obtain the X, Y, and Z components of the optical axis direction. For example, (0.866, 0.0, 0.5) represents a direction that is inclined 30 degrees in the horizontal plane and tilted upward.

[0155] The steps for obtaining the parameter are as follows. The coordinates of the key elements in the scene from the current perspective mainly target elements with stable and unchanging positions such as the screen surface and the device fixing structure. It is necessary to perform three-dimensional scanning or laser ranging on these elements when arranging the scene and register their coordinate values in the scene database. Then, combine these coordinates with the camera internal parameters to form the world coordinate mapping from this perspective.

[0156] The steps for obtaining the parameter are as follows. The camera optical center coordinates of the th perspective. Usually, a laser rangefinder or a special calibration jig is used on-site to repeatedly measure the offset of the camera from the origin of the three-dimensional scene. After measurement, record it as (ox, oy, oz) corresponding to , for example, in an arrangement, a displacement of (10.0, 0.0, 50.0) millimeters is measured, then =(0.01, 0.0, 0.05) meters.

[0157] The steps for obtaining the parameter are as follows. The dynamic space vector direction of the user's screen corresponds to the screen vector direction calculated and recorded in the previous steps. When it is necessary to determine the direction components, the orientation of the edge of the display device can be measured first in the static calibration stage, and then the relative orientation between the user and the display device can be obtained based on real-time monitoring. On this basis, the unit vector components in the X, Y, and Z directions are calculated.

[0158] Calculation process: In an example calculation, select 、 , , , , ;

[0159] First, calculate , where , and take the cross product with to get . The norm ;

[0160] Then, calculate the dot product of vectors and , . The norm . After normalization, , . The norm . After approximate normalization . The dot product of the two . The absolute value ;

[0161] Then, calculate , , . Its norm . Finally, add the three parts together , that is .

[0162] This result shows that there is a certain degree of spatial difference between the user feature points from two perspectives in this example scenario. The larger the value, the more obvious the deviation in position or direction. Subsequently, this result can be compared with a pre-defined spatial tolerance threshold to determine whether the matching requirements are met.

[0163] Based on the spatial position difference value, check the relationship between each pair of user behavior feature points and the key elements of the scene and the camera optical center. It is necessary to give a spatial tolerance threshold during the on-site survey phase to measure whether the coordinate difference is acceptable. This threshold can be measured and recorded according to the size of the site and the activity range of the user. For example, in an indoor site, the position coordinates are restricted within the range of -500 to 500 meters. Combine the previous data statistics to set a threshold range of 40.0 to 60.0 and fix it at 50.0 after multiple tests. When it is detected that a pair of difference values is greater than 50.0, mark this pair of feature points in the comparison list and wait for subsequent verification. If it is lower than or equal to 50.0, it is temporarily classified as within the matching range. Then, perform a consistency check on these point pairs that are within the matching range from each perspective. For the situation where the same feature point satisfies not exceeding the above threshold in multiple perspectives, merge them into the same global coordinate point. If there are problems with relatively large difference values between some perspectives, it is necessary to conduct an in-depth comparison in camera calibration or user movement trajectory recording to avoid deviations in subsequent integration due to inconsistent data. Record all the qualified points item by item and summarize them into the final comparison result index table. Finally, the spatial state consistency determination result can be obtained.

[0164] The above is only a preferred embodiment of the present invention, and it does not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. An unattended vision detection method based on image recognition, characterized in that, It includes the following steps: Collect the image data of the vision detection site, perform scene structure line detection and plane detection to obtain a line segment set and plane parameters, and establish basic geometric description elements; Based on the basic geometric description elements, combine the coordinates of the user's line of sight area obtained by detection and the boundary coordinates of the vision detection screen for spatial integration to obtain the initial scene geometric structure information; Based on the initial scene geometric structure information, track the changes in the camera pose in consecutive image frames, calculate and obtain the displacement and rotation parameters of the camera relative to the initial geometric structure, obtain the camera relative pose sequence, and based on the camera relative pose sequence, construct a three-dimensional space coordinate system and solve the coordinate transformation from the user's line of sight area to the vision detection screen to generate the user screen dynamic space vector; Obtain the multi-view image data of the vision detection site, respectively identify the three-dimensional position information of the user behavior feature points and the key scene elements at each view, gather them into a single-view element positioning table, and based on the unified coordinate system defined by the single-view element positioning table and the user screen dynamic space vector, perform cross-view data mapping and fusion to establish a multi-view fusion scene state point set; Based on the multi-view fusion scene state point set, match the corresponding user behavior feature points and scene elements at different views, calculate the spatial position difference value in the unified coordinate system, compare the spatial position difference value with the preset spatial tolerance threshold for judgment, and obtain the spatial state consistency judgment result.

2. The unattended vision detection method based on image recognition according to claim 1, characterized in that, The steps for obtaining the basic geometric description elements are as follows: Collect the image data of the vision detection site. Based on the collected image data of the vision detection site, perform image edge gradient analysis to calculate the pixel gradient change of the edge points, extract the edge contours at the boundaries of each object in the scene, combine the edge points to form structural line segments, and generate the structural line segment set of the vision detection site; Based on the structural line segment set of the vision detection site, extract the spatial coordinate data of all the line segment endpoints in the line segment set, judge the coplanarity of the line segment endpoint spatial coordinates, screen the coplanar point groups according to the coplanarity of the line segment endpoint coordinates, and fit and calculate the plane equation parameters corresponding to the spatial positions of the coplanar point groups to generate the plane parameters corresponding to the structural line segments; Based on the plane parameters corresponding to the structural line segments, call the line segment endpoint coordinates in the structural line segment set of the vision detection site, perform spatial projection calculations on the line segment endpoint spatial coordinates and the corresponding plane parameters one by one, determine the spatial position relationship between the line segment and the corresponding plane, establish the spatial mapping relationship between the line segment endpoint coordinates and the plane parameters, and generate the basic geometric description elements.

3. The unattended vision detection method based on image recognition according to claim 1, characterized in that The steps for obtaining the initial scene geometric structure information are as follows: Based on the basic geometric description elements, call the line segment endpoint coordinates and plane parameters in the basic geometric description elements, combine the boundary coordinates of the vision detection screen, perform spatial coordinate transformation and projection calculations, determine the spatial position relationship of the vision detection screen in the basic geometric description elements, and generate the screen boundary spatial mapping coordinates; Based on the screen boundary spatial mapping coordinates, obtain the coordinates of the user's line of sight area, perform coordinate system alignment and spatial coordinate transformation, integrate the spatial position relationship between the user's line of sight area and the screen boundary, establish the mapping correspondence relationship between the line of sight area coordinates and the screen boundary coordinates, and generate the line of sight screen integration spatial coordinates; Integrate the spatial coordinates based on the line-of-sight screen to determine the spatial matching relationship between the integrated spatial coordinates of the line-of-sight screen and the basic geometric description elements, and obtain the initial scene geometric structure information.

4. The unattended vision detection method based on image recognition according to claim 1, wherein The steps for obtaining the relative pose sequence of the camera are as follows: Based on the initial scene geometric structure information, obtain the three-dimensional boundary points of the user's line-of-sight area and the corner point coordinates of the vision detection screen boundary. Combine the image acquisition timestamps and camera internal parameters recorded in the consecutive image frames to calculate the matching error and reprojection offset of the corresponding boundary points between the image frames, and generate a consecutive image frame matching error sequence. According to the consecutive image frame matching error sequence, calculate the camera attitude change angle value. According to the camera attitude change angle value, combine the three-dimensional offset of the boundary points of each frame in the consecutive image frame matching error sequence to obtain the three-dimensional rotation axis offset and direction vector change between adjacent frames, and obtain the relative pose sequence of the camera.

5. The unattended vision detection method based on image recognition according to claim 1, characterized in that, The steps for obtaining the dynamic spatial vector of the user screen are as follows: Based on the relative pose sequence of the camera, obtain the rotation matrix and translation vector of each frame. Combine the focal length, image center coordinates, pixel size, and image size parameters in the camera internal parameters to reconstruct the coordinates and spatial vectors of the starting point coordinates of the user's line-of-sight area and the corner point coordinates of the screen in sequence, and generate an initial three-dimensional vector group from the user's line of sight to the screen. According to the initial three-dimensional vector group from the user's line of sight to the screen, calculate the spatial transformation value from the user's line-of-sight vector to the target position on the screen. Based on the spatial transformation value from the user's line-of-sight vector to the target position on the screen, perform vector direction normalization and projection difference correction to obtain the dynamic spatial vector of the user screen.

6. The unattended vision detection method based on image recognition according to claim 1, wherein, The steps for obtaining the single-viewpoint element positioning table are as follows: Obtain multi-viewpoint image data of the vision detection site, extract the pixel information of the user behavior feature area and the scene key element area in the multi-viewpoint image data of the vision detection site, and identify the three-dimensional position information of the user behavior feature points through spatial mapping to generate a set of spatial coordinates of the user behavior feature points. Based on the set of spatial coordinates of the user behavior feature points, identify the contour edges of the scene key elements in the multi-viewpoint image data of the vision detection site, and perform spatial coordinate matching and three-dimensional space fitting to generate a set of three-dimensional spatial position coordinates of the scene key elements. Based on the set of three-dimensional spatial position coordinates of the scene key elements, summarize the three-dimensional spatial position coordinates of the user behavior feature points and the scene key elements in the corresponding viewpoints, and establish a single-viewpoint element positioning table.

7. The unattended vision detection method based on image recognition according to claim 1, characterized in that The steps for obtaining the multi-viewpoint fusion scene state point set are as follows: Based on the single-viewpoint element positioning table, extract the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the scene key elements, the image frame numbers, the camera internal parameter matrix, and the coordinate endpoints of the dynamic spatial vector of the user screen in each viewpoint, and uniformly transform them into the unified coordinate system defined by the dynamic spatial vector of the user screen to generate a set of multi-viewpoint three-dimensional positioning matrices in the unified coordinate system. According to the set of multi-viewpoint three-dimensional positioning matrices in the unified coordinate system, calculate the fusion stable offset of each viewpoint. Based on the fused stable offset, traverse each pair of perspective positioning points, perform spatial position fusion priority sorting in ascending order of the fused stable offset, screen out the point pairs that meet the spatial coincidence degree condition and merge them into a unified spatial point position, and establish a multi-perspective fusion scene state point set.

8. The unattended vision detection method based on image recognition according to claim 1, characterized in that, The steps for obtaining the determination result of spatial state consistency are as follows: Based on the multi-perspective fusion scene state point set, extract the three-dimensional coordinates of the user behavior feature points, the three-dimensional coordinates of the key scene elements, the camera optical center position, the camera optical axis vector, the dynamic spatial vector direction of the user screen, and the orientation vector of the user behavior feature points corresponding to each perspective, and generate a multi-perspective point position direction comparison matrix; According to the multi-perspective point position direction comparison matrix, calculate the spatial position difference value in the unified coordinate system; Based on the spatial position difference value, perform pairwise comparison with the preset spatial tolerance threshold, screen out all matching point pairs within the range of the spatial tolerance threshold, and output a unified judgment result according to the consistent matching situation of the point pairs in all perspectives to obtain the determination result of spatial state consistency.

9. The unattended vision detection system for the image recognition-based unattended vision detection method according to any one of claims 1-8, characterized in that, Including: A data acquisition module, which based on the vision detection on-site image data, performs image acquisition, obtains multi-angle and multi-perspective image data of the vision detection site, performs scene structure line and plane detection, and constructs a scene geometric feature set by extracting the line segment set and plane parameters; A scene modeling module, which based on the scene geometric feature set, performs spatial integration, and obtains the initial scene geometric structure by combining the detected user line-of-sight area coordinates and the vision detection screen boundary coordinates; According to the initial scene geometric structure, track the camera pose change in consecutive image frames, calculate and obtain the relative camera pose, and then construct a three-dimensional space coordinate system, generate a relative camera pose sequence, solve the coordinate transformation from the user line-of-sight area to the vision detection screen, and obtain the dynamic spatial vector of the user screen; A line-of-sight tracking module, which based on the relative camera pose sequence and the dynamic spatial vector of the user screen, performs tracking of the pose change in consecutive image frames, calculates the dynamic changes of the line-of-sight direction and user behavior, and generates a user line-of-sight tracking sequence; A multi-perspective fusion module, which based on the multi-perspective image data of the vision detection site, identifies the user behavior feature points and key scene elements in each perspective, obtains the three-dimensional spatial position data, and aggregates them into a single-perspective element positioning table; combines the user line-of-sight tracking sequence and the single-perspective element positioning table, performs cross-perspective data mapping and fusion, and generates a multi-perspective fusion scene state point set; A spatial consistency module, which based on the multi-perspective fusion scene state point set, matches the user behavior feature points and scene elements in different perspectives, calculates the spatial position difference value, compares the spatial position difference value with the preset spatial tolerance threshold, judges the consistency, and obtains the determination result of spatial consistency.

Citation Information

Patent Citations

  • Unattended group vision screening device

    CN111839453A

  • Multi-view human behavior recognition method and system under edge computing architecture

    CN113743221A