Dynamic vision training method and system based on naked eye 3D imaging, electronic equipment and storage medium
By analyzing user eye movement data in real time, generating viewing angle changes parameters, and using deep learning parallax estimation algorithm to generate pixel-level virtual views, it solves the problem of dynamic changes in visual scenes in dynamic vision training systems, and achieves efficient and immersive visual training effects.
Patent Information
- Application Number
- CN202510342629.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-21
AI Technical Summary
In the dynamic vision training system of naked-eye 3D imaging, how to achieve dynamic changes in visual scenes, quickly and accurately capture user head and eye movement information, and calculate the best parallax parameters based on this information, generate high-quality stereoscopic images, and meet the needs of dynamic vision training.
By obtaining real-time motion information of the user's eyeball, analyzing eye movement data using convolutional neural networks, obtaining view angle change parameters, introducing deep learning parallax estimation algorithm to generate pixel-level virtual views, and rendering them into three-dimensional images.
It realizes dynamic 3D image rendering with low latency and high frame rate, provides an immersive visual training environment, and improves users' visual cognitive ability.
Smart Images

Figure CN119970454A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vision training technology, and in particular to a dynamic vision training method, system, electronic device and storage medium based on naked-eye 3D imaging. Background Art
[0002] In the dynamic visual training system of naked-eye 3D imaging, how to achieve dynamic changes in visual scenes is a key technical issue. Traditional 3D imaging technology usually uses fixed parallax settings, and the depth information of the generated stereo images is limited, which is difficult to meet the needs of dynamic visual training. In order to create realistic stereo dynamic scenes, it is necessary to adjust the parallax in real time according to the user's perspective changes and generate corresponding stereo images. This requires the system to be able to quickly and accurately capture the movement information of the user's head and eyeballs, and calculate the optimal parallax parameters based on this information. At the same time, generating high-quality stereo images also puts higher requirements on computing power. How to achieve low-latency, high-frame-rate dynamic 3D image rendering with limited computing resources is another technical problem that needs to be solved urgently. In addition, in the process of dynamic visual training, it is also necessary to dynamically adjust the training content and difficulty according to user feedback and training progress. This requires the system to be able to analyze user behavior data in real time and adaptively generate training tasks based on certain algorithm models. The collection and analysis of user behavior data, as well as the automatic generation of training tasks, all involve complex artificial intelligence algorithms. How to design and implement these algorithms is also a technical challenge that cannot be ignored. Summary of the invention
[0003] In order to solve the above technical problems, the present invention provides a dynamic vision training method based on naked eye 3D imaging, the method comprising:
[0004] Acquire real-time movement information of the user's eyeball, and analyze the real-time movement information to obtain the user's viewing angle change parameter;
[0005] Based on the perspective change parameters, a disparity estimation algorithm is used to extract and match the input left and right view images to obtain a pixel-level virtual view;
[0006] Perform dynamic rendering based on the pixel-level disparity map to obtain a 3D view;
[0007] A training task is generated, and the 3D view is selected based on the training task to train the user's dynamic vision.
[0008] Preferably, the method for obtaining the viewing angle change parameter includes:
[0009] Acquire real-time eye movement information of the user, perform denoising and outlier cleaning on the real-time motion information, and obtain cleaned eye movement data;
[0010] Using a convolutional neural network model to extract and analyze features of the cleaned eye movement data to obtain a view angle change feature vector;
[0011] The view angle change feature vectors are classified by a clustering algorithm, similar feature vectors are aggregated together, and cluster centers representing different view angle change modes are obtained;
[0012] Calculating the distance between the perspective change feature vector and each of the cluster centers, finding the cluster center with the closest distance, and taking the corresponding perspective change pattern as the perspective change trend of the current user;
[0013] Based on the viewing angle change trend, a corresponding quantization parameter is searched in a viewing angle change parameter mapping table to obtain the viewing angle change parameter.
[0014] Preferably, the method for obtaining the pixel-level virtual view includes:
[0015] According to the left and right view images and the viewing angle change parameter, using a convolutional neural network model to extract features of the left and right view images to obtain feature maps of the left and right view images;
[0016] Matching the feature map using a feature matching algorithm, calculating the position offset of each pixel in the left and right view images, and generating an initial disparity map;
[0017] Denoising the initial disparity map to obtain a smoothed disparity map;
[0018] Combining the smoothed disparity map with the pixel information of the left and right view images, calculating the depth information of the scene through a three-dimensional reconstruction algorithm, and generating a pixel-level depth map;
[0019] The pixel-level depth map is post-processed and fused with the original left view image to obtain the pixel-level virtual view.
[0020] Preferably, the training method comprises:
[0021] According to the training objectives and requirements, select several task templates with high matching degree from the pre-built task template library as basic templates;
[0022] Parsing the basic template, extracting key elements in the template, and generating the training task based on the key elements;
[0023] Selecting the 3D view based on the training task to generate a training scene;
[0024] In combination with the training scenario, the user's dynamic vision is trained based on the training task.
[0025] The present invention also provides a dynamic vision training system based on naked eye 3D imaging, the system is used to implement any of the above methods, including: a viewing angle change parameter acquisition module, a virtual view generation module, a 3D view rendering module and a training module;
[0026] The viewing angle change parameter acquisition module is used to acquire the real-time movement information of the user's eyeball, and analyze the real-time movement information to obtain the user's viewing angle change parameter;
[0027] The virtual view generation module extracts and matches the input left and right view images based on the view change parameters using a disparity estimation algorithm to obtain a pixel-level virtual view;
[0028] The 3D view rendering module performs dynamic rendering based on the pixel-level disparity map to obtain a 3D view;
[0029] The training module is used to generate a training task, select the 3D view based on the training task, and train the user's dynamic vision.
[0030] Preferably, the viewing angle change parameter acquisition module includes: a preprocessing unit, a first feature extraction unit, a clustering unit, a change trend determination unit and a parameter search unit;
[0031] The preprocessing unit is used to obtain real-time movement information of the user's eyeball, denoise and clean up outliers on the real-time movement information, and obtain cleaned eye movement data;
[0032] The first feature extraction unit uses a convolutional neural network model to extract and analyze features of the cleaned eye movement data to obtain a view angle change feature vector;
[0033] The clustering unit classifies the view angle change feature vectors by a clustering algorithm, aggregates similar feature vectors together, and obtains cluster centers representing different view angle change modes;
[0034] The change trend determination unit is used to calculate the distance between the perspective change feature vector and each of the cluster centers, find the cluster center with the closest distance, and use the corresponding perspective change pattern as the perspective change trend of the current user;
[0035] The parameter search unit searches for corresponding quantization parameters from a viewing angle change parameter mapping table based on the viewing angle change trend to obtain the viewing angle change parameter.
[0036] Preferably, the virtual view generation module comprises: a second feature extraction unit, a matching unit, a denoising unit, a three-dimensional reconstruction unit and a fusion unit;
[0037] The second feature extraction unit extracts features from the left and right view images using a convolutional neural network model according to the left and right view images and the viewing angle change parameter to obtain feature maps of the left and right view images;
[0038] The matching unit matches the feature map using a feature matching algorithm, calculates the position offset of each pixel in the left and right view images, and generates an initial disparity map;
[0039] The denoising unit is used to denoise the initial disparity map to obtain a smoothed disparity map;
[0040] The 3D reconstruction unit combines the smoothed disparity map with the pixel information of the left and right view images, calculates the depth information of the scene through a 3D reconstruction algorithm, and generates a pixel-level depth map;
[0041] The fusion unit is used to post-process the pixel-level depth map and fuse it with the original left view image to obtain the pixel-level virtual view.
[0042] Preferably, the training module comprises: a task template confirmation unit, a training task generation unit, a scenario construction unit and a training unit;
[0043] The task template confirmation unit selects a number of task templates with high matching degree as basic templates from a pre-built task template library according to the training objectives and requirements;
[0044] The training task generating unit is used to parse the basic template, extract key elements in the template, and generate the training task based on the key elements;
[0045] The scene construction unit selects the 3D view based on the training task to generate a training scene;
[0046] The training unit trains the user's dynamic vision based on the training task in combination with the training scenario.
[0047] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned dynamic vision training method based on naked-eye 3D imaging when executing the program.
[0048] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, the above-mentioned dynamic vision training method based on naked-eye 3D imaging is implemented.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] The present invention obtains the user's head and eye movement information in real time, uses a convolutional neural network to analyze the eye movement data, obtains the viewing angle change parameters, and then introduces a deep learning disparity estimation algorithm to generate a pixel-level virtual view, and renders it into a three-dimensional image. The rendered three-dimensional image is used to provide an immersive visual training environment. For automatic generation of training tasks, combined with rule templates and learning models, personalized task optimization is achieved. The present invention adopts a modular architecture to design artificial intelligence algorithms, supports online learning and updating, realizes intelligent and personalized visual training, and can effectively improve the user's visual cognitive ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0052] Figure 1 A schematic diagram of a method flow of an embodiment of the present invention;
[0053] Figure 2 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Description of the drawings:
[0055] 1010, processor; 1020, memory; 1030, input / output interface; 1040, communication interface; 1050, bus. DETAILED DESCRIPTION
[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0057] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0058] Embodiment 1
[0059] In this embodiment, if Figure 1 As shown, a dynamic vision training method based on naked eye 3D imaging includes the following steps:
[0060] S1. Obtain real-time movement information of the user's eyeball, and analyze the real-time movement information to obtain the user's viewing angle change parameters.
[0061] The method for obtaining the perspective change parameter includes: obtaining the real-time movement information of the user's eyeball, denoising and cleaning the real-time movement information, and obtaining the cleaned eye movement data; using a convolutional neural network model to extract and analyze the features of the cleaned eye movement data, and obtaining a perspective change feature vector; for the perspective change feature vector, classifying it through a clustering algorithm, aggregating similar feature vectors together, and obtaining cluster centers representing different perspective change patterns; calculating the distance between the perspective change feature vector and each cluster center, finding the cluster center closest to it, and using the corresponding perspective change pattern as the perspective change trend of the current user; based on the perspective change trend, searching for the corresponding quantization parameter from the perspective change parameter mapping table to obtain the perspective change parameter.
[0062] In this embodiment, an infrared camera is used to capture the reflected light of the eyeball, and combined with the head posture sensor data, high-precision eye tracking is achieved to obtain real-time movement information of the user's eyeball; the original real-time movement information often contains noise, such as blinking, head shaking and other interference, and needs to be preprocessed by filtering and outlier detection methods. Specifically, median filtering can be used to remove sudden noise, and then Kalman filtering can be used to smooth the data trajectory; the cleaned eye movement data is input into a pre-trained convolutional neural network model to extract features reflecting changes in viewing angles. The structure of the convolutional neural network model can include multiple convolutional layers and pooling layers for extracting local features and reducing dimensionality, and finally a feature vector is output through a fully connected layer; the extracted features Cluster analysis of vectors can identify different perspective change patterns. Commonly used clustering algorithms such as K-means can cluster similar feature vectors into one category. Assuming that the number of clusters is set to 5, 5 cluster centers can be obtained, representing typical perspective change patterns such as upward, downward, left, right, and gaze. By calculating the Euclidean distance between the current feature vector and each cluster center, the user's current perspective change trend can be determined. A mapping table is established to correspond different perspective change patterns to specific quantitative values. In this embodiment, looking up corresponds to +2, looking down corresponds to -2, looking left and right corresponds to +1 and -1 respectively, and gaze corresponds to 0. This quantification allows the system to accurately measure and respond to the degree of change in the user's perspective. When the quantization value exceeds the preset threshold, the system can trigger the corresponding operation.
[0063] S2. Based on the perspective change parameters, the disparity estimation algorithm is used to extract and match the input left and right view images to obtain a pixel-level virtual view.
[0064] The method for obtaining a pixel-level virtual view includes: extracting features of the left and right view images using a convolutional neural network model according to left and right view images and a viewing angle change parameter to obtain feature maps of the left and right view images; matching the feature maps using a feature matching algorithm to calculate the position offset of each pixel point in the left and right view images to generate an initial disparity map; denoising the initial disparity map to obtain a smoothed disparity map; combining the smoothed disparity map with pixel information of the left and right view images, calculating the depth information of the scene through a three-dimensional reconstruction algorithm to generate a pixel-level depth map; post-processing the pixel-level depth map and fusing it with the original left view image to obtain a pixel-level virtual view.
[0065] In this embodiment, for the left and right view images with a resolution of 640×480, the ResNet50 network name model is used to extract the feature maps of 2048 channels, each channel represents a high-level semantic feature, and the feature maps of the left and right views are obtained; then, the SIFT algorithm is used to detect and describe the key points in the image to establish a correspondence between the left and right views. Assuming that 1000 key points are detected in the left view, 800 pairs of valid matching points can be obtained by matching with the right view, and the position offsets of these matching points constitute the initial disparity map; then, the initial disparity map is denoised by bilateral filtering. For the 8-bit grayscale disparity map, the spatial domain standard deviation can be set to 5 pixels and the value domain standard deviation can be set to 20 gray levels, which effectively removes noise and mismatches in the disparity map while retaining the clarity of the object edge; then, the 3D reconstruction algorithm is used to convert the 2D disparity information into 3D depth information:
[0066] Z=fB / d
[0067] Among them, f is the focal length of the camera, B is the baseline distance, and d is the disparity value. By performing similar calculations on each pixel, a complete pixel-level depth map can be generated. After obtaining the depth map, the median filtering method is used to fill holes and optimize edges of the depth map; finally, the optimized depth map is fused with the original left view image, and a pixel-level virtual view under a new perspective is generated through the viewpoint conversion algorithm according to the perspective change parameters.
[0068] S3. Perform dynamic rendering based on the pixel-level disparity map to obtain a 3D view.
[0069] In this embodiment, in view of the high requirements of dynamic three-dimensional rendering on computing resources, a cloud-based distributed rendering architecture is adopted to distribute rendering tasks to multiple graphics processor server nodes to achieve parallel computing. Specifically, if a dynamic three-dimensional rendering request is received, the rendering task is decomposed into multiple subtasks, and the subtasks are distributed to the corresponding nodes using a load balancing algorithm according to the load of each graphics processor service node. After each graphics processor service node receives the assigned rendering subtask, it uses local computing resources to perform parallel rendering processing and generate corresponding rendering result data. A scheduling server is set in the cloud architecture to track the rendering progress of each node. If the progress of a node lags behind the preset threshold, the task allocation is dynamically adjusted to reallocate some tasks to other nodes with lower loads. After the rendering subtask is completed, the rendering result data is uploaded to the distributed storage system of the cloud, and the result data is indexed by a hash algorithm for subsequent query and call. After all rendering subtasks are completed, the scheduling server summarizes and splices the rendering result data distributed on different nodes according to the task allocation to obtain a complete rendering result, that is, a 3D view.
[0070] S4. Generate a training task, select a 3D view based on the training task, and train the user's dynamic vision.
[0071] The training method includes: selecting several task templates with high matching degree as basic templates from a pre-built task template library according to training objectives and requirements; parsing the basic templates, extracting key elements in the templates, and generating training tasks based on the key elements; selecting 3D views based on the training tasks to generate training scenes; combining the training scenes, and training the user's dynamic vision based on the training tasks.
[0072] In this embodiment, according to the training objectives and requirements, one or more task templates with high matching degree are selected from the pre-built task template library as basic templates, the selected basic templates are parsed, and key elements in the templates, such as attribute features such as task type, difficulty level, knowledge point coverage, etc. are extracted; then, the extracted template attribute features are input into the pre-trained task generation model, a batch of candidate task optimization schemes are generated through model reasoning, and the generated candidate task optimization schemes are screened and sorted by using a rule engine, and non-compliant or redundant schemes are eliminated to obtain a high-quality and available subset of task optimization schemes; from the screened subset of task optimization schemes, the most matching optimization scheme is selected in combination with the trainees' personal learning characteristics, ability level, etc., and the basic template is dynamically adjusted and optimized; the optimized task template is instantiated, and a specific personalized training task is generated according to the task description, parameter indicators, etc. in the template; a 3D view is selected based on the personalized training task to generate a training scene, and the user's dynamic vision is trained based on the training task in combination with the training scene.
[0073] Embodiment 2
[0074] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides a dynamic vision training system based on naked-eye 3D imaging, including: a viewing angle change parameter acquisition module, a virtual view generation module, a 3D view rendering module and a training module.
[0075] The viewing angle change parameter acquisition module is used to acquire the real-time movement information of the user's eyeball, and analyze the real-time movement information to obtain the user's viewing angle change parameters.
[0076] The perspective change parameter acquisition module includes: a preprocessing unit, a first feature extraction unit, a clustering unit, a change trend determination unit and a parameter search unit; the preprocessing unit is used to obtain the real-time movement information of the user's eyeball, denoise the real-time movement information and clean the outliers to obtain the cleaned eye movement data; the first feature extraction unit uses a convolutional neural network model to extract and analyze the cleaned eye movement data to obtain a perspective change feature vector; the clustering unit classifies the perspective change feature vector through a clustering algorithm, aggregates similar feature vectors together, and obtains cluster centers representing different perspective change patterns; the change trend determination unit is used to calculate the distance between the perspective change feature vector and each cluster center, find the cluster center with the nearest distance, and use the corresponding perspective change pattern as the perspective change trend of the current user; the parameter search unit searches for the corresponding quantization parameter from the perspective change parameter mapping table based on the perspective change trend to obtain the perspective change parameter.
[0077] The virtual view generation module uses the disparity estimation algorithm to extract and match the features of the input left and right view images based on the perspective change parameters to obtain a pixel-level virtual view.
[0078] The virtual view generation module includes: a second feature extraction unit, a matching unit, a denoising unit, a three-dimensional reconstruction unit and a fusion unit; the second feature extraction unit uses a convolutional neural network model to extract features of the left and right view images according to the left and right view images and the viewing angle change parameters to obtain feature maps of the left and right view images; the matching unit uses a feature matching algorithm to match the feature map, calculates the position offset of each pixel point in the left and right view images, and generates an initial disparity map; the denoising unit is used to denoise the initial disparity map to obtain a smoothed disparity map; the three-dimensional reconstruction unit combines the smoothed disparity map with the pixel information of the left and right view images, calculates the depth information of the scene through the three-dimensional reconstruction algorithm, and generates a pixel-level depth map; the fusion unit is used to post-process the pixel-level depth map and fuse it with the original left view image to obtain a pixel-level virtual view.
[0079] The 3D view rendering module performs dynamic rendering based on the pixel-level disparity map to obtain a 3D view.
[0080] The training module is used to generate training tasks, select 3D views based on the training tasks, and train the user's dynamic vision.
[0081] The training module includes: a task template confirmation unit, a training task generation unit, a scene construction unit and a training unit; the task template confirmation unit selects several task templates with high matching degree as basic templates from a pre-built task template library according to the training objectives and requirements; the training task generation unit is used to parse the basic template, extract the key elements in the template, and generate training tasks based on the key elements; the scene construction unit selects 3D views based on the training tasks to generate training scenes; the training unit combines the training scenes to train the user's dynamic vision based on the training tasks.
[0082] The system of the above embodiment is used to implement the corresponding dynamic vision training method based on naked-eye 3D imaging in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0083] It should be noted that the above-mentioned dynamic vision training system based on naked eye 3D imaging is embodied in the form of functional units. The term "module" here can be implemented in the form of software and / or hardware, and is not specifically limited to this.
[0084] For example, a "module" may be a software program, a hardware circuit, or a combination of the two that implements the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combined logic circuit, and / or other suitable components that support the described functions.
[0085] Embodiment 3
[0086] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the dynamic visual training method based on naked-eye 3D imaging described in any of the above embodiments is implemented.
[0087] Figure 2 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.
[0088] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0089] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0090] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0091] The communication interface 1040 is used to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB (Universal Serial Bus), network cable, etc.), or through a wireless mode (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).
[0092] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0093] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0094] The system of the above embodiment is used to implement the corresponding dynamic vision training method based on naked-eye 3D imaging in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0095] Embodiment 4
[0096] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the dynamic visual training method based on naked-eye 3D imaging as described in any of the above embodiments.
[0097] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0098] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the dynamic visual training method based on naked-eye 3D imaging as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0099] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.
[0100] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0101] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0102] Therefore, the units of each example described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present application.
[0103] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A dynamic vision training method based on naked eye 3D imaging, characterized in that: The method comprises: Acquire real-time movement information of the user's eyeball, and analyze the real-time movement information to obtain the user's viewing angle change parameter; Based on the perspective change parameters, a disparity estimation algorithm is used to extract and match the input left and right view images to obtain a pixel-level virtual view; Perform dynamic rendering based on the pixel-level disparity map to obtain a 3D view; A training task is generated, and the 3D view is selected based on the training task to train the user's dynamic vision.
2. The method for dynamic vision training based on naked eye 3D imaging according to claim 1, characterized in that: The method for obtaining the viewing angle change parameter includes: Acquire real-time eye movement information of the user, perform denoising and outlier cleaning on the real-time motion information, and obtain cleaned eye movement data; Using a convolutional neural network model to extract and analyze features of the cleaned eye movement data to obtain a view angle change feature vector; The view angle change feature vectors are classified by a clustering algorithm, similar feature vectors are aggregated together, and cluster centers representing different view angle change modes are obtained; Calculating the distance between the perspective change feature vector and each of the cluster centers, finding the cluster center with the closest distance, and taking the corresponding perspective change pattern as the perspective change trend of the current user; Based on the viewing angle change trend, a corresponding quantization parameter is searched in a viewing angle change parameter mapping table to obtain the viewing angle change parameter.
3. The method for dynamic vision training based on naked eye 3D imaging according to claim 1, characterized in that: The method for obtaining the pixel-level virtual view includes: According to the left and right view images and the viewing angle change parameter, using a convolutional neural network model to extract features of the left and right view images to obtain feature maps of the left and right view images; Matching the feature map using a feature matching algorithm, calculating the position offset of each pixel in the left and right view images, and generating an initial disparity map; Denoising the initial disparity map to obtain a smoothed disparity map; Combining the smoothed disparity map with the pixel information of the left and right view images, calculating the depth information of the scene through a three-dimensional reconstruction algorithm, and generating a pixel-level depth map; The pixel-level depth map is post-processed and fused with the original left view image to obtain the pixel-level virtual view.
4. The method for dynamic vision training based on naked eye 3D imaging according to claim 1, characterized in that: The training method comprises: According to the training objectives and requirements, select several task templates with high matching degree from the pre-built task template library as basic templates; Parsing the basic template, extracting key elements in the template, and generating the training task based on the key elements; Selecting the 3D view based on the training task to generate a training scene; In combination with the training scenario, the user's dynamic vision is trained based on the training task.
5. A dynamic vision training system based on naked eye 3D imaging, the system being used to implement the method according to any one of claims 1 to 4, characterized in that: include: View change parameter acquisition module, virtual view generation module, 3D view rendering module and training module; The viewing angle change parameter acquisition module is used to acquire the real-time movement information of the user's eyeball, and analyze the real-time movement information to obtain the user's viewing angle change parameter; The virtual view generation module extracts and matches the input left and right view images based on the view change parameters using a disparity estimation algorithm to obtain a pixel-level virtual view; The 3D view rendering module performs dynamic rendering based on the pixel-level disparity map to obtain a 3D view; The training module is used to generate a training task, select the 3D view based on the training task, and train the user's dynamic vision.
6. The dynamic vision training system based on naked eye 3D imaging according to claim 5, characterized in that: The viewing angle change parameter acquisition module includes: a preprocessing unit, a first feature extraction unit, a clustering unit, a change trend determination unit and a parameter search unit; The preprocessing unit is used to obtain real-time movement information of the user's eyeball, denoise and clean the real-time movement information, and obtain cleaned eye movement data; The first feature extraction unit uses a convolutional neural network model to extract and analyze features of the cleaned eye movement data to obtain a view angle change feature vector; The clustering unit classifies the view angle change feature vectors by a clustering algorithm, aggregates similar feature vectors together, and obtains cluster centers representing different view angle change modes; The change trend determination unit is used to calculate the distance between the perspective change feature vector and each of the cluster centers, find the cluster center with the closest distance, and use the corresponding perspective change pattern as the perspective change trend of the current user; The parameter search unit searches for corresponding quantization parameters from a viewing angle change parameter mapping table based on the viewing angle change trend to obtain the viewing angle change parameter.
7. The dynamic vision training system based on naked eye 3D imaging according to claim 5, characterized in that: The virtual view generation module includes: a second feature extraction unit, a matching unit, a denoising unit, a three-dimensional reconstruction unit and a fusion unit; The second feature extraction unit extracts features from the left and right view images using a convolutional neural network model according to the left and right view images and the viewing angle change parameter to obtain feature maps of the left and right view images; The matching unit matches the feature map using a feature matching algorithm, calculates the position offset of each pixel in the left and right view images, and generates an initial disparity map; The denoising unit is used to denoise the initial disparity map to obtain a smoothed disparity map; The 3D reconstruction unit combines the smoothed disparity map with the pixel information of the left and right view images, calculates the depth information of the scene through a 3D reconstruction algorithm, and generates a pixel-level depth map; The fusion unit is used to post-process the pixel-level depth map and fuse it with the original left view image to obtain the pixel-level virtual view.
8. The dynamic vision training system based on naked eye 3D imaging according to claim 5, characterized in that: The training module includes: a task template confirmation unit, a training task generation unit, a scenario construction unit and a training unit; The task template confirmation unit selects a number of task templates with high matching degree as basic templates from a pre-built task template library according to the training objectives and requirements; The training task generating unit is used to parse the basic template, extract key elements in the template, and generate the training task based on the key elements; The scene construction unit selects the 3D view based on the training task to generate a training scene; The training unit trains the user's dynamic vision based on the training task in combination with the training scenario.
9. An electronic device, characterized in that: It comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the dynamic vision training method based on naked-eye 3D imaging as claimed in any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the dynamic vision training method based on naked-eye 3D imaging as described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Visual training method and system based on naked eye 3D display
CN110099271A
Pupil dynamic tracking method and system based on naked eye 3D
CN116723308A
Naked eye 3D real-time dynamic pupil tracking system based on double cameras
CN119011806A
Visual scene recognition method based on adaptive dynamic feature aggregation
CN119380159A
Naked eye light field 3D display image fusion method and system
CN119399042A
Cited By
Naked eye 3D imaging evaluation method and system based on feedback analysis
CN120455644A
Motion vision training method and system based on naked eye 3D display
CN122140488A