Head-mounted device and computer-readable medium for motion sickness information detection
Through image acquisition by the head-mounted device and visual image processing by the processor, depth and optical flow image data are generated, which solves the problems of reduced virtual reality experience and high computing resource consumption caused by contact devices in the existing technology, and realizes efficient motion sickness detection.
Patent Information
- Application Number
- CN202510694576.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Existing technologies require users to wear dedicated contact devices when detecting motion sickness, which reduces the virtual reality experience and consumes a lot of computing resources.
The image acquisition device in the head-mounted device is used to collect visual image data, which is preprocessed by the processor to generate depth and optical flow image data. The pre-trained motion sickness detection model is used to generate feature maps and temporal feature information to achieve motion sickness detection and reduce the consumption of computing resources.
Through non-contact visual image processing, the user's virtual reality experience loss is reduced, the consumption of computing resources is reduced, and the efficiency of motion sickness detection is improved.
Smart Images

Figure CN120219387B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly, to a head-mounted device and a computer-readable medium for motion sickness information detection. Background Art
[0002] Motion sickness is a common physiological reaction in virtual reality environments. Detecting the severity of motion sickness symptoms in users and providing timely mitigation measures are crucial technologies for improving user experience. Currently, motion sickness detection typically involves monitoring the user's electroencephalogram (EEG), galvanic skin response (GSR), and respiratory rate to capture physiological responses in VR environments and predict motion sickness.
[0003] However, when using the above method to detect motion sickness, the following technical problems often occur:
[0004] Predicting motion sickness by monitoring the user's EEG, galvanic skin response, and respiratory rate requires the user to wear specialized contact devices, which can reduce the user's VR experience. Furthermore, the simultaneous processing of multiple data points, including the user's EEG, galvanic skin response, and respiratory rate, can lead to significant computational resource consumption.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention
[0006] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] Some embodiments of the present disclosure provide a head-mounted device and a computer-readable medium for motion sickness information detection to solve one or more of the technical problems mentioned in the above background technology section.
[0008] In a first aspect, some embodiments of the present disclosure provide a head-mounted device for motion sickness information detection, the head-mounted device for motion sickness information detection comprising: an image acquisition device configured to acquire a visual image data sequence corresponding to a target user; a processor configured to perform the following processing: pre-processing the visual image data sequence to obtain a pre-processed image data sequence; for each pre-processed image data in the pre-processed image data sequence, generating depth image data based on the pre-processed image data; for every two frames of pre-processed image data in the pre-processed image data sequence that meet a preset image condition, generating optical flow image data based on the two frames of pre-processed image data; and generating optical flow image data based on the generated depth image data and the generated depth image data. The method comprises the steps of: generating a horizontal feature map, a vertical feature map, and a depth feature map by combining the optical flow image data and the feature extraction layer of the pre-trained motion sickness detection model, wherein the feature extraction layer includes a horizontal feature extraction network, a vertical feature extraction network, and a depth feature extraction network; inputting the horizontal feature map, the vertical feature map, and the depth feature map into the temporal feature extraction network of the motion sickness detection model to obtain temporal feature information; generating a motion sickness detection result based on the temporal feature information; determining information to be displayed in response to determining that the motion sickness detection result meets a preset detection result condition; a memory is configured to store the motion sickness detection result; and a display is configured to display the information to be displayed.
[0009] In a second aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation of the first aspect is implemented.
[0010] The head-mounted device for motion sickness detection disclosed herein has the following beneficial effects: It can reduce computing resource consumption during processing. Specifically, the high computing resource consumption is caused by the fact that methods for predicting motion sickness by detecting a user's electroencephalogram (EEG), galvanic skin response (GSR), respiratory rate, and other parameters require the user to wear a dedicated contact device, which can reduce the user's virtual reality experience. Furthermore, the detection process requires simultaneous processing of multiple data points, such as the user's EEG, GSR, and respiratory rate, which can lead to high computing resource consumption during processing. Therefore, the head-mounted device for motion sickness detection disclosed herein includes, first, an image acquisition device configured to acquire a sequence of visual image data corresponding to a target user. This allows for the acquisition of raw data to be processed. Second, a processor configured to perform the following processing: First, pre-processing the visual image data sequence to obtain a pre-processed image data sequence. This pre-processing reduces the interference of noise in the raw data with subsequent processing. Second, for each pre-processed image data in the pre-processed image data sequence, depth image data is generated based on the pre-processed image data. In this way, a depth map corresponding to the preprocessed image data sequence can be obtained. Then, for every two frames of preprocessed image data in the preprocessed image data sequence that meet preset image conditions, optical flow image data is generated based on the two frames of preprocessed image data. Thus, an optical flow field map corresponding to the preprocessed image data sequence can be obtained. Then, based on the generated depth image data, the generated optical flow image data, and the feature extraction layer of a pre-trained motion sickness detection model, a horizontal feature map, a vertical feature map, and a depth feature map are generated. The feature extraction layer includes a horizontal feature extraction network, a vertical feature extraction network, and a depth feature extraction network. Thus, each feature map corresponding to the preprocessed image data sequence can be obtained. The horizontal feature map, the vertical feature map, and the depth feature map are then input into the temporal feature extraction network of the motion sickness detection model to obtain temporal feature information. Thus, temporal features corresponding to the preprocessed image data sequence can be obtained. Then, based on the temporal feature information, a motion sickness detection result is generated. Thus, a motion sickness detection result for the user can be obtained using the temporal features. Then, in response to determining that the motion sickness detection result satisfies a preset detection result condition, information to be displayed is determined. Then, the memory is configured to store the motion sickness detection result. Thus, the motion sickness detection result can be stored. Finally, the display is configured to display the information to be displayed. This allows users with more severe symptoms to be provided with relief measures.Furthermore, because motion sickness can be predicted by acquiring and processing a user's visual image data sequence, without requiring the user to wear a dedicated contact device, the likelihood of a poor virtual experience resulting from wearing a dedicated contact device can be reduced. Furthermore, because motion sickness can be directly predicted using a pre-trained motion sickness detection model, without the need to simultaneously process multiple data points such as the user's electroencephalogram (EEG), galvanic skin response, and respiratory rate, the computational resource consumption during processing can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0012] Figure 1 1 is a schematic structural diagram of some embodiments of a head-mounted device for motion sickness information detection according to the present disclosure. DETAILED DESCRIPTION
[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0014] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0015] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0016] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0017] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0018] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0019] Figure 1 A schematic structural diagram of some embodiments of a head-mounted device for motion sickness information detection according to the present disclosure is shown. Figure 1 It includes an image acquisition device 1, a processor 2, a memory 3 and a display 4.
[0020] In some embodiments, the image acquisition device 1 can be configured to acquire a sequence of visual image data corresponding to a target user. Specifically, the image acquisition device 1 can be a device for acquiring the sequence of visual image data. The image acquisition device 1 can include a camera and a posture sensor. The posture sensor can be a sensor capable of monitoring the target user's head posture in a virtual scene in real time. For example, the posture sensor can be an inertial measurement unit. Each visual image data in the sequence of visual image data can be an image corresponding to the virtual scene observed by the user while wearing the head-mounted device. The sequence of visual image data can be a series of frames corresponding to the virtual scene. Each visual image data in the sequence of visual image data has an image width and an image height. Each visual image data in the sequence of visual image data has the same size. The target user can be a user wearing the head-mounted device. The head-mounted device can be a device capable of constructing a virtual scene for the user and detecting the user's degree of dizziness in the virtual scene. The head-mounted device can include the image acquisition device 1, a processor 2, a memory 3, and a display 4. For example, the head-mounted device can be VR glasses. In practice, the image acquisition device 1 can acquire a visual image data sequence corresponding to the target user through a camera.
[0021] In some embodiments, the processor 2 may be configured to perform the following processing:
[0022] First, the visual image data sequence is preprocessed to obtain a preprocessed image data sequence.
[0023] In some embodiments, the processor 2 may preprocess the visual image data sequence to obtain a preprocessed image data sequence. Each preprocessed image data in the preprocessed image data sequence may be preprocessed visual image data. In practice, first, the processor 2 may normalize the visual image data sequence using a normalization algorithm to obtain a normalized visual image data sequence as a normalized image data sequence. Then, the normalized image data sequence may be denoised using an image denoising algorithm to obtain a denoised normalized image data sequence as a preprocessed image data sequence. The normalization algorithm may be an algorithm capable of normalizing an image. For example, the normalization algorithm may be Min-Max Normalization. The image denoising algorithm may be an algorithm capable of denoising an image. For example, the image denoising algorithm may be a Gaussian filter.
[0024] Second, for each pre-processed image data in the pre-processed image data sequence, depth image data is generated based on the pre-processed image data.
[0025] In some embodiments, the processor 2 may generate depth image data for each pre-processed image data in the pre-processed image data sequence based on the pre-processed image data, wherein the depth image data may be a depth map corresponding to the pre-processed image data.
[0026] In some optional implementations of some embodiments, the processor 2 may generate depth image data based on the pre-processed image data by performing the following steps:
[0027] In the first step, feature extraction is performed on the preprocessed image data to obtain a preprocessed feature map group. Each preprocessed feature map in the preprocessed feature map group may be a feature map corresponding to the preprocessed image data. Each preprocessed feature map in the preprocessed feature map group may be a feature map of a different dimension corresponding to the preprocessed image data. In practice, the processor 2 may input the preprocessed image data into a feature extractor to obtain a preprocessed feature map group. The feature extractor may be a variational autoencoder.
[0028] In the second step, for each preprocessed feature map in the above preprocessed feature map group, perform the following steps:
[0029] In the first sub-step, the preset random noise map is downsampled to obtain a downsampled noise map corresponding to the above-mentioned preprocessing feature map. The above-mentioned random noise map can be an image rendered by randomly generated pixel values. The above-mentioned downsampled noise map can be a random noise map that has been downsampled. The above-mentioned downsampled noise map has the same size as the above-mentioned preprocessing feature map. In practice, the above-mentioned processor 2 can reduce the above-mentioned random noise map to the size corresponding to the above-mentioned preprocessing feature map through image pooling technology, so as to downsample the above-mentioned random noise map and obtain the downsampled random noise map as the downsampled noise map. The above-mentioned image pooling technology can be a technology that can perform pooling processing on an image. For example, the above-mentioned image pooling technology can be maximum pooling.
[0030] In the second sub-step, image stitching processing is performed on the downsampled noise map and the preprocessed feature map to obtain a noise stitching feature map. The noise stitching feature map may be a preprocessed feature map with the downsampled noise map added. In practice, the processor 2 may input the downsampled noise map and the preprocessed feature map into a feature stitching function to obtain the noise stitching feature map. The feature stitching function may be a function capable of stitching different feature maps. For example, the feature stitching function may be a concat function.
[0031] In the third step, the obtained noise splicing feature maps are determined as a noise splicing feature map group.
[0032] The fourth step is to generate depth image data based on the noise splicing feature map group. The depth image data may be the depth map corresponding to the preprocessed image data. In practice, the processor 2 may first sample each noise splicing feature map in the noise splicing feature map group using an interpolation sampling technique to obtain each noise splicing feature map after sampling as the sampled splicing feature map group. The sampled splicing feature maps in the sampled splicing feature map group have the same size. The interpolation sampling technique may be a technique capable of scaling the image size. For example, the interpolation sampling technique may be bilinear interpolation upsampling. The number of sampled splicing feature map groups may be determined as the number of feature maps. Furthermore, the reciprocal of the number of feature maps may be determined as feature map ratio data. For each sampled splicing feature map in the sampled splicing feature map group, the product of the sampled splicing feature map and the feature map ratio data may be determined as a sampled ratio feature map using image processing software. The sum of the determined sampled ratio feature maps may then be determined as the target feature map using the image processing software. The image processing software may be software capable of performing operations on images. For example, the image processing software may be HALCON machine vision software. Finally, the target feature map may be input into a convolutional layer to obtain depth image data.
[0033] Third, for every two frames of pre-processed image data that meet the preset image condition in the pre-processed image data sequence, optical flow image data is generated based on the two frames of pre-processed image data.
[0034] In some embodiments, the processor 2 may generate optical flow image data based on every two frames of pre-processed image data in the pre-processed image data sequence that meet a preset image condition. The optical flow image data may be an image representing the optical flow field corresponding to the two frames of pre-processed image data. The preset image condition may be that the two frames of pre-processed image data are adjacent.
[0035] In some optional implementations of some embodiments, the processor 2 may generate optical flow image data based on the two frames of pre-processed image data by performing the following steps:
[0036] In the first step, image conversion processing is performed on the two frames of pre-processed image data to obtain two frames of visual grayscale image data. Each frame of the two frames of visual grayscale image data may be a grayscale image of the corresponding pre-processed image data. In practice, the processor 2 may convert the two frames of pre-processed image data into grayscale images using a grayscale image conversion function to obtain two frames of visual grayscale image data. The grayscale image conversion function may be a function that can convert a color image into a grayscale image. For example, the grayscale image conversion function may be the cvtColor function in OpenCV.
[0037] In a second step, the visual grayscale image data that satisfies a preset image frame condition among the two frames of visual grayscale image data is determined as the first grayscale image data. The image frame condition may be that the pre-processed image data corresponding to the visual grayscale image data is ranked higher in the pre-processed image data sequence.
[0038] In the third step, the visual grayscale image data that does not meet the above image frame condition in the two frames of visual grayscale image data is determined as the second grayscale image data.
[0039] In the fourth step, for each pixel included in the first grayscale image data, perform the following steps:
[0040] The first sub-step is to determine the above pixel point as the first pixel point.
[0041] The second sub-step is to determine the first pixel matrix corresponding to the first pixel point based on the first grayscale image data and the first pixel point. The first pixel matrix can be a third-order matrix composed of the pixel value corresponding to the first pixel point and the adjacent pixel values corresponding to the first pixel point. Each of the adjacent pixel values can be a pixel value corresponding to the pixel point adjacent to the first pixel point. For example, the adjacent pixel values can be 8 pixel values corresponding to the 8 pixel points adjacent to the first pixel point with the first pixel point as the center. In practice, when there is a null value in each of the adjacent pixel values corresponding to the first pixel point, the processor 2 can determine the preset fill value as the element corresponding to the null value. The preset fill value can be a pre-set value. For example, the preset fill value can be 0.
[0042] The third sub-step is to generate a first horizontal spatial gradient and a first vertical spatial gradient based on a preset horizontal operator matrix, a preset vertical operator matrix, and the first pixel matrix. The horizontal operator matrix may be a third-order matrix for detecting horizontal edges of an image. For example, the horizontal operator matrix may be a convolution kernel of a Sobel operator in the horizontal direction. The vertical operator matrix may be a third-order matrix for detecting vertical edges of an image. For example, the vertical operator matrix may be a convolution kernel of a Sobel operator in the vertical direction. The first horizontal spatial gradient may be the gradient amplitude in the horizontal direction corresponding to the first pixel point. The first vertical spatial gradient may be the gradient amplitude in the vertical direction corresponding to the first pixel point.
[0043] In practice, first, the processor 2 may determine the Hadamard product of the horizontal operator matrix and the first pixel matrix as a first horizontal matrix, and then determine the sum of the matrix elements in the first horizontal matrix as a first horizontal spatial gradient.
[0044] In practice, first, the processor 2 may determine the Hadamard product of the vertical operator matrix and the first pixel matrix as a first vertical matrix, and then determine the sum of the matrix elements in the first vertical matrix as a first vertical spatial gradient.
[0045] In a fourth sub-step, the pixel corresponding to the first pixel among the pixels included in the second grayscale image data is determined as the second pixel. For example, when the first pixel is located at an image height of 41 and an image width of 23 in the first grayscale image data, the second pixel is the pixel at an image height of 41 and an image width of 23 in the second grayscale image data.
[0046] In a fifth sub-step, a difference between a pixel value corresponding to the first pixel point and a pixel value corresponding to the second pixel point is determined as pixel difference data.
[0047] The sixth sub-step is to construct optical flow nonlinear information based on a preset custom horizontal vector, a preset custom vertical vector, the first horizontal spatial gradient, the first vertical spatial gradient, and the pixel difference data. The custom horizontal vector may be an unknown number representing the horizontal speed of the first pixel. For example, the custom horizontal vector may be represented by u. The custom vertical vector may be an unknown number representing the vertical speed of the first pixel. For example, the custom vertical vector may be represented by v. The optical flow nonlinear information may be an equation consisting of the custom horizontal vector, the custom vertical vector, the first horizontal spatial gradient, the first vertical spatial gradient, and the pixel difference data.
[0048] In practice, the processor 2 may first determine the product of the custom horizontal vector and the first horizontal spatial gradient as a horizontal term. Secondly, the processor 2 may determine the product of the custom vertical vector and the first vertical spatial gradient as a vertical term. Then, the negative of the pixel difference data may be determined as negative difference data. Finally, the sum of the horizontal term and the vertical term may be set equal to the negative difference data, thereby constructing an equation as the optical flow nonlinear information.
[0049] In a seventh sub-step, each pixel point corresponding to the first pixel point among the pixels included in the first grayscale image data is determined as a target pixel point, wherein each target pixel point may be a pixel point adjacent to the first pixel point.
[0050] The eighth sub-step is to determine the target optical flow nonlinear information corresponding to each target pixel point based on the above-mentioned target pixel points. Among them, each target optical flow nonlinear information in the above-mentioned target optical flow nonlinear information can be the optical flow nonlinear information corresponding to the target pixel point. In practice, for each target pixel point in the above-mentioned target pixel points, the processor 2 can generate the optical flow nonlinear information corresponding to the above-mentioned target pixel point as the target optical flow nonlinear information. Among them, the method of generating the optical flow nonlinear information corresponding to the target pixel point can refer to the specific implementation method of generating the optical flow nonlinear information corresponding to the above-mentioned first pixel point, which will not be repeated here.
[0051] The ninth sub-step is to generate the optical flow horizontal data and the optical flow vertical data corresponding to the first pixel point based on the above-mentioned optical flow nonlinear information and the above-mentioned each target optical flow nonlinear information. The above-mentioned optical flow horizontal data can be the solution of the above-mentioned custom horizontal vector. The above-mentioned optical flow vertical data can be the solution of the above-mentioned custom vertical vector. In practice, the above-mentioned processor 2 can solve the above-mentioned optical flow nonlinear information and the above-mentioned each target optical flow nonlinear information by a polynomial fitting method to obtain the solution of the above-mentioned custom horizontal vector and the solution of the above-mentioned custom vertical vector. The above-mentioned polynomial fitting method can be a least squares method. Then, the solution of the above-mentioned custom horizontal vector can be determined as the optical flow horizontal data corresponding to the above-mentioned first pixel point. Then, the solution of the above-mentioned custom vertical vector can be determined as the optical flow vertical data corresponding to the above-mentioned first pixel point.
[0052] The fifth step is to generate an optical flow level matrix based on the generated optical flow level data. The optical flow level matrix can be a matrix composed of the optical flow level data. In practice, the processor 2 can determine the image matrix corresponding to the first grayscale image data as the first image matrix. Then, the optical flow level data can be combined into an optical flow level matrix based on the position of each pixel point in the first grayscale image data in the first image matrix. For example, when the position of the pixel point corresponding to the optical flow level data in the first image matrix is the third row and second column, the optical flow level data is the element corresponding to the position of the third row and second column in the optical flow level matrix.
[0053] Step 6: Generate an optical flow vertical matrix based on the generated optical flow vertical data. The optical flow vertical matrix may be a matrix composed of the optical flow vertical data. In practice, the processor 2 may combine the optical flow vertical data into an optical flow vertical matrix based on the positions of the pixels in the first grayscale image data in the first image matrix.
[0054] Step 7: Generate optical flow image data based on the optical flow horizontal matrix and the optical flow vertical matrix. In practice, for each element included in the optical flow horizontal matrix, the processor 2 may first determine the element as a target horizontal element. Then, the element corresponding to the target horizontal element among the elements included in the optical flow vertical matrix may be determined as an element to be processed. Then, the square of the target horizontal element may be determined as target square data. Then, the square of the element to be processed may be determined as square data to be processed. Then, the sum of the target square data and the square data to be processed may be determined as amplitude data corresponding to the target horizontal element. Then, the target horizontal element and the element to be processed may be input into a two-parameter inverse tangent function, and the data output by the two-parameter inverse tangent function may be obtained as direction data corresponding to the target horizontal element. Finally, for each determined amplitude data, the amplitude data may be combined into a matrix as an optical flow amplitude matrix according to the position of each target horizontal element corresponding to each amplitude data in the optical flow horizontal matrix. The obtained directional data can be combined into a matrix as an optical flow direction matrix based on the position of the target horizontal element corresponding to each directional data in the optical flow horizontal matrix. The optical flow amplitude matrix can then be input into an amplitude normalization function to normalize the element values in the optical flow amplitude matrix to the range [0, 255]. This results in a normalized optical flow amplitude matrix as a normalized amplitude matrix. The amplitude normalization function can be a function that normalizes the element values in the matrix to the range [0, 255]. For example, the amplitude normalization function can be the Cv2.Normalize function. Each element value in the optical flow direction matrix can then be normalized to the range [0, 180] using an angle normalization function. This results in a normalized optical flow direction matrix as a normalized direction matrix. The angle normalization function can be a function that normalizes the elements in the matrix to the range [0, 180]. For example, the angle normalization function can be the Cv2.cartToPolar function. Then, the normalized amplitude matrix, the normalized direction matrix, and the preset saturation matrix can be input into a merge function to obtain a merged image. The merged image can be an HSV image corresponding to the normalized amplitude matrix and the normalized direction matrix. The merge function can be a function that can convert the normalized amplitude matrix and the normalized direction matrix into an HSV image. For example, the merge function can be the cv2.merge function. The saturation matrix can be a matrix of the same size as the optical flow level matrix, and all element values are preset saturation values. The preset saturation value can be a value between [0, 255].Finally, the merged image can be input into a color space conversion function to obtain a merged image converted to RGB format as the optical flow image data. The color space conversion function can be a function that converts an HSV image to an RGB image. For example, the color space conversion function can be the cv2.cvtColor function.
[0055] In the process of adopting technical solutions to solve the above technical problems, the following problems often arise:
[0056] Methods that predict motion sickness through the user's EEG and other data often require denoising, filtering, data fusion, and other processing of the detected EEG, galvanic skin response, respiratory rate, and other data, which can easily lead to high computational complexity during data processing and, in turn, large computational resources consumed during data processing.
[0057] Faced with the above technical problems, we decided to adopt the following solutions:
[0058] In some optional implementations of some embodiments, the processor 2 may generate optical flow image data based on the two frames of pre-processed image data by performing the following steps:
[0059] In the first step, the pre-processed image data that satisfies a preset image processing condition among the two frames of pre-processed image data is determined as the first pre-processed visual image data, wherein the image processing condition may be that the pre-processed image data ranks higher in the pre-processed image data sequence.
[0060] In the second step, the pre-processed image data of the two frames that do not meet the above image processing conditions are determined as the second pre-processed visual image data.
[0061] The third step is to perform grayscale conversion processing on the first preprocessed visual image data and the second preprocessed visual image data to obtain a first preprocessed visual grayscale image and a second preprocessed visual grayscale image. The first preprocessed visual grayscale image may be a grayscale image corresponding to the first preprocessed visual image data. The second preprocessed visual grayscale image may be a grayscale image corresponding to the second preprocessed visual image data. In practice, the processor 2 may input the first preprocessed visual image data and the second preprocessed visual image data into the grayscale image conversion function, respectively, to obtain a first preprocessed visual grayscale image and a second preprocessed visual grayscale image.
[0062] The fourth step is to perform image transformation processing on the above-mentioned first preprocessed visual grayscale image and the above-mentioned second preprocessed visual grayscale image to obtain a first visual spectrum matrix and a second visual spectrum matrix. The above-mentioned first visual spectrum matrix can be a complex matrix corresponding to the above-mentioned first preprocessed visual grayscale image. The above-mentioned second visual spectrum matrix can be a complex matrix corresponding to the above-mentioned second preprocessed visual grayscale image. In practice, the above-mentioned processor 2 can perform image transformation processing on the above-mentioned first preprocessed visual grayscale image using image transformation technology to obtain a first visual spectrum matrix. The above-mentioned second preprocessed visual grayscale image can be subjected to image transformation processing using the above-mentioned image transformation technology to obtain a second visual spectrum matrix. The above-mentioned image transformation technology can be a fast Fourier transform technology.
[0063] Step 5: Generate a conjugate spectrum matrix based on the second visual spectrum matrix. The conjugate spectrum matrix may be a conjugate matrix corresponding to the second visual spectrum matrix. In practice, the processor 2 may convert the second visual spectrum matrix into a conjugate spectrum matrix using a matrix conversion function. The matrix conversion function may be a function that can convert a complex matrix into a corresponding conjugate matrix. For example, the matrix conversion function may be a conj function.
[0064] In the sixth step, the product of the first visual spectrum matrix and the conjugate spectrum matrix is determined as a conjugate visual spectrum matrix.
[0065] Step 7: Generate a first amplitude matrix based on the first visual spectrum matrix. The first amplitude matrix may be an amplitude spectrum corresponding to the first visual spectrum matrix. In practice, the processor 2 may input the first visual spectrum matrix into an absolute value function to obtain the first amplitude matrix. The absolute value function may be a function capable of generating an amplitude spectrum of a complex matrix. For example, the absolute value function may be an ABS function.
[0066] Step 8: Generate a second amplitude matrix based on the second visual spectrum matrix. The second amplitude matrix may be an amplitude spectrum corresponding to the second visual spectrum matrix. In practice, the processor 2 may input the second visual spectrum matrix into the absolute value function to obtain the second amplitude matrix.
[0067] In the ninth step, the product of the first amplitude matrix and the second amplitude matrix is determined as a visual amplitude matrix.
[0068] In the tenth step, the inverse of the above visual amplitude matrix is determined as the visual amplitude inverse matrix.
[0069] In the eleventh step, the product of the conjugate visual spectrum matrix and the inverse visual amplitude matrix is determined as the visual amplitude spectrum matrix.
[0070] Step 12: Generate a visual phase matrix based on the visual amplitude spectrum matrix. The visual phase matrix can be a real-valued matrix corresponding to the visual amplitude spectrum matrix. In practice, first, the processor 2 can perform an inverse transformation on the visual amplitude spectrum matrix using an inverse transformation technique to obtain a transformed complex matrix as an inverse transformation matrix. The inverse transformation technique can be an inverse fast Fourier transform technique. Then, the inverse transformation matrix can be input into the absolute value function to obtain a real-valued matrix corresponding to the visual amplitude spectrum matrix as a visual phase matrix.
[0071] Step 13, based on the above-mentioned visual phase matrix, determine the first peak data and the second peak data. The above-mentioned first peak data can be the number of rows corresponding to the element with the largest numerical value in the above-mentioned visual phase matrix. The above-mentioned second peak data can be the number of columns corresponding to the element with the largest numerical value in the above-mentioned visual phase matrix. In practice, first, the above-mentioned processor 2 can determine the element with the largest numerical value in the above-mentioned visual phase matrix as the target element, and then, can determine the number of rows corresponding to the above-mentioned target element as the first peak data. Then, the number of columns corresponding to the above-mentioned target element can be determined as the second peak data.
[0072] In the fourteenth step, based on the first peak data and the second peak data, horizontal translation data and vertical translation data corresponding to the first preprocessed visual image data are generated. The horizontal translation data can represent the offset of the first preprocessed visual image data relative to the second preprocessed visual image data in the horizontal direction. The vertical translation data can represent the offset of the first preprocessed visual image data relative to the second preprocessed visual image data in the vertical direction. In practice, first, the processor 2 can determine the image height and image width corresponding to the first preprocessed visual image data as the height data to be processed and the width data to be processed, respectively. Secondly, the ratio of the height data to be processed to the preset value can be determined as the center height data. The ratio of the width data to be processed to the preset value can be determined as the center width data. Then, the difference between the first peak data and the center height data can be determined as the vertical translation data. The difference between the second peak data and the center width data can be determined as the horizontal translation data.
[0073] In step 15, the first preprocessed visual grayscale image is divided into blocks to obtain grayscale image blocks. Each of the grayscale image blocks may be an image block corresponding to the first preprocessed visual grayscale image. In practice, the processor 2 may input the first preprocessed visual grayscale image into an image block division function to obtain the grayscale image blocks. The image block division function may be a function capable of dividing an image into blocks. For example, the image block division function may be a view_as_blocks function.
[0074] In step 16, each of the grayscale image blocks is initialized based on the horizontal translation data and the vertical translation data to obtain each initialized image block. Each of the initialized image blocks may be a grayscale image block whose corresponding motion vector is determined to be the horizontal translation data and the vertical translation data. In practice, for each of the grayscale image blocks, the processor 2 may determine the horizontal translation data and the vertical translation data as the motion vector corresponding to the grayscale image block to initialize the grayscale image block to obtain the initialized image block.
[0075] Step 17. For each of the above-mentioned initialization image blocks, generate image block horizontal data and image block vertical data corresponding to the above-mentioned initialization image block. The above-mentioned image block horizontal data can be the horizontal translation data corresponding to the above-mentioned initialization image block. The above-mentioned image block vertical data can be the vertical translation data corresponding to the above-mentioned initialization image block. In practice, for each of the above-mentioned initialization image blocks, the above-mentioned processor 2 can generate the horizontal translation data corresponding to the above-mentioned initialization image block as the initialization horizontal data. The vertical translation data corresponding to the above-mentioned initialization image block can be generated as the initialization vertical data. The method of generating the horizontal translation data and vertical translation data corresponding to the initialization image block can refer to the specific implementation method of generating the horizontal translation data and vertical translation data corresponding to the above-mentioned first preprocessed visual image data, which will not be repeated here. Then, the above-mentioned initialization horizontal data can be determined as the image block horizontal data. The above-mentioned initialization vertical data can be determined as the image block vertical data.
[0076] In the eighteenth step, optical flow image data is generated based on the generated horizontal data of each image block and the generated vertical data of each image block.
[0077] In practice, first, for each of the aforementioned image block horizontal data, the processor 2 may combine the aforementioned image block horizontal data into a matrix as a local horizontal displacement matrix according to the corresponding positions of the aforementioned initialized image blocks in the first preprocessed visual grayscale image. Then, the vertical data of the aforementioned image blocks may be combined into a matrix as a local vertical displacement matrix according to the corresponding positions of the aforementioned initialized image blocks in the first preprocessed visual grayscale image. Then, for each of the elements included in the local horizontal displacement matrix, the processor 2 may first determine the product of the preset first scale data and the aforementioned element as local horizontal element data. Then, the product of the aforementioned horizontal translation data and the preset second scale data may be determined as horizontal element data. Then, the sum of the aforementioned local horizontal element data and the aforementioned horizontal element data may be determined as horizontal data to be combined. Finally, the determined horizontal data to be combined may be combined into a matrix as an optical flow horizontal component matrix according to the corresponding positions of the aforementioned elements in the local horizontal displacement matrix. Then, an optical flow vertical component matrix corresponding to the aforementioned local vertical displacement matrix may be generated. The method for generating the optical flow vertical component matrix can refer to the specific implementation method of generating the optical flow horizontal component matrix corresponding to the above-mentioned local horizontal displacement matrix, which will not be repeated here. Then, the above-mentioned optical flow horizontal component matrix and the above-mentioned optical flow vertical component matrix can be input into an image conversion function to obtain optical flow image data. The above-mentioned image conversion function can be a function that can convert the optical flow field into an image. For example, the above-mentioned image conversion function can be the Quiver function in Matlab.
[0078] The above-described technical solution and its related contents, as an inventive feature of an embodiment of the present disclosure, address the issue of "high computational resource consumption." Factors contributing to this high computational resource consumption are often as follows: Methods for predicting motion sickness using a user's EEG data, such as a sensory ... In this way, grayscale images of the two frames of preprocessed image data can be obtained. Then, image transformation processing is performed on the first preprocessed visual grayscale image and the second preprocessed visual grayscale image to obtain a first visual spectrum matrix and a second visual spectrum matrix. Thus, the spectrum matrices corresponding to the two frames of preprocessed image data can be obtained. Then, based on the second visual spectrum matrix, a conjugate spectrum matrix is generated. Then, the product of the first visual spectrum matrix and the conjugate spectrum matrix is determined as the conjugate visual spectrum matrix. Then, based on the first visual spectrum matrix, a first amplitude matrix is generated. Then, based on the second visual spectrum matrix, a second amplitude matrix is generated. Then, the product of the first amplitude matrix and the second amplitude matrix is determined as the visual amplitude matrix. Then, the inverse of the visual amplitude matrix is determined as the inverse visual amplitude matrix. Then, the product of the conjugate visual spectrum matrix and the inverse visual amplitude matrix is determined as the visual amplitude spectrum matrix. Thus, the cross-power spectrum between the two frames of preprocessed image data can be obtained. Then, based on the visual amplitude spectrum matrix, a visual phase matrix is generated. Then, based on the above-mentioned visual phase matrix, the first peak data and the second peak data are determined. Then, based on the above-mentioned first peak data and the above-mentioned second peak data, the horizontal translation data and the vertical translation data corresponding to the above-mentioned first preprocessed visual image data are generated. Thus, the displacement of the first preprocessed visual grayscale image relative to the second preprocessed visual grayscale image in the horizontal direction and the vertical direction can be obtained. Then, the above-mentioned first preprocessed visual grayscale image is divided into blocks to obtain various grayscale image blocks. Thus, the first preprocessed visual grayscale image can be divided into blocks. Then, based on the above-mentioned horizontal translation data and the above-mentioned vertical translation data, the above-mentioned various grayscale image blocks are initialized to obtain various initialized image blocks.Thus, each initialization image block can be obtained. Then, for each of the aforementioned initialization image blocks, image block horizontal data and image block vertical data corresponding to the initialization image block are generated. Finally, optical flow image data is generated based on the generated horizontal data and vertical data for each image block. Thus, an optical flow field corresponding to the visual image data can be obtained. Furthermore, because the visual image data sequence corresponding to the user can be processed in blocks to generate each piece of optical flow image data, the user's motion sickness can be predicted using each piece of optical flow image data without the need for fusion or other processing of EEG data. This reduces the computational complexity of data processing and the computing resources consumed during data processing.
[0079] Fourth, based on the generated depth image data, the generated optical flow image data, and the feature extraction layer in the pre-trained motion sickness detection model, a horizontal feature map, a vertical feature map, and a depth feature map are generated.
[0080] In some embodiments, the processor 2 may generate a horizontal feature map, a vertical feature map, and a depth feature map based on the generated depth image data, the generated optical flow image data, and a feature extraction layer in a pre-trained motion sickness detection model.
[0081] The horizontal feature map may be a feature map corresponding to the horizontal motion vector of each optical flow image data. The vertical feature map may be a feature map corresponding to the vertical motion vector of each optical flow image data. The depth feature map may be a feature map corresponding to each depth image data.
[0082] The motion sickness detection model may be a neural network model that takes the depth image data and the optical flow image data as input and outputs a motion sickness data sequence. The motion sickness detection model may include a feature extraction layer, a temporal feature extraction network, and an output layer. The feature extraction layer may include a horizontal feature extraction network, a vertical feature extraction network, and a depth feature extraction network. Each piece of motion sickness data in the motion sickness data sequence may represent a probability distribution of the target user experiencing motion sickness in the virtual scene. The motion sickness data may include symptom severity and probability data. The symptom severity may represent the degree of motion sickness experienced by the target user. Symptom severity may include, but is not limited to, no symptoms, mild symptoms, moderate symptoms, severe symptoms, and extremely severe symptoms. The probability data may represent the probability of occurrence of the corresponding symptom severity. For example, the motion sickness data sequence may be {symptom level: no symptoms, probability data: 10%; symptom level: mild symptoms, probability data: 10%; symptom level: moderate symptoms, probability data: 5%; symptom level: severe symptoms, probability data: 70%; symptom level: extremely severe symptoms, probability data: 5%}.
[0083] In some optional implementations of some embodiments, the motion sickness detection model may be trained by the processor 2 through the following steps:
[0084] The first step is to obtain a sample set. The samples in the sample set include a sample depth image dataset, a sample optical flow image dataset, and a sample motion sickness data sequence. Each sample depth image data in the sample depth image dataset can be depth image data used for model training. Each sample optical flow image data in the sample optical flow image dataset can be optical flow image data used for model training. Each sample motion sickness data in the sample motion sickness data sequence can be standard motion sickness data corresponding to the sample. Each sample motion sickness data in the sample motion sickness data sequence can include a sample symptom degree and sample probability data. The sample symptom degree can be the symptom degree corresponding to the sample motion sickness data. The sample probability data can represent the probability of occurrence of the corresponding sample symptom degree.
[0085] In the second step, based on the sample set, the following training steps are performed:
[0086] In a first sub-step, each sample optical flow image dataset included in at least one sample in the sample set is input into a horizontal feature extraction network included in the feature extraction layer of the initial neural network, thereby obtaining a horizontal feature map corresponding to each sample in the at least one sample. The horizontal feature extraction network may be a neural network that takes each sample optical flow image dataset as input and outputs a horizontal feature map corresponding to each sample optical flow image dataset. For example, the horizontal feature extraction network may be a convolutional neural network.
[0087] In a second sub-step, the sample optical flow image datasets included in the at least one sample are input into a vertical feature extraction network included in the feature extraction layer of the initial neural network to obtain a vertical feature map corresponding to each sample in the at least one sample. The vertical feature extraction network may be a neural network that takes the sample optical flow image datasets as input and outputs the vertical feature maps corresponding to the sample optical flow image datasets. For example, the vertical feature extraction network may be a convolutional neural network.
[0088] In a third sub-step, the depth image datasets of each sample included in the at least one sample are input into a depth feature extraction network included in the feature extraction layer of the initial neural network to obtain a depth feature map corresponding to each sample in the at least one sample. The depth feature extraction network may be a neural network that takes the depth image datasets of each sample as input and outputs the depth feature maps corresponding to the depth image datasets of each sample. For example, the depth feature extraction network may be a convolutional neural network.
[0089] In a fourth sub-step, the horizontal feature map, vertical feature map, and depth feature map corresponding to each sample in the at least one sample are input into the temporal feature extraction network in the initial neural network to obtain temporal feature information corresponding to each sample in the at least one sample. The temporal feature information may be the temporal features corresponding to the horizontal feature map, the vertical feature map, and the depth feature map. The temporal feature extraction network may be a neural network that takes the horizontal feature map, the vertical feature map, and the depth feature map as input and takes the temporal feature information as output. For example, the temporal feature extraction network may be a long short-term memory network.
[0090] In a fifth sub-step, the temporal feature information corresponding to each sample in the at least one sample is input into the output layer of the initial neural network to obtain a motion sickness data sequence corresponding to each sample in the at least one sample. The output layer may be an activation function that takes the temporal feature information as input and outputs the motion sickness data sequence. For example, the activation function may be a Softmax function.
[0091] A sixth sub-step is generating distribution loss data corresponding to each sample in the at least one sample based on the sample motion sickness data sequence and the motion sickness data sequence corresponding to each sample in the at least one sample. The distribution loss data may be a loss value between the sample motion sickness data sequence and the motion sickness data sequence.
[0092] The seventh sub-step is to generate transformation loss data corresponding to each sample in the at least one sample based on preset weight data, the sample motion sickness data sequence corresponding to each sample in the at least one sample, and the motion sickness data sequence. The weight data may be a preset value. The specific setting of the weight data is not limited here. The transformation loss data may be a change value between two adjacent sample motion sickness data in the sample motion sickness data sequence and a loss value between the change values of two adjacent motion sickness data in the motion sickness data sequence.
[0093] In an eighth sub-step, sample loss data is generated based on the distribution loss data and transformation loss data corresponding to each sample in the at least one sample. The sample loss data may be the loss value between each sample motion sickness data sequence and each motion sickness data sequence corresponding to the at least one sample. In practice, for each sample in the at least one sample, the processor 2 may determine the average of the distribution loss data and transformation loss data corresponding to the sample as the average loss data. The average of the average loss data corresponding to the at least one sample may then be determined as the sample loss data.
[0094] In a ninth sub-step, in response to determining that the sample loss data satisfies a preset loss condition, the initial neural network is determined to be a motion sickness detection model. The loss condition may be that the sample loss data is less than a preset loss value. The loss value may be a pre-set value. The specific setting of the loss value is not limited herein.
[0095] In a tenth sub-step, in response to determining that the sample loss data does not satisfy the loss condition, the network parameters of the initial neural network are adjusted, and unused samples are used to form a sample set, and the adjusted initial neural network is used as the initial neural network to perform the training step again. In practice, the processor 2 may adjust the network parameters of the initial neural network using a back propagation algorithm (BP algorithm) and a gradient descent method (e.g., a mini-batch gradient descent algorithm).
[0096] In some optional implementations of some embodiments, the processor 2 may generate distribution loss data corresponding to each sample in the at least one sample based on the sample motion sickness data sequence and the motion sickness data sequence corresponding to each sample in the at least one sample by the following steps:
[0097] For each of the at least one sample above, perform the following steps:
[0098] In the first step, the sample motion sickness data sequence corresponding to the above sample is determined as the target sample data sequence.
[0099] In the second step, for each motion sickness data in the motion sickness data sequence corresponding to the above sample, perform the following steps:
[0100] In a first sub-step, the logarithm of the probability data included in the motion sickness data is determined as motion sickness logarithm data. The logarithm may be a logarithm with base 2.
[0101] In the second sub-step, the target sample data corresponding to the motion sickness data in the target sample data sequence is determined as the sample data to be processed.
[0102] In a third sub-step, the product of the motion sickness logarithmic data and the sample probability data included in the sample data to be processed is determined as motion sickness loss data.
[0103] In the third step, the sum of the determined motion sickness loss data is determined as the distribution loss data corresponding to the above sample.
[0104] In some optional implementations of some embodiments, the processor 2 may generate transformation loss data corresponding to each sample in the at least one sample based on preset weight data, a sample motion sickness data sequence corresponding to each sample in the at least one sample, and a motion sickness data sequence through the following steps:
[0105] For each of the at least one sample above, perform the following steps:
[0106] In the first step, based on the sample motion sickness data sequence corresponding to the above sample, for every two adjacent sample motion sickness data in the above sample motion sickness data sequence, the difference between the two sample probability data corresponding to the above two sample motion sickness data is determined as the sample motion sickness difference data.
[0107] The second step is to generate a sample motion sickness difference data sequence based on the determined sample motion sickness difference data. In practice, the processor 2 may arrange the sample motion sickness difference data into a sample motion sickness difference data sequence in the order corresponding to the sample motion sickness data sequence. As an example, when the sample motion sickness data sequence is {sample motion sickness data 1, sample motion sickness data 2, sample motion sickness data 3}, and sample motion sickness data 1 and sample motion sickness data 2 correspond to sample motion sickness difference data 1, and sample motion sickness data 2 and sample motion sickness data 3 correspond to sample motion sickness difference data 2, the resulting sample motion sickness difference data sequence is {sample motion sickness difference data 1, sample motion sickness difference data 2}.
[0108] In the third step, based on the motion sickness data sequence corresponding to the sample, for every two adjacent motion sickness data in the motion sickness data sequence, the difference between the two probability data corresponding to the two motion sickness data is determined as motion sickness difference data.
[0109] The fourth step is to generate a motion sickness difference data sequence based on the determined motion sickness difference data. In practice, the processor 2 may arrange the motion sickness difference data into a motion sickness difference data sequence in the order corresponding to the motion sickness data sequence.
[0110] Step 5: For each sample motion sickness difference data in the above sample motion sickness difference data sequence, perform the following steps:
[0111] In a first sub-step, the motion sickness difference data corresponding to the sample motion sickness difference data in the motion sickness difference data sequence is determined as the target motion sickness difference data.
[0112] In the second sub-step, the difference between the target motion sickness difference data and the sample motion sickness difference data is determined as the difference data to be processed.
[0113] In the sixth step, the average of the determined difference data to be processed is determined as the average data to be processed.
[0114] In the seventh step, the product of the above average data to be processed and the above weight data is determined as the transformation loss data corresponding to the above sample.
[0115] Fifth, the horizontal feature map, vertical feature map, and depth feature map are input into the temporal feature extraction network in the motion sickness detection model to obtain temporal feature information.
[0116] In some embodiments, the processor 2 may input the horizontal feature map, the vertical feature map, and the depth feature map into a temporal feature extraction network in the motion sickness detection model to obtain temporal feature information.
[0117] Sixth, based on the temporal feature information, motion sickness detection results are generated.
[0118] In some embodiments, the processor 2 may generate a motion sickness detection result based on the time series feature information. The motion sickness detection result may be the degree of motion sickness experienced by the target user. In practice, the processor 2 may first input the time series feature information into the output layer of the motion sickness detection model to obtain a motion sickness data sequence corresponding to the time series feature information. Then, the motion sickness detection result may be determined as the symptom degree with the highest corresponding probability data among the various symptom degrees included in the motion sickness data sequence.
[0119] Seventh, in response to determining that the motion sickness detection result meets a preset detection result condition, determine information to be displayed.
[0120] In some embodiments, the processor 2 may determine information to be displayed in response to determining that the motion sickness detection result satisfies a preset detection result condition. The detection result condition may be that the motion sickness detection result is not asymptomatic. The information to be displayed may be information that needs to be displayed. In practice, the processor 2 may determine preset motion sickness relief information as the information to be displayed. The motion sickness relief information may be information that provides the user with measures to alleviate dizziness. For example, the motion sickness relief information may include "If you experience dizziness or other discomfort, you can reduce the rotation speed to alleviate dizziness."
[0121] In some embodiments, the memory 3 may be configured to store the motion sickness detection result.
[0122] In some embodiments, the display 4 may be configured to display the information to be displayed.
[0123] The head-mounted device for motion sickness detection disclosed herein has the following beneficial effects: It can reduce computing resource consumption during processing. Specifically, the high computing resource consumption is caused by the fact that methods for predicting motion sickness by detecting a user's electroencephalogram (EEG), galvanic skin response (GSR), respiratory rate, and other parameters require the user to wear a dedicated contact device, which can reduce the user's virtual reality experience. Furthermore, the detection process requires simultaneous processing of multiple data points, such as the user's EEG, GSR, and respiratory rate, which can lead to high computing resource consumption during processing. Therefore, the head-mounted device for motion sickness detection disclosed herein includes, first, an image acquisition device configured to acquire a sequence of visual image data corresponding to a target user. This allows for the acquisition of raw data to be processed. Second, a processor configured to perform the following processing: First, pre-processing the visual image data sequence to obtain a pre-processed image data sequence. This pre-processing reduces the interference of noise in the raw data with subsequent processing. Second, for each pre-processed image data in the pre-processed image data sequence, depth image data is generated based on the pre-processed image data. In this way, a depth map corresponding to the preprocessed image data sequence can be obtained. Then, for every two frames of preprocessed image data in the preprocessed image data sequence that meet preset image conditions, optical flow image data is generated based on the two frames of preprocessed image data. Thus, an optical flow field map corresponding to the preprocessed image data sequence can be obtained. Then, based on the generated depth image data, the generated optical flow image data, and the feature extraction layer of a pre-trained motion sickness detection model, a horizontal feature map, a vertical feature map, and a depth feature map are generated. The feature extraction layer includes a horizontal feature extraction network, a vertical feature extraction network, and a depth feature extraction network. Thus, each feature map corresponding to the preprocessed image data sequence can be obtained. The horizontal feature map, the vertical feature map, and the depth feature map are then input into the temporal feature extraction network of the motion sickness detection model to obtain temporal feature information. Thus, temporal features corresponding to the preprocessed image data sequence can be obtained. Then, based on the temporal feature information, a motion sickness detection result is generated. Thus, a motion sickness detection result for the user can be obtained using the temporal features. Then, in response to determining that the motion sickness detection result satisfies a preset detection result condition, information to be displayed is determined. Then, the memory is configured to store the motion sickness detection result. Thus, the motion sickness detection result can be stored. Finally, the display is configured to display the information to be displayed. This allows users with more severe symptoms to be provided with relief measures.Furthermore, because motion sickness can be predicted by acquiring and processing a user's visual image data sequence, without requiring the user to wear a dedicated contact device, the likelihood of a poor virtual experience resulting from wearing a dedicated contact device can be reduced. Furthermore, because motion sickness can be directly predicted using a pre-trained motion sickness detection model, without the need to simultaneously process multiple data points such as the user's electroencephalogram (EEG), galvanic skin response, and respiratory rate, the computational resource consumption during processing can be reduced.
[0124] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. Furthermore, in some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.
[0125] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0126] The above-mentioned computer-readable medium may be included in the above-mentioned head-mounted device for motion sickness information detection; or it may exist independently and not be assembled into the head-mounted device for motion sickness information detection. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the head-mounted device for motion sickness information detection, the head-mounted device for motion sickness information detection: an image acquisition device is configured to acquire a visual image data sequence corresponding to the target user; a processor is configured to perform the following processing: pre-processing the above-mentioned visual image data sequence to obtain a pre-processed image data sequence; for each pre-processed image data in the above-mentioned pre-processed image data sequence, generating depth image data based on the above-mentioned pre-processed image data; for every two frames of pre-processed image data in the above-mentioned pre-processed image data sequence that meet the preset image conditions, generating optical flow image data based on the above-mentioned two frames of pre-processed image data; based on the generated The invention relates to a method for detecting motion sickness, comprising: generating a horizontal feature map, a vertical feature map, and a depth feature map by using depth image data, each of the generated optical flow image data, and a feature extraction layer in a pre-trained motion sickness detection model, wherein the feature extraction layer includes a horizontal feature extraction network, a vertical feature extraction network, and a depth feature extraction network; inputting the horizontal feature map, the vertical feature map, and the depth feature map into a temporal feature extraction network in the motion sickness detection model to obtain temporal feature information; generating a motion sickness detection result based on the temporal feature information; determining information to be displayed in response to determining that the motion sickness detection result meets a preset detection result condition; a memory is configured to store the motion sickness detection result; and a display is configured to display the information to be displayed.
[0127] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0129] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0130] The above descriptions are merely some preferred embodiments of the present disclosure and illustrate the underlying technical principles. Those skilled in the art should understand that the scope of the invention encompassed by the embodiments of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned inventive concept. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A head-mounted device for motion sickness information detection, comprising: an image acquisition device configured to acquire a visual image data sequence corresponding to a target user; A processor configured to perform the following processing: Preprocessing the visual image data sequence to obtain a preprocessed image data sequence; For each pre-processed image data in the pre-processed image data sequence, generating depth image data based on the pre-processed image data; For every two frames of pre-processed image data that meet a preset image condition in the pre-processed image data sequence, generating optical flow image data based on the two frames of pre-processed image data; Generate a horizontal feature map, a vertical feature map, and a depth feature map based on the generated depth image data, the generated optical flow image data, and a feature extraction layer in a pre-trained motion sickness detection model, wherein the feature extraction layer includes a horizontal feature extraction network, a vertical feature extraction network, and a depth feature extraction network. The motion sickness detection model is trained by the following steps: Acquire a sample set, wherein the samples in the sample set include a sample depth image dataset, a sample optical flow image dataset, and a sample motion sickness data sequence, and each sample motion sickness data in the sample motion sickness data sequence includes a sample symptom degree and sample probability data; Based on the sample set, the following training steps are performed: Inputting each sample optical flow image data set included in at least one sample in the sample set into a horizontal feature extraction network included in the feature extraction layer of the initial neural network to obtain a horizontal feature map corresponding to each sample in the at least one sample; Inputting each sample optical flow image data set included in the at least one sample into the vertical feature extraction network included in the feature extraction layer of the initial neural network to obtain a vertical feature map corresponding to each sample in the at least one sample; Inputting each sample depth image data set included in the at least one sample into the depth feature extraction network included in the feature extraction layer of the initial neural network to obtain a depth feature map corresponding to each sample in the at least one sample; Inputting the horizontal feature map, the vertical feature map, and the depth feature map corresponding to each sample in the at least one sample into the temporal feature extraction network in the initial neural network to obtain temporal feature information corresponding to each sample in the at least one sample; Inputting the time series feature information corresponding to each sample in the at least one sample into the output layer of the initial neural network to obtain a motion sickness data sequence corresponding to each sample in the at least one sample, wherein each motion sickness data in the motion sickness data sequence includes symptom severity and probability data; generating distribution loss data corresponding to each sample in the at least one sample based on the sample motion sickness data sequence and the motion sickness data sequence corresponding to each sample in the at least one sample; generating, based on preset weight data, a sample motion sickness data sequence corresponding to each sample in the at least one sample, and a motion sickness data sequence, transformation loss data corresponding to each sample in the at least one sample; generating sample loss data based on the distribution loss data and the transformation loss data corresponding to each sample in the at least one sample; In response to determining that the sample loss data satisfies a preset loss condition, determining the initial neural network as a motion sickness detection model; In response to determining that the sample loss data does not satisfy the loss condition, adjusting network parameters of the initial neural network, using unused samples to form a sample set, using the adjusted initial neural network as the initial neural network, and performing the training step again; Inputting the horizontal feature map, the vertical feature map, and the depth feature map into a temporal feature extraction network in the motion sickness detection model to obtain temporal feature information; generating a motion sickness detection result based on the time series feature information; In response to determining that the motion sickness detection result meets a preset detection result condition, determining information to be displayed; a memory configured to store the motion sickness detection result; The display is configured to display the information to be displayed.
2. The head-mounted device for motion sickness information detection according to claim 1, wherein: The processor is further configured to generate depth image data based on the pre-processed image data by the following steps, including: Performing feature extraction processing on the preprocessed image data to obtain a preprocessed feature map group; For each preprocessing feature map in the preprocessing feature map group, perform the following steps: Downsampling the preset random noise map to obtain a downsampled noise map corresponding to the preprocessed feature map; Performing image stitching processing on the downsampled noise map and the preprocessed feature map to obtain a noise stitching feature map; Determine the obtained noise splicing feature maps as a noise splicing feature map group; Based on the noise splicing feature map group, depth image data is generated.
3. The head-mounted device for motion sickness information detection according to claim 1, wherein: The processor is further configured to generate optical flow image data based on the two frames of pre-processed image data through the following steps, including: Performing image conversion processing on the two frames of pre-processed image data to obtain two frames of visual grayscale image data; Determining the visual grayscale image data that meets a preset image frame condition in the two frames of visual grayscale image data as the first grayscale image data; determining the visual grayscale image data that does not meet the image frame condition in the two frames of visual grayscale image data as second grayscale image data; For each pixel point included in the first grayscale image data, perform the following steps: Determine the pixel point as a first pixel point; determining a first pixel matrix corresponding to the first pixel based on the first grayscale image data and the first pixel; Generate a first horizontal spatial gradient and a first vertical spatial gradient based on a preset horizontal operator matrix, a preset vertical operator matrix, and the first pixel matrix; determining a pixel point corresponding to the first pixel point among the pixels included in the second grayscale image data as a second pixel point; determining a difference between a pixel value corresponding to the first pixel point and a pixel value corresponding to the second pixel point as pixel difference data; constructing optical flow nonlinear information based on a preset custom horizontal vector, a preset custom vertical vector, the first horizontal spatial gradient, the first vertical spatial gradient, and the pixel difference data; determining each pixel point corresponding to the first pixel point among each pixel point included in the first grayscale image data as each target pixel point; Based on the target pixels, determining the target optical flow nonlinear information corresponding to the target pixels; generating optical flow horizontal data and optical flow vertical data corresponding to the first pixel point based on the optical flow nonlinear information and the optical flow nonlinear information of each target; Generate an optical flow level matrix based on the generated optical flow level data; Generate an optical flow vertical matrix based on the generated optical flow vertical data; Optical flow image data is generated based on the optical flow horizontal matrix and the optical flow vertical matrix.
4. The head-mounted device for motion sickness information detection according to claim 1, wherein: The processor is further configured to generate distribution loss data corresponding to each sample in the at least one sample based on the sample motion sickness data sequence and the motion sickness data sequence corresponding to each sample in the at least one sample through the following steps, including: For each sample of the at least one sample, performing the following steps: determining a sample motion sickness data sequence corresponding to the sample as a target sample data sequence; For each motion sickness data in the motion sickness data sequence corresponding to the sample, perform the following steps: determining the logarithm of the probability data included in the motion sickness data as motion sickness logarithmic data; determining the target sample data corresponding to the motion sickness data in the target sample data sequence as sample data to be processed; determining a product of the motion sickness logarithm data and the sample probability data included in the sample data to be processed as motion sickness loss data; The sum of the determined motion sickness loss data is determined as the distribution loss data corresponding to the sample.
5. The head-mounted device for motion sickness information detection according to claim 1, wherein: The processor is further configured to generate transformation loss data corresponding to each sample in the at least one sample based on preset weight data, a sample motion sickness data sequence corresponding to each sample in the at least one sample, and a motion sickness data sequence through the following steps, including: For each sample of the at least one sample, performing the following steps: Based on a sample motion sickness data sequence corresponding to the sample, for every two adjacent sample motion sickness data in the sample motion sickness data sequence, determining a difference between two sample probability data corresponding to the two sample motion sickness data as sample motion sickness difference data; generating a sample motion sickness difference data sequence based on the determined sample motion sickness difference data; Based on the motion sickness data sequence corresponding to the sample, for each two adjacent motion sickness data in the motion sickness data sequence, determining a difference between two probability data corresponding to the two motion sickness data as motion sickness difference data; generating a motion sickness difference data sequence based on the determined motion sickness difference data; For each sample motion sickness difference data in the sample motion sickness difference data sequence, performing the following steps: determining the motion sickness difference data corresponding to the sample motion sickness difference data in the motion sickness difference data sequence as target motion sickness difference data; determining the difference between the target motion sickness difference data and the sample motion sickness difference data as difference data to be processed; Determining the average of each of the determined difference data to be processed as the average data to be processed; The product of the to-be-processed average data and the weight data is determined as the transformation loss data corresponding to the sample.
6. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method performed by the head-mounted device for motion sickness information detection as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Visually induced motion sickness assessment method based on attention mechanism
CN118298352A