Vestibular illusion training type verification method
By performing vestibular illusion type training on preset personnel, using the type recognition model to verify the accuracy of the training data, the problem of inaccurate vestibular illusion training data in the prior art is solved, and the accuracy of training is improved.
Patent Information
- Application Number
- CN202411490995.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2044-10-24
AI Technical Summary
The existing vestibular illusion training data are inaccurate, resulting in insufficient accuracy of vestibular illusion type training.
By obtaining the preset vestibular illusion training type and corresponding training data, the preset personnel are trained for vestibular illusion type, obtain motion parameters, and generate target data through data processing. The target data is identified using the preset type recognition model, the target vestibular illusion recognition type is determined, and compared with the preset vestibular illusion training type to verify the accuracy of the training data.
The accuracy of the training data corresponding to the vestibular illusion training type is improved, thereby improving the accuracy of the vestibular illusion training type.
Smart Images

Figure CN119274746B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vestibular illusion training, and particularly to a method for verifying vestibular illusion training types. Background Art
[0002] Spatial disorientation is one of the problems that almost all pilots have encountered during flight. Among them, vestibular illusions are caused by the physiological limitations of the human vestibular system and are also one of the types of spatial disorientation that pilots find difficult to cope with.
[0003] In existing vestibular illusion training, training data corresponding to each vestibular illusion training type is usually generated based on historical training data or historical experience. However, the generated training data corresponding to the vestibular illusion training type is not accurate, which may lead to inaccurate results when training preset personnel for different vestibular illusion types.
[0004] Therefore, how to improve the accuracy of the training data corresponding to the vestibular illusion training type and thus improve the accuracy of the vestibular illusion type training has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, the present invention provides a method for verifying vestibular illusion training types to solve the problem of how to improve the accuracy of the training data corresponding to the vestibular illusion training type and thus improve the accuracy of the vestibular illusion type training.
[0006] In a first aspect, the present invention provides a method for verifying vestibular illusion training types, the method comprising:
[0007] Obtain a preset vestibular illusion training type and training data corresponding to the preset vestibular illusion training type;
[0008] Conduct vestibular illusion type training on a preset person based on the training data and obtain the motion parameters corresponding to the preset person;
[0009] Perform data processing on the motion parameters to generate target data;
[0010] Based on a preset type recognition model, identify the target data to determine the target vestibular illusion recognition type corresponding to the target data;
[0011] Compare the target vestibular illusion recognition type with the preset vestibular illusion training type;
[0012] Verify the training data according to the comparison result.
[0013] The vestibular illusion training type verification method provided by the embodiments of the present application obtains a preset vestibular illusion training type and training data corresponding to the preset vestibular illusion training type; trains the preset personnel on the vestibular illusion type based on the training data, and obtains the motion parameters corresponding to the preset personnel, so as to ensure the accuracy of the obtained motion parameters. Data processing is performed on the motion parameters to generate target data, ensuring the accuracy of the generated target data. Based on a preset type recognition model, the target data is recognized to determine the target vestibular illusion recognition type corresponding to the target data, ensuring the accuracy of the determined target vestibular illusion recognition type corresponding to the target data. The target vestibular illusion recognition type is compared with the preset vestibular illusion training type, ensuring the accuracy of the generated comparison result. According to the comparison result, the training data is verified, ensuring the accuracy of the verification of the training data. Thus, the accuracy of the training data corresponding to the vestibular illusion training type can be ensured, and further the accuracy of the vestibular illusion type training can be improved.
[0014] In an alternative embodiment, the target data is a target color image; data processing is performed on the motion parameters to generate the target data, including:
[0015] Performing image conversion processing on the motion parameters to generate a grayscale image corresponding to the motion parameters;
[0016] Generating a target color image based on the grayscale image.
[0017] The vestibular illusion training type verification method provided by the embodiments of the present application performs image conversion processing on the motion parameters to generate a grayscale image corresponding to the motion parameters, ensuring the accuracy of the generated grayscale image corresponding to the motion parameters. Generating a target color image based on the grayscale image ensures the accuracy of the generated target color image, so that the accuracy of the target vestibular illusion recognition type corresponding to the target color image determined by recognizing the target color image based on a preset type recognition model can be ensured.
[0018] In an alternative embodiment, the motion parameters include a head motion parameter sequence and a body motion parameter sequence, and the grayscale image includes a head grayscale image and a body grayscale image; performing image conversion processing on the motion parameters to generate a grayscale image corresponding to the motion parameters includes:
[0019] Performing data conversion on the head motion parameter sequence to generate a head grayscale image;
[0020] Performing data conversion on the body motion parameter sequence to generate a body grayscale image.
[0021] The vestibular illusion training type verification method provided by the embodiments of the present application performs data conversion on the head movement parameter sequence to generate a head grayscale image, ensuring the accuracy of the generated head grayscale image. It performs data conversion on the body movement parameter sequence to generate a body grayscale image, ensuring the accuracy of the generated body grayscale image.
[0022] In an alternative embodiment, performing data conversion on the head movement parameter sequence to generate a head grayscale image includes:
[0023] Splitting the head movement parameters to generate multiple non - overlapping head movement parameter subsequences;
[0024] Calculating the correlation integral between the head movement parameter subsequences according to the relationship between the head movement parameter subsequences;
[0025] Calculating the autocorrelation characteristics between the head movement parameter subsequences according to the correlation integral between the head movement parameter subsequences;
[0026] Determining the delay time and embedding dimension corresponding to the head movement parameter sequence according to the autocorrelation characteristics between the head movement parameter subsequences;
[0027] Generating a head phase space corresponding to the head movement parameter sequence based on the delay time and embedding dimension;
[0028] Generating a head grayscale image according to the distances between the head vectors included in the head phase space.
[0029] The vestibular illusion training type verification method provided by the embodiments of the present application ensures the accuracy of the calculated correlation integral and autocorrelation characteristics between the head movement parameter subsequences. Determining the delay time and embedding dimension corresponding to the head movement parameter sequence according to the autocorrelation characteristics between the head movement parameter subsequences. Generating a head phase space corresponding to the head movement parameter sequence based on the delay time and embedding dimension. Generating a head grayscale image according to the distances between the head vectors included in the head phase space, ensuring the accuracy of the generated head grayscale image.
[0030] In an alternative embodiment, the head movement parameter sequence respectively includes the linear acceleration movement parameter sequences of the head on the X, Y, and Z axes and the angular acceleration movement parameter sequences of the head on the X, Y, and Z axes; the head grayscale image includes the linear acceleration grayscale image of the head on the X axis, the linear acceleration grayscale image of the head on the Y axis, the linear acceleration grayscale image of the head on the Z axis, the angular acceleration grayscale image of the head on the X axis, the angular acceleration grayscale image of the head on the Y axis, and the angular acceleration grayscale image of the head on the Z axis.
[0031] In an alternative embodiment, the grayscale image includes a head grayscale image and a body grayscale image. Based on the grayscale image, generating a target color image includes:
[0032] Generating a head color image based on the head grayscale image;
[0033] Generating a body color image based on the body grayscale image;
[0034] Stitching the head color image and the body color image together to generate a target color image.
[0035] The vestibular illusion training type verification method provided by the embodiments of the present application generates a head color image based on the head grayscale image, ensuring the accuracy of the generated head color image. Generating a body color image based on the body grayscale image, ensuring the accuracy of the generated body color image. Stitching the head color image and the body color image together to generate a target color image, ensuring the accuracy of the generated target color image.
[0036] In an alternative embodiment, the head grayscale image includes the head's acceleration grayscale image on the X-axis, the head's acceleration grayscale image on the Y-axis, the head's acceleration grayscale image on the Z-axis, the head's angular acceleration grayscale image on the X-axis, the head's angular acceleration grayscale image on the Y-axis, and the head's angular acceleration grayscale image on the Z-axis. The head color image includes the head's linear acceleration color image and the head's angular acceleration color image. Generating a head color image based on the head grayscale image includes:
[0037] Determining a first red channel from the head's acceleration grayscale image on the X-axis, the head's acceleration grayscale image on the Y-axis, and the head's acceleration grayscale image on the Z-axis, determining a first green channel from the two grayscale images other than the first red channel, and determining the last grayscale image as the first blue channel;
[0038] Generating a head linear acceleration color image according to the determined first red channel, first green channel, and first blue channel;
[0039] Determining a second red channel from the head's angular acceleration grayscale image on the X-axis, the head's angular acceleration grayscale image on the Y-axis, and the head's angular acceleration grayscale image on the Z-axis, determining a second green channel from the two grayscale images other than the second red channel, and determining the last grayscale image as the second blue channel;
[0040] Generating a head angular acceleration color image according to the determined second red channel, second green channel, and second blue channel.
[0041] The vestibular illusion training type verification method provided by the embodiments of the present application ensures that the generated head linear acceleration color image can represent the linear acceleration motion parameter sequences of the head on the X, Y, and Z axes. It also ensures that the generated head angular acceleration color image can represent the angular acceleration motion parameter sequences of the head on the X, Y, and Z axes.
[0042] In an alternative embodiment, based on a preset type recognition model, the target data is recognized to determine the target vestibular illusion recognition type corresponding to the target data, including:
[0043] Input the target color image into the preset type recognition model;
[0044] The preset type recognition model performs feature recognition and feature extraction on the target color image, and based on the extracted features, outputs the target vestibular illusion recognition type.
[0045] The vestibular illusion training type verification method provided by the embodiments of the present application inputs the target color image into the preset type recognition model; the preset type recognition model performs feature recognition and feature extraction on the target color image, and based on the extracted features, outputs the target vestibular illusion recognition type, ensuring the accuracy of the output target vestibular illusion recognition type.
[0046] In an alternative embodiment, the preset type recognition model includes a partitioning layer, a Swin Transformer layer, a pooling layer, a fully connected layer, and a Softmax layer. The preset type recognition model performs feature recognition and feature extraction on the target color image, and based on the extracted features, outputs the target vestibular illusion recognition type, including:
[0047] The partitioning layer partitions the target color image to generate multiple target color image blocks and generates the corresponding marker information for each target color image block;
[0048] The Swin Transformer layer fuses the feature information corresponding to each target color image block to generate a fused feature, performs a linear transformation on the channel data of each pixel in the fused feature to generate a multi-dimensional feature; and performs feature extraction on the multi-dimensional feature to generate a target feature;
[0049] The pooling layer compresses and extracts the target feature to generate a local feature;
[0050] The fully connected layer integrates the local features to generate a global feature and maps the global feature to the output category;
[0051] The Softmax layer converts the output category output by the fully connected layer into a probability distribution and determines the target vestibular illusion recognition type based on the probability distribution.
[0052] The vestibular illusion training type verification method provided by the embodiment of the present application divides the target color image into layers, generates multiple target color image blocks, and generates the corresponding marking information for each target color image block; the Swin Transformer layer fuses the feature information corresponding to each target color image block, generates a fused feature, and performs a linear transformation on the channel data of each pixel in the fused feature to generate a multi-dimensional feature; performs feature extraction on the multi-dimensional feature to generate a target feature, ensuring the accuracy of the generated target feature. The pooling layer compresses and extracts the target feature to generate a local feature, ensuring the accuracy of the generated local feature. The fully connected layer integrates the local features to generate a global feature and maps the global feature to the output category; the Softmax layer converts the output category output by the fully connected layer into a probability distribution and determines the target vestibular illusion recognition type based on the probability distribution, ensuring the accuracy of the determined target vestibular illusion recognition type.
[0053] In an alternative embodiment, the Swin Transformer layer includes four stages. The first stage includes a first normalization layer and a window multi-head self-attention layer. The second stage includes a second normalization layer and a first multi-layer perceptron layer. The third stage includes a third normalization layer and a multi-head self-attention layer based on shifted windows. The fourth stage includes a fourth normalization layer and a second multi-layer perceptron layer; the Swin Transformer layer fuses the feature information corresponding to each target color image block, generates a fused feature, and performs a linear transformation on the channel data of each pixel in the fused feature to generate a multi-dimensional feature; performs feature extraction on the multi-dimensional feature to generate a target feature, including:
[0054] The first normalization layer normalizes the fused feature to generate a first normalized feature;
[0055] The window multi-head self-attention layer performs self-attention calculation on the first normalized feature and performs residual calculation to generate a first attention feature;
[0056] The second normalization layer normalizes the first attention feature to generate a second normalized feature;
[0057] The first multi-layer perceptron layer performs a linear transformation on the channel data of each pixel of the second normalized feature to generate a multi-dimensional feature;
[0058] The third normalization layer normalizes the multi-dimensional feature to generate a third normalized feature;
[0059] The multi-head self-attention layer based on shifted windows performs self-attention calculation on the third normalized feature to generate a second attention feature;
[0060] The fourth normalization layer normalizes the second attention feature to generate a fourth normalized feature;
[0061] The second multi-layer perceptron layer performs a linear transformation on the channel data of each pixel of the fourth normalized feature to generate a target feature.
[0062] In the vestibular illusion training type verification method provided by the embodiments of the present application, the first normalization layer normalizes the fused feature to generate a first normalized feature, making the preset type recognition model more robust to input scale changes. The window multi-head self-attention layer performs self-attention calculation on the first normalized feature and performs residual calculation to generate a first attention feature, so that the multi-head mechanism can be used to simultaneously focus on multiple different representation subspaces, improving the learning ability and generalization performance of the preset type recognition model. The second normalization layer normalizes the first attention feature to generate a second normalized feature; the first multi-layer perceptron layer performs a linear transformation on the channel data of each pixel of the second normalized feature to generate a multi-dimensional feature, increasing the expression ability of the preset type recognition model. The third normalization layer normalizes the multi-dimensional feature to generate a third normalized feature; the multi-head self-attention layer based on shifted windows performs self-attention calculation on the third normalized feature to generate a second attention feature, which helps to better model long-range dependencies and further improve the feature representation ability of the preset type recognition model. The fourth normalization layer normalizes the second attention feature to generate a fourth normalized feature; the second multi-layer perceptron layer performs a linear transformation on the channel data of each pixel of the fourth normalized feature to generate a target feature, ensuring the accuracy of the generated target feature. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0064] Figure 1 is a flowchart of a vestibular illusion training type verification method according to an embodiment of the present invention;
[0065] Figure 2 is a flowchart of another vestibular illusion training type verification method according to an embodiment of the present invention;
[0066] Figure 3 is a flowchart of a method for converting motion parameters to generate a target color image according to an embodiment of the present invention;
[0067] Figure 4 It is a schematic flowchart of another method for verifying the type of vestibular illusion training according to an embodiment of the present invention;
[0068] Figure 5 It is a schematic flowchart of identifying a target color image according to an embodiment of the present invention to determine the type of target vestibular illusion recognition;
[0069] Figure 6 It is a schematic diagram of the Swin Transformer network structure according to an embodiment of the present invention. Detailed implementation manners
[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0071] It should be noted that for the method for verifying the type of vestibular illusion training provided in the embodiments of the present application, the execution subject may be a device for verifying the type of vestibular illusion training. The device for verifying the type of vestibular illusion training may be implemented as part or all of an electronic device through software, hardware, or a combination of software and hardware. Among them, the electronic device may be a server or a terminal. Among them, the server in the embodiments of the present application may be a single server or a server cluster composed of multiple servers. The terminal in the embodiments of the present application may be other intelligent hardware devices such as a smart phone, a personal computer, a tablet computer, a wearable device, and a smart robot. In the following method embodiments, the execution subject is taken as an electronic device for illustration.
[0072] According to an embodiment of the present invention, an embodiment of a method for verifying the type of vestibular illusion training is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0073] In this embodiment, a method for verifying the type of vestibular illusion training is provided, which can be used for the above-mentioned electronic device. Figure 1 It is a flowchart of the method for verifying the type of vestibular illusion training according to an embodiment of the present invention. As Figure 1 shown, the process includes the following steps:
[0074] Step S101, obtain a preset type of vestibular illusion training and training data corresponding to the preset type of vestibular illusion training.
[0075] Specifically, the electronic device can receive the preset vestibular illusion training type input by the user, and then determine the training data corresponding to the preset vestibular illusion training type according to the corresponding relationship between the preset vestibular illusion training type and the training data.
[0076] Among them, the preset vestibular illusion training type is the name of different types of illusions, such as Coriolis-backward roll illusion, Coriolis-forward roll illusion, Coriolis-right roll illusion, somatic gravity illusion, somatic rotation illusion, etc.
[0077] Among them, different preset vestibular illusion training types correspond to different training data, such as setting the rotation axis, rotation direction, rotation speed, movement time, linear movement mode, etc.
[0078] Exemplarily, taking the triggering of the Coriolis-backward roll illusion as an example, the Coriolis illusion-backward roll illusion means that when a preset person rotates uniformly clockwise around the body zb axis, if they suddenly rotate around the body xb axis at the same time, the preset person will have the feeling that the body is rotating backward around the body yb axis. Therefore, the training data fields include {(axis), (turning), (rotation speed), (duration), (starting rotation moment) (axis), (turning), (rotation speed), (duration)}. For example, the corresponding training data includes {(zb), (clockwise), (60° / s), (60s), (30s) (xb), (clockwise), (45° / s), (1s)}.
[0079] Step S102, perform vestibular illusion type training on the preset person based on the training data, and obtain the motion parameters corresponding to the preset person.
[0080] Specifically, the electronic device can drive a six-degree-of-freedom motion platform to perform linear motion and rotational motion based on the training data to conduct vestibular illusion type training on the preset person. The electronic device can obtain the motion parameters corresponding to the preset person based on the motion sensors worn on the preset person.
[0081] Among them, the six-degree-of-freedom motion platform can rotate and perform linear motion along the three axes xp, yp, and zp of the six-degree-of-freedom motion platform coordinate system, and can provide complex six-degree-of-freedom comprehensive motion acceleration stimuli. It is stipulated that the origin Op of the six-degree-of-freedom motion platform coordinate system is located at the center of mass of the six-degree-of-freedom motion platform. The xp axis of the six-degree-of-freedom motion platform coordinate system is within the symmetry plane of the six-degree-of-freedom motion platform, with the forward direction being positive; the yp axis of the six-degree-of-freedom motion platform coordinate system is perpendicular to the symmetry plane of the six-degree-of-freedom motion platform where the xp axis is located, with the direction pointing to the right being positive; the zp axis of the six-degree-of-freedom motion platform coordinate system is perpendicular to the xpOpyp plane, and the positive direction of the zp axis is determined as downward according to the right-hand rule.
[0082] Step S103, perform data processing on the motion parameters to generate target data.
[0083] Specifically, the electronic device can check the integrity and accuracy of the motion parameters, and process the missing values, outliers, and error data in the motion parameters. Then, perform standardization or normalization processing on the motion parameters to make different motion parameters comparable. Extract meaningful features from the original motion parameters. These features may include derivative parameters such as calculated speed, acceleration, and angular velocity. Then, use statistical methods or data mining techniques to analyze the motion parameters. For example, calculate statistics such as mean, variance, and correlation to understand the distribution and characteristics of the data. Then, the electronic device performs data fusion on the motion parameter data from multiple sources, integrating the multi-source motion parameters into a more comprehensive data set, thereby generating target data.
[0084] This step will be introduced in detail below.
[0085] Step S104, based on a preset type recognition model, identify the target data to determine the target vestibular illusion recognition type corresponding to the target data.
[0086] Specifically, the electronic device can input the target data into the preset type recognition model. The preset type recognition model can identify the target data and extract features, and based on the extracted features, determine the target vestibular illusion recognition type corresponding to the target data.
[0087] Among them, the preset type recognition model can be any one of a radial basis function (RBF) network, a feedforward neural network (FFNN), a convolutional neural network (Convolutional neural networks, CNN), a deconvolutional network (Deconvolutional networks, DN), a deep convolutional inverse graphics network (Deep convolutional inverse graphics networks, DCIGN), a generative adversarial network (Generative adversarial networks, GAN), a recurrent neural network (Recurrent neural networks, RNN), a long short-term memory network (Long / short term memory, LSTM), a deep residual network (Deep residual networks, DRN), and an extreme learning machine (Extreme learning machines, ELM). The embodiments of the present application do not limit the preset type recognition model.
[0088] Step S105, compare the target vestibular illusion recognition type with the preset vestibular illusion training type.
[0089] Optionally, the electronic device may compare the target vestibular illusion recognition type with the preset vestibular illusion training type.
[0090] Optionally, the electronic device may also calculate the similarity between the target vestibular illusion recognition type and the preset vestibular illusion training type.
[0091] Step S106, verify the training data according to the comparison result.
[0092] Optionally, if the target vestibular illusion recognition type is consistent with the preset vestibular illusion training type, the electronic device determines that the correspondence between the training data and the preset vestibular illusion training type is accurate, thereby ensuring the accuracy of the training data in the training database and the preset vestibular illusion training type.
[0093] If the target vestibular illusion recognition type is inconsistent with the preset vestibular illusion training type, the electronic device may determine that the correspondence between the training data and the preset vestibular illusion training type is inaccurate. Then, the electronic device may verify the training data to detect whether the training data corresponding to the obtained preset vestibular illusion training type is accurate.
[0094] Specifically, the electronic device may, based on the preset vestibular illusion training type, search in the database for the training data corresponding to the preset vestibular illusion training type, and then compare the training data found in the database with the training data corresponding to the received preset vestibular illusion training type. If the two are inconsistent, it is determined that the received training data is incorrect. Correct the received training data based on the found training data.
[0095] If the found training data is consistent with the training data corresponding to the received preset vestibular illusion training type, it is determined that the training data corresponding to the preset vestibular illusion training type in the database is inaccurate, and then the electronic device may adjust the training data. Optionally, the electronic device may adjust parameters such as the axis and steering included in the test data.
[0096] While ensuring the accuracy of the adjusted training data, save the training data corresponding to the preset vestibular illusion training type to the database, thereby ensuring the accuracy of the training data corresponding to the vestibular illusion training type in the database.
[0097] The vestibular illusion training type verification method provided by the embodiment of the present application obtains a preset vestibular illusion training type and training data corresponding to the preset vestibular illusion training type; performs vestibular illusion type training on a preset person based on the training data, and obtains the motion parameters corresponding to the preset person, so as to ensure the accuracy of the obtained motion parameters. Perform data processing on the motion parameters to generate target data, ensuring the accuracy of the generated target data. Based on a preset type recognition model, identify the target data to determine the target vestibular illusion recognition type corresponding to the target data, ensuring the accuracy of the determined target vestibular illusion recognition type corresponding to the target data. Compare the target vestibular illusion recognition type with the preset vestibular illusion training type, ensuring the accuracy of the generated comparison result. According to the comparison result, verify the training data, ensuring the accuracy of the verification of the training data. Thus, the accuracy of the training data corresponding to the vestibular illusion training type can be ensured, and further the accuracy of the vestibular illusion type training can be improved.
[0098] In this embodiment, a vestibular illusion training type verification method is provided, which can be used in the above-mentioned electronic device. Figure 2 It is a flowchart of the vestibular illusion training type verification method according to an embodiment of the present invention, as Figure 2 shown, and the process includes the following steps:
[0099] Step S201, obtain a preset vestibular illusion training type and training data corresponding to the preset vestibular illusion training type.
[0100] For details, please refer to step S101 of the above embodiment, which will not be elaborated here.
[0101] Step S202, perform vestibular illusion type training on a preset person based on the training data, and obtain the motion parameters corresponding to the preset person.
[0102] For details, please refer to step S102 of the above embodiment, which will not be elaborated here.
[0103] Step S203, perform data processing on the motion parameters to generate target data.
[0104] Specifically, the target data is a target color image; the above step S203 includes:
[0105] Step S2031, perform image conversion processing on the motion parameters to generate a grayscale image corresponding to the motion parameters.
[0106] In an alternative embodiment of the present application, before identifying the motion parameters, the electronic device can unify the sampling frequency of the motion parameters, then perform noise reduction processing, and finally normalize the data using the minimum-maximum method.
[0107] In an alternative embodiment of the present application, the motion parameters include a head motion parameter sequence and a body motion parameter sequence. The motion sensors include a head-mounted motion sensor and a body motion sensor. The head-mounted motion sensor can collect the head motion parameter sequence, and the body motion sensor can collect the body motion parameter sequence. The grayscale images include a head grayscale image and a body grayscale image; the above step S2031 includes:
[0108] Step a1, perform data conversion on the head motion parameter sequence to generate a head grayscale image.
[0109] Specifically, the electronic device can use phase space reconstruction and recurrence plot techniques to perform data conversion on the head motion parameter sequence to generate a head grayscale image, and perform data conversion on the body motion parameter sequence to generate a body grayscale image.
[0110] Specifically, the above step a1 may include the following steps:
[0111] Step a11, split the head motion parameters to generate a plurality of non-overlapping head motion parameter subsequences.
[0112] Step a12, calculate the correlation integral between the head motion parameter subsequences according to the relationship between the head motion parameter subsequences.
[0113] Step a13, calculate the autocorrelation characteristics between the head motion parameter subsequences according to the correlation integral between the head motion parameter subsequences.
[0114] Step a14, determine the delay time and embedding dimension corresponding to the head motion parameter sequence according to the autocorrelation characteristics between the head motion parameter subsequences.
[0115] Step a15, generate a head phase space corresponding to the head motion parameter sequence based on the delay time and embedding dimension.
[0116] Step a16, generate a head grayscale image according to the distances between the head vectors included in the head phase space.
[0117] Specifically, the two key parameters of phase space reconstruction are the delay time and the embedding dimension. For a given time series , the given time series A can represent the head motion parameter sequence, that is, a time series data of a segment of motion parameters.
[0118] In an alternative embodiment, the head motion parameter sequences respectively include the linear acceleration motion parameter sequences of the head in the X, Y, and Z axes and the angular acceleration motion parameter sequences of the head in the X, Y, and Z axes; the head grayscale images include the linear acceleration grayscale image of the head in the X axis, the linear acceleration grayscale image of the head in the Y axis, the linear acceleration grayscale image of the head in the Z axis, the angular acceleration grayscale image of the head in the X axis, the angular acceleration grayscale image of the head in the Y axis, and the angular acceleration grayscale image of the head in the Z axis.
[0119] Specifically, the head-mounted motion sensor is used to detect the head motion of the trainee and record the linear accelerations (hx, hy, hz) and angular accelerations (hrx, hry, hrz) of the trainee's head in three directions. The linear accelerations (hx, hy, hz) and angular accelerations (hrx, hry, hrz) of the head in three directions are relative to the head coordinate system. It is stipulated that the origin Oh of the head coordinate system is located at the center of mass of the head, the xh axis of the head coordinate system is within the head symmetry plane, with the nose direction being positive; the yh axis of the head coordinate system is perpendicular to the head symmetry plane where the xh axis is located, with the right direction being positive; the zh axis of the head coordinate system is perpendicular to the xhOhyh plane, and the positive direction of the zh axis is determined to be downward according to the right-hand rule.
[0120] The body motion sensor is used to detect the motion parameters of the motion platform and record the linear accelerations (bx, by, bz) and three-direction angular positions (bRx, bRy, bRz) of the trainee's body in three directions. The linear accelerations (bx, by, bz) and three-direction angular positions (bRx, bRy, bRz) of the body in three directions are relative to the body coordinate system. It is stipulated that the origin Ob of the body coordinate system is located at the center of mass of the head, the xb axis of the body coordinate system is within the body symmetry plane, with the chest direction being positive; the yb axis of the body coordinate system is perpendicular to the body symmetry plane where the xb axis is located, with the right direction being positive; the zb axis of the body coordinate system is perpendicular to the xbObyb plane, and the positive direction of the zb axis is determined to be downward according to the right-hand rule.
[0121] The electronic device can perform data conversion on the linear acceleration motion parameter sequence of the head in the X axis to generate the linear acceleration grayscale image of the head in the X axis; perform data conversion on the linear acceleration motion parameter sequence of the head in the Y axis to generate the linear acceleration grayscale image of the head in the Y axis; perform data conversion on the linear acceleration motion parameter sequence of the head in the Z axis to generate the linear acceleration grayscale image of the head in the Z axis; perform data conversion on the angular acceleration motion parameter sequence of the head in the X axis to generate the angular acceleration grayscale image of the head in the X axis; perform data conversion on the angular acceleration motion parameter sequence of the head in the Y axis to generate the angular acceleration grayscale image of the head in the Y axis; perform data conversion on the angular acceleration motion parameter sequence of the head in the Z axis to generate the angular acceleration grayscale image of the head in the Z axis.
[0122] Specifically, the phase space reconstruction can be accomplished using phase space B, as shown in Equation (1).
[0123] (1)
[0124] where p is the embedding dimension, τ is the delay time, is the number of phase points in the phase space. By appropriately choosing the embedding dimension and the delay time, the reconstructed phase space becomes a diffeomorphism of the original system.
[0125] Among them, the electronic device can use the C-C method to calculate the delay time τ and the embedding dimension p.
[0126] Specifically, the process is as follows: First, split the time series A into τ non-overlapping subsequences, each with a length of . Among them, the value of τ can start from 1, and the maximum value is the smaller of 40 and (N2 / 16 - 1), where N is the number of data points in the time series A. As shown in Equation (2).
[0127] (2)
[0128] The sampling interval of the time series A is , is the optimal delay time of the time series, is the window of the delay time, which can be calculated according to Equation (3).
[0129] (3)
[0130] The correlation integral is a cumulative distribution function, which represents the probability that the distance between any two points in the phase space is less than a given search radius r. The distance is the distance between points in a high-dimensional space measured by the infinity norm of the vector difference between two points. Therefore, the correlation integral of the reconstructed phase space B is calculated according to Equation (4).
[0131] (4)
[0132] where r is the search radius, and there is , is the infinity norm between two subsequences in the reconstructed phase space B, , is the number of phase points in the phase space. t is
[0133] Define the detection statistic ,which is calculated according to Equation (5).
[0134] (5)
[0135] The actual time series is of finite length and there is a certain correlation between elements, resulting in the actually obtained detection statistic generally not being zero. The detection statistic reflects the autocorrelation characteristics of each sub-time series. The optimal time delay can be taken as the first zero point of, or take the time points with the smallest difference from each other for all radii r. At this time, the points in the reconstructed phase space are closest to a uniform distribution, and the orbit of the reconstructed attractor is fully unfolded in the phase space. Select two radii r with the largest and smallest corresponding values, and define the difference calculated according to formula (6). This time point is either the first zero point of a certain function or the time point with the smallest difference from each other among all radii r. At this time point, the points in the reconstructed phase space are closest to a uniform distribution, and the orbit of the reconstructed attractor is fully unfolded in the phase space.
[0136] (6)
[0137] In formula (6), measures the maximum deviation degree of the detection statistic for all search radii r. So the local maximum time t should be the zero point of and the minimum value of, and the optimal delay time corresponds to the first one among these local maximum times t.
[0138] Then, when the electronic device calculates the delay time and embedding dimension using the C-C method. Exemplarily, take N = 320, p = 2, 3, 4, 5, , is the standard deviation of the time series, where i = 1, 2, 3, 4. Calculate , and respectively according to formulas (7)-(9).
[0139] (7)
[0140] (8)
[0141] (9)
[0142] In formulas (7)-(9), the C-C method takes the first zero point of or the first minimum value of as the optimal time delay , and calculates the time delay τ = . And comprehensively and , take the global minimum value as the time window length of the time series , and finally, according to formula (3), calculate the embedding dimension p. Then, according to formula (3), calculate the delay time using the embedding dimension, and calculate the phase space based on the embedding dimension and the delay time.
[0143] The recurrence plot is a two-dimensional plot composed of black dots, white dots, and two time axes. After reconstructing the time series in the phase space, the phase space is obtained .
[0144] Among them, use to represent the distance between the vector and the vector . If and are less than a certain threshold , then and are considered to be close to each other, and the corresponding position Nij on the recurrence plot is marked as 1 (represented by a black dot); otherwise, it is marked as 0 (represented by a white dot), as shown in formula (10).
[0145] (10)
[0146] In formula (10), , is the distance threshold, is the Heaviside function, which is calculated according to formula (11).
[0147] (11)
[0148] Nij is calculated according to formula (12).
[0149] (12)
[0150] Step a2, perform data conversion on the body movement parameter sequence to generate a body grayscale image.
[0151] Specifically, the body movement parameter sequence respectively includes the linear acceleration movement parameter sequences of the body on the X, Y, and Z axes and the angular acceleration movement parameter sequences of the body on the X, Y, and Z axes. The body grayscale image includes the linear acceleration grayscale image of the body on the X axis, the linear acceleration grayscale image of the body on the Y axis, the linear acceleration grayscale image of the body on the Z axis, the angular acceleration grayscale image of the body on the X axis, the angular acceleration grayscale image of the body on the Y axis, and the angular acceleration grayscale image of the body on the Z axis.
[0152] The electronic device can perform data conversion on the sequence of linear acceleration motion parameter of the body in the X-axis to generate a grayscale image of the linear acceleration of the body in the X-axis; perform data conversion on the sequence of linear acceleration motion parameter of the body in the Y-axis to generate a grayscale image of the linear acceleration of the body in the Y-axis; perform data conversion on the sequence of linear acceleration motion parameter of the body in the Z-axis to generate a grayscale image of the linear acceleration of the body in the Z-axis; perform data conversion on the sequence of angular acceleration motion parameter of the body in the X-axis to generate a grayscale image of the angular acceleration of the body in the X-axis; perform data conversion on the sequence of angular acceleration motion parameter of the body in the Y-axis to generate a grayscale image of the angular acceleration of the body in the Y-axis; perform data conversion on the sequence of angular acceleration motion parameter of the body in the Z-axis to generate a grayscale image of the angular acceleration of the body in the Z-axis.
[0153] Among them, the data conversion method can be referred to the above text and will not be specifically described here.
[0154] Step S2032, generate a target color image based on the grayscale image.
[0155] In some alternative embodiments, the grayscale image includes a head grayscale image and a body grayscale image, and the above step S2032 includes:
[0156] Step b1, generate a head color image based on the head grayscale image.
[0157] Specifically, the head grayscale image includes the grayscale image of the linear acceleration of the head in the X-axis, the grayscale image of the linear acceleration of the head in the Y-axis, the grayscale image of the linear acceleration of the head in the Z-axis, the grayscale image of the angular acceleration of the head in the X-axis, the grayscale image of the angular acceleration of the head in the Y-axis, and the grayscale image of the angular acceleration of the head in the Z-axis. The head color image includes a head linear acceleration color image and a head angular acceleration color image.
[0158] Specifically, the above step b1 may include the following steps:
[0159] Step b11, determine the first red channel from the grayscale image of the linear acceleration of the head in the X-axis, the grayscale image of the linear acceleration of the head in the Y-axis, and the grayscale image of the linear acceleration of the head in the Z-axis, determine the first green channel from the two grayscale images other than the first red channel, and determine the last grayscale image as the first blue channel;
[0160] Step b12, generate a head linear acceleration color image according to the determined first red channel, first green channel, and first blue channel;
[0161] Step b13, determine the second red channel from the grayscale image of the angular acceleration of the head in the X-axis, the grayscale image of the angular acceleration of the head in the Y-axis, and the grayscale image of the angular acceleration of the head in the Z-axis, determine the second green channel from the two grayscale images other than the second red channel, and determine the last grayscale image as the second blue channel;
[0162] Step b14: Generate a head angular acceleration color image based on the determined second red channel, second green channel, and second blue channel.
[0163] Specifically, the electronic device can use the head acceleration grayscale image on the X-axis, the head acceleration grayscale image on the Y-axis, and the head acceleration grayscale image on the Z-axis as the red, green, and blue channels of the head linear acceleration color image respectively to obtain a head linear acceleration color image.
[0164] The electronic device can use the head angular acceleration grayscale image on the X-axis, the head angular acceleration grayscale image on the Y-axis, and the head angular acceleration grayscale image on the Z-axis as the red, green, and blue channels of the head angular acceleration color image respectively to obtain a head angular acceleration color image.
[0165] It should be noted that in the embodiments of this application, no specific limitation is imposed on the correspondence between the head angular acceleration grayscale image on the X-axis, the head angular acceleration grayscale image on the Y-axis, and the head angular acceleration grayscale image on the Z-axis as the red, green, and blue channels of the color image respectively.
[0166] Step b2: Generate a body color image based on the body grayscale image.
[0167] Specifically, the body grayscale image includes the body linear acceleration grayscale image on the X-axis, the body linear acceleration grayscale image on the Y-axis, the body linear acceleration grayscale image on the Z-axis, the body angular acceleration grayscale image on the X-axis, the body angular acceleration grayscale image on the Y-axis, and the body angular acceleration grayscale image on the Z-axis. The body color image includes the body linear acceleration color image and the body angular acceleration color image.
[0168] Specifically, the above step b2 may include the following steps:
[0169] Step b21: Determine the first red channel from the body linear acceleration grayscale image on the X-axis, the body linear acceleration grayscale image on the Y-axis, and the body linear acceleration grayscale image on the Z-axis, determine the first green channel from the two grayscale images other than the first red channel, and determine the last grayscale image as the first blue channel;
[0170] Step b22: Generate a body linear acceleration color image based on the determined first red channel, first green channel, and first blue channel;
[0171] Step b23: Determine the second red channel from the body angular acceleration grayscale image on the X-axis, the body angular acceleration grayscale image on the Y-axis, and the body angular acceleration grayscale image on the Z-axis, determine the second green channel from the two grayscale images other than the second red channel, and determine the last grayscale image as the second blue channel;
[0172] Step b24: Generate a body angular acceleration color image according to the determined second red channel, second green channel, and second blue channel.
[0173] Step b3: Stitch the head color image and the body color image to generate a target color image.
[0174] Specifically, the electronic device can use the body acceleration gray-scale image on the X-axis, the body acceleration gray-scale image on the Y-axis, and the body acceleration gray-scale image on the Z-axis as the red, green, and blue channels of the body linear acceleration color image respectively to obtain a body linear acceleration color image.
[0175] The electronic device can use the body angular acceleration gray-scale image on the X-axis, the body angular acceleration gray-scale image on the Y-axis, and the body angular acceleration gray-scale image on the Z-axis as the red, green, and blue channels of the body angular acceleration color image respectively to obtain a body angular acceleration color image.
[0176] It should be noted that the present application embodiment does not specifically limit the correspondence between the body angular acceleration gray-scale image on the X-axis, the body angular acceleration gray-scale image on the Y-axis, and the body angular acceleration gray-scale image on the Z-axis as the red, green, and blue channels of the color image respectively.
[0177] Step S204: Based on a preset type recognition model, identify the target data to determine the target vestibular illusion recognition type corresponding to the target data.
[0178] For details, please refer to Figure 1 Step S104 of the illustrated embodiment, which will not be elaborated here.
[0179] Step S205: Compare the target vestibular illusion recognition type with a preset vestibular illusion training type.
[0180] For details, please refer to Figure 1 Step S105 of the illustrated embodiment, which will not be elaborated here.
[0181] Step S206: Verify the training data according to the comparison result.
[0182] For details, please refer to Figure 1 Step S106 of the illustrated embodiment, which will not be elaborated here.
[0183] Exemplarily, such as Figure 3As shown, the electronic device respectively converts the linear acceleration motion parameter sequences of the head on the X, Y, and Z axes and the angular acceleration motion parameter sequences of the head on the X, Y, and Z axes to generate a head linear acceleration grayscale image on the X axis, a head linear acceleration grayscale image on the Y axis, a head linear acceleration grayscale image on the Z axis, a head angular acceleration grayscale image on the X axis, a head angular acceleration grayscale image on the Y axis, and a head angular acceleration grayscale image on the Z axis. The electronic device respectively converts the linear acceleration motion parameter sequences of the body on the X, Y, and Z axes and the angular acceleration motion parameter sequences of the body on the X, Y, and Z axes to generate a body linear acceleration grayscale image on the X axis, a body linear acceleration grayscale image on the Y axis, a body linear acceleration grayscale image on the Z axis, a body angular acceleration grayscale image on the X axis, a body angular acceleration grayscale image on the Y axis, and a body angular acceleration grayscale image on the Z axis.
[0184] Then, the electronic device converts the head linear acceleration grayscale image on the X axis, the head linear acceleration grayscale image on the Y axis, and the head linear acceleration grayscale image on the Z axis to generate a head linear acceleration color image; converts the head angular acceleration grayscale image on the X axis, the head angular acceleration grayscale image on the Y axis, and the head angular acceleration grayscale image on the Z axis to generate a head angular acceleration color image; converts the body linear acceleration grayscale image on the X axis, the body linear acceleration grayscale image on the Y axis, and the body linear acceleration grayscale image on the Z axis to generate a body linear acceleration color image; converts the body angular acceleration grayscale image on the X axis, the body angular acceleration grayscale image on the Y axis, and the body angular acceleration grayscale image on the Z axis to generate a body angular acceleration color image.
[0185] The electronic device stitches together the head linear acceleration color image, the head angular acceleration color image, the body linear acceleration color image, and the body angular acceleration color image to generate a target color image.
[0186] The vestibular illusion training type verification method provided by the embodiments of the present application splits the head movement parameters to generate multiple non - overlapping head movement parameter subsequences; according to the relationship between the head movement parameter subsequences, calculates the correlation integral between the head movement parameter subsequences, ensuring the accuracy of the calculated correlation integral between the head movement parameter subsequences. According to the correlation integral between the head movement parameter subsequences, calculates the autocorrelation characteristics between the head movement parameter subsequences, ensuring the accuracy of the calculated autocorrelation characteristics between the head movement parameter subsequences. According to the autocorrelation characteristics between the head movement parameter subsequences, determines the delay time and embedding dimension corresponding to the head movement parameter sequence, ensuring the accuracy of the determined delay time and embedding dimension corresponding to the head movement parameter sequence. Based on the delay time and embedding dimension, generates the head phase space corresponding to the head movement parameter sequence, ensuring the accuracy of the generated head phase space corresponding to the head movement parameter sequence. According to the distances between the head vectors included in the head phase space, generates a head grayscale image, ensuring the accuracy of the generated head grayscale image. Similarly, performs data conversion on the body movement parameter sequence to generate a body grayscale image, ensuring the accuracy of the generated body grayscale image.
[0187] Among them, the head movement parameter sequence respectively includes the linear acceleration movement parameter sequences of the head on the X, Y, and Z axes and the angular acceleration movement parameter sequences of the head on the X, Y, and Z axes; the head grayscale image includes the linear acceleration grayscale image of the head on the X axis, the linear acceleration grayscale image of the head on the Y axis, the linear acceleration grayscale image of the head on the Z axis, the angular acceleration grayscale image of the head on the X axis, the angular acceleration grayscale image of the head on the Y axis, and the angular acceleration grayscale image of the head on the Z axis.
[0188] Determine the first red channel from the head X-axis acceleration grayscale image, the head Y-axis acceleration grayscale image, and the head Z-axis acceleration grayscale image. Determine the first green channel from the two grayscale images other than the first red channel, and determine the last grayscale image as the first blue channel. Generate a head linear acceleration color image based on the determined first red channel, first green channel, and first blue channel, ensuring that the generated head linear acceleration color image can represent the linear acceleration motion parameter sequences of the head in the X, Y, and Z axes. Determine the second red channel from the head X-axis angular acceleration grayscale image, the head Y-axis angular acceleration grayscale image, and the head Z-axis angular acceleration grayscale image. Determine the second green channel from the two grayscale images other than the second red channel, and determine the last grayscale image as the second blue channel. Generate a head angular acceleration color image based on the determined second red channel, second green channel, and second blue channel. Ensure that the generated head angular acceleration color image can represent the angular acceleration motion parameter sequences of the head in the X, Y, and Z axes. Similarly, generate a body color image based on the body grayscale image, ensuring the accuracy of the generated body color image. Stitch the head color image and the body color image to generate a target color image, ensuring the accuracy of the generated target color image.
[0189] In this embodiment, a method for verifying the type of vestibular illusion training is provided, which can be used in the above-mentioned electronic device. Figure 4 It is a flowchart of the method for verifying the type of vestibular illusion training according to an embodiment of the present invention, as Figure 4 shown. The process includes the following steps:
[0190] Step S301, obtain a preset vestibular illusion training type and training data corresponding to the preset vestibular illusion training type.
[0191] For details, please refer to step S201 of the above embodiment, which will not be elaborated here.
[0192] Step S302, perform vestibular illusion type training on a preset person based on the training data, and obtain the motion parameters corresponding to the preset person.
[0193] For details, please refer to step S202 of the above embodiment, which will not be elaborated here.
[0194] Step S303, perform data processing on the motion parameters to generate target data.
[0195] For details, please refer to step S203 of the above embodiment, which will not be elaborated here.
[0196] Step S304, based on a preset type recognition model, identify the target data to determine the target vestibular illusion recognition type corresponding to the target data.
[0197] Specifically, the target data is a target color image. The above step S304 may include the following steps:
[0198] Step S3041: Input the target color image into a preset type recognition model.
[0199] Specifically, the electronic device may input the target color image into a preset type recognition model.
[0200] Step S3042: The preset type recognition model performs feature recognition and feature extraction on the target color image, and based on the extracted features, outputs the target vestibular illusion recognition type.
[0201] Specifically, the preset type recognition model includes a partitioning layer, a Swin Transformer layer (hierarchical vision transformer layer), a pooling layer, a fully connected layer, and a Softmax layer. The above step S3042 may include the following steps:
[0202] Step c1: The partitioning layer partitions the target color image, generates a plurality of target color image blocks, and generates corresponding marker information for each target color image block.
[0203] Specifically, the architecture of the preset image recognition model is as shown in the appendix Figure 5 The whole preset image recognition model includes a block partitioning layer, a Swin Transformer layer, a pooling layer, a fully connected layer, and a Softmax layer.
[0204] The partitioning layer in the preset type recognition model partitions the target color image, generates a plurality of target color image blocks, and generates corresponding marker information for each target color image block.
[0205] Exemplarily, assume that the size of the target color image obtained through phase space reconstruction, recurrence plot, and image stitching is W×H×3, where W and H are the width and length of the target color image respectively, and "3" represents the three color channels of the target color image. First, the partitioning layer in the preset type recognition model partitions the target color image through the block partitioning layer. In the block partitioning layer, the target color image is divided into "blocks", and the block is used as the smallest unit instead of pixels. Each block corresponds to a "marker", and the number of markers is the same as the number of blocks, and the output forms a three-dimensional matrix of H / 4×W / 4.
[0206] Step c2: The Swin Transformer layer fuses the feature information corresponding to each target color image block, generates a fused feature, performs a linear transformation on the channel data of each pixel in the fused feature to generate a multi-dimensional feature; and performs feature extraction on the multi-dimensional feature to generate a target feature.
[0207] Specifically, asFigure 6 As shown, the Swin Transformer layer includes four stages. The first stage includes a first layer normalization (LN) and a window multi-head self-attention (W-MSA). The second stage includes a second layer normalization (LN) and a first multi-layer perceptron (MLP). The third stage includes a third layer normalization (LN) and a shifted window-based multi-head self-attention (SW-MSA). The fourth stage includes a fourth layer normalization (LN) and a second multi-layer perceptron (MLP). The above step c2 includes:
[0208] Step c21: The first layer normalization normalizes the fused features to generate first normalized features.
[0209] Step c22: The window multi-head self-attention performs self-attention calculation on the first normalized features and performs residual calculation to generate first attention features.
[0210] Step c23: The second layer normalization normalizes the first attention features to generate second normalized features.
[0211] Step c24: The first multi-layer perceptron performs a linear transformation on the channel data of each pixel of the second normalized features to generate multi-dimensional features.
[0212] Step c25: The third layer normalization normalizes the multi-dimensional features to generate third normalized features.
[0213] Step c26: The shifted window-based multi-head self-attention performs self-attention calculation on the third normalized features to generate second attention features.
[0214] Step c27: The fourth layer normalization normalizes the second attention features to generate fourth normalized features.
[0215] Step c28: The second multi-layer perceptron performs a linear transformation on the channel data of each pixel of the fourth normalized features to generate target features.
[0216] Specifically, in the first stage, the first normalization layer calculates the mean and variance of each sample feature for the fused features and normalizes the features to generate the first normalized features. During the training process, the parameters of the first normalization layer are adaptively adjusted according to the data distribution. Then, the window multi-head self-attention layer captures the feature dependencies within local windows in the input feature map. It calculates the attention weights between each position and other positions within the window and performs weighted fusion on the features, thereby being able to better model the interactions of local features. Specifically, the input feature map is divided into multiple non-overlapping windows, and self-attention is calculated within each window. Through the multi-head mechanism, multiple attention heads are calculated in parallel to obtain different feature representations. Finally, the outputs of each window are concatenated or merged to generate the first attention features.
[0217] In the second stage, the second normalization layer calculates the mean and variance of each sample feature for the first attention features and normalizes the features to generate the second normalized features. The first multi-layer perceptron layer performs a linear transformation on the channel data of each pixel of the second normalized features to generate multi-dimensional features.
[0218] Among them, the MLP layer usually consists of an input layer, multiple hidden layers, and an output layer. In each neuron, the input signal is linearly combined with the weights and undergoes a non-linear transformation through an activation function. Information is passed between layers, and the weights are adjusted through the backpropagation algorithm to optimize the performance of the model.
[0219] In the third stage, the third normalization layer calculates the mean and variance of each sample feature for the multi-dimensional features and normalizes the features to generate the third normalized features. Similar to W-MSA, the shifted window multi-head self-attention layer first divides the input feature map into windows, but in adjacent layers, the windows are shifted by a certain amount. This allows for some overlap between adjacent windows, enabling the acquisition of more global information. By calculating the self-attention within the shifted windows, the modeling and fusion of features are achieved, generating the second attention features.
[0220] In the third stage, the fourth normalization layer normalizes the second attention features to generate the fourth normalized features. The second multi-layer perceptron layer performs a linear transformation on the channel data of each pixel of the fourth normalized features to generate the target features.
[0221] Exemplarily, as shown in the appendix Figure 5As shown, during the first stage, the number of tokens remains at W / 4 × H / 4, the size of the output feature is set to D, and stage 1 is repeated twice; during the second stage, the number of tokens is reduced to 1 / 4 of the previous one, becoming W / 8 × H / 8, and the output size is 2D, and stage 2 is repeated twice; the third stage and the fourth stage also follow a similar pattern, with the number of output tokens being W / 16 × H / 16 and W / 32 × H / 32 respectively, and the output sizes being 4D and 8D respectively, and stage 3 and stage 4 are repeated 6 times and 2 times respectively.
[0222] Exemplarily, the four stages respectively include 2, 2, 6, and 2 Swin Transformers. The number of channels in the four stages are 96, 192, 384, and 768 respectively. The number of attention heads included in each stage are 3, 6, 12, and 24 respectively.
[0223] Among them, the role of the block fusion layer in the Swin Transformer layer is to perform downsampling, reduce the resolution, and adjust the number of channels to establish a hierarchical structure. The features of each block group are concatenated in the first block fusion layer. Then the concatenated features are input into the linear embedding layer for processing. The linear embedding layer performs a linear transformation on the channel data of each pixel, projects the channel data into any dimension, and the projected dimension is denoted as "D", and then the feature processing is carried out in the Swin Transformer module.
[0224] Specifically, the structure of the SW-MSA layer is similar to that of the W-MSA layer, the difference lies in the feature calculation part, which involves a sliding window operation. The calculations of each part of the network structure are shown in formulas (13)-(16).
[0225] (13)
[0226] (14)
[0227] (15)
[0228] (16)
[0229] Among them, and respectively represent the output features of W-MSA and MLP for l, where l is the index of the partition.
[0230] The use of MSA improves accuracy and loss situation, thus enhancing the generalization ability of the model. However, MSA has high computational complexity. By adopting W-MSA, the computational complexity can be reduced. The standard MSA calculation is shown in formulas (17)-(20). Assume there are i heads and the number of blocks is w×h.
[0231] (17)
[0232] (18)
[0233] (19)
[0234] (20)
[0235] In formulas (17)-(20), is the input matrix, are the query matrix, key matrix, and value matrix respectively, are the query parameter matrix, key parameter matrix, and value parameter matrix respectively. is the transformation matrix, s is the dimension of the query matrix and key matrix, represents the concatenation matrix.
[0236] The self-attention calculation for each window is shown in formula (21).
[0237] (21)
[0238] In formula (21) are the query matrix, key matrix, and value matrix respectively, where is the number of blocks within the current window, s is the dimension of the query matrix and key matrix. F is obtained from the bias matrix obtained.
[0239] The calculation for each window of MSA and W-MSA is shown in formulas (22)-(24).
[0240] (22)
[0241] (23)
[0242] (24)
[0243] Among them, in formulas (22)-(24), D represents the dimension of the feature. P represents the size of the length and width of each window (both the length and width are P).
[0244] Step c3, the pooling layer compresses and extracts the target features to generate local features.
[0245] Specifically, the electronic device can obtain the size and stride of the pooling window in the pooling layer. Then, based on the size and stride of the pooling window, it slides on the input target feature map. For each pooling window position, calculations are performed according to the selected pooling method.
[0246] Optionally, if the pooling layer is max pooling, find the maximum value within the window as the output value at that position. If the pooling layer is average pooling, calculate the average value of all values within the window as the output value at that position.
[0247] Then, after the pooling operation, the obtained output feature map is the extracted local feature. These local features retain the main information of the original target feature, while reducing the dimension of the feature map and the computational amount of subsequent processing.
[0248] Step c4, the fully connected layer integrates the local features, generates global features, and maps the global features to the output categories.
[0249] Specifically, the fully connected layer can perform weighted combination on the local features. Each connection has a corresponding weight. By learning these weights, the fully connected layer can automatically determine how to best integrate the local features to form more representative global features. Then, the generated global features are further mapped to the output categories.
[0250] Step c5, the Softmax layer converts the output categories output by the fully connected layer into a probability distribution, and determines the target vestibular illusion recognition type based on the probability distribution.
[0251] Specifically, the Softmax layer converts the output categories output by the fully connected layer into a probability distribution, and based on the probability distribution, determines the type with the maximum probability distribution as the target vestibular illusion recognition type.
[0252] Step S305, compare the target vestibular illusion recognition type with the preset vestibular illusion training type.
[0253] For details, please refer to step S205 of the above embodiment, which will not be elaborated here.
[0254] Step S306, verify the training data according to the comparison result.
[0255] For details, please refer to step S206 of the above embodiment, which will not be elaborated here.
[0256] The vestibular illusion training type verification method provided by the embodiments of this application inputs the target color image into a preset type recognition model; divides the target color image into layers to generate multiple target color image blocks and generates the corresponding marker information for each target color image block; the first normalization layer performs normalization processing on the fused features to generate the first normalized features, enabling the preset type recognition model to have stronger robustness to input scale changes. The window multi-head self-attention layer performs self-attention calculation on the first normalized features and performs residual calculation to generate the first attention features, so that the multi-head mechanism can be used to simultaneously focus on multiple different representation subspaces, improving the learning ability and generalization performance of the preset type recognition model. The second normalization layer performs normalization processing on the first attention features to generate the second normalized features; the first multi-layer perceptron layer performs a linear transformation on the channel data of each pixel of the second normalized features to generate multi-dimensional features, increasing the expression ability of the preset type recognition model. The third normalization layer performs normalization processing on the multi-dimensional features to generate the third normalized features; the multi-head self-attention layer based on shifted windows performs self-attention calculation on the third normalized features to generate the second attention features, which helps to better model long-range dependencies and further improve the feature representation ability of the preset type recognition model. The fourth normalization layer performs normalization processing on the second attention features to generate the fourth normalized features; the second multi-layer perceptron layer performs a linear transformation on the channel data of each pixel of the fourth normalized features to generate the target features, ensuring the accuracy of the generated target features. The pooling layer compresses and extracts the target features to generate local features, ensuring the accuracy of the generated local features. The fully connected layer integrates the local features to generate global features and maps the global features to the output categories; the Softmax layer converts the output categories output by the fully connected layer into a probability distribution and determines the target vestibular illusion recognition type based on the probability distribution, ensuring the accuracy of the determined target vestibular illusion recognition type.
[0257] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A vestibular illusion training type verification method, characterized in that: The method comprises: Acquire a preset vestibular illusion training type and training data corresponding to the preset vestibular illusion training type; the training data is used to drive a six-degree-of-freedom motion platform; Based on the training data, vestibular illusion type training is performed on a preset person, and motion parameters corresponding to the preset person are obtained; Performing data processing on the motion parameters to generate target data; Based on a preset type recognition model, the target data is recognized to determine a target vestibular illusion recognition type corresponding to the target data; comparing the target vestibular illusion recognition type with the preset vestibular illusion training type; Verifying the training data according to the comparison result; The target data is a target color image; and the step of performing data processing on the motion parameters to generate the target data includes: Performing image conversion processing on the motion parameters to generate a grayscale image corresponding to the motion parameters; Based on the grayscale image, generating the target color image; The motion parameters include a head motion parameter sequence and a body motion parameter sequence, and the grayscale image includes a head grayscale image and a body grayscale image; performing image conversion processing on the motion parameters to generate a grayscale image corresponding to the motion parameters includes: Performing data conversion on the head motion parameter sequence to generate the head grayscale image; The body motion parameter sequence is subjected to data conversion to generate the body grayscale image.
2. The method according to claim 1, characterized in that The step of performing data conversion on the head motion parameter sequence to generate a head grayscale image includes: Splitting the head motion parameter to generate a plurality of mutually non-intersecting head motion parameter subsequences; Calculating the correlation integral between the head motion parameter subsequences according to the relationship between the head motion parameter subsequences; Calculating the autocorrelation characteristics between the head motion parameter subsequences according to the correlation integral between the head motion parameter subsequences; Determine the delay time and embedding dimension corresponding to the head motion parameter sequence according to the autocorrelation characteristics between each of the head motion parameter subsequences; Generate a head phase space corresponding to the head motion parameter sequence based on the delay time and the embedding dimension; The head grayscale image is generated according to the distances between the head vectors in the head phase space.
3. The method according to claim 2, characterized in that The head motion parameter sequence includes the linear acceleration motion parameter sequence of the head in the X, Y, and Z axes and the angular acceleration motion parameter sequence of the head in the X, Y, and Z axes respectively; the head grayscale image includes the acceleration grayscale image of the head in the X-axis line, the acceleration grayscale image of the head in the Y-axis line, the acceleration grayscale image of the head in the Z-axis line, the angular acceleration grayscale image of the head in the X-axis, the angular acceleration grayscale image of the head in the Y-axis, and the angular acceleration grayscale image of the head in the Z-axis.
4. The method according to claim 1, characterized in that The grayscale image includes a head grayscale image and a body grayscale image, and generating the target color image based on the grayscale image includes: Based on the head grayscale image, generate a head color image; Based on the body grayscale image, generate a body color image; The head color image and the body color image are spliced to generate the target color image.
5. The method according to claim 4, characterized in that The head grayscale image includes an acceleration grayscale image of the head on the X axis, an acceleration grayscale image of the head on the Y axis, an acceleration grayscale image of the head on the Z axis, an angular acceleration grayscale image of the head on the X axis, an angular acceleration grayscale image of the head on the Y axis, and an angular acceleration grayscale image of the head on the Z axis; the head color image includes a head linear acceleration color image and a head angular acceleration color image; generating a head color image based on the head grayscale image includes: Determine a first red channel from the grayscale image of the acceleration of the head on the X axis, the grayscale image of the acceleration of the head on the Y axis, and the grayscale image of the acceleration of the head on the Z axis, determine a first green channel from the two grayscale images except the first red channel, and determine the last grayscale image as a first blue channel; generating the head linear acceleration color image according to the determined first red channel, the first green channel and the first blue channel; Determine a second red channel from the grayscale image of the angular acceleration of the head on the X-axis, the grayscale image of the angular acceleration of the head on the Y-axis, and the grayscale image of the angular acceleration of the head on the Z-axis, determine a second green channel from the two grayscale images except the second red channel, and determine the last grayscale image as the second blue channel; The head angular acceleration color image is generated according to the determined second red channel, the second green channel, and the second blue channel.
6. The method according to claim 1, characterized in that The target data is a target color image, and the target data is identified based on a preset type recognition model to determine the target vestibular illusion recognition type corresponding to the target data, including: Inputting the target color image into the preset type recognition model; The preset type recognition model performs feature recognition and feature extraction on the target color image, and outputs the target vestibular illusion recognition type based on the extracted features.
7. The method according to claim 6, characterized in that The preset type recognition model includes a partitioning layer, a SwinTransformer layer, a pooling layer, a fully connected layer and a Softmax layer. The preset type recognition model performs feature recognition and feature extraction on the target color image, and outputs the target vestibular illusion recognition type based on the extracted features, including: The division layer divides the target color image to generate a plurality of target color image blocks, and generates feature information corresponding to each of the target color image blocks; The Swin Transformer layer fuses the feature information corresponding to each of the target color image blocks to generate a fused feature, and performs a linear transformation on the channel data of each pixel in the fused feature to generate a multi-dimensional feature; and performs feature extraction on the multi-dimensional feature to generate a target feature; The pooling layer compresses and extracts the target features to generate local features; The fully connected layer integrates the local features to generate global features, and maps the global features to output categories; The Softmax layer converts the output category of the fully connected layer into a probability distribution, and determines the target vestibular illusion recognition type based on the probability distribution.
8. The method according to claim 7, characterized in that The Swin Transformer layer includes four stages, the first stage includes a first normalization layer and a window multi-head self-attention layer, the second stage includes a second normalization layer and a first multi-layer perceptron layer, the third stage includes a third normalization layer and a multi-head self-attention layer based on a shifted window, and the fourth stage includes a fourth normalization layer and a second multi-layer perceptron layer; the Swin Transformer layer fuses the feature information corresponding to each of the target color image blocks to generate a fused feature, and performs a linear transformation on the channel data of each pixel in the fused feature to generate a multi-dimensional feature; Extracting the multi-dimensional features to generate target features includes: The first normalization layer normalizes the fused features to generate first normalized features; The window multi-head self-attention layer performs self-attention calculation on the first normalized feature, and performs residual calculation to generate a first attention feature; The second normalization layer normalizes the first attention feature to generate a second normalized feature; The first multi-layer perceptron layer performs a linear transformation on the channel data of each pixel of the second normalized feature to generate the multi-dimensional feature; The third normalization layer normalizes the multi-dimensional features to generate third normalized features; The multi-head self-attention layer based on the shifted window performs self-attention calculation on the third normalized feature to generate a second attention feature; The fourth normalization layer normalizes the second attention feature to generate a fourth normalized feature; The second multi-layer perceptron layer performs a linear transformation on the channel data of each pixel of the fourth normalized feature to generate the target feature.
Citation Information
Patent Citations
Vestibular illusion inducing device
CN104524691A
Vestibular tilt illusion simulation method and device and flight illusion simulator
CN113409649A