Point cloud object classification method based on context automatic encoder
By introducing a context-based autoencoder method into the pre-training framework of point cloud mask modeling, the problem of weakening of characterization learning ability is solved, and stronger generalization ability and fast and accurate classification of point cloud objects are achieved.
Patent Information
- Application Number
- CN202510301521.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The existing pre-training framework for point cloud mask modeling has the problem of weakening of characterization learning ability.
The point cloud object classification method based on the context autoencoder is adopted to process point cloud data through the furthest point sampling and KNN algorithm, and a pre-trained network is built and lightweight PointNet is used for embedding and position encoding, and a potential context regressor and potential representation alignment module are trained.
By separating characterization learning and pre-tasks, the generalization ability of point cloud mask modeling is improved, and the rapid and accurate classification of point cloud objects is achieved.
Smart Images

Figure CN120220132A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of three-dimensional computer vision technology, and particularly relates to a point cloud object classification method based on a context autoencoder. Background Art
[0002] Fully supervised learning relies on a large amount of labeled datasets and usually aims to solve specific tasks and datasets. The acquisition and annotation costs of point cloud data are significantly higher than those of two-dimensional images and texts. Therefore, the existing data volume scale is mostly several orders of magnitude different from the datasets of text images. As a result, fully supervised point cloud processing algorithms generally have problems such as high cost, insufficient representation ability, and poor transfer learning effect.
[0003] Point cloud self-supervised learning does not require labels and can well learn point cloud representations and transfer them to downstream tasks.
[0004] Current point cloud self-supervised methods can be roughly divided into two types: the masked modeling paradigm and the contrastive learning paradigm.
[0005] Inspired by the masked modeling paradigm in natural language processing, masked modeling in point clouds is pre-trained through a pre-task of partially occluding the point cloud after sampling and then reconstructing the initial point cloud. Contrastive learning is to apply different perspective transformations to the point cloud, seeking the smallest difference for the same parts and the largest difference for different parts for pre-training.
[0006] Compared with point cloud contrastive learning, point cloud masked modeling does not rely on positive and negative sample pairs, has a more direct learning objective, and better generalization ability. Recently, masked modeling methods such as PointBERT, PointMAE, PointM2AE, and POS-BERT have been proposed and shown great potential. PointBERT divides the point cloud into patches and trains a transformer-based autoencoder to recover the tokens of the masked patches. In contrast, PointMAE directly reconstructs point patches without expensive tokenizer training, using the chamfer distance as the reconstruction loss. PointM2AE introduces multi-scale representation information into the pre-training framework of MAE. POS-BERT improves on Point-BERT and proposes a single-stage pre-training method. However, the current point cloud masked modeling pre-training framework still has the limitation of weakened representation learning ability. Summary of the Invention
[0007] The purpose of the present invention is to provide a point cloud object classification method based on a context autoencoder in order to solve the problem that the current point cloud masked modeling pre-training framework still has weakened representation learning ability.
[0008] The above object of this application is achieved by the following technical solutions:
[0009] S1: Obtain point cloud data and process it through farthest point sampling and the KNN algorithm to obtain point patches P i ;
[0010] S2: Through the point patch P i , obtain the occluded point patch P m and the visible point patch P v ;
[0011] S3: Use lightweight PointNet to embed the visible point patch P v into the visible token T v , and embed the occluded point patch P m into the occluded token T m ;
[0012] S4: Use lightweight PointNet to map the central point coordinates of the point patch to the embedding dimension to obtain the position encoding Pos m of the occluded point patch and the position encoding Pos v of the visible point patch;
[0013] S5: Construct a pre-trained network; train the encoder of the pre-trained network through the occluded token T m , the visible token T v , the position encoding Pos m and the position encoding Pos v ;
[0014] S6: Obtain the point cloud data to be classified; classify the point cloud data to be classified through the trained pre-trained network to complete the classification of point cloud objects.
[0015] Optionally, step S1 includes:
[0016] Perform farthest point sampling on the obtained point cloud data X to obtain s central points C, where i = 1, 2,..., s;
[0017] Through the KNN algorithm, determine the k points closest to each central point C i to generate the point patch P i corresponding to each central point.
[0018] Optionally, step S2 includes:
[0019] Select a preset ratio to occlude the point patch P i to obtain the occluded point patch P m , and the specific steps include: replacing the tokens of the point patch to be occluded with learnable tokens with shared weights;
[0020] The unoccluded point patches are used as the visible point patch P v .
[0021] Optionally, step S3 includes:
[0022] The lightweight PointNet includes: a multi-layer perceptron and a max pooling module;
[0023] S31: Reshape the visible point block P v through the lightweight PointNet, and after the first convolution, use max pooling to obtain the first global feature;
[0024] S32: Concatenate the global feature and the local feature, perform a second convolution on the concatenated feature, then use max pooling to obtain the final global feature, and use the final global feature as the visible token T v ;
[0025] S33: Repeat steps S31 to S32 to obtain the occluded token T m .
[0026] Optionally, step S4 includes:
[0027] Use the learnable multi-layer perceptron of the lightweight PointNet to transform the input three-dimensional coordinates to a high-dimensional space through a linear transformation, that is, map the center point coordinate C m of the occluded point block P m to Pos m , and map the center point coordinate C v of the visible point block P v to Pos v .
[0028] Optionally, step S5 includes:
[0029] The pre-trained network includes: an encoder, a latent context regressor, a latent representation alignment module, a decoder, and a prediction head module;
[0030] The encoder is used to extract point cloud features; the encoder includes: encoder E1 and encoder E2;
[0031] Encode the occluded token T m , the visible token T v , the position encoding Pos m , and the position encoding Pos v through two independent encodings, and the encoded tokens are respectively represented as T ev and T em ;
[0032] The encoded tokens are respectively represented as T ev and T em , combine with the latent context regressor to obtain the encoded keys K and values V of all point blocks, and calculate the similarity of the keys and values;
[0033] When calculating the similarity between keys and values, the position encoding corresponding to each point block is introduced, and the position encoding Pos of the visible point block v and the position encoding Pos of the occluded point block m are connected;
[0034] Using the key K, value V, position encoding Pos v and position encoding Pos m , the latent context regressor iteratively updates the learnable token M and outputs the updated learnable token M as the predicted occluded point block feature R:
[0035]
[0036] where M l-1 is the learnable token of the previous iteration; M1 is the learnable token of this iteration;
[0037] The decoder D consists of Transformer blocks. The decoder uses the predicted occluded point block feature R and the position encoding Pos of the occluded point block m to calculate the predicted occluded point block token T prem ;
[0038] The prediction head module consists of fully connected layers. By projecting the occluded point block token T prem into a vector with the same dimension as the total number of point blocks and reshaping it, the predicted point block P pre is obtained; The predicted point block P pre includes visible and occluded point blocks.
[0039] Optionally, step S5 includes:
[0040] The total loss function L of the pre-trained network = L1 + L2;
[0041] The loss function L1 calculates the difference between the reconstructed points and the real points from the chamfer distance of the predicted point block P pre and the chamfer distance of the actual point block P;
[0042] The loss function L2 uses a coefficient λ multiplied by the mean squared error between the predicted occluded point block feature R and the encoded token T of the occluded point block em ;
[0043]
[0044] L2 = λ × MSE(R, T em )
[0045] L = L1 + L2
[0046] where represents the predicted point block Ppre The chamfer distance; x represents the chamfer distance of the actual point block P; P represents the actual point block; MSE represents the mean squared error function;
[0047] Obtain a point cloud data set without labels; based on the loss function of the pre-trained network, use the point cloud data set without labels to perform self-supervised training on the pre-trained network to obtain a trained encoder.
[0048] An electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device executes a point cloud object classification method based on a context autoencoder.
[0049] A computer-readable storage medium stores instructions that, when executed, perform a point cloud object classification method based on a context autoencoder.
[0050] The beneficial effects brought by the technical solution provided in this application are:
[0051] Using a context autoencoder with alignment constraints to separate representation learning and pre-tasks solves two problems existing in current point cloud mask modeling pre-training: 1. The change of the decoder's representation of the visible part during the pre-training process weakens the encoder's learning ability. 2. The leakage of the position information of the invisible part during the pre-training process. This application has stronger generalization ability compared with the existing methods, and can achieve fast and accurate classification of point cloud objects through the pre-trained network of this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The following will further illustrate this application in conjunction with the drawings. In the drawings:
[0053] Figure 1 is the overall architecture diagram in the embodiment of this application;
[0054] Figure 2 is the schematic diagram of the reconstruction result of the occluded point block in the embodiment of this application;
[0055] Figure 3 is the schematic diagram of the electronic device structure in the embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] In order to have a clearer understanding of the technical features, objectives, and effects of this application, the specific embodiments of this application will now be described in detail with reference to the drawings.
[0057] The embodiment of this application provides a point cloud object classification method based on a context autoencoder.
[0058] Please refer to Figure 1 , Figure 1 which is the overall architecture diagram of a point cloud object classification method based on a context autoencoder in an embodiment of the present application, including:
[0059] S1: Obtain point cloud data and process it through farthest point sampling and the KNN algorithm to obtain point patches P i ;
[0060] S2: Through point patch P i , obtain occluded point patches P m and visible point patches P v ;
[0061] S3: Use lightweight PointNet to embed the visible point patch P v into visible tokens T v , and embed the occluded point patch P m into occluded tokens T m ;
[0062] S4: Use lightweight PointNet to map the center point coordinates of the point patch to the embedding dimension to obtain the position encoding Pos m of the occluded point patch and the position encoding Pos v of the visible point patch;
[0063] S5: Construct a pre-trained network; train the encoder of the pre-trained network through the occluded token T m , the visible token T v , the position encoding Pos m and the position encoding Pos v ;
[0064] S6: Obtain the point cloud data to be classified; classify the point cloud data to be classified through the trained pre-trained network to complete point cloud object classification.
[0065] As an embodiment, for the point cloud object classification task or the point cloud object part segmentation task, set the corresponding task head, fine-tune the training using the trained encoder E, and set the decoder for evaluating the performance metrics.
[0066] Step S1 includes:
[0067] Perform farthest point sampling on the obtained point cloud data X to obtain s center points C, where i = 1, 2,..., s;
[0068] Through the KNN algorithm, determine the k points closest to each center point C i to generate the point patch P i corresponding to each center point.
[0069] Step S2 includes:
[0070] Select a preset ratio to occlude the point block P i to obtain the occluded point block P m , and the specific steps include: replacing the tokens of the point block to be occluded with learnable tokens sharing weights;
[0071] The unoccluded point blocks are used as visible point blocks P v .
[0072] Step S3 includes:
[0073] The lightweight PointNet includes: a multi-layer perceptron and a max pooling module;
[0074] S31: Reshape the visible point block P through the lightweight PointNet v , after the first convolution, use max pooling to obtain the first global feature;
[0075] S32: Concatenate the global feature and the local feature, perform a second convolution on the concatenated feature, then use max pooling to obtain the final global feature, and use the final global feature as the visible token T v ;
[0076] S33: Repeat steps S31 to S32 to obtain the occluded token T m .
[0077] Step S4 includes:
[0078] Use the learnable multi-layer perceptron of the lightweight PointNet to transform the input three-dimensional coordinates to a high-dimensional space through linear transformation, that is, map the center point coordinates C m of the occluded point block P m to Pos m , and map the center point coordinates C v of the visible point block P v to Pos v .
[0079] Step S5 includes:
[0080] The pre-trained network includes: an encoder, a latent context regressor, a latent representation alignment module, a decoder, and a prediction head module;
[0081] The encoder is used to extract point cloud features; the encoder includes: encoder E1 and encoder E2;
[0082] As an embodiment, the encoder consists of standard Transformer blocks, which are composed of a preset number of stacked multi-head self-attention layers and fully connected feed-forward networks.
[0083] Encoding the occluded token T m , the visible token T v , the position encoding Pos m and the position encoding Pos v , and the encoded tokens are respectively denoted as T ev and T em ;
[0084] The encoded tokens are respectively denoted as T ev and T em . Combining with the latent context regressor, the encoded keys K and values V of all point chunks are obtained, and the similarity between the keys and values is calculated;
[0085] As an example, the latent context regressor LCR consists of a series of cross-attention modules. The query is a learnable token M. The initial dimension size and quantity of M are the same as those of the tokens of the occluded point chunks. That is to say, this learnable token replaces the tokens of the occluded point chunks, and the parameters of this learnable token are shared by all occluded point chunks; the key and the value are the encoded tokens of all point chunks, namely T ev and T em :
[0086] K = V = [T ev ; T em + [Pos v ; Pos m
[0087] When calculating the similarity between the keys and values, the position encoding corresponding to each point chunk is introduced, and the position encoding Pos v of the visible point chunks and the position encoding Pos m of the occluded point chunks are concatenated;
[0088] Using the key K, the value V, the position encoding Pos v and the position encoding Pos m , the latent context regressor iteratively updates the learnable token M and outputs the updated learnable token M as the predicted occluded point chunk feature R:
[0089]
[0090] where M l-1 is the learnable token of the previous iteration; M l is the learnable token of this iteration;
[0091] As an example, the potential representation alignment module LRA is embodied in the iterative process of the lower branch encoder E2 and the potential context regressor LCR, which affects the update of the predicted occluded point block feature R through the second loss function L2;
[0092] The decoder D consists of Transformer blocks, and the decoder utilizes the predicted occluded point block feature R and the position encoding Pos of the occluded point block m , and calculates the predicted occluded point block token T prem ;
[0093] The prediction head module consists of fully connected layers. Through the prediction head module, after projecting the occluded point block token T prem into a vector with a dimension equal to the total number of point blocks and reshaping it, the predicted point block P pre is obtained; The predicted point block P pre includes visible and occluded point blocks.
[0094] Step S5 includes:
[0095] The total loss function L of the pre-trained network is L = L1 + l2;
[0096] The loss function L1 calculates the difference between the reconstructed points and the real points from the chamfer distance of the predicted point block P pre and the chamfer distance of the actual point block P;
[0097] The loss function l2 uses a coefficient λ multiplied by the mean squared error between the predicted occluded point block feature R and the encoded token T of the occluded point block em ;
[0098]
[0099] L2 = λ × MSE(R, T em )
[0100] L = L1 + L2
[0101] where represents the chamfer distance of the predicted point block P pre ; x represents the chamfer distance of the actual point block P; P represents the actual point block; MSE represents the mean squared error function;
[0102] Obtain a point cloud data set without labels; Based on the loss function of the pre-trained network, use the point cloud data set without labels to perform self-supervised training on the pre-trained network to obtain the trained encoder.
[0103] As an example, for the point cloud object classification task or the point cloud object part segmentation task, set the corresponding task head, use the trained encoder E for fine-tuning training, and set the decoder for evaluating performance metrics.
[0104] To demonstrate the effectiveness of our pre-training method, we visualized the reconstruction results of the pre-task on the ShapeNet validation set to show the qualitative results of the pre-task. As Figure 2 shown, we occluded the input point cloud at a high ratio of 75% and performed reconstruction. Our method can proficiently predict subsequent point patches and accurately generate the original point cloud, indicating that our encoder can learn deep features.
[0105]
[0106]
[0107] Table 1
[0108] To evaluate the practical application performance of the proposed method on real-world datasets, the present invention applied the pre-trained model to the ScanObjectNN dataset, which contains approximately 15,000 objects extracted from real-world indoor scans. The present invention conducted experiments under three different settings: OBJ-BG, OBJ-ONLY, and PB-T50-RS. Table 1 shows the experimental results. As shown in Table 1, it represents the classification results on the ScanObjectNN dataset. Three settings were used on the ScanObjectNN dataset, and PB-T50-RS is the most challenging setting. Our method outperforms the advanced fully supervised methods (PRANet and above) and is competitive compared to the state-of-the-art self-supervised methods (represents the local reproduction results).
[0109] Self-supervised methods Accuracy Transformer-OcCo
[29] 92.1 Point-MAE
[30] 93.8 Point-MAE(Rep) 93.0 Ours 93.6
[0110] Table 2
[0111]
[0112]
[0113] Table 3
[0114] This study evaluated the pre-trained model on the ModelNet40 dataset, which contains 12,311 clean 3D CAD models covering 40 categories. For fair comparison, the standard voting method was used during testing, and the input point clouds of all comparison methods only contained coordinate information without providing normal information. The experimental results are shown in Table 2, and the method of the present invention achieved the best performance among all methods.
[0115] Methods 5-way,10-shot 5-way,20-shot 10-way,10-shot 10-way,20-shot DGCNN-rand
[36] 31.6±2.8 40.8±4.6 19.9±2.1 16.9±1.5 DGCNN-OcCo
[37] 90.6±2.8 92.5±1.9 82.9±1.3 86.5±2.2 Transformer-rand
[29] 87.8±5.29 93.3±4.3 84.6±5.5 89.4±6.3 Transformer-OcCo
[29] 94.0±3.6 95.9±2.3 89.4±5.1 92.4±4.6 PointBERT
[29] 94.6±3.1 96.3±2.7 91.0±5.4 92.7±5.1 PointMAE
[30] 96.3±2.5 97.8±1.8 92.6±4.1 95.0±3.0 Ours 96.6±2.4 97.7±1.4 92.8±4.5 95.1±2.8
[0116] Table 4
[0117] To demonstrate the generalization ability of the method of the present invention, the present invention conducts few-shot learning experiments on the ModelNet40 dataset. The few-shot learning experiments include four experiments, adopting the setting of n classes and m samples, where n ∈ {5, 10} represents the randomly selected number of classes, and m ∈ {10, 20} represents the number of objects randomly sampled for each selected class. For each setting, 10 independent trials are conducted, and the average accuracy and standard deviation are reported. The results are shown in Table 4. The method of the present invention achieves the best performance among all methods except in the case of 5-way and 20-shot, demonstrating the ability to learn new tasks using limited training data.
[0118]
[0119]
[0120] Table 5
[0121] To demonstrate the representation learning ability of the method of the present invention, the present invention is evaluated on the ShapeNetPart dataset, which includes 16,881 objects of 16 categories. The point cloud is sampled into 2048 points, and the experimental results are shown in Table 5. The present invention performs best in most categories among all comparison methods.
[0122] This application also discloses an electronic device. Referring to Figure 3 , Figure 3 is a schematic structural diagram of an electronic device disclosed in an embodiment of this application. The electronic device 500 may include: at least one processor 501, at least one network interface 504, a user interface 503, a memory 505, and at least one communication bus 502.
[0123] Among them, the communication bus 502 is used to realize the connection and communication between these components.
[0124] Among them, the user interface 503 may include a display screen. Optionally, the user interface 503 may further include a standard wired interface and a wireless interface.
[0125] Among them, the network interface 504 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0126] This application also discloses a computer-readable storage medium, which stores multiple instructions suitable for a processor to load and execute the above-mentioned method for classifying point cloud objects based on a context autoencoder.
[0127] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure.
[0128] This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not recorded in the present disclosure. The description and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A point cloud object classification method based on contextual autoencoder, characterized in that: The method comprises the following steps: S1: Get point cloud data and process it through farthest point sampling and KNN algorithm to get point block P i ; S2: Through point block P i , get the occlusion point block P m and visible point block P v ; S3: Use lightweight PointNet to block visible points P v Embedded in visible token T v , the occluded point block P m Embedded into the occlusion token T m ; S4: Use lightweight PointNet to map the center point coordinates of the point block to the embedding dimension to obtain the position encoding Pos of the occluded point block m And the position code of the visible point block Pos v ; S5: Build a pre-trained network; by blocking the token T m , visible token T v ,Position codePos m And the position code Pos v Train the encoder of the pre-trained network; S6: Obtain the point cloud data to be classified; classify the point cloud data to be classified through the trained pre-trained network to complete the point cloud object classification.
2. A point cloud object classification method based on contextual autoencoder as claimed in claim 1, characterized in that: Step S1 includes: The farthest point sampling is performed on the acquired point cloud data X to obtain s center points C, i = 1, 2, ..., s; Through the KNN algorithm, determine the distance from each center point C i The nearest k points generate a point block P corresponding to each center point i .
3. A point cloud object classification method based on contextual autoencoder as claimed in claim 1, characterized in that: Step S2 includes: Select the preset ratio point block P i Perform occlusion and obtain the occlusion point block P m ,The specific steps include: replacing the tokens of the point blocks that need to be occluded by a learnable marker with shared weights; The unobstructed point blocks are regarded as visible point blocks P v .
4. The point cloud object classification method based on contextual autoencoder according to claim 1, characterized in that: Step S3 includes: Lightweight PointNet includes: multi-layer perceptron and maximum pooling module; S31: Visible point block P through lightweight PointNet v Reshape, after the first convolution, use maximum pooling to obtain the first global feature; S32: The global features are concatenated with the local features, the concatenated features are convolved for the second time, and then the maximum pooling is used to obtain the final global features, and the final global features are used as the visible token T v ; S33: Repeat steps S31 to S32 to obtain the occlusion token T m .
5. The point cloud object classification method based on contextual autoencoder according to claim 1, characterized in that: Step S4 includes: Using a lightweight PointNet learnable multi-layer perceptron, the input 3D coordinates are transformed into a high-dimensional space through linear transformation, that is, the occluded point block P m The center point coordinates C m Mapped to Pos m , the visible point block P v The center point coordinates C v Mapped to Pos v .
6. The point cloud object classification method based on contextual autoencoder according to claim 1, characterized in that: Step S5 includes: The pre-trained network includes: encoder, latent context regressor, latent representation alignment module, decoder and prediction head module; The encoder is used to extract point cloud features; the encoder includes: encoder E1 and encoder E2; The occlusion token T is encoded by two independent m , visible token T v ,Position codePos m And the position code Pos v Encode, and the encoded tags are represented as T ev and T em ; The encoded marks are represented as T ev and T em , combined with the latent context regressor, obtain the encoded keys K and values V of all point blocks, and calculate the similarity between the keys and values; When calculating the similarity between keys and values, the position code corresponding to each point block is introduced, and the position code Pos v And the position code of the occlusion point block Pos m connect; Using key K, value V, position code Pos v and position code Pos m , the latent context regressor iteratively updates the learnable marker M and outputs the updated learnable marker M as the predicted occluded point block feature R: Among them, M l-1 is the learnable marker of the previous iteration; M1 is the learnable marker of this iteration; The decoder D is composed of Transformer blocks. The decoder uses the predicted occlusion point block feature R and the position encoding Pos of the occlusion point block m , calculate the predicted occlusion point block token T prem ; The prediction head module consists of a fully connected layer. The prediction head module transforms the occluded point block token T prem After projecting it into a vector with the same dimension as the total number of point blocks and reshaping it, we get the predicted point block P pre ; Prediction point block P pre Includes visible and occluded patches.
7. A point cloud object classification method based on contextual autoencoder as claimed in claim 6, characterized in that: Step S5 includes: The total loss function of the pre-trained network is L = L1 + L2; The loss function L1 is composed of the predicted point block P pre The chamfer distance of the reconstructed point and the real point is calculated by comparing the chamfer distance of the actual point block P; The loss function L2 uses a coefficient λ to multiply the predicted occlusion point block feature R and the token T encoded by the occlusion point block em The mean square error of L2=λ×MSE(R,T em ) L=L1+L2 in Represents the predicted point block P pre The chamfer distance of the actual point block P; x represents the chamfer distance of the actual point block P; P represents the actual point block; MSE represents the mean square error function; Obtain an unlabeled point cloud dataset; based on the loss function of the pre-trained network, use the unlabeled point cloud dataset to perform self-supervised training on the pre-trained network to obtain a trained encoder.
8. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed by a computer, the method according to any one of claims 1 to 7 is executed.