A semantic communication method, device and equipment for binocular three-dimensional target recognition
Patent Information
- Application Number
- CN202410842679.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-06-27
AI Technical Summary
与此同时,这些传感器数据的失真直接影响着3D检测的准确性
[0043]有益效果:利用两类语义提取、压缩和恢复网络对双目视觉图片中与三维目标检测相关的语义信息进行提取,在保障三维目标检测的准确性的同时,减少了所需传输的数据量,解决了带宽受限下双目视觉图片传输的问题。与此同时,信道编解码网络保障了语义信息在无线信道上的可靠传输,降低了信道噪声对系统的影响,提供了系统在噪声环境下的传输性能。
Smart Images

Figure CN118840743B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of semantic communication technology, specifically relating to a semantic communication method, apparatus, and device for binocular 3D target detection. Background Technology
[0002] Binocular 3D object detection plays a crucial role in fields such as intelligent video surveillance, robot navigation, and autonomous driving. This technology utilizes left and right visual images acquired by binocular cameras to identify and locate objects in a 3D environment, providing necessary detection results for subsequent decision-making by intelligent agents. However, due to limitations in the computing power of the sensors, complex binocular 3D object detection sometimes requires deployment on remote computing devices or in the cloud, posing challenges to data exchange between sensors and remote devices. Therefore, transmitting large amounts of sensor data within limited communication bandwidth has become an important research topic in communications. Simultaneously, the distortion of this sensor data directly affects the accuracy of 3D detection.
[0003] Considering that the accuracy of 3D object detection largely depends on the semantic information in binocular vision images, introducing semantic communication to improve detection performance under limited communication bandwidth is particularly important. Unlike traditional methods, semantic communication considers the meaning or semantics contained in the transmitted data during transmission, i.e., semantic information. It attempts to transmit task-relevant semantic information over noisy channels, rather than pursuing error-free bit-level or symbol-level transmission. By designing specialized semantic and channel codecs, semantic communication systems can selectively extract and compress task-relevant semantic information and effectively overcome channel distortion, thereby improving communication efficiency under limited bandwidth. Therefore, research on semantic communication for binocular 3D object detection is of great significance. Summary of the Invention
[0004] Purpose of the invention: For the task of binocular 3D object detection, a semantic communication framework is proposed, which aims to reduce the loss of semantic information related to 3D object detection while transmitting binocular vision images under limited bandwidth, thereby reducing the performance loss of 3D object detection.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a semantic communication method for binocular 3D target detection, comprising the following steps:
[0006] Step A: Input the acquired left and right visual images into the trained neural network to extract semantic information of key regions and global semantic information;
[0007] Step B: Compress and encode the extracted semantic information at different compression rates to obtain the compressed key regions and global semantic information;
[0008] Step C: Channel coding is performed on the compressed semantic information to obtain channel coded information, which is then transmitted to the receiver through the channel.
[0009] Step D: Decode the received channel coding information. For the key region semantic information and global semantic information that are still in a compressed state after decoding, perform semantic recovery to obtain the key region semantic information and global semantic information.
[0010] Step E: The restored key regions and global semantic information are fused to obtain the restored left and right visual images.
[0011] Furthermore, step A specifically includes the following steps:
[0012] The acquired pair of left and right visual images I r ,I l Input them into the neural network respectively;
[0013] Visual Image I r ,I l Two-dimensional object detection yields a set of bounding boxes for objects in the left and right images. Image information inside the bounding boxes is extracted from the two-dimensional object detection results, and feature extraction is performed on the extracted information to obtain semantic information S of key regions. r ,S l ;
[0014] Visual Image I r ,I l Features are extracted from two sets of convolutional residuals. The extracted left and right image features are directly concatenated along the channel dimension, and joint feature extraction is performed. The fused features are then separated to extract the left and right semantic information F. r ,F l .
[0015] Furthermore, step B specifically includes the following steps: extracting semantic information from key regions. With global semantic information The data are fed into two convolutional neural networks with different compression ratios, and then sequentially processed through a first convolution, pixel re-encoding, and a second convolution to obtain compressed semantic information of the key regions. and global semantic information The compression rate of global semantic information is greater than that of key region semantic information, i.e., n2 > n1.
[0016] Furthermore, the output dimension of the second convolution is the same as the number of channels for key region semantic information or global semantic information.
[0017] Furthermore, the channel coding of semantic information in step C specifically includes the following steps:
[0018] The compressed key region semantic information and global semantic information Each input is fed into a convolutional neural network, and each input is subjected to convolutional coding at least once. The channel coding code rate of the last convolution is C / C″, and the output dimension is C″, where C″≥C and the channel coding code rate is no greater than 1.
[0019] Furthermore, the decoding of the received channel-coded information in step D specifically includes the following steps: the information received after passing through the wireless channel is sequentially subjected to seven convolutions to achieve channel decoding; wherein the input dimension of the first convolution is C″, the output dimension of the last convolution is C, and the compressed key region semantic information after channel decoding is obtained. and global semantic information
[0020] Furthermore, step D, which involves semantic recovery of the key region semantic information and global semantic information that are still in a compressed state after decoding, specifically includes the following steps:
[0021] Key region semantic information after channel decoding First, after one convolution, we get... The feature matrix; then through pixel rearrangement, the feature matrix is... Transformed into Next, the deformed feature matrix is convolved once, with the output dimension of the convolution being the same as the number of channels for the semantic information. Finally, the recovered semantic information is filled into the corresponding regions according to the received bounding box parameters, and the other regions are set to 0, thus obtaining the recovered key region semantic information. global semantic information after channel decoding First, optical flow extraction is performed to extract a pair of optical flow information. in Optical flow information from left to right. This represents the optical flow information from right to left; subsequently, it involves... Residual feature extraction yields a set of feature values f r,1 ,f l,1 ;use For f l,1 By performing optical deformation and a single convolution, the feature f in the right figure is obtained. r,2 Simultaneously utilize For f r,1 Optical deformation is performed and passed through a convolutional layer to obtain the feature f in the left image. l,2 ; will f l,1 and f l,2 The layers are stacked along the channel dimension, and then subjected to a single convolution to obtain f.l,3 At the same time, f r,1 and f r,2 The layers are stacked along the channel dimension, and then subjected to a single convolution to obtain f. r,3 After that, f r,3 and f l,3 After each convolution, a pair of feature matrices are obtained. Then, through pixel rearrangement, the feature matrix is... Transformed into Finally, the deformed feature matrix is convolved once, with the output dimension of the convolution being the same as the number of channels for the semantic information, thus obtaining the restored global semantic information.
[0022] Furthermore, step E specifically includes the following steps:
[0023] The restored semantic information of key regions and global semantic information Divide into two groups according to left and right characteristics. and After directly concatenating the images along the channel dimension, three convolutions are performed to obtain the reconstructed left and right visual images.
[0024] An electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0025] A computer-readable storage medium storing computer instructions for causing a processor to execute the above-described method.
[0026] The training steps for the deep neural network used in semantic communication methods specifically include:
[0027] Step 1: Train the global semantic information extraction network, semantic compression network, global semantic information recovery network, and fusion network without going through the wireless channel layer;
[0028] Step 2: Train the key region semantic information extraction network, semantic compression network, key region semantic information recovery network, and fusion network without going through the wireless channel layer;
[0029] Step 3: Jointly train all semantic information processing networks without going through the wireless channel layer;
[0030] Step 4: Jointly train all semantic information processing networks and channel coding / decoding networks, using the wireless channel layer.
[0031] Furthermore, step 1 specifically includes:
[0032] Without passing through the wireless channel layer, the binocular visual images are sequentially input into a global semantic information extraction network, a semantic compression network, and a global semantic information recovery network to obtain the recovered global semantic information. This global semantic information is then input into a fusion network, and the recovered key region semantic information is replaced with an all-zero tensor to obtain the recovered binocular visual image. Finally, the network is trained using stochastic gradient descent and the Charbonnier loss function. The Charbonnier loss function is as follows:
[0033]
[0034] Among them, I and These are the binocular vision image and the restored binocular vision image, respectively, with ε being a fixed constant.
[0035] Furthermore, step 2 specifically includes:
[0036] Without passing through the wireless channel layer, the binocular visual images are first sequentially input into a key region semantic information extraction network, a semantic compression network, and a key region semantic information restoration network to obtain the restored global semantic information. Simultaneously, based on step 1, the binocular visual images are sequentially input into a global semantic information extraction network, a semantic compression network, and a global semantic information restoration network to obtain the restored global semantic information. The global semantic information and the key region semantic information are then input into a fusion network to obtain the restored binocular visual image. Finally, the key region semantic information processing network and the fusion network are trained using stochastic gradient descent and the Mask Charbonnier loss function, while the parameters of the global semantic information processing network are frozen. The Mask Charbonnier loss function is as follows:
[0037]
[0038] Among them, I mask and The images are the masked binocular vision image and the reconstructed binocular vision image, with the data outside the 2D object detection bounding boxes set to 0. λ is the weight used to balance the reconstruction of key regions and global regions.
[0039] Furthermore, step 3 specifically includes:
[0040] Without passing through the wireless channel layer, the binocular visual images are first sequentially input into a key region semantic information extraction network, a semantic compression network, and a key region semantic information recovery network to obtain the recovered global semantic information. Simultaneously, the binocular visual images are sequentially input into a global semantic information extraction network, a semantic compression network, and a global semantic information recovery network to obtain the recovered global semantic information. Finally, the global semantic information and the key region semantic information are input into a fusion network to obtain the recovered binocular visual image. Finally, all the networks described above are trained using stochastic gradient descent and the same Charbonnier loss function as in step 1.
[0041] Furthermore, step 4 specifically includes:
[0042] Following the steps described in claim 1, the binocular vision image is processed through a complete neural network, including channel encoding / decoding and a wireless channel layer, to obtain the recovered binocular vision image. Finally, all the networks described above are trained using stochastic gradient descent and the same Charbonnier loss function as in step 1.
[0043] Beneficial effects: By utilizing two types of semantic extraction, compression, and recovery networks to extract semantic information related to 3D object detection from binocular vision images, the accuracy of 3D object detection is ensured while reducing the amount of data to be transmitted, thus solving the problem of binocular vision image transmission under bandwidth constraints. Simultaneously, the channel coding and decoding network ensures reliable transmission of semantic information over the wireless channel, reduces the impact of channel noise on the system, and improves the system's transmission performance in noisy environments. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the framework process of an embodiment of the present invention.
[0045] Figure 2 This is a line graph showing the relationship between the accuracy score of three-dimensional detection and SNR on the AWGN channel in an embodiment of the present invention. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] This invention proposes a semantic communication method for binocular 3D target detection. It utilizes a deep neural network to perform semantic extraction and encoding compression at the semantic layer and channel feature extraction at the communication layer, which can demonstrate better efficiency than traditional communication systems under frequency-selective channels.
[0048] The following is an embodiment of the present invention providing a semantic communication method for binocular 3D object detection, including the following steps:
[0049] Step A: For the input left and right visual images, first use the key region semantic information extraction network to extract the key region semantic information related to the target. The key region semantic information extraction network is based on a neural network.
[0050] Step B: For the input left and right visual images, the global semantic information extraction network is used to extract global semantic information related to depth inference. The global semantic information extraction network is based on a neural network.
[0051] Step C: For the extracted semantic information, use a compression coding network to compress and encode it at different compression ratios to obtain the compressed key regions and global semantic information. The compression coding network is based on a neural network.
[0052] Step D: For the compressed semantic information, channel coding is performed using a channel coding network to obtain channel coded information, wherein the channel coding network is based on a neural network;
[0053] Step E: Transmit the channel coding information to the receiver through the channel;
[0054] Step F: For the received channel coding information, decode it using a channel decoding network, which is based on a neural network;
[0055] Step G: For the semantic information of key regions that are still in a compressed state after channel decoding, semantic recovery is performed using a key region semantic recovery network to obtain the semantic information of key regions. The key region semantic recovery network is based on a neural network.
[0056] Step H: For the global semantic information that is still in a compressed state after channel decoding, semantic recovery is performed using a global semantic recovery network to obtain the global semantic information. The global semantic recovery network is based on a neural network.
[0057] Step 1: The restored key regions and global semantic information are fused using a fusion network to obtain the restored left and right visual images. The fusion network is based on a neural network.
[0058] Please refer to Figure 1This embodiment presents a semantic communication system for binocular 3D target detection, comprising the modules shown in the figure in both the transmitter and receiver. In step A, a pair of left and right visual images I... r ,I l The data is input into the 2D object detection module, which generates bounding boxes for objects in the left and right images respectively. Each bounding box has five parameters [u1, v1, u2, v2], where (u1, v1) and (u2, v2) are the coordinates of the top-left and bottom-right corners of the bounding box. Based on the 2D object detection results, information related to 3D detection, i.e., image information inside the bounding boxes, is extracted. The extracted information is then input into at least one convolutional layer for feature extraction to obtain the semantic information S of the key regions. r ,S l At the same time, all bounding box parameters are uniformly encoded and transmitted losslessly to the receiver, enabling the receiver to use the bounding box parameters to determine the spatial location of semantic information of key regions.
[0059] In step B, two sets of convolutional residual networks are first used to analyze a pair of left and right visual images I. r ,I l Features are extracted from the left and right images. Then, the extracted features from the left and right images are directly concatenated along the channel dimension, and a joint feature extraction convolutional layer is used for joint feature extraction. Next, the fused features are input into two feature separation convolutional layers respectively to extract the left and right semantic information F. r ,F l That is, global semantic information.
[0060] Semantic information of key regions extracted in step C With global semantic information The information is fed into two compression coding networks with different compression ratios. Semantic information first passes through a convolutional network to obtain... The feature matrix is denoted by C, where C is the number of channels in the feature matrix. Then, through pixel rearrangement, the feature matrix is... Transformed into Next, the deformed feature matrix is input into a convolutional layer. The output dimension of this convolutional layer is the same as the number of channels of the semantic information, and the compressed semantic information of the key regions is obtained. and global semantic information The compression rate of global semantic information is greater than that of key region semantic information, i.e., n2 > n1.
[0061] In step D, the compressed key region semantic information is... and global semantic information Five convolutional layers are sequentially input to implement channel coding. The output dimension of the last convolutional layer is C″. The channel coding code rate under this structure is C / C″, and to ensure that the channel coding code rate is no greater than 1, C″≥C.
[0062] In step F, the information received after passing through the wireless channel is sequentially input into seven convolutional layers to achieve channel decoding. The input dimension of the first convolutional layer is C″, and the output dimension of the last convolutional layer is C. This yields the compressed semantic information of the key regions after channel decoding. and global semantic information
[0063] Key region semantic information after channel decoding in step G First, it passes through a convolutional network to obtain... The feature matrix. Then, through pixel rearrangement, the feature matrix is... Transformed into Next, the deformed feature matrix is input into a convolutional layer, the output dimension of which is the same as the number of channels for the semantic information. Finally, the recovered semantic information is filled into the corresponding regions according to the received bounding box parameters, and the other regions are set to 0, thus obtaining the recovered key region semantic information.
[0064] Global semantic information after channel decoding in step H First, input the optical flow extraction module to extract a pair of optical flow information. in Optical flow information from left to right. This represents the optical flow information from right to left. Then, a pair of residual networks respectively... Feature extraction yields a set of feature values f r,1 ,f l,1 . use For f l,1 Optical deformation is performed and passed through a convolutional layer to obtain the feature f in the right figure. r,2 Simultaneously utilize For f r,1 Optical deformation is performed and passed through a convolutional layer to obtain the feature f in the left image. l,2 f l,1 and f l,2 Stacking along the channel dimension and passing through a convolutional layer yields f. l,3 At the same time, f r,1 and f r,2 Stacking along the channel dimension and passing through a convolutional layer yields f. r,3 After that, f r,3 and f l,3 After passing through a pair of convolutional layers, a pair of feature matrices are obtained. Then, through pixel rearrangement, the feature matrix is... Transformed into Finally, the deformed feature matrix is input into a convolutional layer whose output dimension is the same as the number of channels of the semantic information, thereby obtaining the recovered global semantic information.
[0065] In step I, the restored semantic information of the key regions will be... and global semantic information Divide into two groups according to left and right characteristics. and After directly concatenating the images along the channel dimension, the images are passed through three convolutional layers to obtain the reconstructed left and right visual images.
[0066] The training process of the deep neural network in the above embodiments is as follows:
[0067] In step 1, without passing through the wireless channel layer, the binocular visual images are sequentially input into the global semantic information extraction network, the semantic compression network, and the global semantic information recovery network to obtain the recovered global semantic information. The global semantic information is then input into the fusion network, and the recovered key region semantic information is replaced with an all-zero tensor to obtain the recovered binocular visual images. Finally, the network is trained using stochastic gradient descent and the Charbonnier loss function. The Charbonnier loss function is as follows:
[0068]
[0069] Among them, I and These are the binocular vision image and the restored binocular vision image, respectively, with ε being a fixed constant.
[0070] In step 2, without passing through the wireless channel layer, the binocular visual images are first sequentially input into the key region semantic information extraction network, the semantic compression network, and the key region semantic information restoration network to obtain the restored global semantic information. Simultaneously, based on step 1, the binocular visual images are sequentially input into the global semantic information extraction network, the semantic compression network, and the global semantic information restoration network to obtain the restored global semantic information. The global semantic information and the key region semantic information are then input into a fusion network to obtain the restored binocular visual image. Finally, the key region semantic information processing network and the fusion network are trained using stochastic gradient descent and the Mask Charbonnier loss function, while the parameters of the global semantic information processing network are frozen. The Mask Charbonnier loss function is as follows:
[0071]
[0072] Among them, I mask and The images are the masked binocular vision image and the reconstructed binocular vision image, with the data outside the 2D object detection bounding boxes set to 0. λ is the weight used to balance the reconstruction of key regions and global regions.
[0073] In step 3, without passing through the wireless channel layer, the binocular visual images are first sequentially input into the key region semantic information extraction network, the semantic compression network, and the key region semantic information restoration network to obtain the restored global semantic information. Simultaneously, the binocular visual images are sequentially input into the global semantic information extraction network, the semantic compression network, and the global semantic information restoration network to obtain the restored global semantic information. Finally, the global semantic information and the key region semantic information are input into the fusion network to obtain the restored binocular visual image. Finally, all the networks described above are trained using stochastic gradient descent and the same Charbonnier loss function as in step 1.
[0074] In step 4, following the steps described in claim 1, the binocular vision image is processed through a complete neural network, including channel encoding / decoding and a wireless channel layer, to obtain the recovered binocular vision image. Finally, all the networks described above are trained using stochastic gradient descent and the same Charbonnier loss function as in step 1.
[0075] The implementation effect of this invention is compared with other traditional communication coding methods, as shown in the figure below. Figure 2 As shown in the figure, the 3D AP (Average Precision) score is an indicator for evaluating the accuracy of 3D detection, ranging from 0 to 100, with higher scores indicating higher recognition accuracy. The figure compares the proposed scheme with traditional communication coding schemes, namely 64QAM modulation with a 2 / 3 code rate and 256QAM modulation with a 1 / 2 code rate. The results show that the proposed scheme achieves higher 3D recognition accuracy for objects of varying recognition difficulty at different signal-to-noise ratios than traditional communication schemes, and approaches the performance limit under channel interference-free transmission.
[0076] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A semantic communication method for binocular 3D target detection, characterized in that... Includes the following steps: Step A: Input the acquired left and right visual images into the trained neural network to extract semantic information of key regions and global semantic information; Step B: Compress and encode the extracted semantic information at different compression rates to obtain the compressed key regions and global semantic information; Step C: Channel coding is performed on the compressed semantic information to obtain channel coded information, which is then transmitted to the receiver through the channel. Step D: Decode the received channel coding information. For the key region semantic information and global semantic information that are still in a compressed state after decoding, perform semantic recovery to obtain the key region semantic information and global semantic information. Step E: The restored key regions and global semantic information are fused to obtain the restored left and right visual images; Step D, which involves semantic restoration of the decoded, still compressed key region semantic information and global semantic information, specifically includes the following steps: Key region semantic information after channel decoding First, after one convolution, we get... The feature matrix; then through pixel rearrangement, the feature matrix is... Transformed into Next, the deformed feature matrix is convolved once, with the output dimension of the convolution being the same as the number of channels for the semantic information. Finally, the recovered semantic information is filled into the corresponding regions according to the received bounding box parameters, and the other regions are set to 0, thus obtaining the recovered key region semantic information. ; global semantic information after channel decoding First, optical flow extraction is performed to extract a pair of optical flow information. ,in Optical flow information from left to right. This represents the optical flow information from right to left; subsequently, it involves... Residual feature extraction yields a set of feature values. ;use right The features shown in the right image are obtained by performing optical deformation and one convolution. Simultaneously utilize right Optical deformation is performed and the data is passed through a convolutional layer to obtain the features shown in the left image. ;Will and Stack them along the channel dimension and then perform a single convolution to obtain... At the same time and Stack them along the channel dimension and then perform a single convolution to obtain... After that, and After each convolution, a pair of feature matrices are obtained. Then, through pixel rearrangement, the feature matrix is... Transformed into Finally, the deformed feature matrix is convolved once, with the output dimension of the convolution being the same as the number of channels for the semantic information, thus obtaining the recovered global semantic information. .
2. The semantic communication method for binocular 3D target detection according to claim 1, characterized in that: Step A specifically includes the following steps: The acquired pair of left and right visual images Input them into the neural network respectively; Visual images Two-dimensional object detection yields a set of bounding boxes for objects in the left and right images. Image information inside the bounding boxes is extracted from the two-dimensional object detection results, and feature extraction is performed on the extracted information to obtain semantic information of key regions. ; Visual images Features are extracted from two sets of convolutional residuals. The extracted left and right image features are directly concatenated along the channel dimension, and joint feature extraction is performed. The fused features are then separated to extract the semantic information of the left and right images. .
3. The semantic communication method for binocular 3D target detection according to claim 1, characterized in that: Step B specifically includes the following steps: extracting semantic information from key regions. With global semantic information The data are fed into two convolutional neural networks with different compression ratios, and then sequentially processed through a first convolution, pixel re-encoding, and a second convolution to obtain compressed semantic information of the key regions. and global semantic information The compression rate of global semantic information is greater than that of key region semantic information, i.e. .
4. The semantic communication method for binocular 3D target detection according to claim 3, characterized in that: The output dimension of the second convolution is the same as the number of channels for the semantic information of the key region or the global semantic information.
5. The semantic communication method for binocular 3D target detection according to claim 1, characterized in that: The channel coding of semantic information in step C specifically includes the following steps: The compressed key region semantic information and global semantic information Each input is fed into a convolutional neural network, and each input undergoes at least one convolutional encoding pass. The channel coding code rate for the final convolution is... The output dimension is ,in The channel coding rate is no greater than 1.
6. The semantic communication method for binocular 3D target detection according to claim 1, characterized in that: The decoding of the received channel-coded information in step D specifically includes the following steps: The information received after passing through the wireless channel is sequentially subjected to seven convolutions to achieve channel decoding; wherein the input dimension of the first convolution is... The output dimension of the last convolution is And obtain the semantic information of the key regions after channel decoding. and global semantic information .
7. The semantic communication method for binocular 3D target detection according to claim 1, characterized in that: Step E specifically includes the following steps: The restored semantic information of key regions and global semantic information Divide into two groups according to left and right characteristics. and After directly concatenating the images along the channel dimension, three convolutions are performed to obtain the reconstructed left and right visual images. .
8. An electronic device, characterized in that: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein... The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions that cause a processor to execute the method of any one of claims 1-7.