Intelligent wearable device image transmission method and intelligent wearable device
Patent Information
- Application Number
- CN202411045166.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2044-07-31
AI Technical Summary
[0004]鉴于上述问题,本发明的目的是提供一种智能穿戴设备图像传输方法即方法,以解决现有图像传输方式存在的无法进行高分辨率传输,影响数据传输及用户体验等问题
[0015] By utilizing the aforementioned image transmission method and smart wearable device, deep learning is used to identify the category or scene of the image to be transmitted. Then, different image processing is performed on the images under different scenes, that is, different pixels and sharpness are adjusted for different images to obtain the transmitted image. The transmitted image is then transmitted to the target smart wearable device for decoding and display. Dynamic compression of the image can be achieved, thereby reducing the latency of image transmission and encoding/decoding, and improving the user experience of using the smart wearable device.
Smart Images

Figure CN121459231B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image data processing technology, and more specifically to an image transmission method for a smart wearable device and a smart wearable device. Background Technology
[0002] Augmented Reality (AR) technology, also known as augmented reality, is a relatively new technology that integrates real-world information with virtual-world information. After the real environment and virtual objects overlap, they can coexist in the same scene and space. In visual augmented reality, users need to wear head-mounted displays to make the real world and computer graphics overlap and be displayed.
[0003] Currently, in the practical use of smart wearable devices and AR devices such as AR glasses, streaming technology is typically used to project content from mobile phones or PCs onto the AR device for display in order to enrich content resources. However, because the overall latency of video codecs (including encoders and decoders) is directly related to the number of bits required for encoding, transmitting high-resolution images from PCs and mobile phones to AR devices is quite challenging. This results in transmission and encoding / decoding latency issues in existing solutions, which cannot meet the transmission requirements of high-resolution images in multiple scenarios. Summary of the Invention
[0004] In view of the above problems, the purpose of this invention is to provide an image transmission method for smart wearable devices, so as to solve the problems of existing image transmission methods that cannot transmit at high resolution, affecting data transmission and user experience.
[0005] The present invention provides an image transmission method for smart wearable devices, comprising: constructing a neural network model and training the neural network model based on an original training dataset until an image classification model is formed; inputting the image to be transmitted into the image classification model and outputting the classification result of the image to be transmitted through the image classification model; adjusting the pixels and sharpness of the image to be transmitted according to the classification result and obtaining the adjusted transmission image; transmitting the transmission image to the target smart wearable device and decoding and displaying it.
[0006] In addition, an optional technical solution is to output the classification result of the image to be transmitted through the image classification model, including: identifying the scene of the image to be transmitted through the image classification model; the scene division is related to the preset requirements of pixels and sharpness; and outputting the type of the image to be transmitted as the classification result according to the scene of the image to be transmitted.
[0007] In addition, an optional technical solution is to identify the scene of the image to be transmitted through the image classification model, including: extracting image features of the image to be transmitted through the image classification model, and selecting representative features related to a preset scene from them; and determining the scene corresponding to the image to be transmitted based on the representative features.
[0008] In addition, an optional technical solution is to adjust the pixels and sharpness of the image to be transmitted when the type of the image to be transmitted is a game, including: performing gaze point compression processing on the image to be transmitted; wherein, the gaze point compression includes: determining the gaze point of the image to be transmitted, and using the gaze point as the origin, dividing the image into cells in a proportional manner in the directions above, below, left, and right of the origin; and sampling each cell after division using a preset sampling rule to form the transmitted image.
[0009] In addition, an optional technical solution is that when the type of the image to be transmitted is a film or television image, the pixel and sharpness of the image to be transmitted are adjusted, including: performing a nearest neighbor similarity comparison on the image to be transmitted, and determining the image transmission method based on the comparison result; wherein, when the difference between the pixel value of any point and the pixel value of its neighboring pixels is less than a preset threshold, the image to be transmitted is transmitted by one eye; otherwise, the image to be transmitted is transmitted by both eyes.
[0010] In addition, an optional technical solution is to adjust the pixels and sharpness of the image to be transmitted when the type of the image to be transmitted is sports or art, including: compressing the image to be transmitted to form the transmitted image; the compression method includes at least one of mjpeg compression ratio, mpeg compression ratio, H264 or H265.
[0011] In addition, an optional technical solution is to adjust the pixels and sharpness of the image to be transmitted when the type of the image to be transmitted is other, including: dividing the image to be transmitted into uniform regions and obtaining the pixel values of the center points of each region after division; obtaining the similarity between the pixel values of adjacent regions; when the similarity is higher than a first preset value, compressing the image to be transmitted using a first compression ratio to form a transmitted image; when the similarity is not higher than the first preset value, compressing the image to be transmitted using a second compression ratio to form the transmitted image.
[0012] In addition, an optional technical solution is that the first compression ratio is not greater than 40M and the second compression ratio is not less than 80M.
[0013] In addition, an optional technical solution is to transmit the transmitted image to the target smart wearable device and decode and display it, including: sending the transmitted image to the target smart wearable device via WIFI or USB.
[0014] On the other hand, the present invention also provides a smart wearable device, which transmits images to the smart wearable device using the above-described smart wearable device image transmission method.
[0015] By utilizing the aforementioned image transmission method and smart wearable device, deep learning is used to identify the category or scene of the image to be transmitted. Then, different image processing is performed on the images under different scenes, that is, different pixels and sharpness are adjusted for different images to obtain the transmitted image. The transmitted image is then transmitted to the target smart wearable device for decoding and display. Dynamic compression of the image can be achieved, thereby reducing the latency of image transmission and encoding / decoding, and improving the user experience of using the smart wearable device.
[0016] To achieve the foregoing and related objectives, one or more aspects of the invention include the features that will be described in detail below. The following description and accompanying drawings illustrate certain exemplary aspects of the invention. However, these aspects indicate only a few of the various ways in which the principles of the invention can be used. Furthermore, the invention is intended to encompass all such aspects and their equivalents. Attached Figure Description
[0017] Other objects and results of the invention will become more apparent and readily understood with reference to the following description taken in conjunction with the accompanying drawings. In the drawings:
[0018] Figure 1 This is a flowchart of an image transmission method for a smart wearable device according to an embodiment of the present invention;
[0019] Figure 2 This is a schematic diagram illustrating the principle of gaze point compression cell division in an embodiment of the present invention.
[0020] Figure 3 This is a schematic diagram of the gaze point compression sampling according to an embodiment of the present invention;
[0021] Figure 4 This is a schematic diagram illustrating the principle of a nearest neighbor similarity comparison in an embodiment of the present invention. Detailed Implementation
[0022] In the following description, numerous specific details are set forth for illustrative purposes and to provide a thorough understanding of one or more embodiments. However, it will be apparent that these embodiments may also be implemented without these specific details. In other instances, well-known structures and devices are shown in block diagram form for ease of description of one or more embodiments.
[0023] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or couplings. The term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.
[0024] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.
[0025] To facilitate understanding of the embodiments of the present invention, the following will provide further explanation and description with reference to the accompanying drawings and several specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.
[0026] Figure 1 A schematic flowchart of an image transmission method for a smart wearable device according to an embodiment of the present invention is shown.
[0027] like Figure 1 As shown, the image transmission method for a smart wearable device according to an embodiment of the present invention includes:
[0028] S100: Construct a neural network model and train the neural network model based on the original training dataset until an image classification model is formed.
[0029] Specifically, the constructed neural network model includes an input layer, convolutional layers, pooling layers, and fully connected layers, and the specific network structure can be flexibly set according to the classification accuracy requirements.
[0030] The original training data can be obtained from publicly available image databases or from self-selected image datasets containing various types of images and involving multiple scenes, such as game images, film and television images, sports images, art images, and other types of images. The categories can also be understood as the application scenarios of the corresponding images. The neural network model is trained with a large amount of original training data until its iteration count meets the preset number of times or the model's loss function meets the preset threshold. At this point, the neural network model is considered to have completed training and formed an image classification model.
[0031] Specifically, the loss function during model training may include Mean Absolute Error (MAE), Mean Squared Error (MSE), Cross-Entropy Function, and Composite Loss Function. The Composite Loss Function is a function that combines at least two loss functions according to a certain coefficient or ratio; among which,
[0032] The formula for the mean absolute error is:
[0033]
[0034] The formula for expressing the mean squared error is:
[0035]
[0036] Where n represents the number of original training data points, y i Let y represent the i-th training data. i p This represents the predicted value of the i-th training data, i.e., the prediction result of image classification.
[0037] S200: Input the image to be transmitted into the image classification model, and output the classification result of the image to be transmitted through the image classification model.
[0038] Once the image classification model is trained, the image data to be transmitted can be input into the model. The model identifies the scene of the image to be transmitted, and the scene classification is related to preset requirements for the image's pixel count and sharpness. For example, if the pixel count and sharpness of the image to be transmitted meet the corresponding preset range requirements, it can be determined as a scene corresponding to the preset range. Alternatively, scene classification can be completed based on the recognition of objects in the image. For example, if objects such as balls or stadiums are identified in the image to be transmitted, the scene is determined to be sports; if objects related to games are identified, the scene is determined to be gaming. The specific scene recognition method can be set through the image classification model and is not limited to the judgment method based on pixels, sharpness, and related objects. For example, the image classification model can extract image features from the image to be transmitted and filter out representative features related to preset scenes to determine the scene corresponding to the image to be transmitted. Finally, the type of the image to be transmitted is output as the classification result based on the scene of the image to be transmitted. For example, the classification result (or type) of the image to be transmitted is output by the image classification model, including game, film and television, sports, art and other categories. The specific type division can be differentiated according to different smart wearable devices. Users can use the default classification form or make their own classification when using smart wearable devices.
[0039] As can be seen, the classification result of the image to be transmitted output by the image classification model is the prediction result of the model. The accuracy of the model will directly affect the accuracy of the classification result. Therefore, the neural network model of the present invention is proposed. By combining the processing of the original training dataset and the control of the number of model iterations, the prediction accuracy of the image classification model can be improved, and the subsequent processing and transmission of the classified image will also be more accurate.
[0040] S300: Adjust the pixels and sharpness of the image to be transmitted based on the classification results, and obtain the adjusted image for transmission.
[0041] Among them, game and sports images may require higher pixel density and clarity, so the compression rate can be slightly lower. Film and television images and art images can reduce pixel density to a certain extent to improve image transmission speed. Other types of images can be flexibly adjusted according to specific content and needs, and are not limited to specific processing methods. The following will describe the pixel and clarity adjustment schemes for each type of image.
[0042] In Example 1, when the type of image to be transmitted is a game, the pixel and sharpness adjustment scheme for the image to be transmitted is to perform foveation point compression processing. Foveation point compression includes: determining the foveation point of the image to be transmitted, and using the foveation point as the origin, dividing the image into cells in a proportional manner in the upward, downward, leftward, and rightward directions from the origin; sampling each cell after division using a preset sampling rule to form the transmitted image. For example, as... Figure 2 The game scene shown is a transmitted image with the gaze point as the origin, in a proportional form a n =a1*q n-1 The size changes divide the cells to the left, right, up, and down. n This represents the nth cell, where a1 represents the origin, and q represents the preset sampling parameters, which can be set according to requirements. The resulting cell structure is as follows: Figure 2 As shown in the grid lines, each cell is then sampled using a different sampling frequency. The sampling frequency 'a' of the nth cell is... n The sampling rate can be set to 1 / q n-1 The transmitted image formed after sampling is as follows: Figure 3 As shown.
[0043] It can be seen that the above-mentioned preset sampling rule can be based on the distance of the divided cells from the origin, according to 1 / q n-1 The changing trend is sampled, where n represents the nth cell in the left, right, up, and down directions from the origin.
[0044] Example 2: When the type of image to be transmitted is video, the pixel and sharpness of the image to be transmitted are adjusted, including: performing a nearest neighbor similarity comparison on the image to be transmitted, and determining the image transmission method based on the comparison result; wherein, when the difference between the pixel value of any point and the pixel value of the neighboring pixels is less than a preset threshold, the image to be transmitted is transmitted by one eye; otherwise, the image to be transmitted is transmitted by both eyes.
[0045] As an example, Figure 4 This illustrates the schematic principle of a single nearest neighbor similarity comparison according to an embodiment of the present invention, such as... Figure 4 As shown, the four pixels a1, a2, a3, and a4 are adjacent. During the comparison process, the difference between each pixel is first obtained: Avg1 = a2 - a1, Avg2 = a3 - a2, Avg3 = a3 - a1, Avg4 = a4 - a3... Then, a preset threshold B is set. If Avg1 < B, Avg2 < B, Avg3 < B, and Avg4 < B, the image to be transmitted can be transmitted by one eye; otherwise, it can be transmitted by both eyes.
[0046] In single-eye transmission, the image aberration between the two lenses of a VR device is relatively small. In this case, the images in the left and right lenses can be obtained through amplitude. However, in binocular transmission, the parallax between the images in the left and right lenses is larger.
[0047] Example 3: When the type of image to be transmitted is sports or art, the pixel and sharpness of the image to be transmitted are adjusted, including compression processing of the image to be transmitted to form the transmitted image; wherein, the compression method includes at least: compressing the image to be transmitted according to any one of mjpeg compression ratio, mpeg compression ratio, H264 or H265, and the specific compression method can be conventional compression, not limited to the above-mentioned mjpeg compression ratio and other methods.
[0048] Example 4: When the type of the image to be transmitted is other, the pixel and sharpness of the image to be transmitted are adjusted, including: comparing the pixel similarity of each pixel in the image to be transmitted, selecting the appropriate compression ratio based on the comparison result, and then compressing and transmitting the image. In this example, the image to be transmitted is first uniformly divided into several small parts of the same size; the smaller the division area, the better. Then, the pixel value of the center point of each divided area is obtained. Next, the similarity between the pixel values of adjacent areas is obtained. Similar to the nearest neighbor pixel comparison method in Example 1, when the similarity of the pixel values between adjacent areas is higher than a first preset value, the image to be transmitted is compressed using a first compression ratio to form the transmitted image; when the similarity is not higher than or lower than the preset value, the image to be transmitted is compressed using a second compression ratio to form the transmitted image, wherein the first compression ratio is less than the second compression ratio.
[0049] Specifically, the first preset value can be flexibly set according to the image transmission requirements or customer requirements. The first compression rate can be set to no more than 40M; the second compression rate can be set to no less than 80M.
[0050] As a specific example, the formula for the pixel compression ratio coefficient used in the compression processing of the image to be transmitted is: CR = (Q0 / Qt) * (B / (B+D)); where CR represents the compression ratio coefficient; Q0 represents the quality of the image to be transmitted, or the original image (which can be the peak signal-to-noise ratio (PSNR) or other quality metrics); Qt represents the quality of the transmitted image or the target image (the target quality is set according to different scenarios, and the metric used is the same as Q0); B represents the target bandwidth (available network bandwidth); and D represents network latency (a factor affecting transmission speed). Based on this formula, the pixel compression ratio coefficient can be determined, and then the pixel compression rate can be determined based on this coefficient to compress the image to be transmitted.
[0051] Specifically, in the compression process of game images and sports images, the parameters in the above pixel compression coefficient formula can be set as follows: Q0 is set to a higher value, such as 40-50dB PSNR, and Qt is also set to a higher value to maintain high image quality. B (higher) adopts a high bandwidth environment, and D (lower) adopts a low latency requirement. Assuming the corresponding compression coefficient is low, usually close to 1, the formula for pixel compression coefficient can be simplified to: CR = Q0 / Qt.
[0052] In the compression of video images, Q0 is set to a moderate value, Qt to a relatively low value, B to a moderate bandwidth environment, and D to a medium latency tolerance. The formula for the pixel compression ratio in this case is: CR = (Q0 / Qt) * (B / (B+D)). However, in the processing of other types of images, the parameters Q0 and Qt can be determined according to specific circumstances, B can be set based on network conditions, and D can be determined based on the image scene. The specific formula for the pixel compression ratio remains unchanged.
[0053] S400: Transmits the image to the target smart wearable device and decodes and displays it.
[0054] In this step, transmitting the image to the target smart wearable device and decoding and displaying it can include: sending the image to the target smart wearable device via various methods such as WIFI or USB, and then decoding and displaying the image through the smart wearable device.
[0055] Corresponding to the above-described image transmission method for smart wearable devices, the present invention also provides a smart wearable device that transmits images to the smart wearable device using the above-described image transmission method for smart wearable devices.
[0056] It should be noted that the embodiments of smart wearable devices can be referred to the description in the embodiments of the image transmission method of smart wearable devices, and will not be repeated here.
[0057] According to the image transmission method and smart wearable device of the present invention, the image to be transmitted is identified by category or scene through deep learning. Then, different image processing is performed on the image under different scenes, that is, different pixels and sharpness are adjusted for different images to obtain the transmitted image. The transmitted image is then transmitted to the target smart wearable device for decoding and display. Dynamic compression of the image can be achieved. High compression rate encoding is performed in some scenes and low compression rate encoding is performed in other scenes. This allows for selective guarantee of image quality or reduction of image transmission and encoding / decoding latency, thereby improving the user experience of using the smart wearable device.
[0058] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0059] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image transmission method for a smart wearable device, characterized in that, include: Construct a neural network model and train the neural network model based on the original training dataset until an image classification model is formed; The image to be transmitted is input into the image classification model, and the classification result of the image to be transmitted is output by the image classification model; Based on the classification results, the pixels and sharpness of the image to be transmitted are adjusted, and the adjusted image to be transmitted is obtained. The transmitted image is transmitted to the target smart wearable device and then decoded and displayed. The image classification model outputs the classification result of the image to be transmitted, including: The scene of the image to be transmitted is identified by the image classification model; the scene division is related to the preset requirements of pixels and sharpness. The type of the image to be transmitted is output as the classification result based on the scene of the image to be transmitted; When the image to be transmitted is a game, the pixel and sharpness of the image to be transmitted are adjusted, including: The image to be transmitted is subjected to gaze compression processing; wherein, the gaze compression includes: Determine the gaze point of the image to be transmitted, and use the gaze point as the origin to divide the image into cells in a proportional manner in the directions above, below, left, and right of the origin; Each cell after division is sampled using a preset sampling rule to form the transmitted image.
2. The image transmission method for a smart wearable device as described in claim 1, characterized in that, Identifying the scene of the image to be transmitted using the image classification model includes: The image features of the image to be transmitted are extracted using the image classification model, and representative features related to a preset scene are selected from them. Based on the representative features, the scene corresponding to the image to be transmitted is determined.
3. The image transmission method for a smart wearable device as described in claim 1, characterized in that, When the type of the image to be transmitted is video or audio, the pixel count and resolution of the image to be transmitted are adjusted, including: A nearest neighbor similarity comparison is performed on the image to be transmitted, and the image transmission method is determined based on the comparison result; wherein, When the difference between the pixel value of any point and the pixel values of its neighboring points is less than a preset threshold, the image to be transmitted is transmitted in one eye. Otherwise, the image to be transmitted is transmitted through both eyes.
4. The image transmission method for a smart wearable device as described in claim 1, characterized in that, When the type of the image to be transmitted is sports or art, the pixel and sharpness of the image to be transmitted are adjusted, including: The image to be transmitted is compressed to form the transmitted image; the compression method includes at least one of mjpeg compression ratio, mpeg compression ratio, H264, and H265.
5. The image transmission method for a smart wearable device as described in claim 1, characterized in that, When the type of the image to be transmitted is other, the pixel and sharpness of the image to be transmitted are adjusted, including: The image to be transmitted is divided into uniform regions, and the pixel values of the center points of each region are obtained. Obtain the similarity between pixel values of adjacent regions; When the similarity is higher than the first preset value, the image to be transmitted is compressed using the first compression ratio to form the transmitted image; When the similarity is not higher than the first preset value, the image to be transmitted is compressed using a second compression rate to form the transmitted image.
6. The image transmission method for a smart wearable device as described in claim 5, characterized in that, The first compression ratio is no greater than 40M, and the second compression ratio is no less than 80M.
7. The image transmission method for a smart wearable device as described in claim 1, characterized in that, Transmitting the transmitted image to the target smart wearable device and decoding and displaying it includes: The transmitted image is sent to the target smart wearable device via WIFI or USB.
8. A smart wearable device, characterized in that, An image is transmitted to the smart wearable device using the image transmission method for a smart wearable device as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Remote sensing image compression algorithm based on deep attention network and scene perception
CN115511983A
Image scene classification method and device, computer equipment and storage medium
CN115908961A
Electronic device, head-mounted display, gaze point detector, and pixel data readout method
US20200412983A1