Audio-video conferencing system and method based on head-mounted device
Patent Information
- Application Number
- CN202610968820.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]MCU类的AR眼镜在算力方面存在限制,无法直接安装和运行大算力的音视频会议软件
[0016]本公开提供了一种基于头戴设备的音视频会议系统,包括头戴设备,用于向近端终端设备发送头戴设备采集到的第一视角下第一分辨率的第一环境图像及针对第一环境图像的第一语音信息;近端终端设备,用于接收第一环境图像和第一语音信息,对第一环境图像执行超分处理得到第二分辨率的第二环境图像,将第二环境图像和第一语音信息发送至远端终端设备,接收远端终端设备发送的针对第二环境图像的至少一个标注信息和与标注图像关联的第二语音信息,根据至少一个标注信息及第二环境图像,生成第三分辨率的标注图像,向头戴设备发送针对第二环境图像的标注图像及与第二语音信息;头戴设备还用于接收并输出标注图像及第二语音信息。基于该系统,算力及显示分辨率受限的头戴设备可借助近端终端设备完成与远端终端设备的音视频会议,这解决了头戴设备算力及显示分辨率不足无法实现音视频会议功能的问题。
Smart Images

Figure CN122824869A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of head-mounted display technology, and more specifically, to an audio and video conferencing system and method based on head-mounted devices. Background Technology
[0002] Currently, AR glasses based on microcontroller units (MCUs) dominate the market for augmented reality (AR) glasses.
[0003] MCU-based AR glasses have limitations in computing power, making it impossible to directly install and run high-performance audio and video conferencing software. Furthermore, AR glasses typically have low display resolutions, making it unable to properly display images with resolutions exceeding those of their own screens, while most images used in audio and video conferencing have resolutions higher than the display's own.
[0004] Therefore, MCU-based AR glasses typically cannot implement audio and video conferencing functions, which limits the intelligence of MCU-based AR glasses. Summary of the Invention
[0005] One objective of this disclosure is to provide a new technical solution for an audio and video conferencing system based on a head-mounted device.
[0006] According to a first aspect of this disclosure, an audio and video conferencing system based on a head-mounted device is provided, comprising: A head-mounted device is used to send a first environmental image at a first resolution from a first viewpoint, acquired by the head-mounted device, and first voice information related to the first environmental image to a near-end terminal device. The near-end terminal device is configured to receive the first environmental image and the first voice information, perform super-resolution processing on the first environmental image to obtain a second environmental image with a second resolution, send the second environmental image and the first voice information to the far-end terminal device, receive at least one annotation information for the second environmental image and second voice information associated with the annotation image sent by the far-end terminal device, generate a third resolution annotation image based on the at least one annotation information and the second environmental image, and send the annotation image for the second environmental image and the second voice information to the head-mounted device. The head-mounted device is also used to receive and output the labeled image and the second voice information.
[0007] Optionally, performing super-resolution processing on the first environmental image to obtain a second environmental image with a second resolution includes: The first environmental image is input into a preset image super-resolution model to obtain a second environmental image with a second resolution output by the image super-resolution model. The image super-resolution model includes a symmetrical encoder and decoder. The outputs of the first and second coding layers in the encoder are connected to corresponding cross-attention modules via the same skip layer. The output of the cross-attention module is connected to the input of the first decoding layer corresponding to the first coding layer. The first coding layer is the coding layer preceding the second coding layer. The cross-attention module is used to fuse the feature maps of corresponding scales output by the first and second coding layers to obtain a fused feature map and output it to the first decoding layer. The encoder is used to perform multi-level downsampling on the input image at a first resolution and output feature maps at different scales. The decoder is used to perform multi-level upsampling based on the fused feature map output by each cross-attention module and the feature map output by the last coding layer to output an output image at a second resolution.
[0008] Optionally, the image super-resolution model further includes a self-attention module located on the input side of the encoder. The self-attention module is used to perform global modeling and weighted fusion of the feature map of the input image to obtain a globally enhanced feature map of the input image, which is then used as the input of the encoder.
[0009] Optionally, the head-mounted device and the near-end terminal device establish a voice transmission channel and an image transmission channel, respectively. Sending the first environmental image at a first resolution from a first viewpoint acquired by the head-mounted device and first voice information relating to the first environmental image to the near-end terminal device includes: The first environmental image at a first resolution from a first viewpoint, acquired by the head-mounted device, is sent to the near-end terminal device through the image transmission channel; The first voice information for the first environmental image is sent to the near-end terminal device through the voice transmission channel. Sending the labeled image and the second voice information to the head-mounted device in relation to the second environmental image includes: The labeled image for the second environment image is sent to the head-mounted device through the image transmission channel; The second voice information is sent to the head-mounted device through the voice transmission channel.
[0010] Optionally, the image transmission channel is a wireless local area network channel, and the near-end terminal device is also used to generate and display an establishment image containing establishment information of the image transmission channel; The head-mounted device is also used to acquire the established image, identify the establishment information of the image transmission channel contained in the established image, and establish an image transmission channel with the near-end terminal device based on the establishment information.
[0011] Optionally, the near-end terminal device is further configured to: A partial image of the size corresponding to the third resolution, centered on the labeled area corresponding to the labeled information, is cropped from the second environmental image according to the labeled information; When the local image is a single frame, the local image is sent to the head-mounted device, and the head-mounted device is also used to store the local image; When the local image consists of at least two frames, the at least two frames of local images are stitched together to obtain a stitched image. The stitched image is then scaled down to the third resolution to obtain a scaled-down stitched image. Each local image and the scaled-down stitched image are then sent to the head-mounted device. The head-mounted device is also used to store the at least two frames of local images and the scaled-down stitched image.
[0012] Optionally, the head-mounted device is also used for: Maintain a hierarchical directory structure, wherein the first-level directory of the hierarchical directory structure is used to describe the image number of the first environmental image; In addition, when a frame of the local image is received, the local image is stored in the second-level directory under the first-level directory corresponding to the image number; Upon receiving at least two partial images and the thumbnail stitched image, the at least two partial images and the thumbnail stitched image are stored in the second-level directory of the first-level directory corresponding to the image number.
[0013] Optionally, the head-mounted device is also used for: Upon receiving the first user input and finding that a partial image is stored in the second-level directory, display the partial image in the second-level directory. And / or, upon receiving a first user input, and if at least two frames of partial images and a thumbnail-stitched image are stored under the second-level directory, the thumbnail-stitched image under the second-level directory is displayed; Upon receiving a second user input regarding the thumbnail stitched image, a partial image from the second-level directory is displayed.
[0014] According to a second aspect of this disclosure, a head-mounted device-based audio and video conferencing method is provided, applied to a head-mounted device in a head-mounted device-based audio and video conferencing system as described in any one of the first aspects, comprising: Sending a first environmental image at a first resolution from a first viewpoint and first voice information for the first environmental image to a near-end terminal device in an audio-visual conferencing system based on a head-mounted device as described in any one of the first aspects; Receive and output the labeled image and second voice information sent by the near-end terminal device.
[0015] According to a third aspect of this disclosure, a head-mounted device-based audio and video conferencing method is provided, applied to a near-end terminal device in a head-mounted device-based audio and video conferencing system as described in any one of the first aspects, comprising: Receive a first environmental image and first voice information sent by a head-mounted device in a head-mounted audio and video conferencing system as described in any one of the first aspects; Perform super-resolution processing on the first environmental image to obtain a second environmental image with a second resolution; Send the second environmental image and the first voice information to the remote terminal device; Receive at least one annotation information for the second environmental image and second voice information associated with the annotation image from the remote terminal device; Based on the at least one annotation information and the second environmental image, a third-resolution annotation image is generated, and the annotation image and the second voice information for the second environmental image are sent to the head-mounted device.
[0016] This disclosure provides an audio and video conferencing system based on a head-mounted device, including a head-mounted device for sending a first environmental image at a first resolution from a first viewpoint, captured by the head-mounted device, and first audio information related to the first environmental image to a near-end terminal device; the near-end terminal device for receiving the first environmental image and the first audio information, performing super-resolution processing on the first environmental image to obtain a second environmental image at a second resolution, sending the second environmental image and the first audio information to a far-end terminal device, receiving at least one annotation information for the second environmental image and second audio information associated with the annotation image sent by the far-end terminal device, generating a third-resolution annotation image based on the at least one annotation information and the second environmental image, and sending the annotation image for the second environmental image and the second audio information to the head-mounted device; the head-mounted device is also used to receive and output the annotation image and the second audio information. Based on this system, head-mounted devices with limited computing power and display resolution can complete audio and video conferencing with far-end terminal devices with the help of a near-end terminal device, which solves the problem that head-mounted devices cannot realize audio and video conferencing functions due to insufficient computing power and display resolution.
[0017] Other features and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments of the present disclosure and, together with their description, serve to explain the principles of the present disclosure.
[0019] Figure 1 This is a schematic diagram of the structure of an audio and video conferencing system based on a head-mounted device provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram illustrating the interaction between a head-mounted device and a near-end terminal device in an audio-visual conferencing system based on a head-mounted device, provided by an embodiment of this disclosure. Figure 3 This is a schematic diagram of the structure of a novel image super-resolution model provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of the encapsulation format of a proprietary protocol data frame provided in an embodiment of this disclosure. Detailed Implementation
[0020] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0021] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0022] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0023] In all the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0024] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0025] This disclosure provides an audio and video conferencing system based on a head-mounted device, wherein the head-mounted device is a low-computing-power and low-resolution head-mounted device that cannot run third-party audio and video conferencing software, such as MCU-based AR glasses. Figure 1As shown, the audio and video conferencing system based on a head-mounted device includes a head-mounted device and a near-end terminal device. The near-end terminal device acts as a relay device between the head-mounted device and the far-end terminal device, establishing communication connections with both. The head-mounted device and the far-end terminal device are the terminal devices for the two parties in the conference, respectively. Either the near-end terminal device or the far-end terminal device can be, for example, a personal computer (PC) or a smartphone.
[0026] The audio-visual system based on a head-mounted device disclosed herein can be applied to scenarios such as industrial maintenance and remote collaboration. For example, the head-mounted device captures first-person view environmental images and corresponding audio information, and sends them to a near-end terminal device. The near-end terminal device transmits the environmental images and audio captured by the head-mounted device to a remote terminal device, where an expert (participating on the remote terminal device's side) analyzes and annotates the received environmental images. The remote terminal device then sends the relevant annotation information and audio information back to the near-end terminal device. After processing the environmental images based on the annotation information, the near-end terminal device feeds back the processed environmental images and the audio information sent by the remote terminal device to the head-mounted device, which then displays the images and plays the audio, thereby achieving remote guidance.
[0027] like Figure 2 As shown, the head-mounted device is used to perform the following step S110.
[0028] Step S110: Send the first environmental image at a first resolution from a first viewpoint, captured by the head-mounted device, and the first voice information for the first environmental image to the near-end terminal device.
[0029] In this disclosure, the head-mounted device is equipped with an image sensor and a microphone sensor. The image sensor is used to acquire an environmental image of a first resolution containing the subjects of a meeting discussion from a first perspective, and the microphone sensor is used to acquire voice information of the wearer discussing the subjects of the meeting in the environmental image. In this embodiment, the environmental image acquired by the image sensor from the first perspective is recorded as a first environmental image, and the voice information of the wearer acquired by the microphone sensor is recorded as first voice information, and the first environmental image and the first voice information are associated. After acquiring the first environmental image and the first voice information, the head-mounted device sends the first environmental image and the first voice information to a near-end terminal device.
[0030] The near-end terminal device is used to perform the following steps S210 to S260.
[0031] Step S210: Receive the first environmental image and the first voice information.
[0032] Step S220: Perform super-resolution processing on the first environmental image to obtain a second environmental image with a second resolution.
[0033] After receiving the first environmental image and first voice information at the first resolution sent by the head-mounted device, the near-end terminal device first performs super-resolution processing on the first environmental image at the first resolution to obtain a second environmental image that allows the participants on the far-end terminal device to clearly view the details of the first environmental image.
[0034] In this embodiment of the disclosure, a preset image super-resolution model is deployed in the near-end terminal device. Specifically, this image super-resolution model is used to super-resolution an image of a first resolution into an image of a second resolution, where the second resolution is greater than the first resolution. Therefore, step S220 is specifically implemented as step S221 below.
[0035] Step S221: Input the first environment image into the preset image super-resolution model to obtain the second environment image with the second resolution output by the image super-resolution model.
[0036] In one embodiment of this disclosure, the image super-resolution model is a traditional image super-resolution model. In another embodiment of this disclosure, a novel image super-resolution model is provided. Figure 3 As shown, the image super-resolution model includes a symmetrical encoder and decoder. The outputs of the first and second coding layers in the encoder are connected to corresponding cross-attention modules via the same skip layer connection. The output of the cross-attention module is connected to the input of the first decoding layer corresponding to the first coding layer. The first coding layer is the coding layer preceding the second coding layer. The cross-attention module is used to fuse the feature maps of corresponding scales output by the first and second coding layers to obtain a fused feature map, which is then output to the first decoding layer. The encoder is used to perform multi-level downsampling on the input image at a first resolution, sequentially outputting feature maps at different scales. The decoder is used to perform multi-level upsampling based on the fused feature map output by each cross-attention module and the feature map output by the final coding layer, outputting an output image at a second resolution. It can be understood that the first coding layer is any coding layer in the encoder except for the final coding layer, and the specific structure of the cross-attention module can be implemented using the traditional cross-attention module structure.
[0037] The new image super-resolution model is specifically an improved U-Net deep learning model. The encoder includes multiple encoding layers, and the decoder, symmetrical to the encoder, includes the same number of decoding layers. The U-Net deep learning model disclosed herein adds at least one cross-attention module to the traditional U-Net deep learning model. This cross-attention module fuses the feature maps of corresponding scales output by the first and second encoding layers to obtain a fused feature map, which is then fed into the corresponding decoding layer. Compared to the traditional U-Net deep learning model, which only uses skip connections to input the feature map output by the first decoding layer into the corresponding first decoding layer, this achieves adaptive association and fusion of shallow and deep features. It solves the problems of feature redundancy and information fragmentation caused by directly inputting the encoding layer feature map into the corresponding decoding layer feature map. This strengthens key effective features for super-resolution, such as image contours, textures, and edges, while suppressing invalid noise and redundant features. Therefore, the improved U-Net deep learning model can achieve accurate super-resolution of the first environmental image.
[0038] Based on the above, the image super-resolution model provided in this disclosure can be exemplarily used as follows: Figure 3 As shown. Among them, Figure 3 The encoder consists of 4 encoding layers, the decoder consists of 4 decoding layers, and there are 3 cross-attention modules. The number of feature map channels in the encoding layers are 16, 64, 128 and 256 respectively, and the number of feature map channels in the decoding layers are 256, 128, 64 and 16 respectively, as shown in the example.
[0039] In another embodiment of this disclosure, such as Figure 3 As shown, the novel image super-resolution model provided in this disclosure further includes a self-attention module, located on the input side of the encoder. The self-attention module is used to globally model and weightedly fuse the feature maps of the input image to obtain a globally enhanced feature map of the input image, which serves as the input to the encoder. It is understood that the specific structure of the self-attention module can be implemented using the structure of a traditional self-attention module.
[0040] In this embodiment, the low-resolution first environment image itself has limited detail information. The self-attention module uses global weighted fusion to associate scattered and weakly correlated features in the first environment image, generating a globally enhanced feature map as input to the encoder. This allows the encoder to retain more semantic information and detail cues related to super-resolution reconstruction during subsequent downsampling, avoiding the loss of key features during downsampling.
[0041] It should be noted that the image super-resolution model in step S221 above is trained using a training sample set according to the traditional model training method. The training sample set includes multiple sets of training samples. Each training sample includes a sample environment image at a first resolution and a corresponding label environment image at a second resolution. The first resolution sample environment image is obtained by downsampling the second resolution label environment image.
[0042] Step S230: Send the second environmental image and the first voice information to the remote terminal device.
[0043] Step S240: Receive at least one annotation information for the second environmental image and second voice information associated with the annotation image from the remote terminal device.
[0044] After super-dividing the first environmental image sent by the head-mounted device into a second environmental image of second resolution, the near-end terminal device sends the second environmental image and the first voice information to the far-end terminal device. At this time, the far-end terminal device can receive the second environmental image and the first voice information sent by the near-end terminal device. Based on this, the far-end terminal device displays the second environmental image and plays the first voice information. Thus, the participants on the far-end terminal device side can view a clear second environmental image and hear the first voice information discussing the second environmental image. Based on the second environmental image and the first voice information, the participants on the far-end terminal device side perform annotation actions on the second environmental image through the far-end terminal device. Annotation actions include adding annotations such as circles, arrows, and text descriptions to the image. While performing annotation operations, the participants on the far-end terminal device side can explain the content and purpose of the annotations via voice. Based on this, the far-end terminal device recognizes the aforementioned annotation actions to obtain annotation information for the second image; simultaneously, the far-end terminal device recognizes the aforementioned voice to obtain the second voice information. Understandably, a complete annotation action corresponds to one annotation information, which is used to reproduce the annotation action and includes parameters such as the coordinate position and color of the identifier corresponding to the annotation action. Further, the remote terminal device sends at least one annotation information for the second image and second voice information to the near-end terminal device. Based on this, the near-end terminal device receives at least one annotation information for the second environmental image and second voice information associated with the annotated image sent by the remote terminal device.
[0045] Step S250: Generate a third-resolution labeled image based on at least one labeled information and a second environmental image.
[0046] The third resolution specifically refers to the display resolution of the head-mounted device, which is smaller than the second resolution.
[0047] Step S260: Send the labeled image of the second environment image and the second voice information to the head-mounted device.
[0048] After receiving at least one annotation information and a second environmental image from the remote terminal device, the near-end terminal device draws annotations added by the participants on the remote terminal device in the second environmental image based on the at least one annotation information, thus obtaining an initial annotation image at a second resolution. Due to the limitation of the head-mounted device's display resolution, the head-mounted device cannot display the initial annotation image at the second resolution. Therefore, the near-end terminal device downsamples the initial annotation image at the second resolution to obtain an annotation image at a third resolution consistent with the display resolution of the head-mounted device. Further, the third-resolution annotation image and the second voice information are sent to the head-mounted device. At this time, the head-mounted device can receive the annotation image and the second voice information. Furthermore, the head-mounted device can correctly display the annotation image that matches its own display resolution and play the second voice information through a speaker sensor. The head-mounted device also performs the following step S120.
[0049] Step S120: Receive and output the labeled image and the second voice information.
[0050] Based on the above step S120, the participants on the remote terminal device side can view the annotations and related voice descriptions of the first environmental image.
[0051] Through the above steps S110, S210 to S260 and S120, head-mounted devices with limited computing power and display resolution can complete audio and video conferencing with remote terminal devices using near-end terminal devices.
[0052] In the audio and video conferencing system based on head-mounted devices provided in this disclosure, the head-mounted device only needs to be responsible for acquiring low-resolution first environmental images and audio information, and displaying labeled images with the same display resolution as itself. The computationally intensive tasks such as image super-resolution processing and labeled image generation are completed by the near-end terminal device, thereby solving the problem that the head-mounted device cannot realize audio and video conferencing functions due to insufficient computing power and display resolution.
[0053] This disclosure provides an audio and video conferencing system based on a head-mounted device, including a head-mounted device for sending a first environmental image at a first resolution from a first viewpoint, captured by the head-mounted device, and first audio information related to the first environmental image to a near-end terminal device; the near-end terminal device for receiving the first environmental image and the first audio information, performing super-resolution processing on the first environmental image to obtain a second environmental image at a second resolution, sending the second environmental image and the first audio information to a far-end terminal device, receiving at least one annotation information for the second environmental image and second audio information associated with the annotation image sent by the far-end terminal device, generating a third-resolution annotation image based on the at least one annotation information and the second environmental image, and sending the annotation image for the second environmental image and the second audio information to the head-mounted device; the head-mounted device is also used to receive and output the annotation image and the second audio information. Based on this system, head-mounted devices with limited computing power and display resolution can complete audio and video conferencing with far-end terminal devices with the help of a near-end terminal device, which solves the problem that head-mounted devices cannot realize audio and video conferencing functions due to insufficient computing power and display resolution.
[0054] In one embodiment of this disclosure, the head-mounted device and the near-end terminal device establish separate voice transmission channels and image transmission channels. In one example, such as... Figure 1 As shown, the voice transmission channel is a Bluetooth Hands-Free Profile (BT HFP) channel, and the image transmission channel is a Wireless Local Area Network (WiFi) channel. Based on this, step S110 is specifically implemented through the following steps S111 and S112.
[0055] Step S111: Send the first environmental image at the first resolution from the first viewpoint, captured by the head-mounted device, to the near-end terminal device through the image transmission channel.
[0056] Step S112: Send first voice information for the first environmental image to the near-end terminal device through the voice transmission channel.
[0057] The above step S260 is specifically implemented through the following steps S261 and S262.
[0058] Step S261: Send the labeled image of the second environment image to the head-mounted device through the image transmission channel.
[0059] Step S262: Send the second voice information to the head-mounted device through the voice transmission channel.
[0060] In this embodiment, image data is large in volume and requires high bandwidth. For example, the image transmission channel of a WiFi network has high throughput capabilities, which can meet the stable uplink and downlink transmission requirements of images. Voice data has stringent requirements for transmission latency and real-time performance. For example, the voice transmission channel of the BT HFP uses a lightweight dedicated call protocol with low transmission latency and lower power consumption. Image data and voice data are transmitted in isolation using independent channels, which can prevent large-volume image data from crowding out voice data transmission resources, prevent voice stuttering and abnormal latency, and reduce the overall power consumption and transmission computing power of the head-mounted device.
[0061] In one embodiment of this disclosure, when the image transmission channel is a wireless local area network channel, the near-end terminal device is further configured to perform the following step S270.
[0062] Step S270: Generate and display an establishment image containing establishment information of the image transmission channel.
[0063] In this embodiment of the disclosure, triggered by a participant on the head-mounted device side, the near-end terminal device first connects to the router and simultaneously generates and displays an establishment image, such as a QR code, which includes the router username and password to which the near-end terminal device is connected, serving as establishment information for the image transmission channel.
[0064] Based on step S270 above, the head-mounted device is also used to perform step S130 below.
[0065] Step S130: Acquire and establish an image, identify the establishment information of the image transmission channel contained in the established image, and establish an image transmission channel with the near-end terminal device based on the establishment information.
[0066] In this embodiment, triggered by a participant on the head-mounted device, the head-mounted device activates its image sensor, uses its head-view sensor to acquire and establish an image, and identifies the establishment information of the image transmission channel within it. Further, based on the establishment information, it connects to the router to which the near-end terminal device is connected. At this time, the head-mounted device and the near-end terminal device are on the same local area network, enabling image transmission.
[0067] It is understood that step S270 is performed before step S210 described above. And step S130 is performed before step S110 described above.
[0068] In one embodiment of this disclosure, in order to facilitate subsequent backtracking, the near-end terminal device is also used to perform the following steps S280 to S2100.
[0069] Step S280: According to the annotation information, crop out a local image of the size corresponding to the third resolution from the second environmental image, centered on the annotation area corresponding to the annotation information.
[0070] In this embodiment of the disclosure, a labeled area corresponding to each labeled information is determined based on the coordinate position of the corresponding identifier included in each labeled information. Further, a local image of a third resolution size centered on each labeled area is cropped from the second environmental image.
[0071] In step S290, if the local image is a single frame, the local image is sent to the head-mounted device.
[0072] When a local image consists of only one frame, the near-end terminal device sends only this one frame of local image to the head-mounted device for storage.
[0073] In one embodiment of this disclosure, the near-end terminal device is further configured to divide the second voice information into sub-voice information associated with each annotation information based on at least one annotation information and the second voice information. Further, the near-end terminal device is specifically configured to associate the sub-voice information with the local image corresponding to the corresponding annotation information, and send the local image and its associated sub-voice information to the head-mounted device after binding them together.
[0074] Based on step S290 above, the head-mounted device is also used to perform step S140 below.
[0075] Step S140: Store the local image.
[0076] Through step S140 above, a partial image of the second environmental image corresponding to the annotation information can be viewed from the head-mounted device. It is understood that the resolution of the partial image is a third resolution, consistent with the display resolution of the head-mounted device, and the head-mounted device can display it correctly.
[0077] In addition to receiving locally bound sub-voice information, the head-mounted device can also store sub-voice information bound to local images. This allows the head-mounted device to also receive sub-voice information about local images.
[0078] In step S2100, if the local image consists of at least two frames, the at least two frames of local images are stitched together to obtain a stitched image. The stitched image is then scaled down to a third resolution to obtain a scaled-down stitched image. Each local image and the scaled-down stitched image are then sent to the head-mounted device.
[0079] Having obtained at least two frames of local images based on step S290 above, all local images are stitched together to obtain a stitched image. It is understood that the resolution of the stitched image is higher than the third resolution, and the head-mounted device cannot display it correctly. Therefore, the near-end terminal device downscales the stitched image to the third resolution to obtain a downscaled stitched image. Further, the near-end terminal device sends the downscaled stitched image and each local image to the head-mounted device for storage.
[0080] Based on the above step S2100, the head-mounted device is also used to perform the following step S150.
[0081] Step S150: Store at least two frames of partial images and thumbnail stitched images.
[0082] Through step S150 described above, the head-mounted device can view the local images in the second environmental image corresponding to the annotation information, as well as a thumbnail-stitched image representing all local images. By viewing the thumbnail-stitched image, the head-mounted device can determine how many frames of local images are present.
[0083] In one embodiment of this disclosure, the head-mounted device is also used to perform the following steps S160 to S170.
[0084] Step S160: Maintain the hierarchical directory structure.
[0085] The first-level directory in the hierarchical directory structure is used to describe the image number of the first environmental image.
[0086] Step S170: When a partial image frame is received, the partial image is stored in the second-level directory under the first-level directory corresponding to the image number.
[0087] Step S180: Upon receiving at least two partial images and a thumbnail stitched image, store the at least two partial images and the thumbnail stitched image in the second-level directory of the first-level directory corresponding to the image number.
[0088] Based on the above, the hierarchical directory structure can be specifically defined as follows: |--Photo (Root Directory) |-- 0001 (First-level directory, image number) |-- 0001.jpg (Second-level directory, partial image) |-- 0002 (Level 1 target, image number) |-- 0001.jpg (Second-level directory, partial image) |-- 0002.jpg (Second-level directory, partial image) |-- 0003.jpg (Second-level directory, partial image) |-- 0004.jpg (Second-level directory, partial image) |-- Thumbnail.jpg (Second-level directory, thumbnail stitched image) The hierarchical directory structure described above allows for quick and convenient viewing of partial images and thumbnail-stitched images.
[0089] In one embodiment of this disclosure, based on the hierarchical directory structure described above, the head-mounted device is further configured to perform the following steps 190 and / or S1100.
[0090] Step S190: Upon receiving the first user input and finding that a partial image is stored in the second-level directory, display the partial image in the second-level directory.
[0091] Step S1100: Upon receiving the first user input and with at least two frames of partial images and a thumbnail-stitched image stored in the second-level directory, display the thumbnail-stitched image in the second-level directory.
[0092] The first user input is the user's request to view images in the second-level directory. Upon receiving this first user input, it is determined that the user needs to view an image in the second-level directory. If a single partial image is stored in the second-level directory, that partial image will be displayed. If both partial and thumbnail-stitched images are stored in the second-level directory, the thumbnail-stitched image will be displayed.
[0093] Step S1110: Upon receiving a second user input for a thumbnail stitched image, display a partial image in the second-level directory.
[0094] The second user input involves selecting a corresponding partial image from the thumbnail-stitched images. Upon receiving the second user input, the partial image in the second-level directory is displayed.
[0095] It should be noted that the specific forms of the first and second user inputs described above are not limited in this embodiment. In one example, the first user input is a left-right swipe operation, and the second input is a click operation.
[0096] In one embodiment of this disclosure, image data is transmitted between the head-mounted device and the near-end terminal device using a proprietary protocol. The encapsulation format of the data frames in the proprietary protocol can be as follows: Figure 4As shown. The HEADER field represents the header of the image data frame; the LENGTH field represents the byte length occupied by the PAYLOAD; the PAYLOAD field represents the transmitted valid data; the CRC field represents the CRC checksum; the TAIL field represents the tail of the image data frame; the FILENAM field in the PAYLOAD represents the file name; the MARK-INDEX field represents the index number of the annotation image in the second image; the INDEX field represents the index number of the data packet for an annotation image when it needs to be sent multiple times; and the DATA field represents the valid data of the annotation image.
[0097] This disclosure also provides an audio and video conferencing method based on a head-mounted device, which is applied to the head-mounted device in the aforementioned audio and video conferencing system based on a head-mounted device, comprising: sending a first environmental image at a first resolution from a first viewpoint and first voice information for the first environmental image acquired by the head-mounted device to a near-end terminal device in the aforementioned audio and video conferencing system based on a head-mounted device; and receiving and outputting an annotated image and second voice information sent by the near-end terminal device.
[0098] It is understood that when the above-described audio and video conferencing method based on head-mounted devices is used in the above-described audio and video conferencing system based on head-mounted devices, it also includes other steps performed by the head-mounted devices in the above-described audio and video conferencing system based on head-mounted devices, which will not be elaborated here.
[0099] This disclosure also provides another audio and video conferencing method based on a head-mounted device, applied to a near-end terminal device in the aforementioned head-mounted device-based audio and video conferencing system, comprising: receiving a first environmental image and first voice information sent by a head-mounted device in the aforementioned head-mounted device-based audio and video conferencing system; performing super-resolution processing on the first environmental image to obtain a second environmental image of a second resolution; sending the second environmental image and the first voice information to a remote terminal device; receiving at least one annotation information for the second environmental image and second voice information associated with the annotation image sent by the remote terminal device; generating a third-resolution annotation image based on the at least one annotation information and the second environmental image; and sending the annotation image for the second environmental image and the second voice information to the head-mounted device.
[0100] It is understood that when the above-described audio and video conferencing method based on head-mounted devices is used in the near-end terminal device of the above-described audio and video conferencing system based on head-mounted devices, it also includes other steps performed by the near-end terminal device in the above-described audio and video conferencing system based on head-mounted devices, which will not be elaborated here.
[0101] This disclosure also provides a computer storage medium storing a computer program thereon, which, when executed by a processor, implements any of the audio and video conferencing methods based on head-mounted devices provided in the above-described method embodiments.
[0102] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0103] This disclosure may be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement any of the methods in the foregoing embodiments of this disclosure.
[0104] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media may include, for example, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), compact disc-read-only memory (CD-ROM), digital versatile disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any combination thereof. The computer-readable storage medium used herein is not to be interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0105] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include one or more of copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to computer-readable storage media in the respective computing / processing device.
[0106] The computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source or object programs written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++, etc.) and conventional procedural programming languages (such as the "C" language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network (e.g., a local area network or a wide area network), or it may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays, or programmable logic arrays, can execute computer-readable program instructions to implement various aspects of the embodiments of this disclosure by utilizing state information from the computer-readable program instructions.
[0107] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0108] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0109] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It should be noted that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are all equivalent.
[0111] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of this disclosure is defined by the appended claims.
Claims
1. An audio and video conferencing system based on a head-mounted device, characterized in that, include: A head-mounted device is used to send a first environmental image at a first resolution from a first viewpoint, acquired by the head-mounted device, and first voice information related to the first environmental image to a near-end terminal device. The near-end terminal device is configured to receive the first environmental image and the first voice information, perform super-resolution processing on the first environmental image to obtain a second environmental image with a second resolution, send the second environmental image and the first voice information to the far-end terminal device, receive at least one annotation information for the second environmental image and second voice information associated with the annotation image sent by the far-end terminal device, generate a third resolution annotation image based on the at least one annotation information and the second environmental image, and send the annotation image for the second environmental image and the second voice information to the head-mounted device. The head-mounted device is also used to receive and output the labeled image and the second voice information.
2. The system according to claim 1, characterized in that, The step of performing super-resolution processing on the first environmental image to obtain a second environmental image with a second resolution includes: The first environmental image is input into a preset image super-resolution model to obtain a second environmental image with a second resolution output by the image super-resolution model. The image super-resolution model includes a symmetrical encoder and decoder. The outputs of the first and second coding layers in the encoder are connected to corresponding cross-attention modules via the same skip layer. The output of the cross-attention module is connected to the input of the first decoding layer corresponding to the first coding layer. The first coding layer is the coding layer preceding the second coding layer. The cross-attention module is used to fuse the feature maps of corresponding scales output by the first and second coding layers to obtain a fused feature map and output it to the first decoding layer. The encoder is used to perform multi-level downsampling on the input image at a first resolution and output feature maps at different scales. The decoder is used to perform multi-level upsampling based on the fused feature map output by each cross-attention module and the feature map output by the last coding layer to output an output image at a second resolution.
3. The system according to claim 2, characterized in that, The image super-resolution model further includes a self-attention module located on the input side of the encoder. The self-attention module is used to perform global modeling and weighted fusion of the feature map of the input image to obtain a globally enhanced feature map of the input image, which is then used as the input of the encoder.
4. The system according to claim 1, characterized in that, The head-mounted device and the near-end terminal device establish a voice transmission channel and an image transmission channel, respectively. Sending the first environmental image at a first resolution from a first viewpoint, captured by the head-mounted device, and first voice information relating to the first environmental image to the near-end terminal device includes: The first environmental image at a first resolution from a first viewpoint, acquired by the head-mounted device, is sent to the near-end terminal device through the image transmission channel; The first voice information for the first environmental image is sent to the near-end terminal device through the voice transmission channel. Sending the labeled image and the second voice information to the head-mounted device in relation to the second environmental image includes: The labeled image for the second environment image is sent to the head-mounted device through the image transmission channel; The second voice information is sent to the head-mounted device through the voice transmission channel.
5. The system according to claim 4, characterized in that, The image transmission channel is a wireless local area network channel, and the near-end terminal device is also used to generate and display an establishment image containing the establishment information of the image transmission channel; The head-mounted device is also used to acquire the established image, identify the establishment information of the image transmission channel contained in the established image, and establish an image transmission channel with the near-end terminal device based on the establishment information.
6. The system according to claim 1, characterized in that, The near-end terminal device is also used for: A partial image of the size corresponding to the third resolution, centered on the labeled area corresponding to the labeled information, is cropped from the second environmental image according to the labeled information; When the local image is a single frame, the local image is sent to the head-mounted device, and the head-mounted device is also used to store the local image; When the local image consists of at least two frames, the at least two frames of local images are stitched together to obtain a stitched image. The stitched image is then scaled down to the third resolution to obtain a scaled-down stitched image. Each local image and the scaled-down stitched image are then sent to the head-mounted device. The head-mounted device is also used to store the at least two frames of local images and the scaled-down stitched image.
7. The system according to claim 6, characterized in that, The head-mounted device is also used for: Maintain a hierarchical directory structure, wherein the first-level directory of the hierarchical directory structure is used to describe the image number of the first environmental image; In addition, when a frame of the local image is received, the local image is stored in the second-level directory under the first-level directory corresponding to the image number; Upon receiving at least two partial images and the thumbnail stitched image, the at least two partial images and the thumbnail stitched image are stored in the second-level directory of the first-level directory corresponding to the image number.
8. The system according to claim 7, characterized in that, The head-mounted device is also used for: Upon receiving the first user input and finding that a partial image is stored in the second-level directory, display the partial image in the second-level directory. And / or, upon receiving a first user input, and if at least two frames of partial images and a thumbnail-stitched image are stored under the second-level directory, the thumbnail-stitched image under the second-level directory is displayed; Upon receiving a second user input regarding the thumbnail stitched image, a partial image from the second-level directory is displayed.
9. A method for audio and video conferencing based on a head-mounted device, characterized in that, The head-mounted device used in any one of the head-mounted device-based audio and video conferencing systems as described in any one of claims 1-8 includes: Sending a first environmental image at a first resolution from a first viewpoint and first voice information for the first environmental image to a near-end terminal device in an audio-visual conferencing system based on a head-mounted device as described in any one of claims 1-8; Receive and output the labeled image and second voice information sent by the near-end terminal device.
10. A method for audio and video conferencing based on a head-mounted device, characterized in that, A near-end terminal device applied in a head-mounted audio and video conferencing system as described in any one of claims 1-8, comprising: Receive a first environmental image and first voice information sent by a head-mounted device in an audio-visual conferencing system based on a head-mounted device as described in any one of claims 1-8; Perform super-resolution processing on the first environmental image to obtain a second environmental image with a second resolution; Send the second environmental image and the first voice information to the remote terminal device; Receive at least one annotation information for the second environmental image and second voice information associated with the annotation image from the remote terminal device; Based on the at least one annotation information and the second environmental image, a third-resolution annotation image is generated, and the annotation image and the second voice information for the second environmental image are sent to the head-mounted device.