Audio and video call adjustment method, adjustment device, AR device and storage medium

By continuously monitoring the network signal quality of AR devices and flexibly selecting the push and pull methods of audio and video streams, the problem of unsmoothness caused by insufficient network bandwidth in remote collaborative calls with AR devices is solved, achieving smooth audio and video transmission and a rich communication experience.

CN117278709BActive Publication Date: 2025-09-12BEIJING DIANJIEZHI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311266472.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2025-09-12
Estimated Expiration
2043-09-27

AI Technical Summary

Technical Problem

During remote collaboration calls, AR devices experience choppy and jammed calls due to insufficient network bandwidth, affecting user experience.

Method used

By continuously monitoring the signal quality of the AR device's access to the network, different audio and video streams are pushed and pulled based on the signal strength, including local audio stream, remote audio stream, local audio and video stream, or remote audio and video stream. This ensures that only audio streams are pushed and pulled when the network signal is weak, and audio and video streams are pushed and pulled simultaneously when the network signal is good.

Benefits of technology

It achieves smooth audio and video transmission under unstable network signal conditions, improves user experience, ensures the continuity of audio communication, and provides a rich audio and video communication experience when the network quality is good.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117278709B_ABST
    Figure CN117278709B_ABST
Patent Text Reader

Abstract

The present disclosure provides an adjustment method, an adjustment device, an AR device and a storage medium for audio and video calls, and relates to the field of real-time call technology. Among them, the adjustment method for audio and video calls includes: in response to the establishment operation of an audio and video call for online guidance, continuously monitoring the quality of the network signal of the AR device accessing the network; if the signal strength of the network signal is detected to be less than a first strength threshold, executing the push operation of the local audio stream and the pull operation of the remote audio stream; if it is detected that the signal strength of the network signal is greater than or equal to a second strength threshold, executing the push operation of the local audio and video stream and the pull operation of the remote audio and video stream, and the second strength threshold is greater than or equal to the first strength threshold. The technical solution of the present disclosure is conducive to achieving smooth switching of audio and video transmission to provide a better audio and video call experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of real-time call technology, and in particular to a method for adjusting an audio or video call, an apparatus for adjusting an audio or video call, an AR device, and a computer-readable storage medium. Background Art

[0002] With the development of virtual reality technology, AR devices are being used more and more widely. However, applications such as remote calls on AR devices rely heavily on wireless networks. If no network coverage is detected, the device will attempt to reconnect. However, if the network bandwidth is low, even if the device is reconnected, the remote collaborative call may experience issues such as lag and freezing, which will affect the user experience.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a method for adjusting audio and video calls, an apparatus for adjusting audio and video calls, an AR device, and a computer-readable storage medium, which can at least to some extent improve the problem of poor user experience of AR devices during remote collaborative calls in related technologies.

[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0006] According to one aspect of the present disclosure, a method for adjusting an audio and video call is provided, comprising: in response to an operation of establishing an audio and video call for online guidance, continuously monitoring the quality of a network signal of the AR device accessing the network; if it is detected that the signal strength of the network signal is less than a first strength threshold, executing a push operation of a local audio stream and a pull operation of a remote audio stream; if it is detected that the signal strength of the network signal is greater than or equal to a second strength threshold, executing a push operation of a local audio and video stream and a pull operation of a remote audio and video stream, the second strength threshold being greater than or equal to the first strength threshold.

[0007] In one embodiment, if the signal strength of the network signal is monitored to be less than a first strength threshold, the local audio stream is pushed and the remote audio stream is pulled, including: during the audio and video call, it is monitored that the signal strength of the network signal drops to less than the first strength threshold, and a first prompt window pops up in the window of the AR device, the first prompt window is used to prompt to stop pulling the remote video stream; in response to the confirmation operation of the first prompt window, the pulled remote audio stream is played.

[0008] In one embodiment, if the signal strength of the monitored network signal is less than a first strength threshold, the local audio stream push and remote audio stream pull operations are executed, which also includes: calling an offline semantic recognition model to identify the pulled remote audio stream to obtain a first guidance semantic recognition result; calling an offline virtual modeling model to model the first guidance semantic recognition result, and generating an augmented reality guidance image based on the modeling result and the image collected in real time by the AR device; and displaying the augmented reality guidance image in the window of the AR device.

[0009] In one embodiment, the calling of the offline virtual modeling model to model the first guidance semantic recognition result, and generating an augmented reality guidance image based on the modeling result and the image collected in real time by the AR device, includes: determining a modeling scene based on the first guidance semantic recognition result and the real-time collected image; determining the guidance intention and guidance requirements based on the first guidance semantic recognition result; performing three-dimensional modeling in the modeling scene based on the guidance intention and the guidance requirements to obtain the modeling result; and mapping the modeling result to the real-time collected image to obtain the augmented reality guidance image.

[0010] In one embodiment, if it is detected that the signal strength of the network signal is greater than or equal to a second strength threshold, the push operation of the local audio and video stream and the pull operation of the remote audio and video stream are executed, which also includes: monitoring that the signal strength of the network signal rises to greater than or equal to the first strength threshold and less than the second strength threshold, starting to push the local video stream; monitoring that the signal strength of the network signal rises to greater than or equal to the second strength threshold and lasts for a duration greater than or equal to the first duration threshold, starting to pull the remote video stream; and calling the remote semantic recognition model to perform real-time recognition of the pulled remote audio stream to obtain a second guiding semantic recognition result; based on the second guiding semantic recognition result, alternatingly playing the pulled remote video stream and the collected local video stream in the window of the AR device.

[0011] In one embodiment, the pulled remote video stream and the collected local video stream are alternately played in the window of the AR device based on the second guidance semantic recognition result, including: if the second guidance semantic recognition result includes a device-related recognition result in the local video stream, playing the local video stream and the remote audio stream; if during the playback of the local video stream, the second guidance semantic recognition result identified in real time does not include the device-related recognition result, when the continuous playback time is greater than or equal to a second time threshold, switching to playing the remote audio and video stream.

[0012] In one embodiment, the alternating playback of the pulled remote video stream and the captured local video stream in the window of the AR device based on the second guiding semantic recognition result also includes: if a playback configuration gesture for the remote video stream and the local video stream is obtained, configuring the playback mode of the remote video stream and the local video stream based on the playback configuration gesture, wherein the playback mode includes any one of continuously playing the remote video stream, continuously playing the local video stream, and playing the remote video stream and the local video stream in parallel.

[0013] In one embodiment, before continuously monitoring the quality of the network signal of the AR device accessing the network in response to the establishment operation of the audio and video call for online guidance, it also includes: joining the remote collaboration meeting and registering a broadcast receiver BroadcastReceiver to continuously monitor the quality of the network signal based on the BroadcastReceiver.

[0014] In one embodiment, in response to the operation of establishing an audio or video call for online guidance, the quality of the network signal of the AR device accessing the network is continuously monitored, and the method also includes: if the network connection is disconnected based on the BroadcastReceiver, a second prompt window pops up in the window of the AR device, and the second prompt window is used to prompt the network disconnection; based on the BroadcastReceiver, network change events are continuously monitored to detect whether network disconnection and reconnection are achieved; if the monitoring time continues to be greater than or equal to a third time threshold and network disconnection and reconnection are not achieved, the network disconnection and reconnection are stopped to remove the AR device from the remote collaboration meeting.

[0015] In one embodiment, if the BroadcastReceiver monitors that the network connection is disconnected, it also includes: obtaining the user identification number of the call target; based on the user identification number, switching the push operation of the local audio stream and the pull operation of the remote audio stream based on Internet voice communication to the push operation of the local audio stream and the pull operation of the remote audio stream based on mobile communication.

[0016] According to another aspect of the present disclosure, a device for adjusting an audio and video call is provided, including: a monitoring module for continuously monitoring the quality of a network signal of the AR device accessing the network in response to an operation of establishing an audio and video call for online guidance; a first execution module for executing a push operation of a local audio stream and a pull operation of a remote audio stream if it is monitored that the signal strength of the network signal is less than a first strength threshold; and a second execution module for executing a push operation of a local audio and video stream and a pull operation of a remote audio and video stream if it is detected that the signal strength of the network signal is greater than or equal to a second strength threshold, the second strength threshold being greater than or equal to the first strength threshold.

[0017] According to another aspect of the present disclosure, an AR device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the method for adjusting the audio and video call described in any one of the first aspect above, or the method for adjusting the audio and video call described in any one of the second aspect above, by executing the executable instructions.

[0018] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any of the above-mentioned methods for adjusting audio and video calls is implemented.

[0019] The adjustment scheme for audio and video calls provided by the embodiments of the present disclosure can timely understand the strength of the network signal by continuously monitoring the quality of the network signal of the AR device accessing the network. According to different signal strengths, it is possible to flexibly choose to push local audio streams, remote audio streams, local audio and video streams or remote audio and video streams, thereby realizing real-time audio and video calls. In the case of weak network signals, only pushing audio streams and pulling audio streams can ensure the continuous passage of audio. In the case of good network signals, audio and video streams can be pushed and pulled at the same time to provide a richer communication experience. By continuously monitoring the network signal strength and selecting appropriate audio and video streams for pushing and pulling according to different signal strengths, while ensuring the real-time nature of data push and its adaptability to network quality, it is also conducive to achieving smooth switching of audio and video transmission, thereby providing a better audio and video call experience.

[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0022] Figure 1 A schematic diagram showing the structure of an audio and video call adjustment system according to an embodiment of the present disclosure;

[0023] Figure 2 A flowchart illustrating a method for adjusting an audio or video call in an embodiment of the present disclosure is shown;

[0024] Figure 3 A flowchart illustrating another method for adjusting an audio or video call in an embodiment of the present disclosure is provided;

[0025] Figure 4 A flowchart illustrating another method for adjusting an audio or video call according to an embodiment of the present disclosure is provided;

[0026] Figure 5 A flowchart illustrating another method for adjusting an audio or video call in an embodiment of the present disclosure is provided;

[0027] Figure 6 A flowchart illustrating another method for adjusting an audio or video call in an embodiment of the present disclosure is provided;

[0028] Figure 7 A schematic diagram illustrating an apparatus for adjusting audio and video calls according to an embodiment of the present disclosure is shown;

[0029] Figure 8 A schematic diagram of an AR device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0030] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0031] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0032] The solution provided by this application can timely understand the strength of the network signal by continuously monitoring the quality of the network signal of the AR device accessing the network. According to different signal strengths, you can flexibly choose to push local audio streams, remote audio streams, local audio and video streams or remote audio and video streams, so as to realize real-time audio and video calls. When the network signal is weak, only pushing audio streams and pulling audio streams can ensure the continuous passage of audio. When the network signal is good, audio and video streams can be pushed and pulled at the same time to provide a richer communication experience. By continuously monitoring the network signal strength and selecting appropriate audio and video streams for pushing and pulling according to different signal strengths, while ensuring the real-time nature of data push and its adaptability to network quality, it is also conducive to achieving smooth switching of audio and video transmission, thereby providing a better audio and video call experience.

[0033] Figure 1 FIG. 1 is a schematic diagram of a computer system according to an exemplary embodiment of the present application. The system includes: a plurality of AR devices 120 , a plurality of terminals 140 , and a server cluster 160 .

[0034] The terminal 140 may be a mobile terminal such as an AR (Augmented Reality) device, a mobile phone, a game console, a tablet computer, an e-book reader, smart glasses, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a smart home device, or a VR (Virtual Reality) device. Alternatively, the terminal 140 may be a personal computer (PC), such as a laptop or desktop computer.

[0035] The AR device 120 and the terminal 140 may be installed with an application for providing adjustment of audio and video calls.

[0036] The AR device 120, the terminal 140, and the server cluster 160 are connected via a communication network to implement real-time communication between the AR device 120 and the terminal 140. Optionally, the communication network is a wired network or a wireless network.

[0037] Server cluster 160 is a single server, or a combination of multiple servers, a virtualization platform, or a cloud computing service center. Server cluster 160 provides backend services for applications that provide statistics on service products. Optionally, server cluster 160 performs primary computing tasks, while AR devices 120 and terminals 140 perform secondary computing tasks. Alternatively, server cluster 160 performs secondary computing tasks, while AR devices 120 and terminals 140 perform primary computing tasks. Alternatively, a distributed computing architecture is employed to enable collaborative computing between AR devices 120, terminals 140, and server cluster 160.

[0038] In some optional embodiments, the server cluster 160 is used to store adjustment program information for audio and video calls.

[0039] Optionally, the application client installed on different AR devices 120 and terminals 140 is the same, or the application clients installed on the two AR devices 120 and terminals 140 are clients of the same type of application on different control system platforms. Based on different terminal platforms, the specific form of the application client may also vary. For example, the application client may be a mobile phone client, a PC client, or a World Wide Web (Web) client.

[0040] Those skilled in the art will appreciate that the number of the AR devices 120 and terminals 140 may be greater or lesser. For example, there may be only one terminal, or there may be dozens, hundreds, or even more terminals. The embodiments of this application do not limit the number and type of terminals.

[0041] Optionally, the system may further include a management device ( Figure 1 (not shown), the management device is connected to the server cluster 160 via a communication network. Optionally, the communication network is a wired network or a wireless network.

[0042] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.

[0043] Hereinafter, the method for adjusting an audio and video call based on interaction between an AR device and a terminal in this exemplary implementation will be described in more detail with reference to the accompanying drawings and embodiments.

[0044] like Figure 2 As shown, a method for adjusting an audio and video call according to an embodiment of the present disclosure is applied to an AR device, including:

[0045] Step S202 : In response to the establishment operation of the audio and video call for online guidance, the quality of the network signal of the AR device accessing the network is continuously monitored.

[0046] Among them, the user uses the AR device to establish an audio or video call with the instructor through the network, which can be achieved through an application on the AR device or specific communication software. The audio and video data between the user and the instructor are transmitted through the network for communication.

[0047] After the audio and video call is established, the AR device can use the built-in camera and sensors to detect the user's line of sight. This can be achieved through the AR device's eye tracking technology or other related technologies. Line of sight detection can determine the direction and target the user is looking at to provide guidance to experts.

[0048] Furthermore, AR devices can use network connection information to monitor the quality of network signals. The detection parameters include but are not limited to detecting network latency, bandwidth, packet loss rate and other indicators. AR devices can use built-in network monitoring tools or APIs related to network connections to obtain information on the quality of network signals.

[0049] AR devices can continuously monitor network access quality and provide feedback to users and instructors based on the monitoring results. Feedback can be provided through the user interface or sound prompts on the AR device. For example, when network access quality deteriorates, the AR device can send warnings or suggestions to users and instructors to improve call quality. Continuous monitoring of network access quality by AR devices during audio and video calls helps provide a better call experience and guidance effect.

[0050] Step S204: If the signal strength of the monitored network signal is less than the first strength threshold, push operations of the local audio stream and pull operations of the remote audio stream are executed.

[0051] Among them, if the signal strength is less than the preset first strength threshold, it means that the network connection is weak. When the network signal strength is detected to be weak, the AR device will perform a push operation of the local audio stream, which means that the AR device will capture and encode the user's audio data from the microphone, and then send it to the expert through the network.

[0052] At the same time, the AR device will also perform remote audio stream pulling operations, that is, the AR device will obtain the audio data sent by the expert from the network, decode and play it. The above operations can provide better audio call quality and improve the user experience when the network connection is unstable or the signal is weak.

[0053] Step S206: If it is detected that the signal strength of the network signal is greater than or equal to the second strength threshold, push operations of the local audio and video stream and pull operations of the remote audio and video stream are executed, and the second strength threshold is greater than or equal to the first strength threshold.

[0054] Among them, the first strength threshold is used to represent the situation where the network signal is weak, and only audio stream is used for calls, while the second strength threshold is used to represent the situation where the network signal is good, and both audio and video streams can be used for calls. The second strength threshold is greater than the first strength threshold, ensuring that users can enjoy a higher quality audio and video call experience when the network signal is good.

[0055] In this embodiment, by continuously monitoring the network access quality of the AR device to the network, the strength of the network signal can be timely understood. According to different signal strengths, it is possible to flexibly choose to push local audio streams, remote audio streams, local audio and video streams, or remote audio and video streams, thereby realizing real-time audio and video calls. In the case of weak network signals, only pushing audio streams and pulling audio streams can ensure the continuous passage of audio. In the case of good network signals, audio and video streams can be pushed and pulled at the same time to provide a richer communication experience. By continuously monitoring the network signal strength and selecting appropriate audio and video streams for pushing and pulling according to different signal strengths, while ensuring the real-time nature of data push and adaptability to network quality, it is also conducive to achieving smooth switching of audio and video transmission, thereby providing a better audio and video call experience.

[0056] In one embodiment, if the signal strength of the monitored network signal is less than a first strength threshold, the local audio stream is pushed and the remote audio stream is pulled, including: during the audio and video call, if the signal strength of the monitored network signal drops to less than the first strength threshold, a first prompt window pops up in the window of the AR device, and the first prompt window is used to prompt to stop pulling the remote video stream; in response to the confirmation operation of the first prompt window, the pulled remote audio stream is played.

[0057] In this embodiment, since the network signal is weak, stable video transmission cannot be guaranteed. When the signal strength of the network signal is monitored to be less than the first strength threshold, a first prompt window pops up in the window of the AR device to prompt to stop pulling the remote video stream. After the user confirms the operation, the pulled remote audio stream can be played. Playing only the audio stream can provide basic voice communication functions and reduce dependence on the network. Even in the case of poor network signal, the voice call function can still be achieved by playing the audio stream.

[0058] In addition, by popping up the first prompt window and playing the audio stream, the user is informed that the network signal is weak, reminding the user to pay attention to the current network status. This helps the user understand the communication quality and may take appropriate measures, such as changing location, adjusting network settings or looking for a better network environment.

[0059] like Figure 3 As shown, according to another embodiment of the present disclosure, a method for adjusting an audio or video call is applied to an AR device, including:

[0060] Step S302: If the signal strength of the monitored network signal is less than a first strength threshold, push operations of the local audio stream and pull operations of the remote audio stream are executed.

[0061] Step S304: calling an offline semantic recognition model to recognize the pulled remote audio stream to obtain a first guiding semantic recognition result.

[0062] Among them, through the offline semantic recognition model, the remote audio stream can be converted into semantic information, which can better understand the other party's language intention and improve the accuracy and efficiency of communication. For example, the voice can be converted into text to facilitate user reading or more accurate replies.

[0063] In addition, through the recognition results of the offline semantic recognition model, the recognized semantic information can be fed back to the user in real time. The user can instantly understand the other party's intentions through text or other forms of display, thereby better communicating.

[0064] Step S306: Calling an offline virtual modeling model to model the first guidance semantic recognition result, and generating an augmented reality guidance image based on the modeling result and the image collected in real time by the AR device.

[0065] Among them, by converting the remote audio stream into semantic information, it can be combined with the real-time image captured by the AR device to generate an augmented reality guidance image, allowing users to observe the virtual guidance information through the AR device, thereby more intuitively understanding the other party's intentions and performing corresponding operations. This interactive method can provide a more comprehensive and rich user experience.

[0066] Step S308: Displaying an augmented reality guidance image in a window of the AR device.

[0067] In this embodiment, under conditions where the network signal is weak, modeling the first guidance semantic recognition result based on the offline virtual modeling model can more accurately understand the semantic information and generate corresponding augmented reality guidance images based on the real-time collected images, which can help improve the communication capabilities and effects, increase interactivity, and improve the accuracy and quality of communication.

[0068] In one embodiment, an offline virtual modeling model is called to model the first guidance semantic recognition result, and an augmented reality guidance image is generated based on the modeling result and the image collected in real time by the AR device, including: determining a modeling scene based on the first guidance semantic recognition result and the real-time collected image; determining the guidance intention and guidance requirements based on the first guidance semantic recognition result; performing three-dimensional modeling in the modeling scene based on the guidance intention and guidance requirements to obtain a modeling result; and mapping the modeling result to the real-time collected image to obtain an augmented reality guidance image.

[0069] In this embodiment, the first guidance semantic recognition result can provide semantic information about the scene, and the real-time collected image can provide actual visual information of the scene. The combination of the two can more accurately determine the modeling scene. By determining the modeling scene, three-dimensional modeling can be performed in the scene to obtain the modeling result. The user can intuitively observe and perceive the modeling result through the augmented reality guidance image, which enhances the interactivity. When the network conditions fail to provide a smooth video stream, more intuitive guidance information can be provided, and corresponding guidance information can be fed back to the user in real time. Three-dimensional modeling is performed in the modeling scene, and the modeling results are mapped to the real-time collected image. Accurate modeling information can be provided, and guidance information can be fed back in real time, thereby enhancing interactivity and intuitiveness, and helping users to better operate and make decisions.

[0070] like Figure 4 As shown, a method for adjusting an audio or video call according to another embodiment of the present disclosure is applied to an AR device, including:

[0071] Step S402: If the signal strength of the monitored network signal is less than the first strength threshold, push operations of the local audio stream and pull operations of the remote audio stream are executed.

[0072] Step S404: When the signal strength of the network signal is monitored to be greater than or equal to the first strength threshold and less than the second strength threshold, push of the local video stream is started.

[0073] Step S406: When the signal strength of the network signal is monitored to be greater than or equal to the second strength threshold and the duration thereof is greater than or equal to the first duration threshold, the remote video stream is started to be pulled.

[0074] Step S408: calling a remote semantic recognition model to perform real-time recognition on the pulled remote audio stream to obtain a second guiding semantic recognition result.

[0075] Step S410 : Based on the second guiding semantic recognition result, the pulled remote video stream and the captured local video stream are alternately played in the window of the AR device.

[0076] In this embodiment, the push and pull of local audio and video streams and remote audio and video streams are realized according to the changes in network signal strength and duration. At the same time, the remote audio stream is recognized in real time through the remote semantic recognition model, and then the remote and local video streams are played alternately in the window according to the semantic recognition results. Dynamic adjustment can be achieved according to the network signal conditions and semantic recognition results, and gradual switching between audio and video can be achieved, thereby improving the transmission and playback effects of audio and video streams.

[0077] In one embodiment, the pulled remote video stream and the captured local video stream are played alternately in the window of the AR device based on the second guidance semantic recognition result, including: if the second guidance semantic recognition result includes a device-related recognition result in the local video stream, the local video stream and the remote audio stream are played; if during the playback of the local video stream, the second guidance semantic recognition result identified in real time does not include a device-related recognition result, when the continuous playback time is greater than or equal to the second time threshold, switch to playing the remote audio and video stream.

[0078] In this embodiment, the switching between playing the local video stream and the remote audio stream is determined based on the second guiding semantic recognition result. If the second guiding semantic recognition result includes the relevant recognition result of the device in the local video stream, the local video stream and the remote audio stream will be played at the same time. If the real-time recognized second guiding semantic recognition result does not include the device-related recognition result, and the continuous playback time is greater than or equal to the second time threshold, it will switch to playing the remote audio and video stream. The playback of the local video stream and the remote audio and video stream can be dynamically determined based on the real-time semantic recognition result and the playback time to achieve a better user experience and more accurate content presentation.

[0079] In one embodiment, the pulled remote video stream and the captured local video stream are played alternately in the window of the AR device based on the second guiding semantic recognition result, and also includes: if a playback configuration gesture for the remote video stream and the local video stream is obtained, the playback mode of the remote video stream and the local video stream is configured based on the playback configuration gesture, wherein the playback mode includes any one of continuously playing the remote video stream, continuously playing the local video stream, and playing the remote video stream and the local video stream in parallel.

[0080] In this embodiment, when playback configuration gestures for remote video streams and local video streams are obtained, the playback modes of the remote video streams and local video streams will be configured according to these gestures. The playback modes include continuous playback of the remote video stream, continuous playback of the local video stream, and parallel playback of the remote video stream and the local video stream. Based on the above operations, the user can freely switch the playback video stream through gesture operations to achieve a better playback experience and personalized configuration. The playback mode of the remote video stream and the local video stream is dynamically determined according to the user's playback configuration gestures to meet the user's personalized needs for video playback.

[0081] like Figure 5 As shown, according to another embodiment of the present disclosure, a method for adjusting an audio or video call is applied to an AR device, including:

[0082] Step S502: Join the remote collaboration conference and register a broadcast receiver.

[0083] Among them, the broadcast receiver BroadcastReceiver is a global listener that can listen to various broadcasts, including system broadcasts and application broadcasts, to achieve communication monitoring between different components.

[0084] Step S504 : In response to the establishment operation of the audio and video call for online guidance, the quality of the network signal is continuously monitored based on the BroadcastReceiver.

[0085] Furthermore, in one embodiment, in response to the establishment of the audio and video call for online guidance, the quality of the network signal of the AR device accessing the network is continuously monitored, further comprising:

[0086] In step S506 , if the network connection is disconnected based on the BroadcastReceiver monitoring, a second prompt window is popped up in the window of the AR device, and the second prompt window is used to prompt the network disconnection.

[0087] Step S508: Continuously monitor network change events based on BroadcastReceiver to detect whether network disconnection and reconnection are achieved.

[0088] In step S510 , if the monitoring duration continues to be greater than or equal to the third duration threshold and the network reconnection is not achieved, the network reconnection is stopped to remove the AR device from the remote collaborative meeting.

[0089] In this embodiment, in a remote collaborative meeting, the quality of the network signal is continuously monitored by registering a broadcast receiver BroadcastReceiver, and corresponding operations are performed according to the monitoring results. When the network connection is disconnected, a prompt window will pop up in the window to remind the user of the network disconnection. At the same time, the network change event is continuously monitored to detect whether the network disconnection and reconnection are achieved. If the continuous monitoring time reaches the third time threshold and the network disconnection and reconnection are not achieved, the network disconnection and reconnection operation is stopped, and the AR device is removed from the remote collaborative meeting. When the network connection is disconnected, the user can be informed in time and take corresponding measures. At the same time, the network changes are continuously monitored to ensure network stability. If the network connection problem cannot be solved, the network disconnection and reconnection operation is stopped, and the AR device is removed from the remote collaborative meeting to prevent waste of memory resources.

[0090] In one embodiment, if the network connection is disconnected based on the BroadcastReceiver monitoring, the method further includes:

[0091] Get the user identification number of the call target.

[0092] The user identification number is the number of the SIM (Subscriber Identification Module) card.

[0093] Based on the user identification number, the push operation of the local audio stream and the pull operation of the remote audio stream based on the Internet voice communication are switched to the push operation of the local audio stream and the pull operation of the remote audio stream based on the mobile communication.

[0094] Among them, the user identification number includes the mobile phone number. Mobile phone number calls are conducted through the mobile communication network and require the use of mobile phone signals for connection, while WeChat calls are Internet-based voice communications and are conducted through network connections.

[0095] In this embodiment, for situations where a stable signal is required, voice calls can be made using a mobile communication-based method. For situations where the network quality is better, voice calls can be made using an Internet-based method. By obtaining the user identity identification number and switching the call operation mode according to its type, a more stable, reliable and adaptable voice communication experience can be provided.

[0096] like Figure 6 As shown, according to another embodiment of the present disclosure, a method for adjusting an audio or video call is applied to an AR device, including:

[0097] Step S602: Establish an online audio and video call conference with a guidance expert.

[0098] In step S604, BroadcastReceiver is used to monitor network change events to detect whether the network is disconnected. If yes, the process goes to step S606; if no, the process goes to step S614.

[0099] Step S606: If a network disconnection is detected, a prompt is given to the user.

[0100] Step S608, check whether the reconnection is successful within 60 seconds, if "yes", go to step S610, if "no", go to step S612.

[0101] Step S610: re-join the network and continue the audio and video call.

[0102] Step S612: The AR device is removed from the audio and video call conference.

[0103] Step S614, check whether the bandwidth is less than 200 kbps, if "yes", go to step S616, if "no", go to step S620.

[0104] Step S616, prompting that video stream cannot be pushed at present, only audio stream is supported.

[0105] Step S618: Push the local audio stream and pull the remote audio stream.

[0106] Step S620: Pushing the local audio and video stream and pulling the remote audio and video stream.

[0107] For example, after joining a remote collaboration conference, the AR device registers a BroadcastReceiver to monitor changes in network status. If the network connection is normal, the network bandwidth is checked every few seconds. If the network bandwidth is less than 200kbps (the first strength threshold), the push of the local video stream and the pull of the remote video stream are stopped, and only the local audio stream is pushed and the remote audio stream is pulled, thereby reducing the network bandwidth requirements for the call process; if the network bandwidth is greater than 200kbps, the pulling and pushing of the video stream are resumed.

[0108] In addition, if the network connection is detected to be disconnected, the local and server sides save the call status. If the connection is successfully reconnected within 60 seconds, the previous call will be continued. If the connection is not successfully reconnected within 60 seconds, the server side will remove the current AR device from the call.

[0109] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0110] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Therefore, various aspects of the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as "circuits," "modules," or "systems."

[0111] Refer to the following Figure 7 The following describes the apparatus 700 for adjusting audio and video calls according to this embodiment of the present invention. Figure 7 The audio and video call adjustment device 700 shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0112] The apparatus 700 for adjusting audio and video calls is implemented as a hardware module. The components of the apparatus 700 may include, but are not limited to: a monitoring module 702 for continuously monitoring the quality of the network signal connecting the AR device to the network in response to the establishment of an audio and video call for online guidance; a first execution module 704 for executing a push operation for the local audio stream and a pull operation for the remote audio stream if the signal strength of the monitored network signal is less than a first strength threshold; and a second execution module 706 for executing a push operation for the local audio and video stream and a pull operation for the remote audio and video stream if the signal strength of the detected network signal is greater than or equal to a second strength threshold, where the second strength threshold is greater than or equal to the first strength threshold.

[0113] Refer to the following Figure 8 The AR device 800 according to this embodiment of the present invention will be described. Figure 8 The AR device 800 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0114] like Figure 8 As shown, the AR device 800 is implemented as a general-purpose computing device. Components of the AR device 800 may include, but are not limited to, the at least one processing unit 810 described above, the at least one storage unit 820 described above, and a bus 830 connecting various system components (including the storage unit 820 and the processing unit 810).

[0115] The storage unit stores program codes, which can be executed by the processing unit 810, so that the processing unit 810 performs the steps according to various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification. For example, the processing unit 810 can perform the following steps: Figure 1 Steps S202 and S206 shown in , as well as other steps defined in the audio and video call adjustment method of the present disclosure.

[0116] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 8201 and / or a cache memory unit 8202 , and may further include a read-only memory unit (ROM) 8203 .

[0117] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, such program modules 8205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0118] Bus 830 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0119] The AR device 800 can also communicate with one or more external devices 860 (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with the AR device, and / or any device that enables the AR device 800 to communicate with one or more other computing devices (e.g., a router, modem, etc.). This communication can occur via an input / output (I / O) interface 840. Furthermore, the AR device 800 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 850. As shown, the network adapter 850 communicates with other modules of the AR device 800 via a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the AR device, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0120] Through the description of the above embodiments, it will be readily understood by those skilled in the art that the example embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes several instructions for enabling a computing device (such as a personal computer, server, terminal device, or network device) to execute the methods according to the embodiments of the present disclosure.

[0121] In exemplary embodiments of the present disclosure, a computer-readable storage medium is also provided, on which is stored a program product capable of implementing the methods described above. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product comprising program code that, when executed on a terminal device, causes the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the "Exemplary Methods" section above.

[0122] According to an embodiment of the present invention, a program product for implementing the above-mentioned method can be a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0123] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0124] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0125] Program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0126] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0127] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0128] Through the description of the above embodiments, it will be readily understood by those skilled in the art that the example embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes several instructions for enabling a computing device (such as a personal computer, server, mobile terminal, or network device) to execute the methods according to the embodiments of the present disclosure.

[0129] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

Claims

1. A method for adjusting an audio or video call, characterized in that: Applied to AR devices, including: In response to establishing an audio or video call for online guidance, continuously monitoring a quality of a network signal of the AR device accessing the network; If the signal strength of the network signal is detected to be less than a first strength threshold, the local audio stream is pushed and the remote audio stream is pulled; When it is monitored that the signal strength of the network signal rises to greater than or equal to the first strength threshold and is less than the second strength threshold, the push of the local video stream is started; when it is monitored that the signal strength of the network signal rises to greater than or equal to the second strength threshold and the duration is greater than or equal to the first duration threshold, the pulling of the remote video stream is started; and the remote semantic recognition model is called to perform real-time recognition of the pulled remote audio stream to obtain a second guiding semantic recognition result; based on the second guiding semantic recognition result, the pulled remote video stream and the collected local video stream are alternately played in the window of the AR device, including: if the second guiding semantic recognition result includes a device-related recognition result in the local video stream, the local video stream and the remote audio stream are played; if during the playback of the local video stream, the second guiding semantic recognition result identified in real time does not include the device-related recognition result, when the continuous playback duration is greater than or equal to the second duration threshold, switch to playing the remote audio and video stream, and the second strength threshold is greater than or equal to the first strength threshold.

2. The method for adjusting audio and video calls according to claim 1, wherein: If it is monitored that the signal strength of the network signal is less than a first strength threshold, executing a push operation of the local audio stream and a pull operation of the remote audio stream, including: During the audio and video call, it is monitored that the signal strength of the network signal drops below the first strength threshold, and a first prompt window pops up in the window of the AR device, the first prompt window being used to prompt to stop pulling the remote video stream; In response to a confirmation operation on the first prompt window, the pulled remote audio stream is played.

3. The method for adjusting audio and video calls according to claim 1, wherein: If the signal strength of the network signal is detected to be less than a first strength threshold, the local audio stream is pushed and the remote audio stream is pulled, further comprising: Calling an offline semantic recognition model to recognize the pulled remote audio stream to obtain a first guided semantic recognition result; Calling an offline virtual modeling model to model the first guidance semantic recognition result, and generating an augmented reality guidance image based on the modeling result and the image collected in real time by the AR device; The augmented reality guidance image is displayed in a window of the AR device.

4. The method for adjusting audio and video calls according to claim 3, wherein: Calling an offline virtual modeling model to model the first guidance semantic recognition result, and generating an augmented reality guidance image based on the modeling result and the image collected in real time by the AR device, including: Determining a modeling scene based on the first guiding semantic recognition result and the real-time acquired image; determining guidance intention and guidance needs based on the first guidance semantic recognition result; Performing three-dimensional modeling in the modeling scene based on the guidance intention and the guidance requirement to obtain the modeling result; The modeling result is mapped to the real-time collected image to obtain the augmented reality guidance image.

5. The method for adjusting audio and video calls according to claim 1, wherein: Alternating between playing the pulled remote video stream and the captured local video stream in the window of the AR device based on the second guidance semantic recognition result, further comprising: If a playback configuration gesture for the remote video stream and the local video stream is obtained, the playback mode of the remote video stream and the local video stream is configured based on the playback configuration gesture, The playback mode includes any one of continuously playing the remote video stream, continuously playing the local video stream, and playing the remote video stream and the local video stream in parallel.

6. The method for adjusting an audio or video call according to any one of claims 1 to 5, characterized in that: Before continuously monitoring the quality of a network signal of the AR device accessing the network in response to establishing an audio or video call for online guidance, the method further includes: Join the remote collaboration conference and register a broadcast receiver BroadcastReceiver to continuously monitor the quality of the network signal based on the BroadcastReceiver.

7. The method for adjusting audio and video calls according to claim 6, wherein: In response to the establishment of an audio or video call for online guidance, the method further includes continuously monitoring the quality of a network signal of the AR device accessing the network: If the BroadcastReceiver detects that the network connection is disconnected, a second prompt window pops up in the window of the AR device, where the second prompt window is used to prompt the network disconnection; Based on the BroadcastReceiver, the network change events are continuously monitored to detect whether the network is disconnected and reconnected; If the monitoring duration continues to be greater than or equal to the third duration threshold and network reconnection is not achieved, the network reconnection is stopped to remove the AR device from the remote collaboration meeting.

8. The method for adjusting audio and video calls according to claim 7, wherein: If the network connection is disconnected based on the BroadcastReceiver monitoring, the method further includes: Obtain the user identification number of the call target; Based on the user identification number, the push operation of the local audio stream and the pull operation of the remote audio stream based on Internet voice communication are switched to the push operation of the local audio stream and the pull operation of the remote audio stream based on mobile communication.

9. An adjustment device for audio and video calls, characterized in that: Applied to AR devices, including: a monitoring module, configured to continuously monitor the quality of a network signal of the AR device accessing the network in response to an operation of establishing an audio or video call for online guidance; A first execution module is configured to execute a push operation of a local audio stream and a pull operation of a remote audio stream if it is detected that the signal strength of the network signal is less than a first strength threshold; The second execution module is used to monitor that the signal strength of the network signal rises to greater than or equal to the first strength threshold and less than the second strength threshold, start pushing the local video stream, monitor that the signal strength of the network signal rises to greater than or equal to the second strength threshold and lasts for a duration greater than or equal to the first duration threshold, start pulling the remote video stream; and call the remote semantic recognition model to perform real-time recognition on the pulled remote audio stream to obtain a second guidance semantic recognition result; based on the second guidance semantic recognition result, alternately play the pulled remote video stream and the collected local video stream in the window of the AR device, including: if the second guidance semantic recognition result includes a device-related recognition result in the local video stream, play the local video stream and the remote audio stream; if during the playback of the local video stream, the second guidance semantic recognition result identified in real time does not include the device-related recognition result, switch to playing the remote audio and video stream when the continuous playback duration is greater than or equal to the second duration threshold, and the second strength threshold is greater than or equal to the first strength threshold.

10. An AR device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the method for adjusting an audio or video call according to any one of claims 1 to 8 by executing the executable instructions.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for adjusting the audio and video call according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Augmented-reality remote guidance method and system

    CN107645651A